Job Summary
We are seeking a Senior AI/ML Engineer to design, deploy, operate, and optimize production-grade AI and Machine Learning solutions at scale. This role focuses on LLM/Generative AI applications, MLOps, cloud platforms, and software engineering excellence, ensuring AI systems are reliable, observable, cost-efficient, secure, and production-ready. You will work closely with data scientists, platform engineers, and product teams to build and support RAG, agentic AI, and ML solutions across the full development lifecycle.
Key Responsibilities
-
Design, build, deploy, and support production AI/ML applications, including LLM-powered, RAG, and agent-based solutions.
-
Develop robust evaluation, testing, observability, and monitoring frameworks for AI systems.
-
Implement and maintain CI/CD pipelines for ML and GenAI workloads.
-
Monitor and optimize model performance, latency, reliability, cost, and operational health.
-
Build and manage AI infrastructure using Infrastructure-as-Code and cloud-native services.
-
Troubleshoot production issues across models, data pipelines, retrieval systems, agents, and integrations.
-
Collaborate with engineering, data science, and platform teams to deliver scalable AI solutions.
-
Drive engineering best practices including code reviews, testing, version control, and documentation.
-
Implement governance, guardrails, tracing, logging, and monitoring to ensure responsible AI deployment.
-
Mentor junior engineers and contribute to technical leadership within the team.
General Qualifications
-
Bachelor's degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or a related discipline.
-
5+ years of experience in Software Engineering, Machine Learning Engineering, MLOps, or AI Engineering roles.
-
Experience designing and supporting production systems in cloud environments.
-
Strong communication, stakeholder management, problem-solving, and mentoring capabilities.
Mandatory Requirements
-
Strong Python programming expertise with experience building and maintaining production-grade applications.
-
Solid software engineering fundamentals, including testing, code reviews, Git/version control, and maintainable code practices.
-
Proven experience delivering and supporting LLM/Generative AI applications in production.
-
Hands-on experience with RAG architectures, AI agents, and/or fine-tuned LLMs.
-
Strong understanding of LLM evaluation, guardrails, observability, latency optimization, and cost management.
-
Experience implementing Infrastructure-as-Code using Terraform or equivalent IaC tools.
-
Production experience on at least one major cloud platform (Azure, AWS, or GCP).
-
Experience with Databricks or comparable lakehouse/MLOps platforms.
-
Hands-on experience with Docker and Kubernetes for containerized AI/ML workloads.
-
Experience building and supporting CI/CD pipelines for ML and AI deployments.
-
Strong knowledge of LLM tracing, logging, telemetry, and observability frameworks.
-
Experience implementing monitoring solutions using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, CloudWatch, or Azure Monitor.
Nice-to-Have Skills
-
Experience with TensorRT-LLM, FlashAttention, or other LLM inference optimization technologies.
-
Knowledge of tensor parallelism and pipeline parallelism for large-scale model deployment.
-
Experience with AI orchestration frameworks such as LangGraph, LlamaIndex, AutoGen, or Semantic Kernel.
-
Familiarity with Model Context Protocol (MCP).
-
Experience with LLMOps tooling, including LiteLLM, model routing/fallback strategies, prompt/version management, and token cost monitoring.
-
Experience with workflow orchestration platforms such as Airflow, Dagster, Kubeflow, or Argo.
-
Relevant cloud, AI, Kubernetes, or Databricks certifications.
-
Previous experience in consulting, professional services, or client-facing delivery environments.