HI
Hope you are doing well!
I have an urgent requirement with one of my clients. Please find the job details below and forward me your updated resume along with your contact details
Senior SRE / AWS DevOps & MLOps Engineer
Location :: Malvern PA( Onsite)
Job Summary
We are seeking an experienced SRE / AWS DevOps & MLOps Engineer to design, build, automate, and operate highly reliable cloud-native platforms supporting applications, machine learning workloads, and Generative AI solutions.
The ideal candidate will have strong hands-on experience with AWS, Kubernetes/EKS, Docker, CI/CD, Infrastructure as Code, observability, MLOps, and AI/LLM monitoring. The role will focus on improving platform reliability, deployment automation, system performance, incident response, and observability across traditional applications and AI/ML workloads.
Experience with Arize AI and AI/ML/LLM observability is highly preferred.
Key Responsibilities
- Design, implement, and maintain highly available and scalable AWS cloud infrastructure.
- Support production environments using EKS, ECS, EC2, Docker, and Kubernetes.
- Build and maintain reliable CI/CD pipelines using GitHub Actions and GitLab CI/CD.
- Automate infrastructure provisioning and configuration using Infrastructure as Code (IaC).
- Develop and maintain AWS Cloud Formation templates and infrastructure automation.
- Implement SRE practices around SLIs, SLOs, SLAs, reliability, availability, and performance.
- Establish monitoring and observability for applications, infrastructure, ML models, and GenAI/LLM workloads.
- Implement AI Observability, ML Observability, and LLM Observability using tools such as Arize AI, OpenTelemetry, Prometheus, Grafana, CloudWatch, and ELK Stack.
- Implement model monitoring, model performance tracking, and model drift detection.
- Support MLOps and GenAI Operations, including monitoring and operationalization of machine learning and AI workloads.
- Work with Amazon SageMaker and related AWS ML services.
- Design and implement monitoring for model quality, latency, throughput, errors, data quality, and inference performance.
- Develop automation and operational tooling using Python and Bash.
- Implement DevSecOps practices across infrastructure and deployment pipelines.
- Apply AWS security best practices including IAM, least privilege, encryption, secrets management, network security, and secure CI/CD.
- Lead incident response, troubleshooting, and Root Cause Analysis (RCA) for production issues.
- Identify reliability risks and implement proactive remediation and performance optimization.
- Establish dashboards, alerts, logging, tracing, and actionable operational metrics.
- Collaborate with development, data science, ML engineering, security, and platform teams.
- Contribute to Platform Engineering initiatives by developing reusable infrastructure, deployment, and observability capabilities.
- Participate in on-call rotations and provide production support for critical systems.
Required Skills
AWS / Cloud
- Strong hands-on experience with AWS.
- Experience with:
- EKS
- ECS
- EC2
- CloudWatch
- SageMaker
- IAM
- AWS networking and security services
- Experience supporting production AWS environments.
Kubernetes & Containers
- Strong experience with Kubernetes and Amazon EKS.
- Hands-on experience deploying and operating containerized applications.
- Strong knowledge of Docker.
- Experience with Kubernetes troubleshooting, scaling, networking, deployments, and production operations.