AI Jobs Map

CISO Aid · Malvern, PA

Senior SRE / AWS DevOps & MLOps Engineer :: Full Time :: Malvern PA(Onsite)

seniorfull timePosted today
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sreawsdevopsmlopsmachine-learninggenerative-aikuberneteseksdockerci/cdobservabilityllmincident-responseartificial-intelligenceecsec2github-actionsgitlab-ciopentelemetryprometheus

HI

Hope you are doing well!

I have an urgent requirement with one of my clients. Please find the job details below and forward me your updated resume along with your contact details

Senior SRE / AWS DevOps & MLOps Engineer

Location :: Malvern PA( Onsite)

Job Summary

We are seeking an experienced SRE / AWS DevOps & MLOps Engineer to design, build, automate, and operate highly reliable cloud-native platforms supporting applications, machine learning workloads, and Generative AI solutions.

The ideal candidate will have strong hands-on experience with AWS, Kubernetes/EKS, Docker, CI/CD, Infrastructure as Code, observability, MLOps, and AI/LLM monitoring. The role will focus on improving platform reliability, deployment automation, system performance, incident response, and observability across traditional applications and AI/ML workloads.

Experience with Arize AI and AI/ML/LLM observability is highly preferred.

Key Responsibilities

- Design, implement, and maintain highly available and scalable AWS cloud infrastructure.

- Support production environments using EKS, ECS, EC2, Docker, and Kubernetes.

- Build and maintain reliable CI/CD pipelines using GitHub Actions and GitLab CI/CD.

- Automate infrastructure provisioning and configuration using Infrastructure as Code (IaC).

- Develop and maintain AWS Cloud Formation templates and infrastructure automation.

- Implement SRE practices around SLIs, SLOs, SLAs, reliability, availability, and performance.

- Establish monitoring and observability for applications, infrastructure, ML models, and GenAI/LLM workloads.

- Implement AI Observability, ML Observability, and LLM Observability using tools such as Arize AI, OpenTelemetry, Prometheus, Grafana, CloudWatch, and ELK Stack.

- Implement model monitoring, model performance tracking, and model drift detection.

- Support MLOps and GenAI Operations, including monitoring and operationalization of machine learning and AI workloads.

- Work with Amazon SageMaker and related AWS ML services.

- Design and implement monitoring for model quality, latency, throughput, errors, data quality, and inference performance.

- Develop automation and operational tooling using Python and Bash.

- Implement DevSecOps practices across infrastructure and deployment pipelines.

- Apply AWS security best practices including IAM, least privilege, encryption, secrets management, network security, and secure CI/CD.

- Lead incident response, troubleshooting, and Root Cause Analysis (RCA) for production issues.

- Identify reliability risks and implement proactive remediation and performance optimization.

- Establish dashboards, alerts, logging, tracing, and actionable operational metrics.

- Collaborate with development, data science, ML engineering, security, and platform teams.

- Contribute to Platform Engineering initiatives by developing reusable infrastructure, deployment, and observability capabilities.

- Participate in on-call rotations and provide production support for critical systems.

Required Skills

AWS / Cloud

- Strong hands-on experience with AWS.

- Experience with:

- EKS

- ECS

- EC2

- CloudWatch

- SageMaker

- IAM

- AWS networking and security services

- Experience supporting production AWS environments.

Kubernetes & Containers

- Strong experience with Kubernetes and Amazon EKS.

- Hands-on experience deploying and operating containerized applications.

- Strong knowledge of Docker.

- Experience with Kubernetes troubleshooting, scaling, networking, deployments, and production operations.

More jobs at CISO Aid