AI Jobs Map

Umanist NA · Pune Division, Maharashtra, India

Senior Site Reliability Engineer (SRE) Engineer

seniorfull timePosted yesterday
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sredevopsobservabilitykubernetesazureawsgcpec2s3ekshelmterraformgitgithubgitlabci/cdpythonbashopentelemetryprometheus

Location: Viman Nagar, Pune – Work From Office

Experience Overall(must have): 8 Years

CTC: Up to ₹25 LPA

Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)

Working Hours: 3:00 PM – 12:00 AM, Monday to Friday

On-Call: 24/7 Production Support – On-Call Rotation Required

Employment Type: Full-Time

About The Role

We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.

The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.

Must-Have Skills & Experience1. SRE & Production Operations

- Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.

- Hands-on experience with 24/7 production support and on-call operations.

- Strong experience in incident management, troubleshooting, RCA, and post-mortems.

- Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.

- Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.

- Ability to improve system availability, performance, scalability, and operational reliability.

- Cloud & Infrastructure

- Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.

- Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.

- Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.

- Experience with:

- Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS

- AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS

- GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring

- Kubernetes & Containerization

- Strong hands-on experience with Kubernetes and containerized workloads.

- Experience with AKS / EKS / GKE or equivalent Kubernetes environments.

- Hands-on experience with Helm deployments.

- Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.

- Infrastructure as Code & DevOps

- Hands-on experience with Terraform / Infrastructure as Code (IaC).

- Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.

- Strong DevOps automation and CI/CD understanding.

- Strong scripting skills in Python and/or Bash.

- Monitoring & Observability

- Strong hands-on experience with OpenTelemetry.

- Experience with monitoring and observability tools such as:

- Prometheus

- Grafana

- Datadog

- Azure Monitor

- AWS CloudWatch

- GCP Cloud Monitoring

- Strong understanding of metrics, logs, distributed tracing, and alerting.

- Experience implementing monitoring based on Golden Signals:

- Latency

- Traffic

- Errors

- Saturation

- Ability to develop symptom-based, user-impact-focused alerting.

- Linux & Networking

- Strong knowledge of Linux system administration.

- Strong understanding of:

- DNS

- TCP/IP

- Load Balancing

- SSL/TLS

- Networking fundamentals

- Experience supporting highly available production environments.

- Incident & Reliability Engineering

- Ability to rapidly diagnose and resolve high-severity production incidents.

- Experience driving MTTR reduction.

- Strong debugging and analytical problem-solving skills.

- Ability to identify recurring issues and implement permanent corrective/preventive solutions.

Good-to-Have Skills

- Experience working across Azure + AWS + GCP in a multi-cloud environment.

- Knowledge of Go (Golang).

- Experience with OpenSearch / ELK Stack.

- Experience supporting AI/ML workloads in production.

- Exposure to Azure AI Services and Azure AI Foundry.

- Experience supporting RAG (Retrieval-Augmented Generation) workloads.

- Experience designing infrastructure for AI/ML platforms.

- Experience building enterprise-wide OpenTelemetry observability frameworks.

- Strong understanding of distributed systems architecture.

- Exposure to advanced cloud-native architectures and reliability patterns.

- Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.

Key ResponsibilitiesProduction & Incident Management

- Participate in the 24/7 on-call rotation.

- Diagnose, mitigate, and resolve production incidents.

- Lead RCA and post-incident reviews.

- Implement corrective and preventive actions.

- Continuously improve MTTR and production stability.

Reliability Engineering

- Define and improve SLIs, SLOs, SLAs, and Error Budgets.

- Identify and eliminate operational toil.

- Conduct reliability and capacity reviews.

- Improve redundancy, failover, disaster recovery, and system resilience.

Cloud & Infrastructure

- Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.

- Manage Kubernetes clusters and containerized applications.

- Implement and maintain Infrastructure as Code using Terraform.

- Support CI/CD and Git-based development workflows.

Observability & Performance

- Build and improve monitoring, logging, metrics, and tracing.

- Implement OpenTelemetry and distributed tracing.

- Establish Golden Signals-based monitoring and alerting.

- Identify and resolve infrastructure and application performance bottlenecks.

Security

- Implement cloud security best practices around IAM, network segmentation, and secrets management.

- Support vulnerability remediation and compliance initiatives.

- Collaborate with Development, Security, and Infrastructure teams.

Ideal Candidate

We Are Looking For Someone With

- Strong SRE mindset and production ownership.

- Excellent troubleshooting and incident-management skills.

- Hands-on expertise in Cloud + Kubernetes + Terraform + Observability.

- Strong understanding of OpenTelemetry and Golden Signals.

- Experience working in highly available, production-critical environments.

- Ability to remain calm and make effective decisions during critical incidents.

- Strong communication and cross-functional collaboration skills.

- Passion for automation, scalability, reliability, and continuous improvement.

Important Hiring Criteria

Must Be

- 7+ years relevant experience

- Immediate joiner

- Willing to work from office in Viman Nagar, Pune

- Comfortable with 3:00 PM – 12:00 AM shift

- Comfortable with 24/7 on-call rotation

- Strong hands-on SRE/DevOps experience

- Strong Cloud + Kubernetes + Observability experience

- Strong production incident management experience

Good To Have

- Multi-cloud: Azure + AWS + GCP

- OpenTelemetry

- AI/ML or RAG production workloads

- Azure AI / AI Foundry

- Go

- OpenSearch / ELK

- Distributed systems

Skills: golang,troubleshooting,error budgets,cloud infrastructure,rca,slo,golden signals,sla,ms azure,production engineering,aws,toil reduction,capacity planning,sre & production operations,opentelemetry,gitlab,github,24/7 production support,terraform,python,mttr,incident management,linux system,kubernetes & containerization,iac,sli,devops,sre,gcp

More jobs at Umanist NA