AI Jobs Map

MetAntz · Bengaluru, Karnataka, India

SRE / Cloud Engineer

Hybridseniorfull timePosted 7 days ago
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

srekubernetesobservabilityawsazuregcpincident-responsemicroserviceseksterraformansiblepythonbashprometheusgrafanaopentelemetrydatadognetwork-securityidentity-and-access-managementlinux

Cloud & Infrastructure Site Reliability Engineer (SRE)

Role Overview

We are looking for a Cloud & Infrastructure Site Reliability Engineer (SRE) to build, operate, automate, and continuously improve highly available cloud-native infrastructure powering SaaS platform. You will work across public cloud, Kubernetes, compute, storage, networking, observability, and infrastructure automation - with a strong focus on reliability, scalability, security, and reducing manual operations. The ideal candidate combines strong infrastructure fundamentals with a software engineering and automation mindset - someone who proactively identifies operational issues and engineers permanent solutions rather than relying on repetitive manual processes.

Key Responsibilities

- Cloud & Infrastructure Operations: Design, deploy, operate, and troubleshoot cloud-native infrastructure across AWS, Azure, and GCP - including compute, storage, networking, and containers. Handle provisioning, upgrades, patching, and capacity management; resolve connectivity, DNS, and load-balancing issues.

- Site Reliability Engineering: Define and improve SLIs, SLOs, monitor for reliability risks, and drive incident response, root-cause analysis, and post-incident reviews. Automate away recurring operational work and improve system resilience through redundancy and automated recovery.

- Kubernetes & Microservices: Deploy, manage, and troubleshoot containerized microservices on Kubernetes (EKS, AKS, GKE), including ingress, storage, secrets, and cluster health.

- Infrastructure as Code & Automation: Automate provisioning and operational workflows with Terraform, Ansible, Python, and Bash; build reusable IaC modules integrated into CI, CD.

- Monitoring & Observability: Implement monitoring, logging, tracing, and alerting (Prometheus, Grafana, OpenTelemetry, Datadog, CloudWatch, etc.) with meaningful dashboards and low-noise alerting.

- Networking & Security: Troubleshoot cloud networking (VPC, VNet, DNS, VPN, firewalls, load balancers); apply IAM, RBAC, secrets, and least-privilege security practices.

- Incident & Production Support: Participate in on-call rotations, own incidents end-to-end, and maintain runbooks and automated remediation procedures.

- Cost & Performance Optimization: Identify underutilized resources and support FinOps initiatives (rightsizing, scheduling, lifecycle management) while balancing cost with reliability.

Required Skills

- Strong hands-on experience with at least one major public cloud - AWS, Azure, or GCP.

- Solid understanding of Linux administration and infrastructure fundamentals.

- Experience with Kubernetes and containerized microservices.

- Experience with Infrastructure as Code (Terraform and, or Ansible).

- Strong automation, scripting skills in Python and Bash (PowerShell a plus).

- Good understanding of networking, DNS, load balancing, firewalls, and cloud routing.

- Experience with monitoring, logging, alerting, and observability platforms.

- Understanding of CI, CD, Git, APIs, and modern DevOps practices.

- Experience troubleshooting distributed production environments.

- Understanding of security, IAM, RBAC, secrets, certificates, and infrastructure security.

- Familiarity with database operations a plus.

Preferred Skills

Experience & Qualifications

- Experience operating hybrid-cloud or multi-cloud environments.

- Experience with VMware, Nutanix, OpenShift, or similar private-cloud technologies.

- Experience building automated remediation or self-healing workflows.

- Familiarity with AIOps, event correlation, anomaly detection, and AI-assisted operations.

- Understanding of FinOps and cloud cost optimization.

- Experience with ITSM platforms such as ServiceNow.

- Cloud, Kubernetes, Linux, or Terraform certifications (AWS, Azure, GCP) are a strong plus. SRE Mindset

- Automate before repeating, and engineer permanent fixes instead of resolving the same incident twice.

- Treat infrastructure and operational configuration as code.

- Design for failure and recovery, and measure reliability using data.

- Build reusable automation and operational standards.

- Reduce operational toil and unnecessary tickets.

- Continuously improve reliability, performance, security, and cost efficiency.

- Bachelor's or Master's degree in Computer Science or a related field, with minimum of 70%.

- SRE Engineer - 2–6 years of relevant Cloud, Infrastructure, DevOps, SRE experience.

- Cloud certifications (AWS, Azure, or GCP) are a strong plus. Roles and responsibilities

- Design, deploy, operate, and troubleshoot cloud-native infrastructure across AWS, Azure, and GCP. Handle provisioning, upgrades, patching, and capacity management.

- Define and improve SLIs, SLOs, monitor for reliability risks, and drive incident response.

- Automate away recurring operational work.

- Deploy, manage, and troubleshoot containerized microservices on Kubernetes.

- Automate provisioning and operational workflows with Terraform, Ansible, Python, and Bash.

- Implement monitoring, logging, tracing, and alerting.

- Troubleshoot cloud networking and apply security practices.

- Participate in on-call rotations and own incidents end-to-end.

- Identify underutilized resources and support FinOps initiatives. Experience and education

- Bachelor's or Master's degree in Computer Science or a related field, with minimum of 70% Location Remote

More jobs at MetAntz