Key Responsibilities
- Design, implement, and maintain highly available, scalable, and reliable infrastructure.
- Manage, monitor, and troubleshoot Kubernetes clusters, workloads, deployments, and cluster-level issues.
- Design and maintain CI/CD pipelines for reliable application build, deployment, release, and rollback.
- Implement deployment strategies including Canary, Blue-Green, and Rolling Deployments.
- Troubleshoot pipeline failures, deployment issues, configuration problems, application failures, and service communication issues.
- Develop automation using Python, Bash, Go, or similar scripting languages to reduce operational toil.
- Monitor system availability, performance, capacity, and reliability using modern observability and monitoring tools.
- Troubleshoot issues across application, infrastructure, Kubernetes, cloud, and networking layers.
- Participate in incident management, root-cause analysis (RCA), and preventive/corrective actions.
- Collaborate with Development, DevOps, Cloud, Platform, and Security teams to improve reliability and operational processes.
Mandatory Skills
- Kubernetes (K8s) - strong hands-on administration and troubleshooting.
- Cloud: AWS, Azure, or GCP.
- CI/CD: Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, Argo CD, or similar.
- Scripting/Programming: Python, Bash, or Go.
- Networking: TCP/IP, DNS, HTTP/HTTPS, routing, subnets/CIDR, load balancing, firewalls/security groups, and service connectivity.
- Containerization: Docker and Kubernetes.
- Strong production troubleshooting and incident management experience.
- Strong experience in automation and operational engineering.
Kubernetes Expertise
Hands-on Experience With
- Kubernetes architecture and control-plane components
- API Server, Scheduler, Controller Manager, etcd, and kubelet
- Pods, Deployments, StatefulSets, DaemonSets, and Jobs
- Services, Ingress, ConfigMaps, and Secrets
- Resource requests/limits and workload scheduling
- Cluster networking and service discovery
- Troubleshooting CrashLoopBackOff, ImagePullBackOff, scheduling, rollout, node health, connectivity, and resource issues
Cloud & Infrastructure
- Hands-on experience with AWS / Azure / GCP infrastructure and native services.
- Knowledge of compute, storage, networking, IAM, load balancing, security, monitoring, and logging.
- Experience with Infrastructure as Code (IaC) using Terraform, CloudFormation, ARM/Bicep, or similar tools.
Observability & Reliability
- Experience with metrics, logs, and traces.
- Exposure to Prometheus, Grafana, ELK/Elastic, Splunk, Datadog, New Relic, or cloud-native monitoring tools.
- Understanding of SLI, SLO, SLA, alerting, incident response, and RCA.
Preferred Qualifications
- Experience supporting production-grade distributed systems and microservices.
- Experience in 24x7 production support / SRE environments.
- Strong Linux/Unix fundamentals.
- Experience with Git and Git-based workflows.
- Experience with Terraform/IaC and configuration management.
- Strong analytical, troubleshooting, communication, and problem-solving skills.
Experience
3+ years of experience in SRE, DevOps, Cloud Infrastructure, Platform Engineering, or related roles.
Mandatory Core Skills
Kubernetes + AWS/Azure/GCP + CI/CD + Python/Bash/Go + Networking + Containerization + Automation + Production Troubleshooting