Key Responsibilities
- Design, implement, and maintain CI/CD pipelines to automate build, test, and deployment of applications across environments.
- Deploy and operate containerised applications on Kubernetes (EKS/AKS/GKE), including rollout strategies, scaling, and upgrades.
- Set up and maintain monitoring, logging, and alerting for applications and platforms, and respond to incidents and performance issues.
- Work closely with development, QA, and operations teams to enforce DevOps best practices, including versioning, branching, and deployment standards.
- Implement security, compliance, and access controls in CI/CD pipelines and Kubernetes/public cloud environments.
- Scheduled upgrade of application packages and tools
- Own and improve release management processes (change management, approvals, deployment windows, rollback procedures).
- Document pipelines, deployment processes, and runbooks for operations and support teams.
Required Skills
- 6–10 years of experience as a DevOps Engineer or similar role in a production environment.
- Solid understanding and strong hands-on experience of CI/CD concepts and tools such as GitHub Actions, Argo CD, Jenkins, or similar.
- Strong hands‑on experience with containerisation (Docker) and deploying applications to Kubernetes (EKS/AKS/GKE or equivalent).
- Practical experience with at least one public cloud provider (AWS, Azure, or GCP) and its core services (compute, networking, storage, IAM).
- Experience managing release cycles, coordinating with multiple teams, and executing deployments with minimal downtime.
- Proficiency in scripting (Bash, Python, or PowerShell) for automation of build, deployment, and operational tasks.
- Knowledge of monitoring and logging tools (Prometheus/Grafana, ELK/EFK, CloudWatch, Azure Monitor, etc.).
- Strong communication, documentation, and problem‑solving skills; comfortable working with cross‑functional teams.
Preferred / Nice-to-have Skills
- Experience with blue‑green, canary, and rolling deployment strategies for Kubernetes applications.
- Experience with incident management, post‑incident reviews, and reliability/SRE practices.
- Relevant certifications (e.g., Certified Kubernetes Administrator, AWS/Azure/GCP associate/professional levels).