Kubernetes Reliability Engineer
Vienna, Austria — Hybrid
Cloud Infrastructure | Kubernetes | Site Reliability Engineering | Platform Engineering
Our client, a growing Cloud Infrastructure company based in Vienna, is looking for a Kubernetes Reliability Engineer to improve the resilience, observability, and operability of large-scale Kubernetes environments.
You'll work at the intersection of Kubernetes Engineering, SRE, and Platform Engineering, building the automation and observability systems that keep critical cloud infrastructure reliable in production.
What You'll Work On
• Improve reliability and resilience across production Kubernetes clusters
• Build platform and operational tooling in Go
• Define and implement SLIs, SLOs, and reliability standards
• Develop observability using Prometheus and OpenTelemetry
• Automate infrastructure provisioning and configuration with Terraform
• Investigate complex Kubernetes and Linux production issues
• Improve alerting, incident detection, and operational visibility
• Automate repetitive operational and recovery workflows
• Analyse capacity, performance, and infrastructure bottlenecks
• Work with engineering teams to improve application reliability on Kubernetes
Core Skills
• 4+ years in SRE, Platform Engineering, Kubernetes Engineering, Cloud Infrastructure, or similar roles
• Strong Kubernetes experience in production environments
• Go
• Prometheus
• OpenTelemetry
• Terraform
• Linux
• Strong understanding of distributed systems and production reliability
Nice to Have
Grafana
Kubernetes Operators / Controllers
Helm
Argo CD / GitOps
AWS / GCP / Azure
eBPF
Service meshes
Incident management and postmortems
Capacity planning and performance engineering
Experience operating multi-cluster or large-scale Kubernetes environments