SRE / Incident Response Engineer
We are seeking a hands-on SRE / Incident Response Engineer to support critical production services in a fast-paced environment. This role is heavily focused on major incident response, service restoration, operational support, and automation.
Key Responsibilities
- Lead and coordinate major incidents, driving issues through to resolution.
- Act as the technical point of coordination during high-severity service outages.
- Collaborate with service owners, engineering teams, and support functions to restore services quickly.
- Perform root cause analysis and contribute to continuous service improvement.
- Support day-to-day operations within a NOC/SRE environment.
- Develop and maintain automation to reduce manual effort and improve reliability.
Essential Skills
- Strong Major Incident Management / Incident Response experience.
- Experience working within a NOC, Operations, or SRE environment.
- Solid Linux administration and troubleshooting skills.
- Automation and scripting experience using Python or similar.
- Experience with Git and Jenkins.
- Strong networking fundamentals and troubleshooting capabilities.
Desirable Skills
- Kubernetes (K8s).
- Cloud platforms: AWS and/or GCP.
- Monitoring, observability, and platform reliability tooling.
Ideal Candidate
A hands-on operational engineer with proven experience managing business-critical incidents, supporting production environments, and using automation to improve service reliability and operational efficiency.