Site Reliability Engineer (SRE, Terraform, AWS, Dynatrace)
Optomi, in partnership with a Fortune 500 digital platform leader, is seeking a Site Reliability Engineer to join their Digital SRE team! In this role, the Site Reliability Engineer will focus on improving the reliability, scalability, and performance of critical customer-facing applications. The ideal candidate will have a strong SRE background, experience with Terraform and AWS, and a passion for observability, automation, and cross-functional collaboration.
What the Right Candidate Will Enjoy!
- A high-impact role within a business-critical Digital SRE team
- Hands-on work with modern DevOps/SRE tools including Terraform, GitLab, and Dynatrace
- Collaborative environment interacting with application teams, sustain partners, and leadership
- Exposure to large-scale systems architecture and service level objective (SLO) management
- Opportunity to contribute to automation initiatives and performance monitoring best practices
Experience of the Right Candidate:
- 3+ years of experience in a Site Reliability Engineer, DevOps, or similar infrastructure-focused role
- Strong experience with Terraform and GitLab
- Hands-on experience with observability and monitoring tools, preferably Dynatrace and Splunk
- Background in AWS or similar public cloud platforms
- Experience working with or supporting APIs and/or customer-facing front-end applications
- Proven ability to lead incident response efforts and create effective runbooks and postmortems
- Familiarity with service level indicators (SLIs), error budgets, and implementing SLO frameworks
- Effective communicator and collaborator who thrives in a cross-functional environment
Responsibilities of the Right Candidate:
- Monitor and manage production environments with a focus on reliability and performance
- Provide on-call support on a rotating basis, ensuring rapid response to production issues
- Implement and maintain observability standards across systems using tools like Dynatrace
- Build automation to eliminate manual operational work and streamline triage processes
- Lead service readiness efforts including documentation, health scoring, and SLO adoption
- Collaborate with sustain partners and application teams to resolve issues and drive stability
- Participate in blameless postmortems and promote continuous improvement practices
- Support cross-functional initiatives to improve platform performance and uptime