Manager, Site Reliability Engineering (SRE)
Location: Remote – US Based
Employment Type: Full-Time
NationsBenefits is seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across production platforms.
This is a player-coach leadership role combining people management with hands-on technical leadership. You will mentor and develop Site Reliability Engineers while actively participating in major incident response, reliability initiatives, automation, observability, and operational reviews.
You will also collaborate closely with SRE leadership in India as part of our global follow-the-sun support model.
Key Responsibilities
Team Leadership
- Lead, mentor, and develop a US-based team of Site Reliability Engineers.
- Conduct 1:1s, performance reviews, and career development discussions.
- Support hiring, onboarding, retention, and team growth.
- Promote a culture of ownership, blameless postmortems, and continuous improvement.
Incident Management & Operational Excellence
- Lead production operations, incident triage, escalation, and resolution.
- Serve as an escalation point and incident commander for major production incidents.
- Drive root cause analysis and problem-management processes.
- Participate in PagerDuty/on-call escalation for critical production issues.
- Monitor and report SLAs, SLOs, availability, MTTR, and incident trends.
Reliability, Observability & Automation
- Improve system reliability, resilience, and observability using Datadog or similar tools.
- Drive automation, self-healing capabilities, and runbook maturity.
- Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
- Contribute hands-on to technical reviews, tooling, scripting, and automation.
Global Collaboration
- Partner closely with SRE leadership in India to support follow-the-sun operations.
- Represent the US SRE team in cross-functional planning and operational reviews.
- Communicate effectively with technical and non-technical stakeholders.
Documentation & Compliance
- Maintain documentation for incidents, postmortems, runbooks, and operational procedures.
- Support adherence to regulated-environment standards including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications
- 5–8+ years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
- 1–2+ years of experience leading, mentoring, or managing engineers.
- Experience operating in a player-coach leadership model.
- Strong hands-on experience with production incident management and escalation.
- Experience with Datadog or similar observability platforms.
- Production experience with Kubernetes and Docker.
- Strong scripting/programming skills using Python, Bash, PowerShell, Java, or C#.
- Experience with Helm, CI/CD pipelines, and deployment automation.
- Working knowledge of ITIL processes and Agile methodologies.
- Experience with SQL, MySQL, or NoSQL databases.
- Strong communication and stakeholder-management skills.
- Willingness to participate in PagerDuty/on-call escalation and a global follow-the-sun operating model.
- Must be based in the United States.
Preferred Qualifications
- Experience with AWS, Azure, or GCP.
- Experience building or scaling SRE teams and on-call programs.
- Experience defining and managing SLIs, SLOs, SLAs, and error budgets.
- Prior experience in healthcare, fintech, or another regulated industry.
- Knowledge of security and compliance frameworks used in regulated environments.
Technical Skills
Kubernetes | Docker | Datadog | Helm | CI/CD | Python | Bash | PowerShell | Java | C# | SQL | MySQL | NoSQL | AWS | Azure | GCP | Incident Management | SRE | DevOps | Observability | Automation
Why Join NationsBenefits?
- Competitive compensation and comprehensive benefits.
- Unlimited PTO.
- Fully remote work environment for US-based employees.
- Opportunity to lead and grow a high-impact SRE team.
- Exposure to modern cloud-native technologies and large-scale reliability challenges.
- Collaborative environment focused on innovation, learning, and continuous improvement.
- Meaningful work supporting healthcare technology at scale.
Ideal Candidate
We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, leading critical incidents, and remaining hands-on with reliability engineering and automation.
Best LinkedIn title:
Manager, Site Reliability Engineering (SRE) – Remote US