Responsibilities
Design and maintain SLO-based monitoring and alerting solutions.
Create and optimize PromQL queries and multi-window burn-rate alerts.
Build and manage Grafana dashboards using configuration-as-code practices.
Develop automation and monitoring configuration generators in Python or TypeScript.
Contribute to Terraform-based observability infrastructure.
Validate monitoring signals and improve alert quality.
Collaborate with engineering and platform teams to enhance reliability and operational visibility.
Skills Must have
Strong experience in Observability, SRE, or Platform Engineering.
Advanced PromQL (or equivalent) expertise, including the ability to identify misleading or incorrect query results.
Hands-on experience designing and implementing SLOs and multi-window burn-rate alerting.
Experience with Grafana provisioning and configuration as code.
Ability to build automation tools and code generators using Python or TypeScript.
Hands-on experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or equivalent. AI-assisted engineering is an expected part of the development workflow.
Strong analytical mindset with a healthy skepticism toward telemetry data.
Experience validating signals through multiple independent sources before operationalizing alerts.
Experience with cloud-native or distributed systems environments.
Nice to have
Groundcover
VictoriaMetrics
OpenTelemetry Collector
AWS EKS
Healthcare or regulated-industry experience
Experience building observability tooling and service health reporting solutions
Domain: US Healthcare (PHI-adjacent)
Okta and observability platform access
AI-assisted engineering is a standard part of the development workflow