Role: Senior Production Support Engineer
Location: Austin, TX (Onsite)
Exp. Level: 8+ yrs
Role Purpose Owns AI-augmented incident triage, runbook execution, and root-cause analysis, and separately owns AI-powered anomaly detection and spend/usage alerting.
Key Responsibilities:
- Own AI-augmented incident triage, runbook execution, and root-cause analysis across the 3 core FinOps applications and 50+ services.
- Build and maintain AI-powered anomaly detection and spend/usage alerting.
- Support the 99.99% uptime target during core PST business hours (8AM-5PM)
- Operate within the support portion of delivery (L1/L2/L3 tiers) within the 35% run-and-support split.
- Contribute to baselining incident response/resolution times post-transition.
Must-Have:
- 8–13 years in production support or SRE, including incident management.
- Experience with monitoring/alerting tooling, ideally GCP-native.
- Demonstrated root-cause analysis discipline at production scale.
Nice-to-Have:
- Experience with AI-based anomaly detection tools.
- Node.js or Python for support automation/scripting.