Our client is seeking a Staff or Principal-level Platform Engineer to take end-to-end ownership of the cloud infrastructure behind its realtime text-to-speech (TTS) and LLM routing products.
You will design, scale, and secure a Kubernetes-based internal platform across multiple cloud providers, supporting low-latency, high-availability AI inference at consumer scale.
Working closely with engineers across the organization, you will shape how the company builds, deploys, monitors, and scales its services, and champion a “you build it, you run it” culture. This is a high-ownership role on a small SRE team of three, with a strong possibility of growing into a team lead position.
Key Responsibilities
- Design, deploy, and maintain reliable, high-performance, and secure cloud infrastructure for the company’s realtime TTS and LLM routing products.
- Manage and scale production Kubernetes clusters, authoring Helm charts and Kustomize manifests for application deployments.
- Own CI/CD pipelines and infrastructure deployments using Terraform, Terragrunt, ArgoCD, GitHub Actions, and related GitOps tooling.
- Partner with engineering teams to deploy and evolve services across Google Cloud Platform, Microsoft Azure, and Oracle Cloud.
- Build the tooling and processes that let teams monitor the reliability, availability, and performance of their own services.
- Lead root cause analysis for critical incidents and deliver automated solutions that prevent recurrence.
- Identify and build AI-powered developer tooling and workflows that increase engineering velocity across the organization.
- Influence org-wide engineering practices and tooling decisions, optimizing the platform for realtime, low-latency AI inference.