AI Jobs Map

Jecona · United States

Site Reliability Engineer

Remoteseniorfull timePosted 2 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sreagentic-aici/cdobservabilitytypescriptpythonkubernetesgcpgitlabhelmterraformdatadogprometheusgrafanagitlab-ci

We are partnering with a company that build a platform that helps contact centers resolve more requests, proactively identify issues, and improve agent performance with AI-powered conversation intelligence and AI agents that act like your best reps. They are looking for a Site Reliability Engineer to join a team that builds - not just supports - an AI-native platform, and they own exciting domains like platform and harness engineering, site reliability, cloud infrastructure, CI/CD, DevEx, observability, incident management, and COGS (e.g. cloud spend visibility). They want an SRE that has opinions about how these domains should work and wants agency in shaping where they go.

Their stack is TypeScript/Node and Python running on Kubernetes - primarily on GCP, with GitLab CI, Helm, Terraform, Datadog, Prometheus, and Grafana.

***Unfortunately, no sponsorship now or in the future is provided for this role***

What You’ll Do

- Contribute to patterns, design, and implementation of our domains; help shape the future of platform engineering.

- Build and improve systems that help reduce toil and enable the company's production infrastructure to remain available and operable under large-scale, real-time conversational AI traffic.

- Extend and iterate our agent harness: Unsupervised AI agents are currently used by about 10% of the dev team - help us grow that number. The agent harness includes CI, sandboxes, guardrails, and validation (e.g. agent-first eval loops).

- Own and improve our CI/CD pipelines and surrounding developer tooling: build and test performance, deployment ergonomics, and paved paths for new services.

- Participate in on-call rotation and incident management to ensure platform uptime and quality. (SRE owns the base infrastructure, not the applications; non-business-hours pages are rare)

What You'll Bring

- 6+ years’ experience in software development enablement roles.

- Solid experience owning CI/CD platforms end to end - including domains like caching, architecture, and developer self-service.

- Effective use of AI tools such as Claude and Cursor for coding, troubleshooting, and reasoning. You pair these skills with a defensible opinion on where to avoid using AI tools.

- Familiarity with Node/TypeScript including making code changes (e.g. exposing new metrics), Python and Terraform for automation, and developing in a Kubernetes/Helm ecosystem.

- Practical experience with observability: logs/metrics/tracing, monitoring/alerting, incident management process, and tooling.

- Experience working in fully remote teams - tell us how you’ve made one work better.

- Bonus:

- Harness engineering experience - building platforms for autonomous agents.

- Production-at-scale experience with GCP.

- Telephony and SIP architectures, FreeSWITCH in particular.

More jobs at Jecona