AI Jobs Map

Ardan Labs · United States

Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY

Remoteseniorfull timePosted 9 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

srekuberneteseksopenshiftawshelmterraformhipaaobservabilityincident-responseistioopentelemetryprometheus

We are looking for a Senior Site Reliability Engineer to join a platform team and take ownership of both the reliability of an internal developer platform and the experience of the application teams building on it.

This is a hands-on senior individual contributor role for an engineer who enjoys solving complex infrastructure problems, operating Kubernetes at scale, and working directly with teams to diagnose and resolve issues.

You will work across Amazon EKS and Red Hat OpenShift on AWS, a curated Helm chart catalog, GitOps pipelines, Terraform, and a multi-account AWS environment built on Control Tower. Our workloads operate under HIPAA requirements, making security hygiene, patch currency, and operational reliability ongoing engineering priorities.

Approximately 60% of your time will focus on reliability engineering and 40% on forward-deployed work with application, security, network, and other technical teams.

What You'll Do

- Own the health and reliability of Kubernetes management and workload clusters across sandbox, development, QA, and production environments.

- Diagnose and permanently resolve GitOps delivery and reconciliation failures.

- Manage version and component adoption to ensure fixes and improvements are consistently deployed across clusters.

- Develop and maintain Terraform and infrastructure pipelines across a multi-account AWS environment.

- Work with cloud identity, resource policies, and key management to diagnose and resolve complex permission issues.

- Partner directly with application teams to troubleshoot manifests, managed resource claims, secrets, certificates, ingress, and other platform issues.

- Build documentation, guides, and guardrails that help application teams become increasingly self-sufficient.

- Own and improve the telemetry path into the observability platform.

- Partner with security, networking, vendors, and other technical teams when the platform is involved in an incident.

- Participate in incident response and drive problems through to durable resolution.

- Lead technical initiatives across teams that do not report to you.

What We're Looking For

- 8+ years of experience in infrastructure, platform engineering, or site reliability engineering.

- At least 3 years of production Kubernetes experience supporting teams beyond your own.

- Strong hands-on experience with production GitOps, including diagnosing conflicts between declared and live state.

- Experience using Terraform across multiple cloud accounts or environments.

- Deep understanding of cloud identity and resource policies, including troubleshooting permissions that appear correct but are still denied.

- Demonstrated ability to lead technical work and influence teams without direct authority.

- Strong troubleshooting, systems thinking, and communication skills.

- Comfortable working directly with application and engineering teams rather than operating solely behind the scenes.

Nice to Have
The following are not required, but would strengthen your application:

- Red Hat OpenShift / ROSA experience

- Experience working in regulated environments

- Production experience with Istio or another service mesh

- Forward-deployed, field engineering, implementation, solutions, or customer-facing engineering experience

- Ownership of OpenTelemetry, Prometheus, or commercial APM platforms

- Go or another compiled language used for infrastructure/tooling

- Published technical writing, such as incident reports, design documents, or technical/customer-facing documentation

Why This Role Is Different
This isn't a role where success is measured by how many tools you know. The platform is highly automated and largely self-healing. The difficult problems are often about figuring out where the problem actually belongs, determining the right durable fix, and building agreement between teams that don't report to one another.

We're looking for an engineer who can operate at both levels: deep technical infrastructure expertise and strong cross-team engagement.

More jobs at Ardan Labs