AI Jobs Map

Reqroute, Inc · Woonsocket, RI

SRE/DevOps/Platform Engineering (W2 Only)

HybridseniorcontractPosted today
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sredevopsgcppythonreactjavaprometheusgrafanaopentelemetrysplunkelasticsearchobservabilitykubernetesetlapache-airflowsqlbigquerypostgresqlllmapache-kafka

W2 | Hybrid Role

Position Title: SRE/DevOps (GCP Exp)

Location: Woonsocket, RI

Duration: 12+ Months Contract

Position Type- W2 Role

Exp Level- 8+Years

Req Skills- SRE/DevOps/Platform Engineering, Incident Commander (IC) for P1 or P2 incidents, Python, React, Java, Prometheus, Grafana, Open Telemetry, (Loki, Splunk, Elasticsearch), Google Cloud Platform (GCP), Rancher K3s

Job Description

- 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility

- Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents — structured leadership updates, not just participant involvement

- Experience tuning and validating time-series anomaly detection models in a production observability context — this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role

- Strong programming proficiency in Python, React, and Java at production quality — capable of writing operational tooling that other engineers will rely on

- Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services

- Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least two of the log aggregation solution (Loki, Splunk, Elasticsearch)

- Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments

- Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.

- Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development.

- Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipeline : Apache Airflow and Tidal.

Preferred Qualifications

- Experience owning Production Readiness Reviews or service launch gates.

- Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics, and reporting frameworks: Google BigQuery, PostgreSQL.

- Hands-on chaos or fault injection experience.

- TIC (Technical Incident Commander) certification or equivalent structured incident command training

- Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact

- LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) — design or implementation experience

- Experience with streaming data platforms: Kafka.

- Experience with service mesh and traffic management: Istio, Envoy.

Infrastructure-as-code proficiency at production scale: Terraform or Ansible

More jobs at Reqroute, Inc