Role: Senior DevOps Engineer
Location: Palo Alto, CA
Duration: 12+ Months Contract
Required Experience:
6–8+ years of experience as a Site Reliability Engineer (SRE) or in a similar DevOps/Platform Engineering role.
Experience working with globally distributed and culturally diverse engineering teams is highly preferred.
Job Description – Site Reliability Engineer (SRE):
Collaborate closely with North America and Japan engineering teams in a multicultural environment.
Demonstrate excellent communication skills with strong cultural awareness, empathy, and cross-functional collaboration experience.
Support and maintain the Weft Platform by ensuring reliability, scalability, and operational excellence.
Triage, troubleshoot, and respond to technical and educational questions from Software Development Engineers (SDEs).
Take end-to-end ownership of production incidents and resolve medium-complexity operational issues.
Build and maintain Infrastructure as Code (IaC) using Terraform (hands-on experience required; expert-level not required).
Develop and manage CI/CD workflows using GitHub Actions.
Write automation and operational scripts using Python.
Work with AWS services, including Lambda, Step Functions, S3, and other data storage services.
Deploy, manage, and troubleshoot applications running on Kubernetes.
Implement and support observability and monitoring solutions using Prometheus, Grafana, Sentry, and the ELK Stack.
Apply security best practices with a solid understanding of AWS IAM and access management.
Participate in incident response, root cause analysis, and continuous service improvement.
Experience with Istio Service Mesh is a strong plus.
Take ownership of incidents, measure operational metrics, and drive service reliability improvements.