Senior SRE (Site Reliability Engineer)
¥10M - ¥20M JPY/year
Tokyo · Onsite-first, partial remote OK · Full-time
A fast-scaling, AI-first e-commerce company bringing Japanese pop culture to a global market is hiring its first dedicated SRE. As traffic and data volume climb, you'll own reliability, availability, and performance across the entire product infrastructure and build the reliability-engineering foundation, and a future SRE team, from the ground up.
This is a rare greenfield ownership seat. You'll set the SLOs, observability, incident culture, and automation standards, defining what AI-augmented infra operations look like in an org where AI-collaborative development (Claude Code, Cursor, MCP) sits at the core.
Environment & Tech Stack
- Cloud & IaC: AWS (ECS/EKS, RDS, CloudFront, Lambda, S3), GCP (partial), Terraform, AWS CDK
- Containers: Docker, ECS/EKS orchestration
- Observability: Datadog (APM, logs, monitors, synthetics)
- CI/CD: GitHub / GitHub Actions
- Incident: PagerDuty, on-call, postmortems
- AI-augmented ops: Claude Code, Cursor, MCP tooling
- Scripting: Go / Python / TypeScript
What We're Looking For
- 5+ yrs as Infrastructure Engineer / SRE / DevOps
- 3+ yrs designing & operating production AWS
- Hands-on IaC (Terraform or equivalent)
- Led incident response, root-cause analysis, and prevention
- Team development & code review with Git/GitHub
- Motivated to leverage AI tools in infrastructure ops
Nice to have: Kubernetes (EKS/GKE), high-traffic ops (10k+ RPS), or launching an SRE function.