We are looking for Principal DevOps Engineer to design and build the AWS infrastructure and delivery platform the new gaming platform will run on. Greenfield: define the cloud foundation — account structure, networking, Kubernetes platform, CI/CD, observability — before the first production workload lands, then scale it to sportsbook peak traffic for millions of players. The platform must hold under live-betting spikes, strict uptime expectations for a money-moving system, and gaming-regulator audit and compliance requirements. Hands-on technical leader: writes infrastructure code daily, sets platform standards, and is the technical authority on how software is built, shipped and operated.
What you will do:
-
Design the AWS foundation from scratch: multi-account architecture (AWS Organizations), landing zone, VPC and networking, IAM strategy, cost governance
-
Build and operate the container platform — Amazon EKS, service mesh, autoscaling tuned for spiky sportsbook load, multi-AZ (and where justified multi-region) resilience
-
Define everything as code: Terraform for all infrastructure, GitOps delivery (Argo CD or similar), paved-road CI/CD pipelines so product teams ship safely and often
-
Establish the observability stack — metrics, logging, tracing, alerting (Prometheus/Grafana, OpenTelemetry, CloudWatch) — and drive an SLO-based reliability practice with error budgets
-
Own production readiness: incident response, on-call design, runbooks, chaos/load testing ahead of major sporting events, blameless postmortems
-
Build the infrastructure side of the migration off the current third-party platform: dual-running environments, data migration pipelines, cutover mechanics
-
Embed compliance into the platform: audit trails, environment segregation, backup/DR, controls that satisfy gaming regulators by construction
-
Partner with the AI coding platform team: provision and operate the infrastructure behind AI-assisted development, and bring AI tooling into DevOps workflows
-
Mentor engineers across teams on cloud-native and operational best practices; set organization-wide standards
Requirements:
Must have:
-
10+ years of DevOps / platform / infrastructure engineering, including 3+ years at staff/principal level with organization-wide influence
-
Deep hands-on AWS: EKS, EC2, RDS/Aurora, networking (VPC, Transit Gateway, Route 53), IAM at scale, multi-account architectures (AWS certifications such as SA Professional / DevOps Professional are a plus)
-
Expert-level infrastructure as code with Terraform
-
Strong Kubernetes operational depth: day-2 operations, upgrades, capacity, cost
-
Track record of building CI/CD and developer platforms engineering teams adopted willingly — golden paths, not gatekeeping
-
Observability and SLO/error-budget practice (Prometheus/Grafana, OpenTelemetry, CloudWatch)
-
Experience operating high-availability, high-throughput production systems with real traffic spikes, with documented playbooks from real incidents
-
Strong scripting/programming (Python, Go or Bash) and comfort reading application code
-
Experience designing backup, disaster recovery and business continuity for systems where data loss is not an option
-
Excellent written and spoken English; communicates platform decisions clearly to engineers and executives
Nice to have:
-
Hands-on use of AI coding tools (Claude Code, Codex) for infrastructure, pipelines and ops automation — a significant plus
-
Security engineering: cloud security posture management, secrets management (e.g. Vault), vulnerability management, SAST/DAST/SCA in CI/CD, ISO 27001 / SOC 2 / PCI DSS
-
iGaming, sports betting, fintech or another regulated, high-transaction-volume domain
-
Migration off a third-party vendor platform
-
Event-streaming infrastructure (Kafka/MSK) and database operations at scale
-
Polish and/or Spanish