About the job
A leading quantitative trading firm is seeking a Senior Site Reliability Engineer to build and evolve the reliability, observability, and automation capabilities powering a highly performance-sensitive trading environment.
Working at the intersection of software and infrastructure, you'll ensure the reliability, scalability, and performance of critical trading and research platforms. The team works across Linux, distributed systems, observability, automation, and platform engineering to solve complex production challenges at scale. You'll play a key role in designing, building, and optimising the infrastructure that underpins a best-in-class trading environment.
What You'll Bring
- Strong software engineering experience with one or more languages such as Python, Go, or C++
- Deep understanding of Linux systems and production infrastructure
- Experience building and operating highly available production environments
- Strong knowledge of distributed systems, observability, monitoring, and incident management
- Experience improving the reliability, scalability, and efficiency of critical systems and services through automation
- Excellent troubleshooting and fire-fighting ability within complex production environments
Nice to Have
- Experience with Kubernetes, Docker, Helm, OpenShift, EKS, GKE, AKS, or other container orchestration platforms
- Experience with observability tooling such as Prometheus, Grafana, OpenTelemetry, Datadog, Splunk, or ELK
- Background working within a quantitative trading firm, leading technology company, or high-growth startup
On Offer
- Market-leading compensation and benefits
- Growth-focused, high-performance engineering culture
- Opportunity to make a direct impact on business-critical trading infrastructure
Apply or get in touch for more information!