We’re looking for mid‑level Site Reliability / Application Engineers to join a high‑impact team building and operating the firm’s core AI platform — powering model access, workflow runtimes, the enterprise AI portal, and experimentation environments used across investment research.
If you love solving deep infrastructure problems, scaling AI systems, and building reliable automation, this role is for you.
🔧 What You’ll Do
- Build and operate the firm’s AI platform and supporting infrastructure
- Troubleshoot production Kubernetes issues (EKS preferred)
- Implement and extend Terraform modules across multi‑environment setups
- Drive observability, monitoring, and SLO discipline
- Support release engineering, including data migration planning and verification
- Partner with platform, infra, and research teams to ensure reliability and scalability
⭐ Must‑Have Skills
- Degree in Computer Science / IT or related fields
- Strong hands‑on experience with production Kubernetes (EKS ideal)
- Solid experience with Terraform and IaC workflows
- Practical experience with AWS / Azure / GCP core services
✨ Good‑to‑Have Skills
- Datadog (APM, SLOs, alert tuning)
- Helm chart authoring
- Data migration experience
- Chaos engineering / fault injection (AWS FIS or similar)
- Operating OpenSearch or Temporal
- Enterprise identity & access provisioning
- AWS cost optimization
- Experience operating AI / LLM systems