About the Company
We are representing a rapidly growing financial technology company building a next-generation trading platform. Our client is developing a consumer-focused platform that combines real-time trading with social discovery, giving users a streamlined way to discover market activity, follow traders, receive real-time insights, and execute trades across a growing range of financial products. Behind the consumer product is a sophisticated, highly distributed backend platform supporting real-time market data, trading activity, social features, financial data, and user-facing services across multiple regions. With significant growth ahead, the engineering organization is expanding its core infrastructure capabilities and is looking for a Staff Distributed Systems Engineer to take ownership of reliability, scalability, and performance across the platform.
About the Role
This is a hands-on Staff-level position with direct ownership of critical production infrastructure. You will own shared systems including datastores, caches, messaging infrastructure, and regional application services. The focus will be on designing systems that remain predictable and resilient through traffic surges, dependency failures, infrastructure changes, and partial regional outages. You will also lead the development of new failover and disaster-recovery capabilities, including defining recovery objectives and implementing the systems and testing required to safely recover services and data.
Responsibilities
- Design and operate high-throughput, multi-region backend services.
- Own the reliability, scalability, and performance of critical shared infrastructure.
- Improve datastore and cache performance, capacity, replication, and failure handling.
- Implement backpressure, concurrency limits, load shedding, rate limiting, circuit breakers, and bounded retries.
- Reduce cross-region latency and improve data locality.
- Design and test service, datastore, and regional failover procedures.
- Build disaster-recovery systems covering backup restoration, replication, and regional recovery.
- Define and validate RTO/RPO objectives for critical services and data.
- Identify production bottlenecks, capacity constraints, and failure modes before they become incidents.
- Help architect new features so they can operate reliably at scale from day one.
- Improve observability, debugging, and operational tooling across the platform.
- Establish strong engineering practices around distributed systems and reliability.
- Mentor engineers and raise the team's technical bar around operating systems at scale.
Qualifications
- 10+ years of experience in backend, platform, infrastructure, or distributed systems engineering, or equivalent practical experience.
- Deep experience designing, operating, and debugging distributed, high-throughput production systems.
- Strong PostgreSQL experience, including query performance, indexing, connection pooling, replication, transaction contention, and database failure modes.
- Strong experience with Redis or Redis-compatible systems such as Valkey, Dragonfly, or KeyDB, including sharding, replication, memory management, hot keys, and failure handling.
- Experience operating production services on AWS, ideally using ECS, RDS, and ElastiCache.
- Strong infrastructure-as-code experience, preferably Terraform.
- Proficiency in Go, TypeScript/Node.js, or another systems-oriented programming language.
- Hands-on experience designing and testing failover and disaster-recovery systems.
- Experience with backup restoration, replication, regional failover, and validating RTO/RPO objectives.
- Strong production engineering instincts and a track record of owning systems through real-world scale and failure scenarios.
Required Skills
- Experience with NATS JetStream, Kafka, or another durable messaging system.
- Familiarity with Datadog APM, AWS Performance Insights, or comparable observability tooling.
- Experience performing live datastore or cache topology migrations.
- Experience operating systems with highly variable or bursty traffic.
- Background in financial technology, trading, cryptocurrency, gaming, or another high-throughput environment.
- Experience working on systems where availability, latency, and data correctness are critical.
Why This Role
This is not a purely architectural Staff position. You will be deeply involved in production systems, making decisions around databases, caching, replication, regional architecture, failure recovery, traffic management, and system performance. The role is particularly well suited to an engineer who has spent years dealing with the realities of distributed systems: overloaded services, database contention, hot keys, failed deployments, network failures, dependency degradation, regional outages, and unpredictable traffic. It is an opportunity to have broad technical ownership over the distributed systems layer of a rapidly scaling financial platform, while helping shape the architecture and engineering practices as the platform grows.