Our client is an established proprietary trading firm trading its own capital across global markets. No external investors, no clients; research goes straight into production, and performance is the only feedback loop that matters.
They are building out a deep learning research function and are actively interviewing now. We're partnering with them confidentially on a small number of senior hires. Firm details shared on a call.
The role
This is a frontier-scale deep learning research seat, not an applied-LLM or platform role. You'll be doing the modelling work itself, architecture, objectives, training dynamics, on models at the largest scale, with the compute and data to match.
What you'll do
- Research and design deep learning models and architectures at scale
- Work on pre-training: architecture design, objectives and loss design, data mixture as a modelling decision, optimization and training stability, scaling behavior
- Take ideas from hypothesis through large-scale training runs to production
- Push on the parts of the stack that actually move model quality — attention variants, long context, mixture-of-experts routing, tokenization, curriculum
- Work alongside researchers, traders and infrastructure engineers who handle the plumbing so you can stay on the modelling
What they're looking for
- 5+ years of LLM research experience
- Experience at genuine scale. You've worked on frontier-class models. Think the largest publicly known model generations and above, and can talk concretely about parameter counts, token budgets, cluster size and what broke.
- A pre-training background is the strongest fit. Post-training profiles are welcome where the work is heavily modelling RL algorithm design, reward modelling, objective design, distillation research — rather than data pipeline ownership or infrastructure
- Deep fluency in PyTorch (or JAX) and genuine command of the theory, not API-level familiarity
- Strong publication record is a plus, not a requirement — frontier-run experience speaks for itself
Trading experience is not required and not expected. They hire from big tech, frontier AI labs, and research groups doing serious work at scale.
Why it's worth a conversation
- A genuine research seat with serious compute — no product roadmap, no launch cycle, no committee between your idea and a training run
- Immediate, measurable feedback on every model you ship
- Small team, high trust, no organizational drag
- Compensation for this skillset is genuinely open-ended at the top end — well beyond standard big-tech senior bands