- Do you have experience in AI research across LLMs, agents, evaluations, benchmarking, AI safety, alignment, agent understanding, RL environments or post-training?
- Have you built practical AI/ML research projects, benchmarks, evaluations or agent-related systems through academic research, open-source work, internships or industry?
- Do you have hands-on experience with post-training methods such as RLHF, DPO, distillation or high-stakes model evaluation, ideally beyond coursework?
- Are you a strong hands-on coder or full-stack technical generalist who enjoys building, testing and shipping rather than only writing papers?
- Can you operate autonomously on ambiguous technical problems and move quickly from research through to task creation and delivery?
If you answer yes to the above, then this confidential role could be the one for you.
This is a rare opportunity to join a very early-stage AI infrastructure company working at the intersection of research, data, evaluations, post-training and frontier model improvement. The successful candidates will help build high-impact benchmarks, datasets and environments that improve model reasoning and agent performance. Fresh PhD graduates can be considered.
Salary on offer is $200,000 to $400,000, with more potentially available for an exceptional candidate.