AI Research Engineer – Frontier Models & Benchmarking
Dublin | 5 Days Onsite
We’re working with a well-funded AI company building training and evaluation environments for some of the world’s leading AI labs.
They’re growing a small research team in Dublin and are looking for people who are genuinely deep into frontier AI models.
This is a deliberately specialised role.
If your LLM experience is mainly building RAG applications, chatbots, agent orchestration or integrating models into existing products, this probably isn’t the right role for you.
They’re looking for people who study the models themselves.
You’ll be working on problems like:
- Designing and building benchmarks for frontier LLMs
- Comparing how different models perform and behave
- Building difficult tasks and datasets to expose model strengths and weaknesses
- Adversarially testing models and finding where they break
- Designing rigorous evaluation methodologies
- Working with synthetic training and evaluation data
- Investigating model reasoning, reliability and unexpected behaviour
- Working on post-training and reinforcement learning
- Running experiments across multiple frontier models
The people we particularly want to hear from have experience in one or more of:
- Building or contributing to recognised/public LLM benchmarks
- LLM benchmarking or model capability evaluation
- Adversarial LLM testing or red teaming
- AI safety / model safety research
- LLM post-training or reinforcement learning
- Synthetic data generation for model training/evaluation
- Designing evaluation frameworks, metrics or model judges
- Research into LLM behaviour and failure modes
- Working deeply across multiple models such as Claude, GPT, Gemini, Llama or DeepSeek
-
Strong research credentials are highly valued. That could mean a PhD in a relevant area, significant research experience, publications at conferences such as NeurIPS, ICML or ICLR, or demonstrably strong work building benchmarks and evaluation systems.
Most importantly, we’re looking for people who have actually done this work, rather than simply used the terminology.
If your experience is predominantly:
- RAG
- LangChain
- Vector databases
- Prompt engineering
- Building chatbots
- Agent orchestration
- Calling LLM APIs
- Adding GenAI features to existing products
...without deeper model evaluation, benchmarking or research experience, this role is unlikely to be a fit.
This is also 5 days per week onsite in Dublin city centre. It’s a small, ambitious team that moves quickly and works hard, so it’ll suit someone who actively wants that kind of start-up environment rather than a traditional 9-to-5.
If you’re the sort of person who sees a new frontier model released and immediately wants to test it, break it, compare it and understand why it behaves differently from the others, we’d like to hear from you.