AI Jobs Map

Imperial College London · South Kensington

Research Associate in Adaptive and Efficient LLM Architectures

full time$68,908 – $80,794 / yearPosted 2 days ago
Apply on IndeedIndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmgenerative-aitransformersagentic-aideep-learningpytorchhugging-facecudaartificial-intelligencenlpvllm

Job number

ENG04037

Faculties

Faculty of Engineering

Departments

Department of Computing

Salary or Salary range

£50,733 - £59,484 per annum

Location/campus

South Kensington Campus - On site only

Contract type work pattern

Full time - Fixed term

Posting End Date

30 Sept 2026

About the role

We are looking for creative and passionate researchers to join the ERC project AToM (Adaptive Tokenization and Memory in Foundation Models) at Imperial College London, led by Dr. Edoardo Ponti, in a fully funded postdoctoral role to lead transformative research in adaptive and efficient architectures for AI models.

Overview. The Department of Computing, is seeking highly motivated and talented Postdoctoral Research Associates, who have conducted cutting-edge research and/or have extensive experience in frontier AI labs. The position is fully funded by the ERC project with a focus on designing, implementing, and publicly releasing LLMs with efficient architectures (including adaptive memory, latent tokenization, sparse attention, multi-token prediction, adaptive depth, among others).

The recent revolution in generative AI has been driven by the rapid growth in training and inference compute for Foundation Models (FMs). This scaling paradigm, however, is characterised by high energy demand, high latency, and environmental impact. AToM sets out to reverse this trend by targeting a fundamental inefficiency in dominant FM architectures such as Transformers (including SWA/SSM hybrids). Currently, models store, access, and convert data into sequences of internal representations, whose length bottlenecks both training/prefill (compute-bound) and decode (memory-bandwidth bound). Yet the granularity of the representations is largely determined upfront by input segmentation (tokenisation), which typically remains fixed across layers, and by the memory update mechanism, which accumulates most tokens in the key–value cache.

AToM aims to prototype new classes of FM architectures that learn, end-to-end, to compress their sequences of internal representations, effectively redefining the model’s “atomic units” for processing and memorising information. To accelerate adoption, we will retrofit state-of-the-art open-weight FMs into adaptive variants. In the first year of the project, we have already successfully released Qwen 3 with adaptive memory in collaboration with NVIDIA, and OLMo 3 with latent tokenization in collaboration with AI2.

This project will lead not only to substantial gains in efficiency (several orders of magnitude speedups without accuracy degradation) but also to the emergence of new capabilities: adaptive FMs can operate over broader effective horizons, as they can perceive longer inputs and generate longer outputs under a fixed budget. This enables (1) lifelong learning via a permanent, sub-linearly growing memory, (2) inference-time hyper-scaling for reasoning and agentic tasks (e.g., advanced maths, science, coding), and (3) world modelling for multimodal planning and simulation. Adaptive FMs thus open a path towards more efficient and more capable generative AI.

What you would be doing

Within the project, you will conduct original research in the new and exciting field of efficient and adaptive LLM architectures and explore its applications across long-context understanding and reasoning (for code, maths, and agentic workflows) as well as long-horizon multimodal world modelling. We will strive to release new, more efficient and capable AI models and to publish in top-tier conferences and journals. Specifically, you will work closely with Dr. Edoardo Ponti in:

- Designing and developing novel architectures for AI models

- Performing retrofitting, post-training, and evaluation of SOTA open-weight models

- Conducting independent and collaborative research within the group

- Publishing results in top-tier conferences and journals (e.g., NeurIPS, ICML, ICLR, *ACL, EMNLP, Nature)

- Implementing research ideas using modern deep learning frameworks (PyTorch/JAX), model/dataset libraries (Huggingface transformers/diffusers), and efficient kernels (Triton/CUDA)

- Contributing to research projects on adaptive memory, latent tokenization, sparse attention, long-context understanding and reasoning, agentic, and multimodal world modelling

- Contributing to the life and development of Edoardo Ponti’s Lab, including weekly meetings, presentations, maintaining the website and other resources).

- Delivering tutorials at conferences and summer schools.

- Mentoring PhD and MSc students.

- Presenting research at international conferences and workshops.

- Contributing to grant proposals and collaborative research initiatives.

What we are looking for

- Self-driven and motivated individuals with genuine love for research.

- The applicant is also expected to have a strong track record in top conferences and journals in the fields of AI/ML/NLP, such as NeurIPS, ICML, ICLR, *ACL, EMNLP, etc. Candidates with research experience as part of frontier AI labs are also welcome.

- Excellent skills in coding, strong foundations in mathematics (especially information theory, linear algebra, calculus), and knowledge in deep learning.

- Practical experience in a broad range of techniques including LLM training, evaluation, RLVR, PEFT, quantisation, tensor/data parallelism.

- Ideally, the candidate should have familiarity with CUDA kernels and/or Triton, and inference engines (vLLM, SGLang, et cetera).

- Experience coding with deep learning libraries such as Pytorch/JAX is essential.

- Fluent written and spoken English skills as well as contributions to the group culture are expected.

- Applicants must hold a PhD in computer science or equivalent experience.

Please see job description for a full list of requirements.

What we can offer you

- Extensive funding for conference travel (2 international conferences per year)

- Access to extensive compute via the AToM project GPUs (B200s and cloud credits) and GPUs from the Department of Computing and the College of Engineering (A100s and H200s)

- The opportunity to continue your career at a world-leading institution and be part of our mission to continue science for humanity.

- Grow your career: gain access to Imperial’s sector-leading dedicated career support for researchers as well as opportunities for promotion and progression.

- Sector-leading salary and remuneration package (including 43 days off a year and generous pension schemes).

- Be part of a diverse, inclusive and collaborative work culture with various staff networks and resources to support your personal and professional wellbeing.

Further information

Full-time, fixed term 2-year contract to start around winter 26/27 (the start date is flexible).

- Candidates who have not yet been officially awarded their PhD will be appointed as Research Assistant within the salary range £45,399 - £48,876 per annum.

In addition to completing the online application candidates should attach:

- A full CV with a list of all publications

- A research statement (max 2 pages) indicating what you see are the most interesting research questions relating to adaptive and efficient AI architectures and why your expertise is relevant.

Informal enquiries related to the position should be directed to Dr Edoardo Ponti:

[email protected]

For queries regarding the application process contact Jamie Perrins:

[email protected]

Available documents

Attached documents are available under links. Clicking a document link will initialize its download.

Please note that job descriptions are not exhaustive, and you may be asked to take on additional duties that align with the key responsibilities mentioned above.

We reserve the right to close the advert before the stated closing date, should we receive a high volume of applications. It is therefore advisable that you submit your application as early as possible to avoid disappointment.

If you encounter any technical issues while applying online, please don't hesitate to email us at [email protected]. We're here to help.

About Imperial

Welcome to Imperial, a global top-ten university where scientific imagination leads to world-changing impact.

Join us and be part of something bigger. From global health to climate change, AI to business leadership, here at Imperial we navigate some of the world’s toughest challenges. Whatever your role, your contribution will have a lasting impact.

As a member of our vibrant community of 22,000 students and 8,000 staff, you’ll collaborate with passionate minds across nine London campuses and a global network.

This is your chance to help shape the future. We hope you’ll join us at Imperial College London.

Our Culture

We work towards equality of opportunity, to eliminating discrimination, and to creating an inclusive working environment for all. We encourage applications from all backgrounds, communities and industries, and are committed to employing a team that has diverse skills, experiences and abilities. You can read more about our commitment on our web pages.

Proud signatory of the Armed Forces Covenant. We welcome applications from the Armed Forces community.

Our values are at the root of everything we do, and everyone in our community is expected to demonstrate respect, collaboration, excellence, integrity, and innovation.

More jobs at Imperial College London