Hands-on Research Engineering Manager to establish the R&D team that builds an ๐๐ ๐ฎ๐ด๐ฒ๐ป๐ ๐ต๐ฎ๐ฟ๐ป๐ฒ๐๐ and evaluation regime for ๐ฐ๐ผ๐ป๐๐ฟ๐ผ๐น๐น๐ถ๐ป๐ด ๐๐ฒ๐ฏ ๐ฏ๐ฟ๐ผ๐๐๐ฒ๐ฟ๐ with open-weight LLMs. Greenfield.
THE TECHNICAL CHALLENGE
We build agents on open-weight LLMs that operate web applications through the browser. The engineering work is the harness around the model: the abstraction that lets a pipeline or another agent drive a browser agent from the terminal, and the evaluation regime across levels of abstraction, from single actions ("find the element matching this description") to whole tasks, run against simulated applications before a live one. Agents maintain the evaluation suite; humans define its schema. Progress is measured against benchmarked results. You lead the team that builds this with the ability to write code.
KEY RESPONSIBILITIES
โข Set the technical direction of the harness and the interfaces between its levels of abstraction
โข Establish the evaluation regime and the fortnightly score, and hold the team to measured results
โข Run the R&D team on a two-week rhythm: planning, review with working software, retrospective, written post-mortems for failed experiments and incidents
โข Coach, level and hire the engineers; interface with product management on the backlog and with the platform team on what the harness exposes
DESIRED QUALIFICATIONS
โข Managed both a research-cadence team and a sprint-cadence team, and can articulate the interface between them
โข Built an agent harness or an evaluation harness for computer-use systems, agents or reinforcement learning, with an evaluation culture recognised outside the team, and can say which abstraction was right and which was wrong
โข Has taken apart the harnesses of coding agents (OpenHands, SWE-agent, Aider, DeepSeek Harness or comparable) and can explain how they work in plain terms
โข Has fine-tuned an LLM, from data preparation to evaluation on a held-out set
EXPECTED QUALIFICATIONS
โข T-shaped: deep in one domain, with working breadth in a neighbouring one
โข Structures a large, incompletely specified problem and drives it to a working result independently
โข Managed a team building on LLMs or machine-learning models in production, with evaluation pipelines as a management requirement and a clear account of how quality was gated
โข Turns research or prototype code into reference implementations others reuse, using AI coding tools daily and verifying their output
โโ
HOW WE WORK
โข Product engineering: we own what we build and run it in production
โข Small teams, two-week cycles, working software at every review
โข AI coding tools are part of the standard workflow
WHAT WE OFFER
โข Founding-team scope with a direct line to the Product CTO
โข A supported track to the UAE Golden Visa (ten-year residency) for AI and technology talent, and relocation support including the UAE AI Specialist visa
โข MacBook Pro M5 Max with 64 GB RAM or more, gear, Nvidia B200s and a serious budget for AI tokens, credits and tooling
โข Up to six weeks per year working remotely from anywhere
PROCESS
โข Introductory call with our recruiter
โข Technical conversation with the Product CTO
โข Practical session; the format is agreed with you
โข In exercises, AI tools are allowed and expected (no LeetCode)
WHO WE ARE
New product organisation backed by a semi-government in Abu Dhabi. International, ex-FAANG team. Completely greenfield, with a modern tech stack.
โโ
REQUIREMENTS TO BE CONSIDERED
โข Clear written and spoken English; has reported technical progress to non-technical executives on a fixed cadence
โข 6+ years in engineering, of which 2+ leading a team of engineers where work was accepted on measured results
โข Still coding, able to review TypeScript and Python; own code operated in production at a product company or a startup
โข Bachelor's degree in any field, or self-taught with a track record of open-source contributions
RELATED TECHNOLOGIES AND CONCEPTS
โข Agent harnesses and frameworks: OpenHands, SWE-agent, Aider, DeepSeek Harness, Hermes Agent, LangGraph, smolagents
โข Programmatic prompt and pipeline optimisation: DSPy, TextGrad, Ax
โข Browser automation and computer use: Playwright, Chrome DevTools Protocol, Stagehand, Browser Use, Steel Browser
โข Evaluation and observability: evaluation harnesses, LLM-as-judge, trajectory evaluation, Inspect, Langfuse, OpenTelemetry; benchmarks such as Online-Mind2Web, WebArena, OSWorld
โข Models and serving: open-weight LLMs, vLLM, SGLang
- โข State and orchestration: state machines, event sourcing, XState, Trigger.dev, Temporal, Kubernetes, TypeScript, Python