AI Jobs Map

Stealth Startup ยท Abu Dhabi Emirate, United Arab Emirates

Engineering Manager, AI Agent Harness (Product R&D)

Remotedirectorfull timePosted 2 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmtypescriptpythonplaywrightobservabilityopentelemetryvllmkubernetesreinforcement-learning

Hands-on Research Engineering Manager to establish the R&D team that builds an ๐—”๐—œ ๐—ฎ๐—ด๐—ฒ๐—ป๐˜ ๐—ต๐—ฎ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€ and evaluation regime for ๐—ฐ๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น๐—น๐—ถ๐—ป๐—ด ๐˜„๐—ฒ๐—ฏ ๐—ฏ๐—ฟ๐—ผ๐˜„๐˜€๐—ฒ๐—ฟ๐˜€ with open-weight LLMs. Greenfield.

THE TECHNICAL CHALLENGE

We build agents on open-weight LLMs that operate web applications through the browser. The engineering work is the harness around the model: the abstraction that lets a pipeline or another agent drive a browser agent from the terminal, and the evaluation regime across levels of abstraction, from single actions ("find the element matching this description") to whole tasks, run against simulated applications before a live one. Agents maintain the evaluation suite; humans define its schema. Progress is measured against benchmarked results. You lead the team that builds this with the ability to write code.

KEY RESPONSIBILITIES

โ€ข Set the technical direction of the harness and the interfaces between its levels of abstraction

โ€ข Establish the evaluation regime and the fortnightly score, and hold the team to measured results

โ€ข Run the R&D team on a two-week rhythm: planning, review with working software, retrospective, written post-mortems for failed experiments and incidents

โ€ข Coach, level and hire the engineers; interface with product management on the backlog and with the platform team on what the harness exposes

DESIRED QUALIFICATIONS

โ€ข Managed both a research-cadence team and a sprint-cadence team, and can articulate the interface between them

โ€ข Built an agent harness or an evaluation harness for computer-use systems, agents or reinforcement learning, with an evaluation culture recognised outside the team, and can say which abstraction was right and which was wrong

โ€ข Has taken apart the harnesses of coding agents (OpenHands, SWE-agent, Aider, DeepSeek Harness or comparable) and can explain how they work in plain terms

โ€ข Has fine-tuned an LLM, from data preparation to evaluation on a held-out set

EXPECTED QUALIFICATIONS

โ€ข T-shaped: deep in one domain, with working breadth in a neighbouring one

โ€ข Structures a large, incompletely specified problem and drives it to a working result independently

โ€ข Managed a team building on LLMs or machine-learning models in production, with evaluation pipelines as a management requirement and a clear account of how quality was gated

โ€ข Turns research or prototype code into reference implementations others reuse, using AI coding tools daily and verifying their output

โ€”โ€”

HOW WE WORK

โ€ข Product engineering: we own what we build and run it in production

โ€ข Small teams, two-week cycles, working software at every review

โ€ข AI coding tools are part of the standard workflow

WHAT WE OFFER

โ€ข Founding-team scope with a direct line to the Product CTO

โ€ข A supported track to the UAE Golden Visa (ten-year residency) for AI and technology talent, and relocation support including the UAE AI Specialist visa

โ€ข MacBook Pro M5 Max with 64 GB RAM or more, gear, Nvidia B200s and a serious budget for AI tokens, credits and tooling

โ€ข Up to six weeks per year working remotely from anywhere

PROCESS

โ€ข Introductory call with our recruiter

โ€ข Technical conversation with the Product CTO

โ€ข Practical session; the format is agreed with you

โ€ข In exercises, AI tools are allowed and expected (no LeetCode)

WHO WE ARE

New product organisation backed by a semi-government in Abu Dhabi. International, ex-FAANG team. Completely greenfield, with a modern tech stack.

โ€”โ€”

REQUIREMENTS TO BE CONSIDERED

โ€ข Clear written and spoken English; has reported technical progress to non-technical executives on a fixed cadence

โ€ข 6+ years in engineering, of which 2+ leading a team of engineers where work was accepted on measured results

โ€ข Still coding, able to review TypeScript and Python; own code operated in production at a product company or a startup

โ€ข Bachelor's degree in any field, or self-taught with a track record of open-source contributions

RELATED TECHNOLOGIES AND CONCEPTS

โ€ข Agent harnesses and frameworks: OpenHands, SWE-agent, Aider, DeepSeek Harness, Hermes Agent, LangGraph, smolagents

โ€ข Programmatic prompt and pipeline optimisation: DSPy, TextGrad, Ax

โ€ข Browser automation and computer use: Playwright, Chrome DevTools Protocol, Stagehand, Browser Use, Steel Browser

โ€ข Evaluation and observability: evaluation harnesses, LLM-as-judge, trajectory evaluation, Inspect, Langfuse, OpenTelemetry; benchmarks such as Online-Mind2Web, WebArena, OSWorld

โ€ข Models and serving: open-weight LLMs, vLLM, SGLang

- โ€ข State and orchestration: state machines, event sourcing, XState, Trigger.dev, Temporal, Kubernetes, TypeScript, Python

More jobs at Stealth Startup