Mission: Build the engineering harness that allows a small specialist model to operate reliably as an end-to-end software engineer.
Responsibilities
- Design and build the agent/harness architecture around specialist models.
- Build repository understanding and codebase retrieval.
- Connect models to Git, compilers, debuggers, test runners and development environments.
- Build planning → execution → testing → evaluation → correction loops.
- Integrate architecture/dependency analysis.
- Integrate security and static-analysis tooling.
- Build automated performance/load-testing workflows.
- Design model verification and critic/evaluator systems.
- Develop tool-use and function-calling infrastructure.
- Build mechanisms that prevent models from modifying code before understanding the task/context.
- Develop observability and evaluation infrastructure for agent behavior.
- Determine which capabilities should be handled by the model and which should be deterministic system components.
- Work closely with the Model Research Engineer to feed real-world failures back into model training.
Ideal background
- Strong AI agent / ML systems engineering experience.
- Experience building production coding agents or developer tools.
- Strong Python/TypeScript and backend engineering.
- Experience with tool calling, orchestration and agent frameworks.
- Strong understanding of software architecture.
- Experience with RAG, code retrieval and knowledge graphs.
- Familiarity with Docker, Kubernetes, CI/CD and cloud infrastructure.
- Experience with evaluation and observability systems.
Bonus
- Compiler/toolchain experience.
- Static analysis.
- Code security.
- Performance engineering.
- SWE-bench or repository-level coding agents.
- LangGraph, Claude Agent SDK, OpenAI Agents SDK or similar.
What success looks like
Take a small specialist model and build a system around it that allows it to reliably understand, plan, implement, test, debug and validate real software-engineering tasks.