Research Engineer – Evaluation
AI Infrastructure Company (Series A, open-source core)
San Francisco | Onsite. Permanent.
$250k–$350k base + Equity & Benefits
This role will suit a research engineer with end-to-end experience designing and independently owning evaluation systems for LLM, extraction or search output on open web data at scale – a business whose product converts arbitrary web pages into structured, model-ready data for AI companies and agents, and whose quality claims are published as benchmarks.
Metric design, evaluation harnesses, gold-set and synthetic dataset creation, LLM-as-a-judge calibration, statistical validation, release gating. Practical knowledge of measuring output quality where no ground truth exists, across millions of heterogeneous web pages, and turning the results into model and product decisions.
The Role:
- Ownership of the metrics that define good output for page conversion, structured extraction, web search, agent task completion and document parsing.
- Evaluation harnesses that run thousands of live URLs per change and gate releases.
- Human-annotated, synthetic and adversarial datasets, with agreement statistics.
- Judge design and calibration against human labels; drift and bias control.
- Feedback loop from evaluation deltas into model, routing, prompt and product decisions.
- Benchmarks published alongside product launches, with open datasets where appropriate.
- Delivery directly with the Head of Evaluations and the founders. One hire. No product-management layer.
Tech Estate
- Python, PyTorch/JAX-adjacent tooling. LLM APIs and in-house models. TypeScript/Node API and workers, Redis and PostgreSQL queues, Playwright browser services, Rust document parser. Hugging Face datasets. CI-gated evaluation.
- Coverage, precision, recall and F1 on extracted text; recall@k and MRR for retrieval; task-success and hallucination rate for agents; inter-annotator agreement, stratified sampling and bootstrap confidence intervals.
Requirements
- Built an evaluation system from scratch for LLM, extraction or search output – not run someone else's benchmarks.
- Designed a metric where none existed, validated against human judgement.
- Evaluated at web scale (millions of documents), not a bounded corpus.
- A result of yours that gated a release or reversed a model decision.
- Statistical reasoning that holds under 45 minutes of questioning.
- Python; 3–4+ years in ML, research engineering or data-heavy backend work with evaluation as the primary responsibility. Published or open-source evaluation work considered in place of years.
LLM Evaluation
Benchmarking
Metric Design
Dataset Curation
Statistical Analysis
Information Extraction
Web Scraping
Search Relevance
Python (Programming Language)
PyTorch