We are looking for a hybrid Systems Engineer and AI Researcher to lead the development of our agent evaluation framework and post-training data pipelines. You will design sandboxed execution environments, high-throughput reinforcement learning feedback loops, and robust evaluation setups that handle non-deterministic agent outputs.
Core Responsibilities
-
Build Stateful Agent Environments: Design deterministically verifiable, stateful sandboxes (web, OS, API, database) where agents can execute 50+ step action trajectories safely.
-
Scale Post-Training & RL Pipelines: Implement high-throughput post-training infrastructure (RLHF, Direct Preference Optimization, Process-Supervised Reward Models) for dynamic policy optimization.
-
Design Enterprise Benchmarks: Formulate evaluation metrics and automated grading harnesses that catch agent drift, hallucination, and loops in realistic enterprise environments.
-
Systems Optimization: Keep latency low and compute efficiency high across distributed GPUs and sandboxed runtime environments.
Requirements
-
Production AI Experience: Track record of deploying evaluations, RL loops, or sandboxed agent environments into production.
-
Systems Polyglot: Deep systems background (Python, Rust, C++, Go, TypeScript, CUDA)—you pick up new tools and frameworks within days.
-
San Francisco On-Site: In-person collaboration at our SF office to iterate quickly with founders and domain experts.
-
Pragmatic Problem Solver: Comfortable navigating raw paper implementations, undocumented SDKs, and custom distributed training setups.