About the Role
This Staff Engineer role sits at the core of an applied agentic AI product that automates complex, multi-step workflows across desktop engineering tools used by hardware engineers every day. Reporting directly to the CTO, you will own the agent intelligence layer end-to-end and lead a small, focused team of AI engineers, a user researcher, and domain expert contractors. The work you do here directly determines how much real-world value the product delivers to enterprise customers.
What You'll Do
-
Lead development of the core agent intelligence layer that executes multi-step workflows across complex desktop engineering software.
-
Own the full product loop: define agent capabilities from user stories, build implementations, and benchmark against real workflows.
-
Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating on completion metrics.
-
Set and enforce per-task token budgets and track cost per completed workflow to ensure commercial viability.
-
Build rigorous, reproducible evaluation infrastructure grounded in validated user stories.
-
Lead user story mapping and validation through engineer interviews and close collaboration with domain experts.
-
Translate validated user stories into testable evals, closing the loop between user research and agent benchmarking.
-
Own agent architecture decisions including tool-calling strategies, state management, error recovery, model routing, and context management.
-
Act as a player-coach: write production code, review designs, unblock the team, and raise engineering standards.
-
Collaborate cross-functionally with integrations, product, and customers during POCs to align agent behavior with real-world usage.
What We're Looking For
-
7+ years of software engineering experience, including at least 2 years building LLM-based agents that take real-world actions.
-
Deep experience designing LLM application architectures: model selection, context and window management, retrieval, and orchestration patterns.
-
Proven ability to build evaluation and benchmarking frameworks measuring task completion, cost efficiency, and failure modes.
-
Strong Python skills and hands-on familiarity with LLM tooling including function calling, tool APIs, observability and tracing, and evaluation frameworks.
-
Experience shipping AI or LLM tooling on top of proprietary engineering data or desktop engineering software, such as agents or MCP servers over CAD, PLM, or simulation platforms.
-
Technical leadership experience setting direction for small teams of 3 to 6 engineers while continuing to write and review production code.
-
Experience with desktop automation or programmatic control of applications such as COM or similar interfaces.
-
Domain familiarity with mechanical engineering, CAD, CAE, PLM, or adjacent engineering software industries.
-
Understanding of enterprise deployment constraints on locked-down corporate workstations.
-
Comfort operating in a fast-paced, high-intensity early-stage environment with significant customer demand.
Compensation & Benefits
Base salary range: $160,000 to $250,000 USD annually, plus equity. Visa sponsorship is not available for this role.
Location
On-site in San Francisco, California, United States.