Staff Platform Engineer — Agentic AI
Harrison Clarke is partnering with a fast-scaling vertical AI company — building autonomous agents that don't just assist, but execute — to hire a Staff Platform Engineer who will design the internal platform that lets this company ship agents into production faster than anyone else in the space.
This isn't "platform engineering for a company that happens to use AI." This is platform engineering where the workload is non-deterministic, the runtime is an LLM, and the deployment target is a system that makes real decisions in your customers' businesses with real consequences.
What You'll Own
- Internal developer platform — designing and building the foundational abstractions, tooling, and self-service workflows that product engineers use to define, test, deploy, and operate autonomous agents across verticals
- Agent runtime infrastructure — building the execution layer that orchestrates multi-step, stateful agent workflows — handling branching logic, tool invocation, retry semantics, human-in-the-loop checkpoints, and graceful failure across long-running asynchronous processes
- LLM gateway and model management — designing the abstraction layer that sits between product code and model providers (OpenAI, Anthropic, internal fine-tunes) — routing, fallback, rate limiting, cost attribution, prompt versioning, and response caching
- Observability for non-deterministic systems — building tracing, logging, and evaluation infrastructure purpose-built for agent workflows where traditional request/response observability doesn't capture what matters. You'll instrument chain-of-thought reasoning, tool call sequences, decision branching, and outcome quality — not just latency and error rates
- CI/CD and testing for agents — creating deployment pipelines and testing frameworks for systems where "does it work" isn't a binary — including regression detection for prompt changes, eval harnesses for agent behavior, and canary deployment strategies for non-deterministic outputs
- Multi-tenancy and isolation — designing the platform primitives for securely running agents across customer environments with strict data isolation, credential management, and per-tenant resource controls
What You Bring
- 7+ years of software engineering experience, with significant depth in platform engineering, infrastructure, or developer experience — you've built internal platforms that other engineers genuinely want to use
- Strong production experience with Kubernetes and container orchestration — you've operated clusters at scale, built custom operators or controllers, and understand the platform's sharp edges
- Deep familiarity with event-driven and workflow orchestration systems — Temporal, Argo, Step Functions, Inngest, or equivalent — and an opinion on where they break for AI workloads
- Hands-on experience with observability infrastructure — you've designed and operated tracing (OpenTelemetry, Jaeger), structured logging, and custom instrumentation for complex distributed systems
- Production exposure to LLM-powered applications — you understand the operational realities of working with model APIs: token economics, prompt management, non-deterministic outputs, latency variance, and provider reliability
- Proficiency in Go, Python, or TypeScript — ideally comfortable across more than one, because the platform serves teams writing in all of them
Nice to Have
- Direct experience building infrastructure for LLM-based agents, copilots, or autonomous workflow systems — you've already hit the walls and know where the conventional tools stop working
- Background in LLMOps or ML platform engineering — model routing, prompt registries, evaluation frameworks, A/B testing for generative outputs
- Familiarity with agent frameworks (LangGraph, CrewAI, AutoGen, or custom) and an informed opinion on what belongs in the framework vs. the platform
- Experience designing plugin or tool-use architectures — systems where agents dynamically invoke external capabilities with structured inputs and outputs
Why This Opportunity
- The platform IS the product moat — in agentic AI, the company that ships reliable, observable, controllable agents fastest wins the vertical. The platform you build is the single biggest lever for that velocity
- Genuinely novel problem space — platform engineering for autonomous agents is being invented right now. There's no off-the-shelf Helm chart for "deploy an agent that negotiates contracts on behalf of a customer." You'll define patterns that the rest of the industry adopts later
- Vertical focus, horizontal complexity — the company goes deep in a specific industry, which means you build against real, demanding use cases with paying customers — not a horizontal demo looking for a problem
- Staff-level ownership and influence — you'll set the technical direction for the platform, make build-vs-buy decisions with real budget behind them, and mentor a growing engineering team. Your architecture choices compound for years
- Well-capitalized and scaling — backed by top-tier investors with the runway to build this properly, not cut corners