AI Jobs Map

Haven AI · Boston, MA

Founding AI Engineer - Agent Platform & Intelligence

entry_levelfull timePosted today
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmobservabilityagentic-ai

Company Description:

___________________

Haven is building the first continuously learning cognitive layer that preserves institutional memory and powers full AI transformation for organizations. Applicants joining Haven will be truly moving the needle in AI infrastructure, contributing to a research-grade technology that helps power a future where companies can unlock reliable, context-rich AI at scale. We have a goal to improve how every company works with AI, and to have fun doing it.

Role Description:

_______________

You will work directly alongside the Cofounder, CTO to own the Haven systems that transform enterprise signals into persistent organizational understanding and allow agents to safely reason and act on top of that understanding. You’ll be responsible for improving the quality of Haven's intelligence for our users, so you should be deeply comfortable with LLM/agent observability systems. You'll actively use the traces, datasets, experiments, and scores it produces to help improve the platform.

Qualifications and Experience:

___________________________

Observability and evaluation

- Debug agent behavior using traces, spans, sessions, generations, prompt and version history, datasets, and experiment results.

- Define quality metrics for accuracy, groundedness, hallucination, retrieval, graph, memory, entity resolution, tool use, and trajectory quality.

- Create and curate golden datasets and benchmark scenarios for the agent and intelligence capabilities that matter most.

- Inspect agent trajectories to catch wrong tool choices, stale context, bad retrieval, unnecessary calls, incorrect sub-agent selection, or failure to stop.

- Own the semantic interpretation of eval results and decide what counts as an acceptable regression.

- Partner closely with other engineers who owns our observability and evaluation infrastructure: telemetry pipelines, experiment execution, dashboards, and the operational Langfuse integration.

Agent harness and runtime

- You've built or owned an agent runtime: orchestration and execution loops, planning, tool calling, a tool registry, MCP support, agent state, and model routing.

- You're comfortable with the hard parts of production agents: parallel agents and sub-agents, handoffs, long-running jobs, checkpoints and resumability, human approvals, interruptions, retries, and failure recovery.

- You treat safety as part of the runtime: permission-aware tools, sandboxed execution, agent identity, tenant-aware authorization, read/write boundaries, execution limits, network restrictions, and full auditability.

Memory, provenance, and contradiction handling

- You've designed persistent profiles for customers, products, employees, projects, teams, and the organization as a whole, and know when short-term, long-term, episodic, semantic, entity, and historical memory each earn their place.

- You believe every fact should carry its provenance: source, timestamp, original and derived evidence, confidence, and the evidence that supports or contradicts it.

- Experience with contradiction detection, evidence weighting, source reliability, confidence scoring, and temporal reasoning.

Retrieval and context construction

- You can decide what context actually matters: how much to retrieve, which memories to use, which graph paths to traverse, and when raw evidence is required.

- You build the minimum high-quality context an agent needs to reason correctly, rather than dumping large stores into an LLM.

- Experience with model and provider abstraction, model routing, fallbacks, and task-, cost-, and latency-aware model selection.

Email the team at [email protected]!

Similar roles in Boston