AI Jobs Map

AKKA · United States

Applied AI Research Engineer: Benchmarking & Performance Economics

Remotefull timePosted yesterday
Apply on IndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

machine-learningagentic-aiopenaistatisticsvllmllmpythonsystem-designa/b-testing

Description

This is an applied research role, not a machine learning science role. You will not be advancing the state of the art in model architecture. You will be producing decision-grade numbers: benchmarks, cost models, and calculators that determine what we build, what we tell customers, and what we are willing to claim in public.

The work sits at the intersection of three disciplines that rarely overlap in one person — systems engineering, inference-stack depth, and honest experimental design. If you have ever read a vendor benchmark and immediately knew which confound made it meaningless, this is the seat for you.

What you'll own

-
The token economics of agentic execution. What it actually costs to run an autonomous agent to task completion, and how that compares to LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK, Ray, and whatever exists six months from now. Tokens per call is the wrong unit and you know it, the unit is cost per completed task at a fixed success rate, and getting there means measuring steps, retries, tail latency, success rate, and run-to-run variance.

- Routing and model-selection economics. Where each routing strategy sits on the cost/quality frontier, what the router costs to run, and how a 5% misroute rate compounds into a 40% task failure rate across a thirty-step trajectory. Routing quality gets evaluated at the trajectory level here, not the request level.

- Small language model economics. The crossover analysis: at what volume, task narrowness, and quality tolerance does a fine-tuned 3B beat a frontier API call? You own the full stack: data curation, training, evaluation, serving, and the ongoing cost of drift and retraining, and the judgment to say "don't train this" when the numbers say so.

- GPU and serving-stack modeling. A parameterized calculator that takes model size, quantization, batch size, context distribution, KV cache footprint, and concurrency, and returns throughput, memory headroom, latency percentiles, and cost per million tokens. Validated against measured ground truth, not derived from spec sheets. It accounts for prefix caching, because agentic workloads resend the same system prompt hundreds of times and a model that ignores that is wrong by a multiple.

- The cost of governance. What policy evaluation, guardrails, evaluator calls, and human-in-the-loop suspension actually cost in latency and tokens. Enterprise buyers assume the number is bad. Being able to state it precisely turns an objection into a differentiator.

- The benchmark harness as a product. Pinned versions, controlled cache state, variance reported rather than averaged away, full configuration captured with every result, reproducible by a third party. Most benchmark work in this industry does not survive scrutiny. Yours will be the asset that does.

What We're Looking For

-
You are constitutionally unwilling to publish a number you can't defend. When the result is inconvenient, you report the result.

- Fair to the competition, to the point of discomfort. You will implement a rival framework properly, in its own idiom, and you will spend the extra two days doing it well. An unfair benchmark is worse than no benchmark, it's a liability the moment someone reproduces it.

- You have been wrong in public and corrected it. We consider this a qualification, not a blemish. Someone who has never had a result challenged has never had a result that mattered.

- Scoping discipline. Most of this work arrives as an ill-posed question. Turning "is our routing better?" into a measurable experiment with a defined success criterion is the core skill, and it happens before any code is written.

- Suspicious of your own instruments. You assume the harness is lying until you've proven otherwise. You measure the measurement.

- Writer. The output of this role is arguments supported by evidence. A correct result that can't be explained to an engineer, an architect, and a buyer has delivered a fraction of its value.

- Builder, not administrator. Akka is small enough that you write this playbook rather than inherit it. That should energize you.

- Real experimental design and statistics. Confidence intervals, confound control, sample sizing, and the judgment to know when a difference is noise. This is the most common gap in otherwise strong candidates and the one we're least able to flex on.

- Inference stack depth, hands-on. vLLM, SGLang, or TensorRT-LLM in anger. Quantization formats and their quality trade-offs. Continuous batching, paged attention, prefix caching, GPU memory arithmetic.

- Strong systems engineering. Python for the harness, plus enough comfort in a systems language to read and profile serving code. What you build here is infrastructure others extend. It needs to be maintainable, not notebook-shaped.

- Evaluation design for non-deterministic systems. LLM-as-judge and its failure modes, task-completion rubrics, trajectory-level evaluation. You know that a benchmark measuring the wrong thing precisely is worse than one measuring the right thing roughly.

- Cost modeling that survives interrogation. Comfort building models a finance-literate stakeholder will go through line by line: amortization, utilization assumptions, marginal versus fully-loaded cost.

- Hands-on time across multiple agent frameworks, sufficient to build the same task idiomatically in each.

- Distributed systems intuition, concurrency, backpressure, failure modes, tail latency behavior.

- Published benchmarks, teardowns, or analyses under your own name that held up to challenge.

Familiarity with event-driven architectures or actor-model platforms.

-

A PhD. A publication record. Novel architecture or training-methods research. Deep theoretical ML. We are not screening for any of these, and candidates should not self-select out for lacking them.

We weigh rigor and approach above years-of-experience checkboxes. The interview loop is built to surface:

- A benchmark you have actually built (published or internal) and how you handled variance, what confounds you controlled for, and what you got wrong.

- A work sample: we hand you a specific performance or cost claim from a competitor's marketing page and ask you to design the experiment that would confirm or refute it in a day. We are watching for how fast you find the unstated assumptions and how honestly you scope what you'd leave unmeasured.

- How you reason about fairness when the comparison makes us look worse than we'd like.

- How you'd explain your most technical result to someone evaluating our platform against three alternatives.

A candidate with a strong research pedigree and no hands-on serving experience is not a fit for this seat, regardless of institution. A candidate who has spent three years profiling inference workloads, has shipped a fine-tuned model, and can explain exactly why their last benchmark was subtly wrong probably is.

Benefits

-
Competitive salary with performance-based incentives.

- Comprehensive health and wellness benefits.

- Opportunities for professional development and continuous learning.

- Flexible remote working environment.

- Collaborative, inclusive, and innovative company culture.

- A transparent, distributed work environment with a strong focus on work-life balance.

- Challenging work that interacts with innovative applications used by millions.

- A collaborative culture that attracts the "brightest minds" in the technology community.

About Akka

Akka builds the platform that enterprises use to develop, deploy, and operate agentic AI systems. Our customers run mission-critical workloads on top of Akka, many of them in regulated industries (financial services, healthcare, public sector, defense-adjacent) where the cost of an outage, a data leak, or a botched upgrade is measured in regulatory exposure, not just churn. They are not buying SaaS. They are buying a managed runtime they can stake their business on.

The agentic AI market is being sold on assertion. Vendors publish throughput numbers with no methodology, cost comparisons with no workload definition, and benchmarks built by people who learned the competing framework the afternoon before they measured it. We intend to be the company whose numbers hold up when someone tries to prove them wrong.

Our Vision:

Distributed systems that feel local.

Our Mission:

To make it simple to build, run, and evaluate agentic systems.

Our Values:

- We’re Authentic: We value transparency and genuine communication, without politics or games. We're honest and assume good intentions, cultivating trust and accountability within our organization and in our interactions with the community. Authentic defines who we want to work with.

- We’re Customer-focused: We value customer outcomes above all else. By prioritizing our customers' interests, and meeting them where they are today, we help ensure their success. We are dedicated to deeply understanding our customer’s needs, anticipating challenges, navigating time constraints, and striving to exceed expectations. Customer-focused defines how we decide where we spend our time.

- We’re Nonconventional: We value fearless innovation by challenging the status quo and embracing alternative approaches. Continuous learning and a growth mindset aimed at improving ourselves, our company, and our products, drives us to push boundaries and explore new solutions. Guided by a bias for action, we leverage industry and customer insights to inspire fresh ideas, enabling optimal future offerings. Nonconventional defines how we approach and test a strategy.

- We’re Persistent: We value excellence through continuous experimentation and courageous problem-solving. We recognize that achieving success often demands approaching challenges with tenacity and taking calculated risks to achieve leading-edge solutions. Persistent defines how we want to be perceived by others.

More jobs at AKKA