MLOps Engineer — AI Platform & Infrastructure
AHL + Saaf AI — Mortgage Lending, Reimagined
Saaf AI is building the future of mortgage lending by combining cutting-edge AI with proven lending operations. Saaf AI is part of American Heritage Lending, a top-10 private lender processing billions in loan volume, with 15+ years of growth. We are backed by some of the largest asset managers and funds and are growing fast.
We’re not experimenting with AI. We’re deploying it. From underwriting to document processing to borrower experience, all are shaped by AI. If you want to work somewhere that uses AI as a core building material — you’re in the right place.
The Role
We’re hiring an MLOps Engineer to own the infrastructure our AI runs on.
Our AI systems are already in production and already touching real loans. The next chapter is making them scale, deploy safely, and change quickly — infrastructure that absorbs real production load, a promotion path that makes shipping boring, and internal tooling that lets the team improve AI behaviour without an infrastructure engineer in the loop for every change.
This is not a role where you inherit a mature platform and keep the lights on. You’ll take AI services that were built to prove the idea and turn them into systems that hold up under load, deploy predictably across environments, and give the rest of engineering a safe, fast path to production. You own the arc from “this works on one machine” to “this is a platform the whole company builds on.”
This role is right for you if:
-
You’ve taken AI or ML systems from a single deployment to infrastructure that scales — and you’ve been on call for the result
-
You treat infrastructure as a product with users, not a ticket queue
-
You think deploys should be boring, reversible, and frequent — and you’ve built the systems that make that true
-
You’d rather remove yourself from the critical path than be the person everyone has to ask
What You’ll Build
Scalable AI Serving Infrastructure
Our AI workloads are moving from early-stage deployment to production scale. You’ll design what they run on:
-
Re-architect how AI services are deployed and run — from single-host setups to horizontally scalable, orchestrated infrastructure that absorbs traffic spikes without degrading
-
Design for the specific realities of LLM-backed workloads: long-running requests, streaming responses, bursty concurrency, expensive downstream calls, and upstream rate limits
-
Own capacity, autoscaling, and unit economics — you should be able to say what we spend per unit of work, and why
Deployment & Environment Promotion
Shipping AI changes should be a routine, reviewable event — not a coordinated risk:
-
Build the promotion path from development through pre-production to production, with environments that are consistent and reproducible rather than each one a special case
-
Make everything version-controlled and reviewable — infrastructure, configuration, and application logic on the same rails
-
Ship CI/CD that gives engineers fast deploys with real rollback, staged rollout, and change history you can audit
Internal Platform & Self-Serve Tooling
The highest-leverage thing you’ll build is the thing that lets other people ship without you:
-
Build internal tooling that lets engineers — and technically-minded teammates outside engineering — define, modify, and test AI workflow logic without touching deployment plumbing
-
Design the guardrails that make that safe: validation, versioning, review, staged promotion, and a clear path back when something is wrong
-
Treat internal users as customers. Success is measured in how many changes ship correctly without you being involved
Reliability, Observability & Operations
You’ll care about whether the system actually holds — not just whether it deployed:
-
Instrument the AI stack end to end: latency, throughput, failure modes, cost, and output-quality signals
-
Build alerting that catches real degradation rather than noise, and incident practice that changes the system instead of assigning blame
-
Own secrets, access control, and environment isolation in a regulated industry where those things carry real consequences
What We’re Looking For
Must-Have
-
Production infrastructure ownership: You’ve deployed and operated containerized services in production under real traffic. You’ve dealt with rollouts, autoscaling, failure isolation, resource limits, and rollback — not just a first deploy.
-
Infrastructure as code: You build environments declaratively and reproducibly. Manually-configured infrastructure makes you uncomfortable, and you know how to migrate away from it without a big-bang rewrite.
-
CI/CD depth: You’ve built pipelines that others depend on daily — automated testing, promotion between environments, safe rollout, and fast recovery.
-
Strong Python: Our AI stack is Python. You should be able to read and change application code, profile it, and fix it — not just package and deploy it.
-
Operational judgment: You can debug a production problem across application, network, and infrastructure boundaries, and you know the difference between a fix and a workaround.
-
Ownership without guardrails: You drive work end to end — design, implementation, deployment, monitoring, and the follow-up when it doesn’t behave.
-
Fast iteration: You ship incrementally and improve. You’re impatient with process that doesn’t reduce risk.
Strong Preferences
-
Experience operating LLM or ML workloads in production — where cost, latency, and non-deterministic output are all live concerns
-
Cloud deployment experience (AWS preferred) — you can containerize, deploy, scale, and operate the systems you own
-
Experience in fintech, lending, insurance, or another regulated industry — secrets management, access control, audit trails, PII handling
-
Built internal developer platforms or self-serve tooling that non-infrastructure engineers actually used
-
Full-stack comfort — you can build a lightweight UI or internal tool when the platform needs one, rather than waiting for someone else to
Nice-to-Have
-
Experience designing multi-environment promotion pipelines where configuration and logic move together with code
-
Familiarity with agent orchestration frameworks and what it takes to run them reliably
-
Observability for non-deterministic systems — tracing, evaluation signals, quality monitoring alongside standard telemetry
-
Inference optimization: model serving, batching, caching, and cost reduction strategies
-
Exposure to workflow automation tooling and low-code builders
Why Saaf
Mortgage AI is still wide open. Unlike ad tech or e-commerce, where AI optimization is a rounding error, in mortgage lending an AI system that works changes whether a family gets their home. The problems are complex, the data is rich, and the solutions don’t exist yet.
The infrastructure problems here are real ones. Non-deterministic workloads, strict regulatory constraints, real financial stakes, and a company shipping AI changes faster than most infrastructure can absorb. This is not maintaining someone else’s platform — it’s designing the one everything else will be built on, at the moment when that decision matters most.
AI-first means AI-first. Every engineer here uses AI to build AI. Claude Code, agentic workflows, AI-assisted code review — we don’t add AI to our process, AI is our process. The team you’ll join is already operating at the frontier.
Direct leverage on everything we ship. Every AI feature this company builds runs on what you build. When deploys get faster and safer, every engineer here gets faster. When the platform scales, every loan we process moves quicker.
Compensation & Benefits
-
Competitive compensation
-
Unlimited PTO
-
Remote-first with flexible hours
-
$2,000/year professional development budget
-
Home office setup stipend