AI Jobs Map

FutureFit AI · United States

Staff ML Platform Engineer (MLOps)

Remoteseniorfull time$172,000 – $215,000 / yearPosted yesterday
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

mlopsllmrailsa/b-testingobservabilitydata-engineeringdata-scienceci/cdartificial-intelligencesqlpythonapache-airflowdbtpostgresqlredshiftmongodbmachine-learningawspytorchs3

Come join our Data team!

High velocity, high trust, and high impact with a will to win.

If that resonates deeply with you, this could be your next career move. We're seeking someone who leads with humility, pursues audacious goals, and is motivated by meaningful impact on people and the world.

At FutureFit AI, our core mission is to help more people get to better jobs faster and cheaper, with a specific focus on those facing barriers to opportunity. Our work helps resolve the growing issue of economic inequality, ensuring that no one is left behind in the future of work. Our AI-powered platform brings efficiency and insight to workforce development, replacing outdated systems and unlocking human potential at scale.

Ready to make an impact? Apply today.

Important note: Data shows that men typically apply when meeting 3/10 requirements, while women often wait until it's 10/10. We encourage you to apply if you see a strong (not necessarily perfect) fit.

The Opportunity

We're seeking a Staff ML Platform Engineer (MLOps) to build the platform our ML and LLM-powered products run on. Our ML footprint has grown fast, but the layer underneath it has not kept pace. You'll own that layer end to end: how models get built, deployed, evaluated, and served; how compute and environments get provisioned and managed; how our LLM calls get routed and optimized for cost; and how we know quickly when a recommender goes down or goes off the rails.

This is a build role and an operate role: when a model regresses or a recommendation looks wrong, you can trace it back to the inputs that produced it and help fix it.

Your Role

Our ML footprint has grown quickly: batch models, real-time recommendation models, LLM-powered features, and daily pipelines processing every available job across the US and Canada. What we haven't built is the platform underneath it: consistent compute and environments, a disciplined path from experiment to production, cost-aware routing across LLMs, and the monitoring that tells us fast when something breaks.

You'll assess our current pipelines and ML workflows with clear eyes, decide what to build and in what order, then build it. This is greenfield platform work with direct influence on production models and how our ML team operates, and it comes with real operational ownership: you will be close enough to the running systems to debug them, not one step removed.

What You'll Own

- Assessment and plan: Evaluate our current pipelines, data architecture, and ML workflows, and produce a prioritized, opinionated plan for what needs to change.

- Platform foundations and optimization: Own compute provisioning and environment management, keep training and serving environments reproducible, keep frameworks and packages current across services and model images, and tune latency, throughput, and spend, all without destabilizing production.

- LLM infrastructure and smart routing: Build the layer our LLM features run on, including smart routing that sends each request to the cheapest model that can handle it well, plus the prompt and response evaluation needed to prove quality holds when we route down.

- Experimentation and safe rollout: Give us a real discipline for A/B testing models before they are fully ramped: shadow deploys, canaries, holdouts, and success criteria agreed in advance, so a model earns its way into production instead of being switched on.

- Observability and traceability: Know within minutes when a recommender goes down or starts drifting, and be able to explain why: model and data monitoring, alerting, regression detection, lineage, and enough traceability to reproduce a questionable recommendation on demand or trace a prediction back to the inputs that produced it (we currently use Braintrust; comparable tooling counts too).

- Hands-on operations: Stay close enough to the running systems to operate them. You will work with the team to keep models online, but when something breaks in production you can dive in and help fix it, including models other people built.

- Data and feature infrastructure: Own how features are computed, stored, and served consistently between training and inference, and take on the data engineering the team needs along the way.

- Standards, not sole ownership: Establish the deployment and monitoring standards the rest of the team can run with. You are building shared ownership, not becoming the only person who keeps things online.

Required Experience

- Staff-level, hands-on experience in MLOps, ML platform, or ML infrastructure (we're also open to Data Platform Engineer, ML Infrastructure Engineer, or Data Scientist backgrounds with strong platform ownership: the title on your last resume matters less than what you actually built)

- Experience standing up MLOps practice end to end: CI/CD for models, experiment tracking, model registries, deployment workflows, and monitoring

- Production experience with LLM-based systems: serving, prompt and response evaluation, routing across models and providers, and managing cost and latency tradeoffs. If you've done this with traditional ML systems and can show you pick up LLM tooling fast, that counts too

- Experience operating models in both batch and real-time serving contexts

- Hands-on with compute provisioning and environment management: containers, reproducible training and serving environments, and keeping frameworks and packages current on a fast-moving stack without breaking production

- Experience running controlled model experiments in production: A/B tests, shadow or canary deploys, holdouts, and the judgment to set success criteria before ramping

- Depth in observability and traceability for production ML: drift and regression detection, alerting that catches a recommender going down or going off the rails, lineage, and the ability to trace a prediction back to the inputs that produced it and reproduce it after the fact

- Hands-on operational experience: you have carried the pager or its equivalent, debugged production ML incidents under time pressure, and fixed systems you did not originally build

- Comfort doing the data engineering the platform needs: pipelines, feature computation and storage, and keeping training and serving features consistent

- A track record of walking into complex, fast-grown systems, diagnosing the real problems, and materially improving them

- Strong systems design ability: you can translate product needs into durable architecture and stay close enough to the code to build it yourself

Bonus Points

- Interest in growing into model development yourself. This role starts on the platform side, but the line between platform and modeling is thin here, and we would rather hire someone who wants to cross it

- Feature store experience. We are early here, so you would be shaping it rather than inheriting it

- Experience evaluating AI/ML observability or LLM evaluation vendors, with judgment on when to buy versus build

- Background in mission-driven, workforce, or government-adjacent data environments

- Comfort mentoring a small data and engineering team while you build

Our Tech Stack for Data

- Languages: SQL, Python

- Data orchestration and transformation: Airflow, dbt

- Data storage and warehousing: PostgreSQL, Redshift, MongoDB

- Machine learning and model serving: AWS SageMaker (PyTorch models, artifact upload to S3, model registration), serving real-time and batch inference

- Visualization and reporting: Looker, Quicksight

- Infrastructure: AWS (S3, Redshift), GitHub Actions for CI/CD

Your Education

Your alma mater isn't our focus. Your grit, hunger, and drive are. If you learn continuously, tackle challenges head-on, and know your strengths and gaps intimately, you're our person.

Location

Remote (CA/US). Toronto-based candidates are welcome to work from our office at 325 Front St West if they prefer, but it's optional, not a hybrid requirement.

Travel Expectations

Approximately 2-3 trips per year, including our company off-site in August.

Compensation

Pay Range: USD $172,000-$215,000 (United States) / CAD $172,000-$220,000 (Canada)

As a remote-first company, we benchmark to the national market for comparable roles at institutionally-funded startups, targeting the middle of market. Bands reflect applied experience, with room to grow.

If establishing the standards a growing ML org runs on, rather than inheriting a finished one, is the kind of problem you want, let's talk.

Hiring Journey

At FutureFit AI, our hiring process is designed to help you assess whether this role and our culture are the right fit based on your unique skills, mindset, and experiences. We move fast and work with intensity, so we want you to get a real sense of that from the start.

Each journey includes a mix of interviews and a performance challenge. For this role, that might look like:

- Online Application

- Initial Screen with Talent Acquisition

- Interview with Hiring Manager

- Performance Challenge

- Final 1:1 Interviews

- Final Decision

Generally, this entire process takes around 6 weeks, although the timing can vary due to specific candidate circumstances.

More jobs at FutureFit AI

  • FutureFit AI · United States

    6 days ago

    Software Engineer

    Remotesenior$150,000 – $185,000system-designreacttypescriptnode.js+4
  • FutureFit AI · Canada

    6 days ago

    Director of Product

    Remotedirector$150,000 – $185,000agentic-ai
  • FutureFit AI · New York, NY

    8 days ago

    Brand & Marketing Manager

    Hybrid$80,000 – $130,000artificial-intelligence
  • FutureFit AI · New York, NY

    13 days ago

    Strategic Account Manager

    Remote$125,000 – $160,000artificial-intelligence
  • FutureFit AI · New York, NY

    13 days ago

    Technical Project Manager

    $125,000 – $165,000oauthgraphqletlgithub+4