AI Jobs Map

Arcadia · Greater Toronto Area, Canada

Staff AI Engineer

Hybriddirectorfull timePosted today
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

agentic-aiobservabilitysystem-designllmmachine-learningpythontypescriptkubernetesdockerawsgcpvllmapache-airflowprefectfine-tuningci/cd

Staff AI Engineer

Location: Toronto, ON (Hybrid - x3 per week in downtown Toronto office)

Type: Full-Time

Comp: $230-300K CAD Base + Benefits

We're partnering with a highly technical AI organization building the infrastructure that powers production AI systems at massive scale. Operating as a AI innovation startup within a much larger global technology business, the team combines the pace, ownership, and greenfield engineering opportunities of an early-stage company with the resources and reach of an established platform serving hundreds of millions of users.

This is not an AI research role. We're looking for a Staff-level software engineer who can define the architecture and technical direction for large-scale AI infrastructure across inference, agentic systems, distributed platforms, and cloud-native backend services.

You’ll help build the infrastructure used to deploy and operate text, voice, vision, code, and domain-specific models, while architecting the runtime, orchestration, safety, and developer tooling required for autonomous AI agents to operate reliably in production.

This is a highly hands-on Staff position. You’ll solve complex engineering problems, lead major technical initiatives, establish platform standards, and influence how production AI systems are built across the organization.

What You'll Do

- Define the technical architecture and direction for production AI platforms across inference and agentic systems

- Architect and build the runtime, orchestration, and developer tooling required for autonomous AI agents

- Design multi-agent coordination systems that enable agents to reason, collaborate, use tools, and execute complex workflows

- Build multi-model serving infrastructure across text, voice, code, vision, and domain-specific models

- Own the complete model lifecycle, including deployment, serving, monitoring, updating, routing, and model swapping

- Optimize inference performance across latency, throughput, reliability, and cost using batching, caching, quantization, and intelligent routing

- Build secure tool-use infrastructure that allows agents to interact safely with APIs, databases, and internal services

- Develop guardrails covering permissioning, sandboxing, prompt injection, data leakage, and human-in-the-loop oversight

- Build evaluation, observability, and monitoring frameworks that measure agent behaviour, detect regressions, and diagnose non-deterministic failures

- Design scalable backend services, APIs, event-driven systems, and durable workflows supporting production AI applications

- Develop SDKs, APIs, and platform capabilities that allow internal teams to build and deploy AI agents quickly and safely

- Lead complex, cross-functional technical initiatives in partnership with Product, ML, Infrastructure, and Security teams

- Establish engineering standards and best practices for agent design, model serving, tool calling, evaluation, and production reliability

- Mentor senior engineers and raise the technical bar across the broader engineering organization

What We're Looking For

- 8+ years of software engineering experience, including significant experience building large-scale backend or distributed systems

- 3+ years of experience building production AI systems, LLM applications, agentic platforms, or machine learning infrastructure

- Demonstrated experience owning the architecture and delivery of complex, business-critical technical initiatives

- Strong understanding of LLM-based agent architectures, including tool use, memory, planning, multi-step workflows, and multi-agent coordination

- Experience building highly reliable distributed systems using event-driven architectures, task queues, state management, and durable workflows

- Experience evaluating production LLM systems, building automated evaluations, detecting regressions, and debugging non-deterministic failures

- Strong programming experience with Python and/or TypeScript, with the ability and willingness to work across both

- Experience with Kubernetes, Docker, AWS or GCP, and modern cloud-native deployment practices

- Experience working with commercial LLM APIs, open-source models, or model-serving technologies

- Understanding of inference optimization techniques such as quantization, batching, caching, routing, and GPU utilization

- Strong understanding of the security risks associated with agentic systems, including prompt injection, privilege escalation, and data leakage

- Exceptional system-design and software-engineering fundamentals

- Strong written and verbal communication skills, with the ability to influence technical direction across teams

- Comfortable operating in an ambiguous, fast-moving environment with substantial ownership and autonomy

- Passion for building production software and infrastructure rather than purely research-focused AI

Nice to Have

- Experience with model-serving technologies such as vLLM, TensorRT-LLM, or Triton

- Experience with Temporal, Airflow, Prefect, or similar workflow-orchestration platforms

- Familiarity with Model Context Protocol (MCP) or other agent communication standards

- Experience with model fine-tuning, LoRA, or quantization

- Experience building AI infrastructure within fintech, healthcare, or another regulated industry

- Experience working with multimodal, voice, or edge-inference systems

- Experience designing human-in-the-loop approval and oversight systems

- Experience building developer platforms, internal SDKs, or CI/CD automation for AI workloads

Why Apply?

- Take Staff-level ownership over the architecture and technical direction of major AI platforms

- Build inference and agentic infrastructure used across a global technology organization

- Work on genuinely greenfield engineering problems spanning LLMs, autonomous agents, distributed systems, security, and cloud infrastructure

- Remain deeply hands-on while influencing engineering standards and mentoring a high-calibre technical team

- Join a startup-style environment with significant autonomy, backed by the scale and resources of an established global platform

More jobs at Arcadia