Founding AI Quality & Release Engineer | AI/LLM | Python / TypeScript | Agentic Testing / CI/CD | Must have startup experience
🚨 Only candidates with hands-on AI-native quality engineering experience - testing or evaluating agentic/LLM-based systems within the last 12 months - along with startup experience will be considered.
Location: Fully Remote (US or Canada) - optional office space for those who want it
Package: $180,000 - $200,000 + meaningful equity - Can go to $220,000 if you tick every box
Eligibility: Must have full work authorization in the US or Canada - No visa sponsorship or transfer support available
🚨 Please only apply if you have commercial experience with ALL of the following 🚨
- 7+ years as a quality, release, or platform engineer
- Hands-on AI-native quality engineering (testing or evaluating agentic / LLM-based systems) within the last 12 months
- Strong GitHub Actions and GitHub-based CI/CD (container builds, pipeline automation)
- Python and/or TypeScript
- Startup / scale-up experience
- Happy as a pure hands-on IC who writes code most days
Join an early-stage, venture-backed AI company building agentic systems that are transforming how technical teams work. This is the first dedicated quality-infrastructure hire - a true founding role where you own the entire verification and release system end to end.
This isn't traditional QA. The agents run the tests here - your job is to build the systems that verify every build, gate releases automatically, and make shipping safe and predictable from merge to customer.
What You'll Own:
- AI-powered verification - The trustworthiness of the test signal: agent-run harnesses, LLM graders, regression coverage, and automated release gates. Closing the gaps that let bugs escape to customers.
- Release and environment engineering - The mechanics of shipping: release pipelines, CI health, artifact publication, and the reliability of demo and staging environments.
What this role is NOT:
- It is not manual test execution - you fix the underlying system, not staff the symptom
- It is not a release gatekeeper approving builds on intuition - you build systems that make readiness visible and stop unsafe releases automatically, with evidence
- It is not a hidden 24/7 on-call role - business-hours escalations only, with production on-call staffed separately as it matures
Required Background:
- 7 - 15 years across quality, release, or platform engineering
- Proven experience building, testing, or evaluating systems where AI agents do the actual work
- Deep GitHub Actions / GitHub CI/CD (container builds, pipeline automation)
- Python and/or TypeScript (flexible on language for the right person)
- Experience with containerized infrastructure and modern CI/CD practices
- Track record of taking full ownership of a process or system - deciding and executing independently
Must-Haves:
- Startup mindset with the autonomy to make and implement decisions without close guidance
- Comfortable owning trustworthiness of test signal in an AI-native environment
- Strong software engineering depth - this is an engineering role, not traditional QA
Bonus Experience:
- Familiarity with a major cloud or data platform (AWS, Azure, Snowflake, Databricks)
- Modernizing legacy QA / DevOps practices inside a larger org
- Building internal AI-assisted tooling (test generation, log analysis, agent workflows)
This is a rare shot to define quality and release from day one at a company moving fast. If you tick every box above and want to own the system that makes shipping safe, get in touch for a fast response.