Summary
Build the quality and evaluation engine for next-generation, compliant AI systems! If you are a senior quality and test automation engineer with 1 to 5 years of experience, eager to pioneer testing frameworks for probabilistic LLMs and agentic workflows, our client wants to hear from you. This ground-floor role makes you the first dedicated QA leader in a high-velocity startup. Step into an environment backed by top-tier investors where your evaluation pipelines, CI/CD gates, and rigorous testing standards will directly ensure the safety and reliability of life-changing enterprise platforms.
Description
Our client is a startup redefining pharmaceutical marketing with compliant, reference-backed automation, partnering with top life sciences manufacturers and dozens of major brands. In this early-stage environment, you'll step in as the foundational owner of software quality and AI evaluation.
- The Challenge: Build novel evaluation frameworks and automated test suites for non-deterministic AI agents where traditional pass/fail assertions do not apply.
- The Difference: Solve cutting-edge technical problems at the intersection of machine learning, compliance, and product reliability in a fast-growing startup.
- Core Responsibilities: Design LLM eval systems, establish CI/CD quality gates, build end-to-end UI suites with Playwright, and execute hands-on exploratory testing.
- Why Apply: Enjoy ground-floor equity, incredible ownership over the quality function, and the opportunity to define how reliable AI systems are built.
Skills
- 1 to 5 years of professional software engineering and testing experience, featuring strong Python skills and real-world implementation of CI/CD pipelines.
- Demonstrated expertise in end-to-end and UI test automation using Playwright, paired with a rigorous manual and exploratory QA discipline.
- Proven experience—or a deep appetite—for testing non-deterministic, ML, or LLM-based systems using golden datasets, LLM-as-judge, and statistical variance analysis.
- Strong instinct for error analysis, tracing system failures, and translating complex agent behaviors into permanent regression tests.