We’re working with a fast-growing AI company looking for a QA Engineer to own how the quality of its AI products is measured, tested and improved. This isn’t traditional QA — you’ll work at the intersection of AI evaluation, software engineering and product quality, building the frameworks, datasets and automated tests that determine whether LLM-powered features are genuinely accurate, reliable and ready for production. You’ll work closely with engineering and product teams to identify where AI systems fail, understand why, and turn those findings into measurable improvements.
Must Haves
Strong software engineering or technical background, ideally with hands-on experience working with AI/ML systems
Experience building testing, evaluation or quality frameworks for AI-powered products
Strong Python or similar programming experience, with the ability to build evaluation tooling and automation
Good understanding of LLMs and Generative AI, including how model outputs should be tested and evaluated
Experience creating evaluation datasets, test cases and regression suites for complex AI features
Ability to analyse model outputs, identify failure patterns and root causes, and turn findings into measurable improvements
Strong understanding of accuracy, reliability and consistency when evaluating AI systems
Excellent analytical and problem-solving skills, particularly when dealing with ambiguous or difficult-to-measure problems
Comfortable working closely with Engineering and Product teams and communicating technical findings clearly
Nice to Haves
Experience evaluating RAG / retrieval systems, including the quality and relevance of retrieved information
Experience testing AI agents, tool-calling or multi-step AI workflows
Knowledge of different LLM evaluation methodologies, metrics and benchmarking approaches
Experience implementing automated AI regression testing within development or CI/CD workflows
Experience with human-in-the-loop evaluation alongside automated testing
Knowledge of prompt evaluation and comparing changes across different models or prompts
Experience building internal AI evaluation tooling, dashboards or monitoring
Previous experience working with AI products where accuracy and reliability are particularly important