AI Jobs Map

Ikigaienablers · Singapore

AI - Quality Assurance Specialist (Senior / Lead individual contributor)

director$38,030 – $240,000 / yearPosted today
Apply on IndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmagentic-aipythoncybersecurityanthropiclinuxbashgitrest-apitest-automationci/cdgenerative-aiapi-testingragplaywrightpytestseleniumdockerperformance-testing

Job Title: AI Quality Assurance Specialist – Senior / Lead

Location: Singapore

About the Role

We are looking for a Senior / Lead AI Quality Assurance Specialist to build and operate an evidence-backed quality assurance capability for LLM-based conversational AI systems.

This is not a conventional QA execution or prompt-writing role. You will evaluate whether an AI system not only works technically, but whether its responses are correct, relevant, grounded, faithful, useful and trustworthy.

You will work with an advanced LLM through an agentic coding environment to test and evaluate another LLM-based conversational system, combining automation, exploratory investigation, independent verification and human judgment.

What You Will Do

- Build and evolve an AI assurance and quality strategy for LLM-based conversational systems.

- Design quality gates, test coverage, evidence standards and release-quality criteria.

- Perform single-turn and multi-turn conversational testing.

- Test realistic, ambiguous, multilingual and challenging user scenarios.

- Develop repeatable Python/API-based automation and adaptive exploratory tests.

- Evaluate AI responses for:

- Correctness

- Groundedness

- Faithfulness

- Relevance

- Hallucinations

- Conversational understanding

- Usefulness

- Workflow behaviour

- Verify AI responses against authoritative sources, product data and observed system behaviour.

- Build and maintain regression coverage from exploratory findings.

- Evaluate and calibrate LLM-as-a-Judge approaches and understand the risks of one LLM evaluating another.

- Preserve technical evidence and create traceable findings for developers.

- Produce clear quality and maturity recommendations for product and senior management.

- Identify potentially significant guardrail, misuse or prompt-injection behaviour and work with Cybersecurity/governance teams when required.

Technical Environment

You should be comfortable working from the terminal and using an agentic coding CLI such as Codex CLI, Claude Code or an equivalent tool.

Hands-on experience with:

- Python

- Linux / Bash

- Git

- REST APIs

- JSON

- SSE / streamed APIs

- Browser/network evidence

- Dependency management

- API automation

- Test automation

- CI/CD

Required Experience

- 5+ years of QA, test automation, software testing or quality engineering experience.

- Hands-on experience testing or evaluating Generative AI / LLM-based systems.

- Strong Python programming and automation experience.

- Experience with API testing and automation.

- Experience testing chatbots, conversational AI, RAG or multi-turn LLM systems.

- Practical understanding of hallucination, groundedness, faithfulness, relevance and response quality.

- Experience building or substantially improving a QA/test/evaluation capability, rather than only executing an existing test suite.

- Strong analytical and investigative skills.

- Ability to distinguish between a successful API call, a fluent response and a semantically correct response.

- Strong written and verbal communication skills.

Preferred Experience

- RAG evaluation

- LLM-as-a-Judge

- Golden/reference datasets

- Evaluator calibration

- DeepEval

- Ragas

- Promptfoo

- LangSmith / Langfuse

- Playwright

- PyTest

- Selenium

- Docker

- CI/CD

- Agentic AI / AI agents

- Prompt-injection and guardrail testing

- Performance testing

- Evidence dashboards / QA reporting

What We Value

We are particularly interested in candidates who can:

- Challenge the AI system rather than simply follow scripted test cases.

- Decide when to automate, investigate or manually verify.

- Question the output of an evaluator LLM rather than treating it as unquestionable.

- Establish sources of truth in an unfamiliar business domain.

- Turn exploratory discoveries into repeatable regression tests.

- Connect raw evidence → technical finding → developer action → management recommendation.

Pay: $3,169.23 - $20,000.00 per month

Benefits:

- Additional leave

- Health insurance

- Parental leave

- Professional development

- Promotion to permanent employee

Experience:

- QA/QC: 4 years (Required)

- AI: 2 years (Required)

- GenAI: 1 year (Required)

- RAG: 1 year (Required)

Location:

- Singapore (Required)

Work Location: In person

More jobs at Ikigaienablers