ai-application-testing.service · golden set loaded

AI Application Testing Services

For products built on LLMs, agents and copilots — where "correct" is a distribution, not a boolean. We evaluate output quality, catch prompt regressions, measure hallucination, and verify agent behavior — wired into your CI like any other test.

▶ Catch a hallucination
golden-set evalshallucination trackingprompt regression in CIagent E2E suites

what it is

Your model changed its mind overnight. Did anyone test that?

AI application testing verifies products where the same input can produce different outputs — and "passing" means staying inside a quality distribution, not matching a fixed string. We score outputs against rubrics, run prompts against golden sets to catch silent regressions, measure how often your RAG system invents facts, and verify that your agents take the right actions. Traditional tests check if the app works. AI tests check if the intelligence behaves.

graded · not matchedrubric-scored quality, thresholds instead of exact strings
regression-safeevery prompt & model change diffed against your golden set
hallucination %measured, baselined, and gated in CI — not hoped about

is this you?

Check what's true. The page will be honest back.

Tick anything that sounds familiar — we'll tell you honestly whether AI testing is your next move.

how we work

Four phases. An artifact at the end of each.

PHASE · WEEK 1

Curate

We build your golden set — representative prompts, edge cases, adversarial inputs and the ground-truth answers — and define the quality rubric that matters for your product.

→ artifact: versioned golden set + scoring rubric, owned by you
PHASE · WEEK 2–3

Evaluate

Automated evals score faithfulness, relevance, safety and tone across the set; RAG answers are checked for grounding; agents are run through multi-step scenarios with tool-use verification.

→ artifact: baseline scorecard — quality, hallucination rate, agent pass rate
PHASE · WEEK 3–4

Gate

Evals wire into your CI as a quality gate: every prompt edit and model bump is diffed against the golden set before it ships. Regressions block the merge, not the user.

→ artifact: CI eval gate + prompt-diff reports on every PR
PHASE · ONGOING

Watch

Production sampling tracks drift over time; the golden set grows as new failure modes appear; hallucination and quality trends are reported like any other reliability metric.

→ artifact: drift dashboard + weekly AI-quality report

see it run

Watch an eval catch what a human reviewer would miss.

A 20-second simulation of an eval run across a golden set — including the moment that matters: a confident, fluent, completely wrong answer, caught and scored.

qacraft@eval — run golden-set.jsonlIDLE
▶ press run — 8 graded prompts, one confident lie
simulation · your real evals report exactly like this

the toolkit

Eval frameworks, your models, your CI.

OEOpenAI Evals RgRagas DEDeepEval PWPlaywright GHGitHub Actions AWAWS Bedrock

We're model-agnostic — OpenAI, Anthropic, Google, open-weights, your own fine-tunes. The evals live in your repo and run against your endpoints, so quality travels with the product, not with us.

deliverables

The AI-quality system your product is missing.

The golden set

Curated, versioned, ground-truthed — the asset every future prompt change is judged against.

Scoring rubrics

Faithfulness, relevance, safety, tone — defined for your domain, applied consistently.

CI eval gate

Prompt and model changes diffed against the golden set on every PR, in your pipeline.

Hallucination tracking

Grounding and citation accuracy measured and baselined, trended over time.

Agent scenario suites

Multi-step, tool-using workflows tested end to end, including failure recovery.

AI-quality dashboard

Drift, quality and hallucination trends your whole team reads — reliability for intelligence.

proof

Caught in eval, not in a screenshot.

SAMPLE DATA — replace with verified client story
9% → 1.2%hallucination rate · RAG support assistant, US SaaS

A golden set of 240 real support questions exposed a 9% hallucination rate on the customer-facing RAG assistant. Grounding fixes and a CI eval gate brought it to 1.2% and blocked three regressions before release in the following quarter — each of which would have shipped a confidently wrong answer to customers.

"QACraft found numerous defects with our SaaS software, as well as making good suggestions for improvements to functionality and usability."
Matt Williams · New Zealand · SaaS quality engagement

pricing logic

Build the eval system once. Gate forever.

fixed-price

Eval Framework Build

Golden set, rubrics and CI eval gate stood up for your AI features — one scope, 3–6 weeks, you own it all.

the AI-quality foundation, installed
continuous

AI Quality Pod

Ongoing eval curation, agent testing and drift monitoring inside a dedicated pod — quality that keeps pace with your model.

priced by scope · scale monthly

The Escaped-Defect Response Guarantee applies to AI too. If a hallucination or regression escapes a gate we certified: 4-hour senior swarm, root cause, new eval cases to harden the golden set — at no additional cost, until it's covered.

straight answers

Asked on every AI call. Answered here.

What is AI application testing?

It verifies software built on LLMs and agents: rubric-scored output quality, prompt regression against golden sets, hallucination and grounding measurement for RAG, and end-to-end agent testing including tool use — wired into your CI like any other quality gate.

How do you test something non-deterministic?

By testing distributions, not single outputs. Golden-set evaluation across many runs, scored rubrics with pass thresholds instead of exact matches, and statistical regression detection that flags when a change shifts behavior beyond tolerance.

How do you measure hallucination?

For RAG systems, every answer is checked for grounding in its retrieved sources, citations are verified, and unsupported claims are flagged. The rate is baselined so each prompt or model change is judged against a known number — not a vibe.

Can you test agents that take actions?

Yes — multi-step plan execution, correct tool selection, recovery from failed calls, and guardrails against unsafe actions, all captured as repeatable scenario suites that run in CI.

We use OpenAI / Anthropic / our own model — does that matter?

No. We're model-agnostic and the evals run against your endpoints, so switching or upgrading models becomes a measured decision instead of a leap of faith.

Is this separate from our normal app testing?

It's a layer on top. The app shell still needs functional, performance and security testing; AI testing adds the intelligence layer. Most clients run both — often as one pod.

Ship AI features you can actually trust.

Scope your AI test plan in 60 seconds — or bring your flakiest prompt or newest agent to a 30-minute call and leave with an eval strategy and one number.

Book a Scoping Call