AI-Assisted QA & Release Assurance
AI Evaluation & Testing
Test the AI inside your product, not just the app around it. We evaluate answer quality, grounding, permissions and failure handling in agents, chatbots and LLM features, and set clear release criteria.

Who needs AI evaluation
Your AI feature looks good in a demo, but nobody can yet say how it behaves across real questions, hostile inputs and missing data, or whether the next model or prompt change makes it worse.
- Teams about to put a chatbot or assistant in front of customers
- Product owners changing model, provider or prompts who need a before-and-after comparison
- Companies letting an agent read internal records or trigger actions in other systems
How each behavior gets evaluated
Illustrative plan for a customer-facing assistant; methods shift with your use case and risks.
| Automated evals | Expert review | Red-team testing | Production signals | |
|---|---|---|---|---|
| Answer accuracy | Primary method | Primary method | Rarely used here | Supporting or sampled |
| Grounding and citations | Primary method | Supporting or sampled | Rarely used here | Supporting or sampled |
| Permission boundaries | Supporting or sampled | Primary method | Primary method | Supporting or sampled |
| Prompt injection | Supporting or sampled | Supporting or sampled | Primary method | Supporting or sampled |
| Refusals and fallbacks | Primary method | Supporting or sampled | Supporting or sampled | Primary method |
- Primary method
- Supporting or sampled
- Rarely used here
Testing AI features is a different job
A chatbot can pass every functional test and still give a confident wrong answer, reveal data it shouldn't, or follow instructions hidden in an uploaded file. AI features need their own evaluation: datasets of real and difficult cases, checks on answer quality and grounding, tool and data permissions, prompt-injection exposure, failure handling, and regression checks whenever a model or prompt changes. That's different from testing an ordinary app that AI helped build, which Software QA & Testing covers.
AI helps grade at scale; experts set the bar
How AI assists
- Generates variations of real user questions you approve, including paraphrases, edge cases and adversarial inputs.
- Grades large batches of answers with model-based scoring, calibrated against answers labeled by experts.
- Clusters failures by cause, such as missing context, a wrong tool call or an ignored instruction.
- Compares runs across model, prompt or retrieval changes and flags where behavior got worse.
What our experts own
- Experts define what a correct, safe and useful answer is, together with your domain specialists.
- Model-based grades are spot-checked by people, and a grader that disagrees with experts is fixed or dropped.
- People decide which tools and data an agent may use, and where a human must approve or take over.
- Release criteria and the go/no-go call stay with your team and ours, not with a score alone.
What you receive
Evaluation you can rerun on every change
Evaluation dataset
Real and synthetic cases labeled with expected behavior: common questions, edge cases, out-of-scope requests and past failures.
Answer quality and grounding
Accuracy, relevance and tone, and whether answers are supported by your sources, with citations checked and hallucinations flagged.
Tool and data permission tests
What an agent can read, change or trigger, tested against least privilege, approval points and the paths for human takeover.
Prompt-injection exposure
Direct and indirect injection through messages, uploaded files, web pages and retrieved documents, rated by what an attacker could reach.
Failure handling and fallbacks
Behavior when the model times out, a provider is down, retrieval finds nothing or confidence is low: fallbacks, clear messages, hand-off to people.
Regression checks and release criteria
Automated evaluations that rerun when models, prompts or data change, with agreed criteria that decide whether a change ships.
How an AI evaluation runs
- 01
Define good behavior
Agree use cases, risks, quality criteria and permissions with your team and domain experts, and collect real examples you allow us to use.
- 02
Build the dataset
Assemble and label evaluation cases, including difficult and adversarial ones, and set the release criteria each change must meet.
- 03
Evaluate and probe
Run automated evaluations with expert review, red-team permissions and prompt injection, and trace failures to prompts, retrieval, tools or the model.
- 04
Gate and monitor
Add evaluations to your release process, report results with recommended fixes, and keep the dataset current as the product changes.
Two ways to work with AI tools
AI helps draft tests and investigate defects. Choose where it may process your code and test data.
- Private / Local AI Engineering
Privately hosted models inside infrastructure you control or an agreed isolated environment.
Discuss with this package - Claude Code / OpenAI Codex Engineering
Claude Code and/or OpenAI Codex with cloud settings your organization approves.
Discuss with this package
Not sure? We'll recommend one during scoping. Compare the packages in detail
Typical evaluation requests
Typical scenarios we scope, not client case studies.
Help-center assistant before public launch
A team wants an assistant that answers from its help articles. We turn real support questions into a labeled set, grade answers and citations, add out-of-scope and adversarial cases, and agree release criteria before launch.
Moving a feature to a lower-cost model
A product owner wants to switch a feature to a cheaper model or provider. We run the same evaluation set on both, show where answers get better or worse and how latency changes, and leave the decision with them.
An agent that can change records
A company lets an internal agent update orders. We red-team its tool permissions and approval points, plant instructions in documents it reads to test indirect injection, and recommend where a person must approve before it acts.
Where AI evaluation stops
- Changing prompts, retrieval or the model itself isn't part of evaluation; we recommend fixes, and your team makes them or we scope them under LLM Integration.
- Attacks on the servers, APIs and cloud account behind the feature belong to Penetration Testing; here we probe the model's instructions, tools and data access.
- Latency, rate limits and model costs at peak traffic are measured in Performance Testing.
- Legal or regulatory sign-off on your AI use isn't included; results are evidence for your own risk and compliance review.
How AI evaluation connects to design, QA and operations
Design: supervision people can use
Designers shape how users see sources, uncertainty and errors, and how your staff review, correct or take over from an agent.
QA: the rest of the product too
Software QA covers the app around the AI, such as sign-in, payments, roles and integrations, tested against requirements like any release.
Operations: quality in production
Production logs, user feedback, latency and cost feed monitoring and new evaluation cases, with alerts when quality starts to drift.
Ongoing: re-evaluate every change
Model upgrades, prompt edits and new data sources are evaluated before release, as part of AI operations agreed with you.
FAQ
Frequently asked questions
How is this different from testing an app that AI helped build?
An app written with AI coding tools still behaves like ordinary software, so Software QA & Testing covers it. AI Evaluation is for products where a model generates answers or takes actions at runtime. Outputs vary, so we measure quality across many cases, test permissions and plan for failure. Many products need both, and we scope them together.
Which models and AI stacks can you evaluate?
Features built on hosted models such as OpenAI or Claude, open-weight models such as Llama or Mistral, retrieval pipelines and agent frameworks. The method doesn't depend on the provider, so you can also use it to compare models or providers on your own cases.
Can evaluation stop the AI from ever being wrong?
No. Models can still make mistakes. Evaluation shows where and how they fail, so you can set release criteria, add guardrails and fallbacks, design human review where errors are costly, and notice quickly when quality changes.
Where are our prompts, logs and evaluation data processed?
Your feature's own model calls go to the provider you already use. AI tools we use in the work, such as model-based grading, run within the boundary you choose: Private / Local AI Engineering on infrastructure you control, or Claude Code / OpenAI Codex Engineering under agreed data-handling terms. Personal data in logs is masked before use, as agreed with you.
Related reading
- LLM Evals You Can Show a Client
Golden sets, regression dashboards, and how we stop prompt tweaks from silently breaking last month’s wins.
- RAG Pipelines That Do Not Hallucinate Your Docs
Chunking, citations, and evals we use when a client wants an LLM that actually stays on their knowledge base.
- Observability for AI Features in Production
Traces, token budgets, and user-visible fallbacks when the model or the vector store blinks.
See how your AI behaves before you ship it
Tell us what your AI feature does, which model it uses and what worries you. We'll suggest an evaluation plan and release criteria.


