An agent can return a fluent, correct-sounding reply and still file the wrong ticket, double-charge a card, or burn the budget on retries. Reading transcripts will never catch that, because the damage lives in the tool calls and state changes a transcript never shows.
Fluent answers can still fail production
An AI agent can produce a polished response and still create the wrong ticket, call an unnecessary API, retry a payment action, or exceed budget. Transcript review catches only part of production behavior: answer, decision sequence, tool calls, external state changes, latency, and cost.
Regression testing treats a release as a behavioral contract, not a prompt comparison. It asks whether the agent completed the task, used permitted tools correctly, avoided prohibited effects, stayed within limits, and can be stopped safely. A small runnable suite around these questions is more useful than subjective “good answer” examples.
The suite need not mirror all traffic. Cover expensive mistakes, high-volume workflows, known failures, and decisions affected by the change. Its purpose is blocking customer incidents, not certifying universal intelligence.
A release contract beats transcript review
Define each case with a scenario, initial state, available tools, and pass/fail rubric. Expected text can be part of the rubric, but rarely all of it. A support agent replacing a damaged order might need to verify eligibility, create one replacement request, explain the next step, and avoid an unsupported refund.
Record separate expected outcomes:
- Task result: the requested business outcome exists and meets constraints.
- Tool trace: required calls have valid arguments; forbidden calls do not occur; ordering respects dependencies.
- External state: the expected record exists, with no duplicate or stray records.
- Response quality: the user receives an accurate, policy-compliant explanation matching the action.
- Resource envelope: latency, model use, and paid tool activity stay within budget.
A final answer is not completion proof. An agent may claim a case was updated after a failed tool call, or update it correctly while telling the user the wrong resolution. Both fail differently.
Make rubrics runner-evaluable. “Helpful and accurate” helps reviewers but is not a gate; use observable checks: required fields, a named API method called once, no forbidden method, and a response containing the approved disposition. Where language judgment is unavoidable, use a bounded evaluator rubric with explicit failure labels and retain failures for human review.
One passing run is weak evidence
Agent behavior is partly stochastic even with fixed workflow, tools, and input. Sampling settings, model updates, retrieval ranking, service latency, and retries can alter a trace. One success proves success was possible, not that the release is dependable for its exposure.
Repeat every nondeterministic scenario and store every trace, not just aggregate scores. Set repetition count before running based on action risk: a read-only research agent can use a lighter check than one modifying customer records. Specify both a success requirement and blocking failure pattern. One prohibited side effect can be a hard failure even if every other repetition succeeds.
Task Success Rate = successful task runs / attempted task runs
This needs a companion measure. Nine correct outcomes and one duplicate cancellation are not “90% good” when duplication is unacceptable. Track hard separately from recoverable failures, and inspect clusters by input type, tool, model route, or retry branch.
Repeated runs also reveal fixture-based false confidence. Identical cached retrieval and a mock API that never times out measure a narrow path. Keep a stable baseline, then add controlled variants: missing records, ambiguous identifiers, delayed tool responses, malformed results, and duplicate user requests. The goal is not random chaos, but evidence that known production conditions do not let the agent escape its authority.
Tool traces reveal the hidden failure
A tool-enabled agent is a transaction coordinator with a language interface. Scrutinize its trace and response: each call’s intent, arguments, ordering, result handling, and side effects. Tool names alone are insufficient; create_case with the wrong account ID is not a near miss.
For every state-changing action, assert it was allowed for the task, preconditions were met, parameters came from validated data, it occurred the permitted number of times, and resulting state matches the claimed outcome. Record correlation IDs so failed tests connect traces to downstream logs.
Test idempotency separately. Agents retry after timeouts, but a timeout does not show that a downstream service missed the request. Simulate an ambiguous tool response and require an idempotency key, a state read before retry, or review routing. Blindly retrying a write can create duplicate orders, duplicate emails, or conflicting account updates.
Verify side effects where real outcomes are visible. Mocks support deterministic branch coverage but cannot prove field mapping or deduplication in an integration. Use sandbox accounts or isolated test tenants for end-to-end checks. For irreversible actions, test the authorization boundary and handoff path rather than performing the action.
Score outcomes, traces, and spend separately
A compact scorecard ties release discussion to observable behavior and prevents cost or safety regressions from hiding in a blended score.
| Measure | Calculation | Release question |
|---|---|---|
| Task Success Rate | successful runs / attempted runs | Did it complete intended work? |
| Trace Compliance Rate | compliant traces / attempted traces | Did it respect tool and policy constraints? |
| Side-Effect Error Rate | runs with invalid state changes / attempted runs | Did it cause forbidden or duplicate changes? |
| Cost per Successful Task | total model and tool cost / successful runs | Is it affordable at expected volume? |
| Tail Latency | chosen high-percentile completion time | Will slow paths breach user experience or worker timeout? |
Compare a candidate with a versioned baseline using the same test set, tool configuration, and evaluation window. Lower cost is not a win if success falls; higher success may be suspect if a more expensive fallback model handled every case.
Cost gates include model tokens, external tool charges, retrieval requests, and retries. Four search-endpoint calls before an answer can pass quality while hurting unit economics. For varying complexity, compare groups rather than one mean: a small group of long failed runs can be hidden by many easy requests.
Do not treat evaluator scores as ground truth. Automated graders can miss policy violations, reward confident language, or share the agent’s ambiguity. Use them for repeatable, clear criteria; audit disagreements; and reserve deterministic assertions for consequential actions.
Risk changes the test shape
A read-only internal research agent can start with representative prompts, citation checks, and budget limits. A customer-facing agent with account access needs stronger identity, permission, and trace checks. An agent writing to operational systems needs state verification, idempotency scenarios, and an immediate kill switch. Strictness should rise with authority, not merely traffic.
Maturity changes useful regressions. Early failures often involve missing instructions, poor tool schemas, and incomplete event logging. Later cases involve routing changes, tool-version drift, retrieval-corpus changes, model-provider updates, and fallback-path interactions. A suite frozen after launch becomes a museum of old bugs.
Use incidents and near misses as test sources. Convert each into a minimized scenario with sanitized data, define expected trace and final state, and retain it in the regression set. This preserves concrete failure memory without making every release review forensic.
Rollback must be tested before release day
A passing suite reduces, not removes, risk. Production has inputs, timing conditions, and dependency failures no pre-release set fully reproduces. Releases need a reversible control plane: versioned prompts and tool schemas, feature flags, traffic routing, audit logs, and a documented owner able to disable the capability.
Set measurable rollback criteria before expanding traffic: a prohibited tool action, sustained breach of the task-success floor, abnormal cost per successful task, or a trace indicating permission bypass. Define who decides, which switch they use, and whether in-flight work may finish. “We can roll back quickly” is not a control until exercised.
Use canaries when isolation by tenant, workflow, or traffic slice is possible. Watch leading signals such as tool errors and trace violations, plus lagging signals such as customer corrections and downstream reconciliation failures. If a write cannot be safely reversed, use shadow evaluation or human approval until evidence supports broader authority.
The durable suite stays small and opinionated
The strongest production suite is not a leaderboard. It is a short set encoding non-negotiable behavior: tasks the agent must complete, actions it must never take, resource limits it must respect, and stop conditions.
Keep every case runnable; version fixtures and rubrics; save traces from every run; and treat failures as engineering evidence, not model folklore. An agent is ready for wider exposure when its gate detects a bad answer, tool call, state change, expensive detour, and failed rollback path. A suite that checks only the final answer leaves tool behavior, downstream state, and rollback untested.
