How to Verify an AI Agent Fix After Deployment

Published Sources reviewed October 6, 20267 min read

To verify an AI agent fix after deployment, count the specific failure the fix targeted in production conversations before and after the release, using the same detection rule, enough conversations on each side, and a check for anything else that changed in between.

The answer is one of four: verified, not fixed, confounded, or not enough evidence yet. This guide shows how to get to an honest one. It is part of the AI agent optimization hub.

The short version

A fix is verified by the failure it targeted, measured the same way before and after, with nothing else changing the agent in between. Anything less is a hypothesis, and an honest inconclusive result beats a false win.

1. Decide what will count before you ship

Write down four things before the release: the failure pattern the fix targets, the rule that detects it, the window you will compare, and what result counts as success or failure. Deciding afterward invites picking whichever metric moved.

Measure the targeted pattern, not an overall quality score. A fix for wrong-specialist routing should be judged on wrong-specialist routing; an overall score can move for reasons that have nothing to do with the fix.

2. Build a baseline with the same detector you will use after

Count the targeted pattern over a fixed window before the release, for example seven days. Use the exact detector that will measure the after window: the same labeling rules, and if an LLM judge does the labeling, the same judge version and prompt.

If the detector changes mid-measurement, the comparison is between two rulers, not two agents. Converra pauses a verdict when its failure detector's version changed inside the window and waits for same-version evidence.

3. Mark exactly what shipped and when

Record the deployment time and the version of the prompt or configuration that went live. That marker is what separates before from after. Without it, conversations from a slow rollout or a cached prompt end up on the wrong side.

Converra groups conversations by the prompt version that actually served them where it can, so a conversation counts as after only if the new version handled it.

4. Compare like with like

Traffic changes on its own. A product launch, a marketing campaign, a holiday, or a new customer segment can change which requests arrive, and with them how often a failure occurs. Before reading a drop as a fix, check whether the mix of request types moved between the two windows, and compare within the segments the fix targets.

Converra does not automatically detect traffic-mix shifts today, so this check is yours. When it is feasible, a concurrent split of live traffic between the old and new version removes most time-based differences; see production A/B testing for agents.

5. Rule out other changes in the same window

List everything that could have changed the agent between the two windows: other prompt edits, other merged pull requests, a model version change, a new tool or API behavior, or a change to retrieval data. If any of them could affect the targeted failure, the result cannot be credited to the fix alone.

Converra marks a result confounded when the agent's prompt was edited outside the deployed fix during the window, or when another merged change touched the same prompt. It does not automatically flag model-version or tool changes for a prompt fix, so check release notes and your own deploy log. Model swaps are verified separately, on quality and cost at the same prompt.

6. Check that nothing else got worse

A fix can remove the targeted failure and create another one. Compare the overall failure rate and the other top patterns across the same windows. Converra applies this guardrail to merged changes: if the targeted patterns improved but the overall failure rate rose significantly, the result is reported as regressed.

Regression cases tested before release lower this risk but do not remove it, because production users find paths that test suites miss. The guide to regression testing covers the pre-release side.

7. Read the result honestly

Small samples move a lot by chance. A common rule of thumb is to treat a difference smaller than about two standard errors as no detectable change; the NIST engineering statistics handbook explains how sample size, significance, and the size of the change you want to detect interact.

The table shows the outcomes Converra reports, as of October 2026. Fixes Converra ships are judged by recurrence of the targeted failure. Changes your team merges outside Converra are judged by comparing seven days before and after the merge.

Outcome
Not enough evidence
Fixes Converra ships
Monitoring until at least 5 diagnosed conversations after the change
Changes you merge yourself
Monitoring until the 7-day after window closes; insufficient data below 10 diagnosed conversations per side
Outcome
Verified
Fixes Converra ships
At least 15 conversations after the change and zero recurrences of a failure that occurred before it
Changes you merge yourself
Not used; a measured drop is reported as improving
Outcome
Improving or likely fixed
Fixes Converra ships
Fewer recurrences than before, or zero recurrences with fewer than 15 conversations
Changes you merge yourself
Targeted failure share fell by at least 1 percentage point or 25%
Outcome
Not fixed or flat
Fixes Converra ships
The targeted failure recurred at least as often as before
Changes you merge yourself
Targeted failure share moved by less than 1 point and less than 25% (flat); without targeted patterns, a change under two standard errors is flat
Outcome
Regressed
Fixes Converra ships
A previously verified fix whose failure returned
Changes you merge yourself
Targeted failure rose, or the overall failure rate rose significantly
Outcome
Confounded
Fixes Converra ships
The prompt was edited outside the fix during the window
Changes you merge yourself
Another merged change touched the same prompt, or targeted patterns moved in opposite directions

Worked example: reading a before and after

The numbers in this example are illustrative, not customer data. A fix targets refund promises made after an empty order lookup. In the seven days before release, the detector finds the pattern in 38 of 600 diagnosed conversations, 6.3%. In the seven days after, it finds 9 of 580, 1.6%.

The drop is 4.8 percentage points. The standard error of the difference is about 1.1 points, so the drop is more than four standard errors: not chance. Nothing else touched the prompt, the request mix was stable, and the overall failure rate did not rise. Under Converra's rules for a fix it shipped, this reads as improving rather than verified, because 9 recurrences remain. The honest next step is to look at those 9 conversations, not to declare victory.

A real example: in the Salespeak case study, unsupported pricing, VAT, and infrastructure claims did not recur on live traffic after the April 23 deploy, and routing failures fell 74% after the April 25 deploy. Limits: one customer and one agent; the case study publishes percentages rather than conversation counts; and the two changes shipped two days apart, so each was measured against its own target pattern.

What to do with an inconclusive or negative result

Not enough evidence: wait for more conversations rather than lowering the bar, and avoid shipping another change to the same behavior in the meantime. Confounded: either wait for a clean window or ship the next change on its own so it can be measured. Not fixed: return to diagnosis; the cause was probably in a different layer than the fix changed. The guide to turning production traces into improvements covers that loop.

A not-fixed result is useful. It tells you the diagnosis was wrong while the evidence is fresh, instead of letting a persistent failure hide behind a closed ticket.

Frequently asked questions

How long should I wait before judging a fix?

Long enough to collect a comparable number of diagnosed conversations before and after. Converra waits for a 7-day window after a merged change and at least 5 conversations before judging a fix it shipped; low-traffic agents need longer.

Can I use an LLM judge to detect the failure?

Yes, if the judge is the same version and prompt in both windows and you have checked its labels against human review. A judge that changes between windows makes the comparison meaningless.

Is an A/B test better than a before-and-after comparison?

When you can run one, yes. Splitting live traffic between versions at the same time removes most time-based confounders such as campaigns or seasonality. A before-and-after comparison is still useful when a split is not possible, as long as you check for other changes.

What if the model provider updated the model during the window?

Treat the result as confounded unless you can show the update did not affect the targeted behavior. Pinning model versions in production avoids this problem.

Does a fix that passed simulation still need production verification?

Yes. Simulation chooses the best candidate against simulated users. Only production traffic shows whether real users stopped hitting the failure.

Oren Cohen, founder of Converra

Written by

Founder of Converra. Previously founded Buildup (acquired by Stanley Black & Decker) and, as VP Product Growth at Totango, owned AI end-to-end from design through production.

Get a verdict on every fix you ship

Converra diagnoses the failure, tests a bounded fix against your current agent, ships what you approve, and measures the targeted failure on live traffic: verified, not fixed, or confounded.