How to Verify an AI Agent Fix After Deployment
To verify an AI agent fix after deployment, count the specific failure the fix targeted in production conversations before and after the release, using the same detection rule, enough conversations on each side, and a check for anything else that changed in between.
The answer is one of four: verified, not fixed, confounded, or not enough evidence yet. This guide shows how to get to an honest one. It is part of the AI agent optimization hub.
The short version
A fix is verified by the failure it targeted, measured the same way before and after, with nothing else changing the agent in between. Anything less is a hypothesis, and an honest inconclusive result beats a false win.
1. Decide what will count before you ship
Write down four things before the release: the failure pattern the fix targets, the rule that detects it, the window you will compare, and what result counts as success or failure. Deciding afterward invites picking whichever metric moved.
Measure the targeted pattern, not an overall quality score. A fix for wrong-specialist routing should be judged on wrong-specialist routing; an overall score can move for reasons that have nothing to do with the fix.
2. Build a baseline with the same detector you will use after
Count the targeted pattern over a fixed window before the release, for example seven days. Use the exact detector that will measure the after window: the same labeling rules, and if an LLM judge does the labeling, the same judge version and prompt.
If the detector changes mid-measurement, the comparison is between two rulers, not two agents. Converra pauses a verdict when its failure detector's version changed inside the window and waits for same-version evidence.
3. Mark exactly what shipped and when
Record the deployment time and the version of the prompt or configuration that went live. That marker is what separates before from after. Without it, conversations from a slow rollout or a cached prompt end up on the wrong side.
Converra groups conversations by the prompt version that actually served them where it can, so a conversation counts as after only if the new version handled it.
4. Compare like with like
Traffic changes on its own. A product launch, a marketing campaign, a holiday, or a new customer segment can change which requests arrive, and with them how often a failure occurs. Before reading a drop as a fix, check whether the mix of request types moved between the two windows, and compare within the segments the fix targets.
Converra does not automatically detect traffic-mix shifts today, so this check is yours. When it is feasible, a concurrent split of live traffic between the old and new version removes most time-based differences; see production A/B testing for agents.
5. Rule out other changes in the same window
List everything that could have changed the agent between the two windows: other prompt edits, other merged pull requests, a model version change, a new tool or API behavior, or a change to retrieval data. If any of them could affect the targeted failure, the result cannot be credited to the fix alone.
Converra marks a result confounded when the agent's prompt was edited outside the deployed fix during the window, or when another merged change touched the same prompt. It does not automatically flag model-version or tool changes for a prompt fix, so check release notes and your own deploy log. Model swaps are verified separately, on quality and cost at the same prompt.
6. Check that nothing else got worse
A fix can remove the targeted failure and create another one. Compare the overall failure rate and the other top patterns across the same windows. Converra applies this guardrail to merged changes: if the targeted patterns improved but the overall failure rate rose significantly, the result is reported as regressed.
Regression cases tested before release lower this risk but do not remove it, because production users find paths that test suites miss. The guide to regression testing covers the pre-release side.
7. Read the result honestly
Small samples move a lot by chance. A common rule of thumb is to treat a difference smaller than about two standard errors as no detectable change; the NIST engineering statistics handbook explains how sample size, significance, and the size of the change you want to detect interact.
The table shows the outcomes Converra reports, as of October 2026. Fixes Converra ships are judged by recurrence of the targeted failure. Changes your team merges outside Converra are judged by comparing seven days before and after the merge.
- Outcome
- Not enough evidence
- Fixes Converra ships
- Monitoring until at least 5 diagnosed conversations after the change
- Changes you merge yourself
- Monitoring until the 7-day after window closes; insufficient data below 10 diagnosed conversations per side
- Outcome
- Verified
- Fixes Converra ships
- At least 15 conversations after the change and zero recurrences of a failure that occurred before it
- Changes you merge yourself
- Not used; a measured drop is reported as improving
- Outcome
- Improving or likely fixed
- Fixes Converra ships
- Fewer recurrences than before, or zero recurrences with fewer than 15 conversations
- Changes you merge yourself
- Targeted failure share fell by at least 1 percentage point or 25%
- Outcome
- Not fixed or flat
- Fixes Converra ships
- The targeted failure recurred at least as often as before
- Changes you merge yourself
- Targeted failure share moved by less than 1 point and less than 25% (flat); without targeted patterns, a change under two standard errors is flat
- Outcome
- Regressed
- Fixes Converra ships
- A previously verified fix whose failure returned
- Changes you merge yourself
- Targeted failure rose, or the overall failure rate rose significantly
- Outcome
- Confounded
- Fixes Converra ships
- The prompt was edited outside the fix during the window
- Changes you merge yourself
- Another merged change touched the same prompt, or targeted patterns moved in opposite directions
Worked example: reading a before and after
The numbers in this example are illustrative, not customer data. A fix targets refund promises made after an empty order lookup. In the seven days before release, the detector finds the pattern in 38 of 600 diagnosed conversations, 6.3%. In the seven days after, it finds 9 of 580, 1.6%.
The drop is 4.8 percentage points. The standard error of the difference is about 1.1 points, so the drop is more than four standard errors: not chance. Nothing else touched the prompt, the request mix was stable, and the overall failure rate did not rise. Under Converra's rules for a fix it shipped, this reads as improving rather than verified, because 9 recurrences remain. The honest next step is to look at those 9 conversations, not to declare victory.
A real example: in the Salespeak case study, unsupported pricing, VAT, and infrastructure claims did not recur on live traffic after the April 23 deploy, and routing failures fell 74% after the April 25 deploy. Limits: one customer and one agent; the case study publishes percentages rather than conversation counts; and the two changes shipped two days apart, so each was measured against its own target pattern.
What to do with an inconclusive or negative result
Not enough evidence: wait for more conversations rather than lowering the bar, and avoid shipping another change to the same behavior in the meantime. Confounded: either wait for a clean window or ship the next change on its own so it can be measured. Not fixed: return to diagnosis; the cause was probably in a different layer than the fix changed. The guide to turning production traces into improvements covers that loop.
A not-fixed result is useful. It tells you the diagnosis was wrong while the evidence is fresh, instead of letting a persistent failure hide behind a closed ticket.
Frequently asked questions
How long should I wait before judging a fix?
Long enough to collect a comparable number of diagnosed conversations before and after. Converra waits for a 7-day window after a merged change and at least 5 conversations before judging a fix it shipped; low-traffic agents need longer.
Can I use an LLM judge to detect the failure?
Yes, if the judge is the same version and prompt in both windows and you have checked its labels against human review. A judge that changes between windows makes the comparison meaningless.
Is an A/B test better than a before-and-after comparison?
When you can run one, yes. Splitting live traffic between versions at the same time removes most time-based confounders such as campaigns or seasonality. A before-and-after comparison is still useful when a split is not possible, as long as you check for other changes.
What if the model provider updated the model during the window?
Treat the result as confounded unless you can show the update did not affect the targeted behavior. Pinning model versions in production avoids this problem.
Does a fix that passed simulation still need production verification?
Yes. Simulation chooses the best candidate against simulated users. Only production traffic shows whether real users stopped hitting the failure.

Written by
Oren CohenFounder of Converra. Previously founded Buildup (acquired by Stanley Black & Decker) and, as VP Product Growth at Totango, owned AI end-to-end from design through production.
Related reading
What is AI agent optimization?
The hub: the full loop from diagnosis to a production verdict.
Production verification in Converra
How Converra issues verified, not fixed, and confounded verdicts.
Why loops stop before verification
Why a shipped change remains a hypothesis until production evidence returns.
Get a verdict on every fix you ship
Converra diagnoses the failure, tests a bounded fix against your current agent, ships what you approve, and measures the targeted failure on live traffic: verified, not fixed, or confounded.