Did My AI Agent Fix Work? Fix Verification Calculator
Enter how often the targeted failure happened before and after your fix. The calculator shows whether the change is bigger than chance, how many conversations you need to detect a change of that size, and how Converra would label the fix.
It runs in your browser; nothing you enter is sent anywhere. The method and its limits are explained below the calculator.
Enter your counts
Count the specific failure the fix targeted, with the same detection rule in both windows.
- Failure rate
- 6.3% → 1.6%
- Change
- -4.8 pts
- Standard errors
- 4.3
The drop is larger than two standard errors, so chance alone is an unlikely explanation. Before crediting the fix, rule out other changes and traffic shifts in the same window.
To detect a change this size reliably (5% significance, 80% power), you need about 259 conversations in each window.
How Converra would label this fix: Improving: fewer recurrences than before, but not zero.
How to get the right counts
Count one specific failure, the one the fix was meant to remove, not an overall quality score. Use the same detection rule, and if an LLM judge does the labeling, the same judge version, in both windows.
Use windows of similar length on either side of the release, and count only conversations the new version actually handled after it shipped. The guide to verifying an AI agent fix after deployment covers baselines and release markers in detail.
How the calculation works
The calculator compares two failure rates with the standard normal approximation for independent proportions. The standard error of the difference is the square root of p1(1 − p1)/n1 + p2(1 − p2)/n2. A difference of more than about two standard errors is unlikely to be chance alone at the usual 5% level.
The sample-size estimate uses the two-proportion formula from the NIST/SEMATECH engineering statistics handbook, at 5% two-sided significance and 80% power. It answers: if the true rates were what you observed, how many conversations per window would you need to detect the difference reliably?
What the calculator cannot tell you
A real drop is not proof that your fix caused it. Other prompt edits, a model version change, a new tool behavior, or a shift in the kind of requests users sent can all move a failure rate. Rule those out before crediting the fix.
The math assumes each conversation is an independent observation. Repeated conversations from the same user, or a bug that fails in bursts, make the true uncertainty larger than the calculator shows. When you can, a concurrent split of live traffic between the old and new version is a stronger test than a before-and-after comparison.
How Converra labels a fix
For fixes it ships, Converra judges recurrence of the targeted failure rather than a significance test. It waits for at least 5 diagnosed conversations after the change. A fix is verified when at least 15 post-fix conversations show zero recurrences of a failure that occurred before it, improving when recurrences fell but did not reach zero, and not fixed when the failure recurred at least as often as before. With zero recurrences and fewer than 15 conversations, it is likely fixed.
Converra also marks a result confounded when the prompt was edited outside the fix during measurement. The calculator cannot see your other changes, so it does not show confounded. See production verification for how verdicts are reported.
Frequently asked questions
How do I know if my AI agent fix actually worked?
Count the specific failure the fix targeted in comparable windows before and after the release, with the same detection rule. If the drop is larger than about two standard errors and nothing else changed the agent in that window, the fix is the likely cause.
How many conversations do I need to verify a fix?
It depends on how common the failure is and how big the change is. Detecting a drop from 10% to 5% at 5% significance and 80% power takes about 435 conversations in each window; bigger drops need fewer.
Why does the calculator say my drop is too noisy?
With few conversations or rare failures, chance alone can produce sizable swings. Wait for more conversations, or measure over a longer window, rather than lowering the bar.
Is my data sent anywhere?
No. The calculation runs in your browser and nothing you enter is stored or transmitted.

Written by
Oren CohenFounder of Converra. Previously founded Buildup (acquired by Stanley Black & Decker) and, as VP Product Growth at Totango, owned AI end-to-end from design through production.
Get this verdict automatically on every fix
Converra diagnoses the failure, tests a fix against your current agent, ships what you approve, and measures the targeted failure on live traffic: verified, not fixed, or confounded.