Most AI benchmarks end before the hard part begins
How we measure production AI agent reliability
A static score can tell you an agent passed a test. It cannot tell you whether the agent keeps doing the right thing after a prompt change, a model update, or a week of real users. We follow the entire chain from production failure to deployed fix, and we do not compress that chain into one flattering number.
The operating question
Did the agent get better where people actually use it?
Production reliability is not uptime, one model score, or a pile of passing evals. It is the ability to find a recurring behavior failure, reproduce it, improve it without breaking something else, and verify the result after deployment.
If any link in that chain is missing, the result says so.
One finding, five linked receipts
Every claim must survive the full evidence chain
Observe
Did a recurring failure happen in production?
Start with an eligible set of real interactions. Publish the failure rule, numerator, denominator, exclusions, and time window together.
Detect
Can the failure be identified consistently?
Apply a declared rubric, preserve the evidence behind each finding, and send ambiguous cases to review instead of forcing a label.
Reproduce
Can the behavior be recreated on demand?
Freeze the current agent version and run an applicable probe family that varies wording and context without changing the target behavior.
Improve
Does a targeted change beat the current agent?
Compare baseline and candidate on the same eligible probe pairs. Unpaired runs do not select a winner, and regressions can block a lift.
Verify
Did the improvement survive production?
After an authorized deployment, measure the same failure pattern again and issue a verdict only when the production evidence supports one.
Production outcome
Three verdicts. No victory lap for missing data.
The post-deployment measure gets one of three verdicts. If the required measure is absent, the result remains unobserved and no verdict is issued.
Verified
The targeted failure decreased in a comparable production measure.
Not fixed
The problem persisted or worsened after the deployed change.
Confounded
Other changes make the observed movement impossible to attribute.
Rules before results
The methodology is designed to make inconvenient facts hard to hide.
Consent is specific
Permission to analyze an agent is not permission to publish its name, users, prompts, screenshots, quotes, or results. Each right is recorded separately.
Comparisons stay paired
A baseline and candidate must complete the same applicable probe. Missing or ineligible runs never become convenient zeroes.
Review can disagree
Automated scoring may surface a candidate finding. Reviewers can uphold it, reject it, or leave it unresolved. Unresolved cases do not enter a claimed rate.
Corrections stay visible
A changed denominator, rubric, inclusion rule, or finding creates a new version and a readable correction record instead of rewriting history.
The standard for a fix
Only ship fixes that beat what's running
We test the proposed fix and the current agent against the same failure cases. The fix advances only when it improves the target behavior without introducing a regression. Agents doing different jobs are never compared on one shared ranking.
- No universal reliability score that hides different tasks and rubrics.
- No ranking of agents that do different jobs.
- No industry failure rate built from a small, self-selected pilot.
- No production claim based only on a simulation result.
- No filled-in estimate when the required measure is missing.
Finding record
No findings published
The empty state is intentional. Rules are public before results so the cohort, denominator, scoring, and exclusions cannot be quietly optimized around a preferred story.
- Eligible production sample and failure definition
- Frozen baseline and candidate versions
- Paired probe outcomes and regression results
- Reviewer agreement and unresolved cases
- Deployment reference, measurement windows, and confounders
- A plain statement of what the evidence does not establish
Who is behind this work
Converra has a stake in the outcome.
Converra authors and sponsors this methodology, would choose the pilot cohort, and would operate the analysis. We also build software for diagnosing, testing, and verifying changes to production AI agents. Favorable findings could benefit us commercially.
This is not a third-party-independent benchmark. Before any pilot result is called externally reviewed, an outside method reviewer must be named with their scope, conflicts, review date, and disposition.
Questions worth asking before the first result
Is this an independent benchmark?
No. Converra authors, sponsors, and would operate it. The narrower promise is an inspectable protocol, frozen rules, disclosed conflicts, and outside method review before any result is called externally reviewed.
Are failure rates comparable across agents?
Only when the task boundary, eligible opportunity, probe family, rubric, missing-run rule, and analysis conditions are materially equivalent. Otherwise results stay within each agent family.
Does a better simulation score prove the fix worked?
No. Simulation supports a deployment decision. A production claim requires a separate post-deployment measure of the same target behavior.
Where are the results?
There are none yet. This page publishes the rules first so future inclusion, scoring, and analysis choices cannot be quietly changed after seeing outcomes.