Research methodology

Production-agent reliability methodology

This is a pre-registered protocol for a future production-agent reliability benchmark. It sets the conditions for comparable, consented publication before any benchmark result is considered.

No benchmark results are published here

Converra has not published participant results, pilot metrics, rankings, rates, or customer evidence on this page. The protocol is intentionally public first: a future report can only add evidence after its inclusion, rights, normalization, redaction, and adjudication gates are met.

Converra's role and the publication boundary

Converra sponsors and operates this proposed benchmark. That role is disclosed so readers can evaluate the protocol and any future report with the appropriate context. Converra does not treat private customer material or a product audit as consent to publish benchmark evidence.

Inclusion and scope

A future participant must have written approval for the tested agent, its owner, the task scope, the version under test, and the publication scope before any benchmark material is collected. A participant that cannot satisfy those conditions is excluded rather than approximated.

Publication rights

Future publication requires a page-level grant that covers the named organization, tested scope, findings, excerpts, screenshots, term, withdrawal process, and takedown contact. A benchmark entry is not publishable from a private audit, a token-linked report, or an informal verbal approval.

Normalization

Comparable claims require a documented common task family, dated agent version, probe contract, scoring contract, and collection window. If those conditions differ materially, the work may be described as a separate case but not used for a cross-agent rate or ranking.

Redaction and security review

Future materials are reviewed for personal data, secrets, proprietary content, and coordinated-disclosure risks before publication. Any unresolved item blocks the affected material; redaction is recorded as a limitation rather than presented as complete evidence.

Adjudication

A future report will state its scoring rubric, reviewer roles, disagreement process, and unresolved cases. Reviewers do not adjudicate material they authored, and disputed classifications remain disputed until the published record explains their resolution or exclusion.

Corrections and withdrawals

Any future public report will carry its version, collection window, last-verified date, correction log, and a contact for rights, factual, privacy, or security concerns. A valid rights withdrawal or unresolved security concern removes or noindexes the affected material while the correction is reviewed.

Limitations that remain part of the record

  • This is a Converra-sponsored and Converra-operated methodology, not an independent industry census.
  • A future result would be point-in-time evidence for the disclosed scope and version; it would not establish general agent reliability.
  • Different task mixes, traffic, instrumentation, scoring rules, or collection windows can make two agent results non-comparable.
  • Redactions, exclusions, and rights limits can narrow what is publishable and must remain visible in any future report.

What a future report must show

If this methodology produces a publishable report, that report must identify its tested scope and version, collection window, evidence contract, publication rights, redactions, limitations, and correction status. Until then, this page remains a methodology and no-results record.