How to Turn Production Agent Traces Into Improvements
To turn production agent traces into improvements, label failures at the step where they happen, group them into counted patterns, pick the costly pattern a change can fix, make one bounded change, test it against the current agent, and measure the same pattern again after release.
Reading traces one at a time finds examples. The improvement comes from turning examples into counts, because a count is what tells you which fix to make first and whether it worked. This guide is part of the AI agent optimization hub.
The short version
A trace is an example; a pattern with a count is a target. Every step below exists to turn examples into a counted target, change it with one bounded fix, and count it again after release.
1. Collect traces that can show a failure
Start with complete records. A useful trace for improvement work includes the full conversation, every model call, every tool call with its arguments and result, and any handoff between agents. The OpenAI Agents SDK, for example, records model generations, tool calls, handoffs, and guardrails as spans in one trace.
Do not sample only errors. Many agent failures throw no exception: the agent answers fluently and wrongly. Include conversations with negative signals, such as a user repeating a request, asking for a human, or abandoning the flow, and a random slice of ordinary traffic so you can see how common a pattern really is.
2. Label the failing step, not the whole conversation
A label like "bad conversation" cannot be fixed. Label the first step that went wrong and what kind of failure it was: the router picked the wrong specialist, retrieval returned the wrong document, a tool was called with the wrong argument, the answer ignored a tool result, or the reply broke a policy in the prompt.
Use a short, fixed taxonomy so labels can be counted across weeks. The taxonomy of agent failure modes is one starting point, and step-level diagnosis shows how Converra records the failing step and failure type for each conversation.
3. Group labels into patterns and count them
Group conversations that share the same failing step and failure type into one pattern, then count it over a fixed window. Platforms can help: LangSmith's Insights feature, for example, clusters traces into usage patterns and failure modes.
Counting is the easiest step to skip and the one that matters most. Google's site reliability engineers describe analyzing postmortems in aggregate so they can target systemic root causes instead of the most recent incident. The same applies to agents: the most recent bad trace is rarely the most frequent failure.
4. Prioritize by frequency, cost, and fixability
Rank patterns by three things: how many conversations they affect, what a failure costs (a wrong refund promise costs more than a clumsy greeting), and whether a change you control can fix them. A pattern caused by a broken upstream API is an engineering ticket, not a prompt fix.
The table below is illustrative, not customer data. It shows how the same week of traces can produce a clear first target.
- Pattern
- Claims a refund after an empty order lookup
- Conversations affected
- 38 of 600
- Cost when it happens
- High: false promise to a customer
- Fixable by
- Prompt instruction
- Decision
- Fix first
- Pattern
- Routes billing questions to the technical specialist
- Conversations affected
- 52 of 600
- Cost when it happens
- Medium: one extra handoff
- Fixable by
- Routing instructions
- Decision
- Fix second
- Pattern
- Order API times out
- Conversations affected
- 21 of 600
- Cost when it happens
- High
- Fixable by
- Engineering, not the agent
- Decision
- File a ticket
- Pattern
- Greets returning users as new
- Conversations affected
- 90 of 600
- Cost when it happens
- Low
- Fixable by
- Prompt instruction
- Decision
- Defer
5. Find the layer that caused the failure
Before writing a fix, decide which layer to change. The same symptom can come from the prompt, retrieval, a tool definition, routing, the model, or the data the agent was given. Changing the prompt to compensate for a broken tool hides the problem until the next model change exposes it again.
The guide to AI agent root cause analysis by layer walks through how to separate these causes from trace evidence.
6. Make one bounded change, with a stated target
Write the change as a hypothesis: which pattern it targets, what it changes, and what result would count as success or failure. One pattern, one change. If you rewrite the system prompt to fix the refund claim, you cannot tell afterward whether the routing pattern moved because of it.
Converra generates prompt or configuration variants aimed at a diagnosed pattern as small incremental edits, and records which pattern each one targets.
7. Test the change against the current agent
Run the candidate and the current agent on the same conversations: cases that reproduce the pattern and regression cases the agent already handles well. A candidate that fixes the refund claim but starts refusing valid refunds is not an improvement.
The guide to building a regression suite from production traces covers turning traces into test cases without treating logs as ground truth, and simulation testing covers multi-turn conversations with simulated users.
8. Ship with approval, then count the same pattern again
Have the person accountable for the agent approve the change, and record when it shipped. Then count the targeted pattern in new production conversations, with the same labels and window length you used for the baseline.
Three outcomes are possible: the pattern stopped (verified), it recurred at least as often (not fixed), or another change to the same agent shipped in the window and the result cannot be attributed (confounded). The guide to verifying an AI agent fix after deployment explains how Converra decides each verdict.
Worked example: from routing traces to a measured fix
In the Salespeak case study, an orchestrator agent decided which specialist agent should handle each user. Converra analyzed production conversations and labeled two recurring patterns at the step where they occurred: users routed to the wrong specialist, and the orchestrator stating pricing, VAT rules, and infrastructure details that were not true.
Each pattern got its own fix, aimed at routing behavior and at unsupported claims rather than a general rewrite. Variants were tested in multi-turn simulation before any change reached production, and Salespeak's CTO reviewed and applied the winning changes. On live traffic, routing failures fell 74% after the April 25 deploy, and the hallucinated claims did not recur after the April 23 deploy.
Limits: one customer and one agent, so this is not a rate to expect elsewhere. The published case study reports percentages rather than conversation counts. Because the two changes shipped two days apart, each was measured against the pattern it targeted rather than as one combined result.
Frequently asked questions
How many traces do I need before I can prioritize?
Enough to count the patterns you care about, not a fixed number. A pattern that appears in a handful of conversations a week can still be worth fixing if each failure is costly, but you will need a longer window to see whether a fix changed it.
Should I use an LLM to label traces?
It can speed up labeling, but check its labels against human review on a sample before you trust the counts. A labeler that is wrong in one direction will make one pattern look larger than it is.
What if a failure has no error in the trace?
Most agent failures do not. Look for user signals such as repeated requests, requests for a human, or abandonment, and compare the agent's answer with its tool results and the rules in its prompt.
Can I fix several patterns at once?
You can ship several changes, but give each one its own target and measure each pattern separately. If two changes touch the same behavior in the same window, the result for either cannot be attributed.
How is this different from building a regression suite?
A regression suite is one output of this workflow: traces become test cases that protect behavior. This workflow goes further, from choosing which pattern to fix to measuring whether the fix worked in production.

Written by
Oren CohenFounder of Converra. Previously founded Buildup (acquired by Stanley Black & Decker) and, as VP Product Growth at Totango, owned AI end-to-end from design through production.
Let the loop do the counting
Converra reads your production traces, diagnoses failures at the step, generates and simulates a bounded fix, ships what you approve, and returns a production verdict: verified, not fixed, or confounded.