The people who teach AI evals to product teams keep giving the same first instruction, and it has nothing to do with tooling. Read your traces. Hamel Husain and Shreya Shankar call error analysis the most important activity in evals in their AI evals FAQ. They write that they have spent 60 to 80 percent of their development time on error analysis and evaluation.
Any UX researcher who reads their method will feel at home. You collect real sessions, write open notes on each one, group the notes into themes and count them. Researchers call that coding and affinity mapping. Eval practitioners call it error analysis. The two groups have been doing the same work in different rooms.
That split costs teams. AI product managers who treat error analysis as user research get evals that measure what users care about. They also get a research practice that can keep pace with a product whose behaviour changes each time the model does.
The method evals borrowed from qualitative research
The FAQ lays out error analysis in four steps.
- Build a dataset of traces. A trace is the full record of one interaction: the user's message, the context the agent pulled in, each tool call and the final answer.
- Open coding. A domain expert reads traces and writes a free-text note about the first thing that went wrong in each. Husain and Shankar suggest at least 30 traces before any automation gets involved.
- Axial coding. You group the notes into failure categories and count how often each one appears. The FAQ calls this the most important step.
- Keep going until nothing new appears. They put saturation at around 100 diverse traces.
Swap "trace" for "interview transcript" and you have a textbook description of qualitative synthesis. Open and axial coding come from grounded theory, the same tradition that gave UX research its affinity walls. The eval community rediscovered an old discipline and pointed it at machine output.
Traces show intent in the user's own words
Interviews tell you what people remember and are willing to say. Analytics tell you what they clicked. Traces sit between the two: they record what a person typed to get something done, in their own words, followed by everything the system did in response.
Picture a customer-service agent for an online retailer whose customers write in two languages, often in the same message. Read 40 of its traces and a few codes surface fast:
- "User mixed two languages, agent answered a different question"
- "Agent quoted a return window that the policy does not contain"
- "User asked the same thing twice, agent gave the same answer twice"
- "User typed 'ok tq' and left with the order issue unresolved"
None of those codes would show up on a dashboard. The third and fourth are usability findings: the user did not understand the answer, or gave up. A researcher spots those patterns because spotting them is the job. An engineer reading the same traces may file them under "working as intended", since no tool call failed.
Put the PM and the designer in the same trace review
Husain and Shankar recommend one domain expert as a "benevolent dictator" who makes the final call on quality. The FAQ also notes that engineers tend to catch technical failures while PMs catch product failures. The strongest sessions include both, with a designer or researcher in the room to name the comprehension and trust problems the others step past.
A format for a first session:
- Before: the engineer exports 50 recent traces into a spreadsheet, one row per trace, with the full conversation readable in a cell.
- Solo reading, 45 minutes: everyone reads the same 30 traces and writes one short note per trace on the first failure they see. No categories yet.
- Clustering, 30 minutes: read the notes aloud, group them and name each group with a plain sentence a stakeholder would understand.
- Counting, 15 minutes: tally each group. The biggest cluster becomes next sprint's problem.
The session costs an afternoon. It ends weeks of debate about whether the agent is "good enough", because the team now shares a list of the ways it fails and how often.
If your team has never run qualitative synthesis, start with our overview of UX research methods that drive product decisions. The skills transfer one to one.
Turning codes into evals
Each failure category from your review is a candidate eval. The FAQ gives precise advice on how to write them.
Use binary pass or fail. A 1 to 5 scale lets annotators drift toward the middle and turns the difference between a 3 and a 4 into a matter of mood. "Did the agent quote a policy that exists? Yes or no" forces a clear definition.
Let code decide where code can. A check that the quoted return window matches the policy document is a string comparison. Save LLM judges for judgement calls, such as whether an answer addressed the question the user asked.
Validate every LLM judge against human labels. Measure how many real failures the judge catches (true positive rate) and how many good answers it passes (true negative rate). Eugene Yan makes the broader point in An LLM-as-Judge won't save the product: evals are a process of looking at data, and a judge nobody has checked is one more unreviewed opinion.
Your eval suite ends up as a written record of what your users need from the product. That makes it a research artefact, and it should live next to the rest of your research.
Where evals stop and usability testing starts
Evals score outputs. They cannot see what happens in the user's head after the output arrives.
NN/g's research on explainable AI in chat interfaces found that people seldom click citations, so a citation can create confidence the content has not earned. An answer can pass a "cites a source" eval and still mislead the person reading it. You find that by watching people use the product.
METR's 2025 study gives a sharp example of the gap between feeling and fact. In a randomised trial with experienced open-source developers, participants took 19 percent longer with AI tools while believing they had been faster. METR's 2026 update found selection problems in the design and called its new estimates inconclusive, which is its own lesson in reading one study with care. Self-report and measurement disagree, and a product team needs both.
The research design that fits is what NN/g calls explanatory sequential mixed methods: quantitative data first, then qualitative work to explain it. Your eval failure clusters are the quantitative strand. Interviews and usability sessions with the users who hit those failures tell you why, and whether they noticed at all. Our guide to mixed methods UX research in the AI era sets out the full study design. For the interface patterns that follow from those findings, see designing for AI: UX patterns for intelligent interfaces.
Your first month of error analysis
For a PM on a team with an AI feature in beta or production:
- Week 1: get trace access. If your team cannot export full conversations with tool calls, fix that first. Our PM's guide to the AI stack covers the observability to ask for.
- Week 2: run the trace review above. Pick one person as the final judge of quality.
- Week 3: write binary evals for the top two failure categories and run them on every change to prompts, context or model.
- Week 4: recruit five users who hit the biggest failure and interview them. Compare what they say with what the traces show.
Then repeat. Models change underneath you, and so do your users.
Research teams will own the traces
Agent engineering is heading toward systems that improve themselves from production data. LangChain's Sydney Runkle describes a hill-climbing loop that mines production traces to improve the agent's harness, with humans making the judgement calls at each level. Those judgement calls are research decisions: which failures matter to which users, and what "good" means for them.
That puts research operations in a new position. Someone has to decide who may read customer conversations. Someone also has to handle consent and data protection duties under GDPR, PDPA or whichever law applies, for transcripts that contain names, phone numbers and order details. Name that owner before the repository grows, since retrofitting consent onto a year of stored conversations is far harder than setting it up on day one.
If you want help setting up trace reviews and evals for an AI feature, see our product management consulting or get in touch. Teams that want to build the skill in-house can join the AI Product Management cohort, which teaches trace review and judge validation hands-on.