Skip to content

UX for AI & Product Design

Mixed Methods UX Research in the AI Era

Mixed methods UX research for AI products: which design to pick, where traces and evals fit, how to use AI in synthesis, and where synthetic users belong.

Published
Read
9 min read

AI products need two kinds of evidence

An AI feature does not behave like a form or a checkout flow. The same prompt can produce two different answers. A user can trust it on Monday and abandon it on Wednesday after one bad reply. Your analytics dashboard shows a drop in usage, and it cannot tell you whether people left because the answers were wrong, slow, or phrased in a way that felt off.

Mixed methods research gives you both halves of the picture. Quantitative data tells you how often something happens and how many people it affects. Qualitative research tells you why it happens and what it means to the person on the other side of the screen. For AI products, you need both, because the failure you can count and the failure you can feel are often two different failures.

The practice is also shifting under researchers' feet. User Interviews' State of User Research 2025 (485 responses, fielded July to August 2025) found that 80% of researchers now use AI in their work, up 24 points in a year. In the same survey, 41% said AI's current impact on research is negative and 32% said positive. Teams are adopting the tools faster than they are agreeing on how to use them. A clear mixed-methods plan is the best guard against that drift.

Three designs, one product decision

Nielsen Norman Group's Rachel Banawa describes three standard designs in Mixed-Methods Research: Combining Qualitative and Quantitative Data (July 2025). Each one fits a different moment in an AI product's life.

Explanatory sequential: count first, then ask

You start with numbers and follow up with conversations to explain them. NN/g's example is an A/B test followed by usability testing to learn why one version won.

For an AI feature, the numbers can come from a new place: production traces and evaluation results. Say your support assistant fails its quality checks far more often on refund questions than on anything else. That count tells you where to look. Interviews and session reviews with the customers who asked those questions tell you whether they lost trust and what they did next.

Exploratory sequential: talk first, then measure

You start with qualitative work to map an unfamiliar problem, then test what you learned at scale. Use this before you build.

An insurer exploring an AI claims assistant might begin with contextual interviews with claims staff and customers. It could then turn the patterns into a survey in each language its customers use to size them. The interviews generate the hypotheses. The survey tells you which ones hold across your customer base.

Convergent parallel: both at once

You run both strands in the same window and compare them at the end. This suits a beta, when time is short and the methods can run without affecting each other. A two-week diary study with twenty pilot users can run alongside a review of every trace those same users produced. At the end, you lay the diary entries next to the logs and look for the places they disagree.

Deciding if a study needs both strands

Mixed methods cost more time and coordination than a single study. NN/g's advice is to pick each method to serve the research goal. In practice, combine qualitative and quantitative work in four situations:

  • The decision is expensive to reverse. Rolling an AI agent out to every customer, or letting it take actions without approval, deserves evidence from more than one angle.
  • The numbers show a pattern without a cause. A drop in completion rate or a cluster of failed evals is a starting point for research, never the conclusion.
  • Self-report and behaviour may diverge. People describe AI tools with confidence that their behaviour does not always support.
  • Stakeholders trust one kind of evidence. A finance director may want numbers. A product lead may want to watch users struggle. Giving each the evidence they trust shortens the debate.

If one method answers the question, run one method. A usability test on five people is enough to find a confusing approval screen.

The METR lesson on self-report

The strongest recent argument for mixed methods came from a study about developers. In a randomised controlled trial published by METR in July 2025, 16 experienced open-source developers took 19% longer to finish tasks when they could use AI tools. They believed the tools had made them about 20% faster.

METR's February 2026 update complicates the picture. The team redesigned the study after finding selection bias: some developers refused to work without AI. The new estimates were inconclusive. That caveat matters, and it reinforces the lesson. Perception and measurement can point in opposite directions, and a study that collects only one of them will miss the gap. If your AI pilot's success case rests on a satisfaction survey, add a behavioural measure before you scale it.

Triangulating contradictory findings

Sooner or later your interviews and your data will disagree. NN/g's guidance on interpreting contradictory research findings and triangulation treats the contradiction as a finding in its own right. Check whether the two studies measured the same population and the same task. Check whether one method was too small or too leading. Often the disagreement points to a segment you had not separated, such as new users who love the assistant and experienced staff who route around it.

Traces and evals are your new quantitative strand

AI products generate a record of every interaction: the user's input, what the system retrieved, what the model produced and what the user did next. Engineering teams read these traces to debug. Researchers should read them too.

Hamel Husain and Shreya Shankar's AI evals FAQ puts error analysis on real traces first, before any automated scoring. They recommend binary pass or fail judgements over 1 to 5 scales, and they recommend that a single domain expert own the quality call. That process looks a lot like qualitative coding. You read a sample and label what went wrong. Then you group the labels into themes.

This gives you a natural explanatory sequential loop:

  1. Sample 50 to 100 traces from the last two weeks and label each one pass or fail, with a short note on the failure.
  2. Group the failures and count them. This is your quantitative strand.
  3. Pick the two largest clusters and recruit users who hit them for interviews or a moderated session.
  4. Bring the quotes and clips back to the product team next to the counts.

A designer who sits in on trace review learns things a usability test cannot show, such as how often users paste in half a document or ask two questions in one message. We cover the interface side of this in Designing for AI: UX patterns for intelligent interfaces.

AI-assisted synthesis, with an evidence trail

AI tools save hours in transcription, translation and first-pass clustering. The risk sits in the step after that: a tidy summary that nobody can trace back to what a participant said.

Set one rule for your team. Every theme in a report must link to the quotes or clips that support it. If an AI tool proposes a theme and nobody can find three participants who said it, the theme does not go in the report. NN/g's guidance on accelerating research with AI supports AI for notes and transcripts and warns against using it to moderate usability tests.

AI-moderated interviews sit in between. In a January 2026 NN/g article on AI-moderated interviews, Maria Rosala recommends them for structured product feedback and multilingual studies. She advises against them for discovery research and topics that need domain expertise. The multilingual case matters for any product with users in several languages. In Malaysia, where we are based, one structured feedback study can reach Bahasa Malaysia, English, Chinese and Tamil speakers without four moderators. Discovery research, where you do not yet know which questions to ask, still needs a person in the room.

One step teams skip: consent. If recordings and transcripts pass through an AI service, say so in the consent form. Store the files in line with the data protection law that covers your participants, such as the GDPR in Europe or the PDPA in Malaysia.

Synthetic users: rehearsal yes, decisions no

Synthetic users are AI-generated personas that answer research questions in place of people. NN/g's 2024 review accepts them for desk research and for drafting interview guides. It rejects them for validation. A 2025 follow-up by Raluca Budiu looked at three studies and found that simulated participants showed less variation than humans and carried demographic bias. Its conclusion: complement, never replace.

Researchers agree. User Interviews' State of Synthetic Users report (150 researchers, May 2026) found 47% skeptical, 17.3% opposed and 3.3% enthusiastic. The top concerns were quality and accuracy (88%) and stakeholders over-trusting AI-generated findings (79%). And 62.7% said their organisation has no guidance on synthetic users at all.

If your team has no policy, write a short one. Allow synthetic users to pressure-test a discussion guide or generate hypotheses before fieldwork. Forbid them as evidence in any decision about what to build, ship or remove. Label any synthetic output in a report so nobody mistakes it for a participant.

A mixed-methods plan for next month

For your next AI feature decision, fill in one page before any fieldwork starts:

  • Decision: the specific call this research informs, with the name of the person who makes it and a date.
  • Quantitative strand: the numbers you will collect (analytics, eval results or a survey) plus the threshold that would change your mind.
  • Qualitative strand: the sessions you will run, plus the people you will recruit for them.
  • Integration point: the meeting where both strands sit on one page, and who resolves disagreements.

That last line is where most mixed-methods studies fail. Two good reports that never meet produce two opinions instead of one decision. The structure in UX research methods that drive product decisions applies here too: start from the decision and work backwards.

Research that keeps pace with the product

AI features change every time the model, the prompt or the retrieval source changes. Research for these products runs as a loop that reads traces every sprint and talks to users every month, with one shared place where the two meet.

If you want help setting that loop up, or a second pair of eyes on a study plan, our UX research consulting covers discovery research and usability testing for AI features, or you can get in touch. For teams who want to build the skills in-house, see our training programmes.