Skip to content

UX for AI & Product Design

AI Error Handling UX: When the AI Gets It Wrong

Your AI feature will be wrong in front of users. Design recovery first: cheap correction, scoped answers, honest uncertainty and a clean human handoff.

Published
Read
9 min read

A wrong answer looks like a right one

Traditional software fails in the open. A form rejects an invalid ID number, a payment page shows a red banner, a server returns an error page. The user knows something broke.

AI fails in silence. A wrong answer arrives in the same confident font as a right one. The chatbot gives last year's public holiday dates. The assistant summarises a contract and drops the termination clause. Nothing on screen tells the user to look twice.

That is why error handling for AI is a design problem before it is an engineering one. You cannot remove every mistake from a probabilistic system. You can decide what happens next. Does the user notice, and how much does the fix cost them? Their answer decides whether they trust the product tomorrow.

Google's People + AI Guidebook has a whole chapter on errors and graceful failure, and Microsoft's Guidelines for Human-AI Interaction group five of their 18 guidelines under "when wrong." This article turns both into patterns you can put in a design file this week.

Name the kind of wrong you are designing for

"The AI made a mistake" covers several different failures, and each needs a different response. The PAIR chapter separates them. Four matter most for product teams:

  • Context errors. The system works as built, but the user sees a failure because it made a poor assumption or did not explain itself. A food delivery assistant that suggests pork dishes to a user whose order history is all halal has a context problem. The fix is usually better defaults and a visible explanation.
  • Failstates. The system cannot answer because of a hard limit, such as missing data or a request outside its scope. The fix is to say so and offer a path forward.
  • Relevance errors. The output is low confidence or misses what the user needed. The fix is to narrow the answer, ask a clarifying question, or show alternatives.
  • Background errors. Neither the user nor the system notices the mistake. These are the most dangerous, and no interface pattern catches them alone. You find them through evaluation, covered below.

Map your feature's expected failures to these types before you design the screens. The exercise tends to show that the generic "Something went wrong, please try again" state covers almost none of them.

The recovery toolkit

Make correction cheaper than starting over

The HAX guideline to support efficient correction asks products to make it easy to edit, refine or recover when the AI is wrong. In practice:

  • Let people edit the output in place instead of rewriting their prompt.
  • Let them regenerate one part, such as a paragraph, a line item or a field, without losing the rest.
  • Keep their edits when the AI runs again. An agent that overwrites a correction teaches the user to stop correcting.

A user who fixes one wrong figure in a generated quotation in five seconds will keep using the tool. A user who has to start the quotation again will go back to the spreadsheet.

Make dismissal free

Some AI suggestions are wrong in a harmless way: an irrelevant autocomplete, an unwanted summary. The HAX guidelines ask for efficient dismissal. One tap should do it, with no confirmation dialog and no guilt-trip copy asking the user to explain why.

Narrow the scope when unsure

The HAX guideline to scope services when in doubt describes two moves: ask a clarifying question, or fall back to a smaller answer. Both beat a confident guess.

If a user asks an HR assistant "How many days of leave do I have?" and the system cannot tell which leave type they mean, one short question ("Annual leave or medical leave?") costs less than a wrong number the user acts on.

If the system has no good answer, say so. PAIR recommends making failure safe and boring, with a clear route back to manual control. An honest "No answer found in your policy documents. Here is the HR contact." is a working feature.

Explain the specific failure with care

The HAX guidelines ask products to make clear why the system did what it did. Be careful with how. NN/g's research on explainable AI in chat interfaces found that step-by-step reasoning displays are often written after the fact and do not reflect how the model produced its answer. A plausible but invented explanation for an error makes things worse.

Explain what you can verify: the documents the answer came from, the date range the data covers, the input the system could not read. "This answer uses the 2025 policy. The 2026 policy was not found." is honest and actionable.

Put a checkpoint before actions

Agents that act raise the cost of error. A wrong summary wastes a minute. A wrong email to a customer, or a wrong journal entry, takes an afternoon to repair.

Emily Campbell's The Shape of AI collects patterns for this under Governors, including action plans, draft mode and verification. Show the plan before the agent executes. Stage outputs as drafts until a person releases them. Offer undo wherever the underlying system allows it, and say so in plain words when it does not.

The HAX guidelines also ask products to convey the consequences of user actions. On an approval screen that means a preview of what will happen, with the actual email, entry or change on screen. A summary of the agent's intent is too thin to judge. How much checkpointing you need depends on the task's autonomy level, which we cover in Levels of Autonomy for AI Agents: Set the Dial.

Hand off to a person with the context attached

Some errors need a human. The handoff is part of the AI experience, and it often decides how the customer remembers the whole interaction.

Pass along the conversation history and a note on why the AI stopped. A customer who explained their problem to a bot on WhatsApp should never have to type it again for the agent who picks it up. Tell the customer what happens next and when.

Write error copy people read

NN/g's research found that disclaimers work when they are plain, specific and placed near the input. Generic warnings in the footer go unread. The same applies to error messages.

Two rules for AI error copy:

  1. Name the limit and the next step. Say what the system could not do, then give the user an action they can take.
  2. Drop the persona. NN/g also found that human-like language inflates trust. "Sorry, I got confused!" invites the user to treat the system as a person. Neutral, factual copy sets the right expectation.
  • Weak: "Oops! Something went wrong." Stronger: "This invoice could not be read. Upload a clearer image or enter the total by hand."
  • Weak: "I might make mistakes." Stronger: "Dates and amounts may be wrong. Check them against your bank statement."
  • Weak: "Maaf, saya keliru." Stronger: "Jawapan ini mungkin tidak tepat. Sila semak tarikh pada penyata bank anda."

Write error and uncertainty copy in every language your users type in. The third example above is Bahasa Malaysia, a language Malaysian speakers often mix with English. If someone asks in Malay and the error comes back in English, the product has failed twice.

Find your failures before users do

The best error handling starts upstream. Hamel Husain and Shreya Shankar's evals FAQ puts error analysis first. Read real traces and label each failure. Group the failures into types before you write automated checks. They recommend pass or fail judgments over 1 to 5 scores.

That is qualitative research on machine output, and designers belong in it. A weekly session where the designer and PM read 30 real conversations together surfaces the failure types that deserve a designed state. It also catches background errors that no user will report.

You can test error states before the AI exists. In a Wizard of Oz study, a researcher plays the AI and gives a planted wrong answer at a set point. Watch whether the participant notices and how they try to fix it. Then see whether they trust the next answer. Run a follow-up session a week later to see how that trust holds up. Our UX research consultation runs studies like this, and our post on UX research methods that drive product decisions covers how to plan them.

Measure recovery

Thumbs up and thumbs down tell you little about errors. Track signals tied to recovery:

  • Correction rate. How often users edit AI output, and which fields they edit most.
  • Abandonment after error. Whether users leave the feature after a failure state.
  • Escalation rate and resolution time. How often the AI hands off, and how long the human takes to close the case.
  • Return after failure. Whether a user who hit an error comes back to the feature within a week.

These numbers show whether your recovery design works. Pair them with eval pass rates and you can see both how often the AI is wrong and how much each mistake costs your users.

Start with the failure flow

On your next AI feature, design the error states before the happy path. Sketch a response for each of the four error types, then test at least one with users before launch. For the full set of UX-for-AI principles, read UX for AI: Designing AI Products People Trust, and for earlier patterns see designing for AI: UX patterns for intelligent interfaces.

If you want help designing recovery flows for an AI product, our UX and product design consultation can review your current states and prototype better ones, or you can book a call to start with one feature. To build the skill in-house, see the training programmes Marcus teaches.