Skip to content

Lab Notes

Lab Notes: Design the States the AI Demo Skips

An AI demo shows one good answer. Real users meet the empty, slow, unsure, wrong and corrected states. Put them in the spec before anyone builds.

Published
Read
4 min read

One answer, perfect every time

The demo of an AI feature has one path. Someone types a clear request and the model returns a good answer. The room nods. Generators build to that path by default because the prompt described it and nothing else.

A real user arrives with a blank account, a vague question and a slow connection. The model is sometimes unsure and sometimes wrong. The user wants to fix the answer, undo the action or reach a person. None of that shows up in the demo, and all of it decides whether the feature gets used a second time.

Google's People + AI Guidebook gives errors and graceful failure a full chapter. Microsoft's Guidelines for Human-AI Interaction group 18 guidelines by phase, and one phase is labelled "when wrong". Both treat failure as a normal part of the product.

The state list

We write this list into the spec before anyone prompts a prototype. Each line needs a designed screen or a written decision to skip it.

Empty

This is the screen before the user has asked anything or added any data. A good empty state shows one example of a request that works and one way to start. A blank text box with a blinking cursor asks the user to guess what the feature can do.

Working

This covers the wait while the model thinks. For anything longer than a second or two, show what the system is doing, in the user's terms: "Reading your three invoices" tells them more than a spinner. For long agent runs, show the plan and the step it has reached.

Unsure

This covers low model confidence and requests outside the feature's scope. Scoping the answer, or saying it cannot help with this one, costs less trust than a confident wrong answer. The HAX guidelines call this scoping services when in doubt.

Wrong

This covers how the user finds out the answer is wrong, and what the product does next. On screen, a wrong answer looks the same as a right one. This state depends on what you show alongside it: the sources and inputs the model used, plus a visible way to flag a problem.

NN/g's research on explainable AI in chat interfaces adds a warning here. People rarely click citations, so a row of sources can create confidence without checking anything. Put the evidence where the decision happens.

Corrected

This is the fix. Edit the output in place, regenerate with a hint, or undo an action the agent took. Correction should cost less effort than starting over, or users will start over, then stop.

Handed off

This is the point where a person takes over. For customer-facing features, that means a route to a human with the conversation attached. For agents, it means an approval gate before anything irreversible: sending, paying, deleting. Levels of autonomy helps you decide which steps need the gate.

How to use the list

Put the states in the PRD as a section of their own, one line each: what triggers the state, what the user sees, what they can do next. Our free AI Feature PRD Template has the section ready, with a prompt that drafts a first pass for you to edit.

Then prototype the unhappy states first. A generator will build the happy path without being asked. The wrong and corrected states need a deliberate prompt, and they are the screens a user test will lean on.

You can test most of them before a model exists. NN/g's Wizard of Oz method has a person play the AI behind the interface, which lets you watch a user hit a wrong answer and try to recover. That session tells you more about the product than another round of prompt tuning.

What changes in review

Once the list exists, design review gets a concrete question for every AI screen: which state is this, and where are the other five? A screen set that only shows the success path goes back for the rest.

The same list feeds the evals. Each state that depends on model behaviour, unsure and wrong in particular, becomes a case your team checks in real traces. LLM error analysis is user research explains how designers and PMs read those traces together. AI error handling UX goes deeper on the wrong and corrected states.

This list forms the AI UX row of our AI Slop Test. If your team is shipping an AI feature and the spec stops at the happy path, UX for AI consulting starts with this list.