The next AI products will act for people
For most people, AI still means a chat box. Gartner expects that to shift soon: in 2025 it predicted that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from less than 5% in 2025. Treat the figure as a forecast. The direction matches what product teams are shipping.
The next generation of AI features will draft the invoice, reply to the customer, book the site visit and reconcile the accounts. Once software starts acting on someone's behalf, the model stops being the hard part. The hard part is the moment a person decides whether to let it act, and what they do when it gets something wrong.
That moment is the territory of UX for AI.
UX for AI, defined
UX for AI is the design of every surface where a person meets an AI system. It covers what the system says it can do and how sure it is. It also covers the moment things go wrong, when a person needs to correct the system or take over.
Engineers now talk about the "harness" around a model: the tools, rules and checks that keep an agent on track. Birgitta Böckeler's essay on harness engineering makes a point designers should borrow. A good harness does not aim to remove people from the loop. It directs human input to "where our input is most important."
UX for AI is that idea seen from the person's side. Agent experience asks what context and tools the model needs. User experience asks what the customer service lead, the finance clerk or the patient needs in order to trust the result the right amount. Both matter. This article is about the second.
NN/g's State of UX 2026 names trust as a major design problem for AI this year, and lists transparency, control, consistency and support when the system fails as the ingredients. The five jobs below turn those ingredients into work a product team can plan.
The five jobs of UX for AI
1. Calibrate trust, don't maximise it
Google's People + AI Guidebook frames the goal as calibrated trust. People should rely on the system when it is reliable and use their own judgment when it is not. Too much trust leads to people approving wrong answers. Too little leads to what the guidebook calls algorithm aversion, where people ignore a system that would have helped them.
Calibration is a design output. You shape it with the words the product uses, the confidence it shows and the limits it admits up front. Two practical rules from the guidebook:
- Show confidence where it changes a decision, in a form people can read. Categories such as high, medium and low with a clear next step beat a percentage most users cannot interpret.
- Say what data the system uses and where that data is thin. A property valuation tool trained on big-city transactions should say so before someone in a rural town relies on it.
Watch your tone as well. NN/g's research on explainable AI in chat interfaces found that human-like language inflates trust. "I've checked this for you" promises more than "Checked against your March statement."
2. Make transparency faithful
Transparency features can mislead. The same NN/g study found that step-by-step "reasoning" displays are often written after the fact and do not reflect how the model reached its answer. Citations can be hallucinated, and people seldom click them, so a row of source links can create confidence the product has not earned.
Faithful transparency shows what the system did. Emily Campbell's pattern library The Shape of AI groups these features as Governors: action plans, citations, draft mode and verification steps that keep a person in charge. The test for each one is simple to state. Does it help someone catch a mistake before it costs them? If it only decorates the answer, cut it.
For citations, NN/g recommends placing each source next to the claim it supports and linking to the exact section.
3. Match control to risk
An AI that suggests a subject line and an AI that transfers money need different controls. Researchers Kevin Feng, David McDonald and Amy Zhang propose five levels of autonomy defined by the user's role: operator, collaborator, consultant, approver and observer. Their key claim is that autonomy is a design decision, separate from what the model is capable of.
This gives teams a shared question for every feature: which role should the person play here? We cover the five roles and how to design the controls for each in Levels of Autonomy for AI Agents: Set the Dial.
4. Design the wrong-answer path first
Every AI feature will be wrong in front of a user. Microsoft's Guidelines for Human-AI Interaction, based on research by Saleema Amershi and colleagues, give a whole group of their 18 guidelines to this situation. Make the system easy to invoke and easy to dismiss. Make correction efficient. Narrow the scope when the system is unsure. Let people see why it did what it did.
Teams tend to design the happy path in Figma and leave errors to a generic toast. For AI, the recovery flow deserves the same care as the main flow, because it is where trust is won or lost. Our guide to AI error handling UX walks through the patterns.
5. Evaluate with real users and with evals
AI products need two kinds of evidence. Evals measure output quality at scale. User research measures whether people understand, trust and recover. Neither replaces the other.
On the eval side, Hamel Husain and Shreya Shankar's evals FAQ recommends starting with error analysis. Read real traces and label what went wrong before you write automated checks, and make those checks pass or fail. That practice looks a lot like qualitative research, and designers should be in the room for it.
On the research side, you can test before the model exists. In a Wizard of Oz study, a person behind the scenes plays the AI while a participant uses the interface. You learn how people phrase requests, what they expect back and where they stop trusting the output, all before an engineer writes a prompt.
Do not rely on what users say about the tool. METR's 2025 study of experienced open-source developers found that they took 19% longer with AI tools while believing they were about 20% faster. METR's 2026 update says newer estimates are inconclusive, and the gap between perception and measurement is the lesson that holds. Watch behaviour and measure outcomes.
Constraints that change the design
The frameworks cover the principles. Real products add constraints the frameworks leave out.
Language. Many users write in more than one language. In Malaysia, where Produlogi is based, a single message often mixes Bahasa Malaysia and English. Your disclaimers, error messages and clarifying questions should answer in the language the person used. An English-only fallback on a Malay query tells the user the system did not understand them, even when it did.
Governance. If your organisation has no written AI policy, the product interface becomes the policy. If the screen lets a junior staff member approve an AI-drafted refund with one tap, that is the rule, whatever the handbook says.
Regulated sectors. Finance, insurance and healthcare products need audit trails and human review built into the flow. We look at this in more depth in AI in financial services: designing compliant, trustworthy products.
A starting checklist for your next AI feature
Use this before a feature reaches design review:
- Write one sentence on what the AI can do and one on what it cannot. Put both in the interface.
- Decide the user's role for each action: operator, collaborator, consultant, approver or observer.
- Mark every irreversible action and add an approval step with a clear preview of the consequence.
- Design three failure states: wrong answer, no answer and partial answer. Give each a next step.
- Replace human-like phrasing with plain, factual copy.
- Place citations beside the claims they support, or remove them.
- Run a Wizard of Oz session with five target users before building.
- Book a weekly trace review where the designer and PM read 30 real conversations together.
- Write error and disclaimer copy in every language your users type in.
Our earlier piece on UX patterns for intelligent interfaces covers streaming, confidence indicators and feedback controls in more detail.
Where UX for AI is heading
The chat box is a transitional interface. As agents take on longer tasks, people will spend less time reading answers and more time reviewing plans, approving actions and checking logs. The core design skill shifts from arranging outputs to designing supervision: showing a person enough to make a good call in the few seconds they will give it.
Addy Osmani's essay on loop engineering warns that "a loop running unattended is also a loop making mistakes unattended." The teams that design good supervision will be the ones allowed to hand their agents more work.
If you are planning an AI feature and want a second pair of eyes on the experience, our consultation services for UX and product design start with these questions. You can also book a call to talk through yours. To build the skills inside your team, see the training programmes Marcus teaches.