The pilot worked. Then nothing happened.
The demo went well. The chatbot answered product questions and the management team nodded. Someone said "let's roll it out." Six months later, the bot still sits on the website and a handful of customers use it. Nobody can say whether it saves time or money.
Research from several markets suggests this pattern is common, and it shows up close to home for us too. In Malaysia, where Produlogi is based, Amazon Web Services' Unlocking Malaysia's AI Potential 2026 study found that 67% of AI adopters focus on basic applications such as public chatbots, and only 19% have a formal strategy for scaling AI across functions. (We break down that data in AI adoption in Malaysia 2026.)
Moving an AI pilot to production depends on the work around the model. You choose the right job and design for the people who use it, then keep measuring after launch. This guide covers that work.
Read the "95% of AI pilots fail" figure with care
A single statistic dominates conversations about AI pilots. In mid-2025, MIT's NANDA initiative published The GenAI Divide, reported in the press as finding that 95% of enterprise AI pilots deliver no measurable return.
Before you quote it in a board paper, look at how it was built. The report drew on 52 interviews, 153 survey responses from senior leaders and a review of over 300 public AI initiatives. It used a narrow definition of success, measurable profit-and-loss impact, and it was not peer reviewed. Analysts such as Futuriom have challenged its methods.
The headline number is weak. The reasons the report gives for failure are more useful. It found that most generative AI systems do not retain feedback or improve over time. It also found that brittle workflows and poor fit with day-to-day operations stop pilots from scaling. Those are design and operations problems, and you can fix them.
Gartner adds a forward-looking warning. In June 2025 it predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, citing rising costs, unclear business value and weak risk controls. It also warned about "agent washing", where vendors rebrand chatbots and scripts as agents. Treat it as a prediction, and as a checklist of risks to design against.
Four reasons pilots stall
From the research above and the patterns practitioners describe, most stalled pilots share one or more of these causes:
- The pilot tested the demo. It proved the model could answer questions. It never tested whether a specific person's job got faster, cheaper or better.
- Nobody measured the starting point. Without a baseline for handling time or error rate, the team cannot prove value, so the budget conversation stalls.
- The design ignored trust. Nielsen Norman Group's State of UX 2026 names trust as the central design challenge for AI, built on transparency, control, consistency and support when the AI fails. A pilot that hides its sources and offers no way to correct it loses users after the first wrong answer.
- Nobody owns it after launch. AI output drifts as models, prompts and data change. A system without an owner and a review routine degrades in silence.
Five moves from pilot to production
1. Anchor on one job and one baseline
Pick a single workflow with a person attached to it: the claims officer processing a form, the sales coordinator preparing quotations, the support agent answering refund questions. Watch them do the job. Record how long it takes and how often it goes wrong today.
That baseline becomes your success measure. If you cannot name the person and the number, you are not ready to scale.
2. Choose the simplest shape that works
Anthropic's often-cited guide Building effective agents advises starting with the simplest solution. Many use cases need a fixed workflow with one or two model calls, and an autonomous agent adds cost, latency and new ways to fail.
For most small and mid-size businesses, a well-designed workflow beats an agent: classify the incoming email, draft a reply from approved templates, route anything unusual to a person. Save agents for tasks where the steps vary from case to case. Our practical guide to building AI agents covers the build decisions once you get there.
3. Design for trust and for the wrong answer
Google's People + AI Guidebook frames the goal as calibrated trust: people should rely on the system when it is reliable and use their own judgement when it is not. Microsoft's Guidelines for Human-AI Interaction include a full set for the moment the AI is wrong: make it easy to dismiss, correct and recover.
Decide how much autonomy each action gets. Researchers Feng, McDonald and Zhang describe five levels of autonomy defined by the user's role, from operator to observer. A draft reply to a customer can go out after one click of approval. A refund or a change to a customer record may need a named approver. Put the approval step where a mistake would cost money or trust, and nowhere else, so staff do not learn to click through it.
Test in the languages your customers use, including the way they mix them. In Southeast Asia, a single support message can switch between Bahasa Malaysia and English mid-sentence. An assistant that handles clean English and stumbles on that message will lose the customers who write that way. For interface patterns, see Designing for AI: UX patterns for intelligent interfaces.
4. Measure with evals and with users
Two kinds of testing catch two kinds of failure. Evaluations check output quality at scale. Hamel Husain and Shreya Shankar's AI evals FAQ recommends starting with error analysis: label real conversations pass or fail, then automate the checks. Usability testing with real users catches the failures evals miss: people who misread a confident answer, ignore a citation or cannot find the undo.
Run both before you scale. Five usability sessions with the people who will use the system, plus a labelled sample of one hundred real interactions, will tell you more than a month of dashboards. Our piece on mixed methods UX research in the AI era shows how to combine the two.
5. Give it an owner and a review rhythm
Name one person accountable for the system's quality after launch. Give them access to traces, the logs of what the system received and produced. Book a short review of failures every week. Ask for traces in the OpenTelemetry format, an open standard that observability tools such as Langfuse accept, so you are not locked into one vendor's dashboard.
This is also where the MIT report's "learning gap" gets closed. A system improves when someone reads its mistakes and changes the prompts, data or workflow in response.
Build, buy or partner
Many companies buy their AI capability from outside. In the AWS Malaysia study, 57% of businesses source AI through external providers first. For any company in that position, vendor choice becomes one of the biggest decisions in the move from pilot to production. Ask any provider, including us, these questions before you sign:
- Can you show me traces of what the system did for a sample of real requests?
- How do you evaluate quality, and can we see the eval set and results?
- What happens when the system is wrong, and how does a user correct it?
- Who owns the prompts, workflows and data when the engagement ends?
A provider who answers with a demo instead of evidence is selling a pilot.
A 90-day plan
Here is a realistic path from stalled pilot to a production workflow:
- Weeks 1 to 2: discovery. Observe the workflow, interview the people who do it, record the baseline and write down the success threshold.
- Weeks 3 to 6: prototype and test. Build the simplest version. Test it with five to eight users, using a person behind the scenes where the model is not ready yet, which NN/g calls the Wizard of Oz method.
- Weeks 7 to 10: limited rollout. Release to one team or one customer segment. Label a sample of real interactions every week and fix the largest failure cluster first.
- Weeks 11 to 13: decide. Compare results with the baseline. Decide to scale it or stop it. Write down why.
Stopping a pilot at week 13 with evidence is a better result than running it for a year without any.
Design is the next advantage
Over the next two years, access to models will stop being a differentiator for most businesses. Every competitor will have the same tools. The difference will come from teams that choose the right jobs and design AI their staff and customers trust. Those teams keep improving it after launch.
If you have a pilot that stalled, or a use case you want to get right from the start, see how we work with product teams on choosing, designing and measuring AI that reaches production. Or tell us about your pilot.