Skip to content

Consulting · Product management and AI PM

Product management consulting for the AI era

Product management for the AI era: discovery, AI PRDs, evals and roadmaps, as a sprint or a fractional AI PM embedded in your team.

The job of a product manager is moving outward from the model. A year ago the conversation was prompts. Now practitioners talk about context, the harness around the model and the loops that run agents without a person watching each step. Each of those layers holds product decisions: what good looks like, when a person approves, what the team gets to see when something goes wrong.

We help product teams make those decisions with evidence. That covers classic product work from discovery to roadmap. It also covers the newer AI PM work of evals, traces and agent operations.

Built for

  • A head of product with AI features in production and no way to measure them. Users complain and the team tweaks prompts. Nobody can say whether this week's version beats last week's.
  • A startup adding its first agent. You need someone to write the PRD and choose between a workflow and an agent, with evals in place before launch.
  • A PM team moving into AI product management. You want a working method your PMs can adopt, with real artefacts from your own product.

Deliverables

  • AI PRD with an eval plan. The feature, its users, the definition of good output and the checks that prove it.
  • Error analysis report. We read real traces with your PM and designer, cluster the failures and rank them by user impact.
  • Eval suite, first version. Binary pass/fail checks tied to the failure clusters, with any LLM judges calibrated against human labels before you trust them.
  • Observability requirements. The traces, spans and fields your team or vendor must capture so you can see what the agent did, in line with the OpenTelemetry conventions for generative AI.
  • Agent runtime decision memo. Criteria and a recommendation for where your agents run, covering session state, tracing, cost and data residency.
  • AI roadmap. Discovery findings turned into a sequenced roadmap, with the eval gate each item must pass before it ships.

The AI PM stack, in product terms

Evals and error analysis

Good eval practice starts with reading traces, before any tooling. We run error analysis as a research session: PM and designer read real conversations and agent runs, name the failures in plain language and decide which ones matter to users. Those names become the first evals. LLM error analysis is user research walks through the method.

Harness and loop engineering

The harness is everything around the model: instructions, tools, tests, guardrails and the checks that let an agent correct itself. Loop engineering designs the cycles that run agents on a schedule or on events. Both decide where a human reviews the work. We map your harness and loops with your engineers and mark the points where a person must stay in charge. Harness engineering: a PM's guide to the stack covers the vocabulary.

Agent runtime and observability

An agent runtime is the managed place where long, stateful, tool-using sessions run. The product questions come before the vendor comparison. Can you see what the agent did, step by step? Can you judge it with evals? Can you stop it? We write those requirements first and pick the runtime against them.

Discovery and roadmaps

The classic PM work stays: interviews, opportunity sizing, prioritisation and a roadmap your leadership can fund. We add one rule for AI items. Each item names the eval it must pass and the user signal that proves it worked.

Working together

  1. Audit. A clarity session or kickoff to review your product, your current AI features and any traces you already collect.
  2. Read the traces. Error analysis sessions with your team.
  3. Define good. The PRD, the eval plan and the observability requirements.
  4. Build the gate. The first eval suite, wired into your release process with your engineers.
  5. Hand over or embed. Your team runs it, or Marcus stays on as a fractional AI PM.

See the engagement formats and pricing. We work remote and async first from Kuala Lumpur (UTC+8).

The craft difference

  • Evals and user research share one loop. Eval failure clusters tell us where to interview. Interviews tell us which failures users care about. We run both as one loop.
  • A human stays where the risk is. We design approval points around irreversible actions and leave routine steps to the agent. Levels of autonomy for AI agents explains the dial.
  • Claims with sources. Roadmap cases cite data from your product. Where the data does not exist yet, the roadmap says so.

Proof

Marcus teaches the orchestration day of Mastering Claude and Multi-Agent Orchestration as AiTraining2U faculty, covering Claude Code, plugins, skills and MCP. He is also the named instructor for AI Vibe Coding, where every build starts from an AI-assisted PRD. More on the about page.

TBC: first consented case study for this track.

If the eval failures point to the interface, the UX, UI and product design track takes over. PMs who want to build these skills themselves can apply for coaching on the PM path or join the AI Product Management cohort.

Questions people ask

We have no AI features yet. Is this track too early for us?

No. Discovery and roadmapping for a first AI feature is core to this track. The AI PM work on evals and observability starts once you have something running, even a prototype.

Do you write the evals, or teach our team to?

Both, in that order. We run the first rounds of error analysis with your PM and designer, write the first pass/fail checks with you, and leave your team able to extend the suite.

Can you work as our fractional AI PM?

Yes. In the fractional format Marcus joins your team on a monthly retainer and owns agreed outcomes, such as the eval suite or the AI roadmap. He takes part in your planning and reviews. TBC: minimum term.

Which agent runtimes and observability tools do you recommend?

It depends on your stack, data residency needs and budget. We write the decision criteria with you and compare two or three options against them. The one fixed requirement: you must be able to see a full trace of what the agent did.

Do you work with teams outside Malaysia?

Yes. We are based in Kuala Lumpur and work remote with product teams worldwide, async first, with a weekly live call where our time zones overlap.

How do you handle time zones?

Updates are written or recorded so you can read them in your morning. We place the live call in the overlap window, or fix one slot at the edge of both days for teams in the Americas.

Next step

Tell us what you're building

Based in Kuala Lumpur, working with teams worldwide.

Book a call about product management