Skip to content

AI Product Management

Harness Engineering: A PM's Guide to the Stack

Prompt, context, harness and loop engineering explained for AI product managers, plus the agent runtime and observability questions to ask your team.

Published
Read
8 min read

The vocabulary of AI product work has spent the past year moving outward from the model. Anthropic writes about context engineering. Thoughtworks and OpenAI write about harness engineering. Addy Osmani and LangChain write about loop engineering. Each term names a wider ring around the model, and each ring holds decisions that belong to the product manager as much as to the engineer.

If you manage an AI product, your job in this stack is to know which layer a problem lives in. The layer tells you who fixes it and how long the fix will take. Below are the four layers from the inside out, followed by the two topics underneath them that PMs get asked about most: the agent runtime and observability.

Prompt engineering: the instructions

The prompt is the text that tells the model who it is and how to behave. It is the smallest ring and the cheapest to change. It still carries product decisions: the agent's tone, what it refuses, how it says "I don't know", and which language it answers in when a customer switches languages halfway through a sentence.

Prompt fixes are fast, which makes them tempting. When the same failure returns after three prompt edits, the problem sits in a wider ring.

Context engineering: what the model knows at each step

Anthropic's Applied AI team defines context engineering as curating and maintaining the best set of tokens the model sees while it works. Their core argument is that context is a finite attention budget. As the window fills, the model's recall of any one detail degrades, an effect they call context rot.

For long tasks the team describes four techniques:

  • Compaction: summarise the history when the window fills, keeping decisions and open issues.
  • Structured notes: the agent writes progress to a file outside the window and reads it back later.
  • Sub-agents: hand focused tasks to agents with clean contexts that return short summaries.
  • On-demand retrieval: give the agent tools to fetch what it needs instead of loading everything up front.

The PM decisions at this layer concern knowledge. Which documents should the agent see, and which version wins when two conflict? What should it remember about a customer between sessions, and what must it forget? If your policy handbook exists in English and half your customers write in Spanish, which one does the agent retrieve? Those questions belong in the product spec, and they shape answer quality more than most prompt edits do.

Harness engineering: everything except the model

Birgitta Böckeler's article on harness engineering at martinfowler.com gives the clearest definition: the harness is everything in an agent except the model. She splits its controls into two kinds.

  • Guides act before the agent does. They steer it toward a good first attempt: instruction files, examples, templates and scoped permissions.
  • Sensors act after. They observe the result and let the agent correct itself: tests, linters, type checks, or a second model reviewing the output.

Each control can be computational (deterministic code, fast and reliable) or inferential (an LLM judging, slower but able to assess meaning). Böckeler notes that the harness for functional correctness, which she calls the behaviour harness, is the least developed of the three kinds she describes.

Mitchell Hashimoto, co-founder of HashiCorp, describes the working habit in his account of adopting AI: each time the agent makes a mistake, change its environment so that mistake cannot happen again. OpenAI's own harness engineering write-up sums up its internal experiment as humans steering while agents execute.

For a PM, the harness is where product policy becomes enforceable. "The agent must not approve refunds above $500" is a guide if you write it into instructions. It becomes a sensor when code blocks the action and routes it to a person. Böckeler argues that a good harness should direct human input "to where our input is most important." Deciding where that is falls to product and design.

Loop engineering: the system that runs the agent

Osmani's post on loop engineering describes a change of role. You stop being the person who prompts the agent and design the system that does it for you. His building blocks include:

  • scheduled automations that find and triage work
  • isolated worktrees so several agents can work in parallel
  • reusable skills that hold project knowledge
  • connectors to the tools your team already uses
  • sub-agents that separate doing the work from checking it

Progress lives in external state, such as a markdown file or a project board, because the model forgets between runs.

Sydney Runkle at LangChain frames the same idea as nested loops. The tool-calling loop sits inside a verification loop that grades output against rubrics. Event-driven loops, triggered by webhooks or schedules, wrap those. The outermost loop mines production traces to improve the harness. Humans stay at each level for judgement calls.

Osmani also names the risk. He warns about "cognitive surrender", the point where the people overseeing a loop stop exercising judgement and accept whatever it produces. So the PM questions at this layer are about attention. Which loops run unattended, and which wait for a person? What triggers each run? Who reads the output, and how often? A loop that drafts customer replies overnight for a human to approve in the morning is a different product from one that sends them.

The agent runtime underneath

"Agent runtime" now tends to mean a managed place to run long, stateful, tool-using sessions. Amazon Bedrock AgentCore is one example. It hosts agents built with several frameworks, including the Claude Agent SDK, with session isolation and support for long-running work.

Engineers pick the runtime. PMs should check that the choice answers four product questions:

  • Where does customer data sit, and does that meet your data protection law and any sector rules, such as a banking regulator's?
  • Can one customer's session leak into another's?
  • How long can a task run, and what does the user see while it does?
  • Can you see what the agent did?

Observability: what a PM should be able to see

The last question carries the most weight, because every other layer depends on it. You cannot improve context, harness or loop without traces.

OpenTelemetry, the open standard many engineering teams already use for tracing, now has semantic conventions for generative AI. They standardise spans for model calls and token counts, with optional capture of prompt and response content. Operation names such as invoke_agent and execute_tool mark each step; Greptime's walkthrough covers them in detail. Most of these conventions are still experimental. Claude Code and OpenAI Codex can already emit OpenTelemetry, and platforms such as Langfuse accept OTLP traces, so you are not tied to one vendor's SDK.

Questions to ask any team or vendor building your agent:

  1. Can I open one user's conversation and see each model call and tool call in order, with the documents it retrieved?
  2. Can I export 100 traces to review with my team?
  3. Do traces record cost and latency per step?
  4. Are traces emitted in OpenTelemetry format, or kept in a proprietary store?
  5. Who can read traces that contain customer data, and how long do you keep them?

Once you have traces, the next step is reading them with your designer. We cover that method in LLM error analysis is user research.

Start with the smallest ring that works

Anthropic's Building effective agents is a useful counterweight to agent hype. Many use cases need a single model call or a fixed workflow, and agents trade cost and latency for flexibility. Gartner predicts that over 40 percent of agentic AI projects will be cancelled by the end of 2027, pointing to rising costs and unclear business value among other reasons.

Adoption data from mid-size markets points the same way. An AWS study of Malaysian businesses, reported by Xinhua in August 2026, found that 67 percent of AI adopters stick to basic uses like public chatbots. For teams at that stage, the next gain sits in context and harness: better retrieval over their own documents, and clear rules for what needs a human's approval. Loops come after the team can see and judge what a single agent run does.

The PM's job in this stack

All four layers point to one idea: the product is the system around the model. Most of that system consists of decisions about what people see, approve and correct. Those are product and design decisions, and an AI PM who can place each one in the right layer will ship agents that hold up after launch.

For a longer view of how these layers fit a full build, read the full stack AI product and our practical guide to building an AI agent. If you are scoping an agent and want a second opinion on which layer to invest in first, see our product management consulting or get in touch.