Skip to content
MKT Studios

· Technologist · 8 min read

No Authority, Full Context

How a hidden product manager builds every AI response

This is the first essay in a series of AI 101 posts that I’ll be writing. In each of these, I’ll break down topics and concepts related to Artificial Intelligence into digestible chunks with a product manager’s twist.

Today, when you enter a prompt into an AI like Claude, ChatGPT, or Gemini from providers like Anthropic, OpenAI, or Google, you’ll see it think for a bit and return a response. In reality, what you’re typing into is just a thin client that doesn’t do anything significant by itself. Like most user interfaces, it’s a facade that abstracts the technical complexities underneath. The response actually comes from several complex layers beyond what you see, and how it comes about looks a lot like how a product gets built.

The harness

The next layer your prompt goes to is the Harness that lives hidden from plain sight on your device, or on the servers of your organization or provider.

It is similar to a product manager in many ways. A product manager’s responsibilities include:

  • Understanding explicit and implicit customer and business needs.
  • Producing supplementary artifacts that better clarify the problem.
  • Drafting the product requirements document (PRD) including the problem and the artifacts.
  • Communicating the PRD to a cross-functional team for development.
  • Ensuring that the product gets clearance for shipping.
  • Shipping the finished product to users.

The harness does all of these.

The primary role of the harness is to draft the Context (PM analogy: PRD).

To do this, it gathers the following:

  • Your newest message and conversation history. (Customer needs)
  • A set of high-level instructions pre-baked into the harness by the provider, called System Prompt. (Business needs)
  • A dictionary of tools that the harness can invoke, including ones from third parties created via the Model Context Protocol (MCP). (Supplementary artifacts)
  • Documents that the harness can retrieve to augment the generated response (RAG). (Supplementary artifacts)

And while gathering, the harness optimizes context for ideal relevance and length by trimming or summarizing older history and retrieving relevant chunks of information, all of which have a direct impact on quality of output and performance of the system. It also indicates which portion of the context is cacheable to the provider, further boosting performance. And when this is done, it passes the context to the Model (cross-functional team) via its API.

The harness has no real authority over the model and purely influences its output with a powerful context. The model is also stateless. It retains no part of the context and requires the entire context—including the content marked cacheable—to be passed from scratch again. Think of this as a product manager updating requirements mid-development, except that, in this case, the cross-functional team always requires a full PRD rewrite for every iteration.

Once the model processes the context, it generates a response and passes it back to the harness. The harness ensures that the entire process stays within the guardrails set by the provider (clearance). Think of guardrails as caps like content safety concerns, security permissions, maximum tool calls, time consumed, or usage credits. And finally, when the harness determines that the response is ready, it presents the response back to you via the interface where it all started (shipping product).

The spectrum of prompt complexity

Your prompts to an AI are not always simple. They could fall anywhere across a spectrum of increasing sophistication as in the following examples:

  1. “What’s Tokyo like to visit in November?”
  2. “What’s the current exchange rate from USD to JPY?”
  3. “Find me a flight to Tokyo under $800 for the second week of November, check if it conflicts with anything on my calendar, and book it if it’s clear.”
  4. “Plan my Tokyo trip: research flights, find a hotel near Shibuya under $200/night, and put together a 5-day itinerary based on my interest in food and temples — then combine it all into one plan.”
  5. “Plan and book my entire Tokyo trip: first find flights. Only search hotels near whichever area the chosen flight actually arrives at. If total cost of flight plus hotel exceeds $1,500, drop to economy-tier hotel options and re-search. Once both are confirmed, check for calendar conflicts, and only finalize if there are none. Log every booking attempt.”

The most basic chatbot

Simple prompts like (1) are completed in a single run of the harness-model interaction. The harness passes the context to the model; the model returns a response that the harness surfaces to you. These are how Chatbot harnesses work. They can answer simple questions that just require a single harness-model roundtrip but can do nothing more.

The chatbot that does one more thing

Slightly more complex prompts like (2) require a little more than fetching a response from the model. These require invoking tools available to the harness such as the ability to search the internet, check the weather, execute some code, or some third party capability made available via the Model Context Protocol. So, in this flow, the harness reaches out to the model with the context. The model responds asking the harness to run a specific tool and to use the tool’s output in the response to you.

The need for agents

Significantly more complex prompts, such as (3) through (5) above, require much more work. These not only include tool runs (2) but also Loops, Orchestration, and Graphs.

Loops are iterative harness-model interactions (3). Orchestration involves managing multiple parallel but independent iterations (4). And graphs are multiple sequential iterations where the next step is conditional on the previous step’s outcome (5). All of these require one or more agents.

Agent is the name of a harness in a loop—drafting, sending, receiving, and revising—repeating its core process multiple times to arrive at the ideal response. Just like how product managers can’t always complete a PRD in a single run, harnesses too might require multiple runs (PRD iteration) to get to the best possible response.

In each loop, the harness updates the context, feeds the model, receives the response, and updates the context again until it thinks the response is ready for you or it hits a guardrail set by the provider. This loop is called the ReAct loop or the “Thought → Action → Observation, repeating”1 cycle.

The agent can update the context in several ways. It can use the output that comes from executing tools. Or it can also spawn multiple sub-agents that each update the context.

Sometimes, these sub-agents work independently with the parent agent orchestrating the operation. Think of this as a product leader who manages a portfolio of unrelated products with a team of product managers.

Otherwise, these sub-agents work sequentially and based on some specific condition being met. These related sub-agents comprise a graph. Think of this as a large enterprise product that has product managers owning individual components. Product updates could require changes to multiple components. This work sometimes has to be done sequentially; for example, update the foundational platform framework before its dependent clients can implement. And the components that need rework would vary based on findings from the prior set of changes. These related components form the graph that lives under the hood of a large enterprise product.

The spectrum of agent autonomy

Consider two other prompts:

  1. “Find flights to Tokyo under $800 for the second week of November. Don’t check my calendar or do anything else without asking me first — confirm before each step.”
  2. “Plan and book my entire Tokyo trip — flights, hotel, and a 5-day itinerary — based on my usual budget and preferences. Handle everything end to end; I don’t need to review anything unless something goes wrong.”

Prompt (6) explicitly requires the harness to check in prior to each step. Prompt (7), however, gives the agent full autonomy. There are several prompts that would sit in between, granting varying levels of autonomy to the agent. Generally, the need for agent autonomy/human involvement is determined either by the prompter, the provider, or the model itself.

Explicit prompting

If the prompt clearly specifies check-in points, the harness follows that instruction and seeks direction for next steps. So in prompt (6) above, the agent does what it is asked to do and checks in before proceeding. This is the simple case and is similar to a product manager iterating on the PRD with further input from the user.

Harness policy

Sometimes the model’s response could require certain actions from the harness, like accessing and updating a file on the system or running some command via a tool, that the harness is restricted from carrying out. These restrictions are explicitly built in as a guardrail by the provider. In such cases, the harness checks in with the prompter, seeks appropriate permissions, and proceeds based on the consent received. Consider these similar to stakeholder reviews seeking sign-offs for key product decisions.

Model’s judgment

Even if the prompt and the harness allow for full autonomy, the model’s response could require the harness to seek explicit permissions. While explicit prompting and harness policy are deterministic—they are based on explicit conditions—the model’s judgment is not guaranteed; it is probabilistic, as with most things concerning the model.

The PM parallel, if I have to, admittedly, stretch the analogy, would be the cross-functional team seeking user feedback on a build because they are conflicted on some decision.


Every response you get from an AI goes through one or more product development cycles. The harness uses the context to influence the model without authority, and updates the context using user and stakeholder guidance amid ambiguity. This is all too similar to how a product manager uses a PRD to build a product with a cross-functional team.

Most of the time, the shipped product is only as good as the final PRD that stands behind it. Writing a complete PRD end-to-end requires heart: you need to be able to fully empathize with your customers; it requires mind: you need to think of all the use cases the product must cover and to help your cross-functional team see what you see. And most importantly, it requires a lot of time—the most valuable resource for us humans. I once wrote a PRD for software that recorded inbound inventory at a grocery store. I had to work with a global engineering team, store operations, and my end users—the store workers and managers themselves. I had to fully understand how the stores operated physically as well as how the current software worked, including its design and architecture. I just had to write the PRD once and make occasional updates based on new findings as the team went through redesign. And that entire process took me months. The harness writes its PRD, the context, from scratch for every iteration of the response, and it does that significantly faster, often in minutes. And that fascinates me. I hope it uses AI.

Footnotes

  1. Yao et al., 2022, Princeton/Google Research ↩

← Back to all posts