Scoping AI Agents: A PM's Framework

For Technical product managers · Based on Rajeev Kanth Agentic AI Architecture Stack

// TL;DR

Technical product managers can use the Agentic AI Architecture Stack to scope AI agent projects correctly before engineering starts. The framework's use-case-first principle forces you to define what friction you're removing and what non-value-added work you're eliminating — the two questions that justify any agent. It also gives you a shared vocabulary for evaluating proposals: is this genuinely agentic, or just generative AI? Use it to right-size model spend, decide single versus multi-agent, ensure all four guardrail layers are planned, and set clear test criteria so you can measure whether the agent is production-ready.

Why should product managers start with the use case, not the technology?

Because building an agent without knowing why is the most common and most costly mistake. Before any model or framework is chosen, answer two questions: what friction am I removing — which multi-step human process am I collapsing — and what non-value-added activity am I releasing humans from, meaning work the organisation does that customers won't pay for. Document these first. If you can't answer why the agent exists, stop. These two answers become the acceptance criteria you'll test against later and the basis for ROI justification to stakeholders.

How do I tell if a proposed 'AI agent' is actually an agent?

Audit it against the four core traits. Autonomy: does it detect errors in its own output and self-correct through loops without human intervention? Tool use: is it connected to external tools like web search, databases, or email to take real-world actions? Memory: does it retain short-term conversation and long-term organisational context? Reflection: does it evaluate its output against user intent before finalising? If a proposal is just an LLM taking an input and returning an output, it's generative AI — not an agent — and should be scoped and priced accordingly.

How do I keep agent projects from over-spending on models and complexity?

Two rules protect your budget. First, model selection must be justified by use case requirements, token billing at projected scale, and projected ROI — never by defaulting to the largest or most prestigious model. High-volume moderate-reasoning use cases run fine on mid-tier models. Second, if a single agent can solve the use case, don't approve multi-agent complexity. Multi-agent architectures increase failure surface, debugging cost, and time to ship. Require the team to demonstrate single-agent viability before adding orchestrators and sub-agents.

How do I make sure the agent is safe before it ships?

Require all four guardrail layers in the plan: Input (which inputs the agent will and won't respond to), LLM (behavioural constraints on tone, scope, and escalation), Tool (what actions tools are permitted to take), and Output (format, safety, and quality on every response). Guardrails are not optional in production — omitting any one creates unpredictable behaviour. For agents that mutate real systems, insist on tight tool guardrails, such as restricting code writes to non-main branches.

How do I define done for an AI agent?

Define done as passing the use case criteria set in Step 1, measured on realistic cases. In a logistics example, the team tests on 50 historical cases and measures accuracy and escalation rate before deploying. Set concrete metrics — resolution rate, escalation rate, loop count, accuracy — and treat failure to meet them as a trigger to loop back to the relevant layer and iterate, not to ship anyway.

Next step: Draft a one-page use case brief answering the friction and non-value-added questions, list the required tools and memory sources, and define your acceptance metrics. Hand that to engineering as the scope contract before any architecture decisions are made.

// FREQUENTLY ASKED QUESTIONS

How do I justify an agent's ROI to leadership?

Quantify the friction removed — how many steps in a multi-step process the agent collapses — and the non-value-added activity eliminated, meaning work the organisation does that customers won't pay for. Then tie the LLM choice to token cost at projected scale and the value of solving the use case. This makes cost and benefit explicit before development begins.

What questions should I ask a vendor claiming to sell an 'AI agent'?

Ask whether the system self-corrects its own errors through loops (autonomy), what external tools it uses to take actions (tool use), how it handles short-term and long-term memory, and whether it evaluates its output against user intent (reflection). If it can't demonstrate all four, it's generative AI, not an agent, and should be evaluated as such.

How do I set acceptance criteria for an agent build?

Derive acceptance criteria directly from your Step 1 use case: the specific friction removed and non-value-added work eliminated. Convert these into measurable metrics like accuracy, escalation rate, and loop count, then require testing on realistic historical or sandboxed cases before deployment. If the agent fails the criteria, the team iterates on the relevant layer rather than shipping.