How PMs Scope AI Features and Choose the Right Approach

For technical product managers · Based on Outcome School AI Engineering Stack Framework

// TL;DR

Technical product managers can use the AI Engineering Stack Framework to scope AI features accurately and pick the right technique before writing a single spec. The framework gives you a shared vocabulary — LLM, RAG, MCP, fine-tuning, quantization — and a decision path to determine whether a feature needs live data, tool routing, domain styling, or hardware optimization. This lets you write requirements engineers can build directly, push back on over-engineering, and explain to stakeholders exactly why a technique fits a scenario. You'll stop conflating RAG with fine-tuning and stop assuming the LLM can fetch data it can't.

Why do AI feature specs so often miss the mark?

Because PMs frequently misunderstand what an LLM can and can't do. A common failure is assuming the LLM can fetch live data itself — it can't. It's a closed room: model.py plus parameters.bin, working offline with no network access. Another failure is confusing RAG with fine-tuning, or assuming every AI feature requires training a model. This framework gives you a precise mental model so your specs match technical reality and your engineers don't have to reverse-engineer your intent.

How do I scope whether a feature needs RAG?

Ask one question: can the LLM answer this from general knowledge it was trained on? If yes — like explaining a widely-known concept — a direct LLM call suffices, and your spec should say so. If the answer requires real-time data, private company documents, or anything after the model's training cutoff, the feature needs RAG. In your spec, name the exact source for each knowledge gap: a live API, a database, or a set of documents. This clarity tells engineers precisely what to retrieve and augment.

When should I spec MCP or fine-tuning in a feature?

Spec MCP when a feature touches multiple data sources and would otherwise force engineers into brittle routing logic. MCP lets the LLM decide which tool to call from plain-English descriptions, so include a note that each source should expose clear tool metadata. Spec fine-tuning only when the requirement is about style, tone, or domain depth — for example, an AI tutor that must explain concepts in your platform's specific teaching style rather than generic internet phrasing. Crucially, never spec fine-tuning to keep the model current on facts; that's a RAG requirement. Mixing these up is the single most common PM mistake.

Consider a coding education platform wanting an AI tutor that explains Python in its own simplified style. The LLM already knows Python but explains it generically. That's a fine-tuning requirement: curate the platform's existing lessons and blog posts as a dataset, run further training on the Foundation Base Model, and the model now teaches in the platform's voice. If the tutor also needs live docs, you'd layer MCP servers on top. Naming both requirements correctly in the spec prevents wasted engineering.

How do I handle features that run on user devices?

When a feature must run on-device — say, inside a mobile app on mid-range phones — quantization becomes the primary constraint. A 7B model at 32-bit is ~28GB, impossible on a 4GB device. Engineers will quantize down to 8-bit or 4-bit to reach a compatible size, accepting some accuracy loss. As a PM, your job is to flag the hardware constraint early and set a realistic quality bar, because on-device features trade accuracy for deployability. External data still flows through the app's backend, never the on-device model.

What's my next step?

For your next AI feature, write a one-page decision doc using the framework: state whether you're using or building a model (almost always using), list what the LLM already knows, name each data gap and its source, decide RAG vs direct call, note whether MCP is needed for multiple sources, and flag any fine-tuning or quantization requirement with its reason. Share it with engineering as the spec foundation. You'll cut clarification cycles and ship features that were scoped right the first time.

// FREQUENTLY ASKED QUESTIONS

How do I explain to stakeholders why we chose RAG over fine-tuning?

Explain that RAG solves knowledge freshness by retrieving live or private data and adding it to the model's input, while fine-tuning changes the model's style and domain depth but not its facts. If the feature needs current or private information, RAG is the correct and cheaper choice. Fine-tuning would neither keep data current nor be cost-justified for that requirement.

Can I write AI feature specs without a deep ML background?

Yes. AI engineering is about using existing models, not building them, so you don't need ML research knowledge. Focus on the decision path: what the model already knows, what data gaps exist, and which technique fills each gap. The framework's shared vocabulary lets you communicate precisely with engineers without understanding model internals.

How do I estimate effort for an AI feature during planning?

Break the feature into framework components. A direct LLM call is lightest. RAG adds retrieval and a Vector DB. MCP adds tool servers and the agent loop. Fine-tuning adds dataset curation and a training run. Quantization adds testing across precision levels. Count the components each feature genuinely needs — over-scoping happens when teams assume all of them apply.