Outcome School AI Engineering Stack Framework
Map any real-world AI product requirement onto the correct combination of LLM, RAG, MCP, Agents, Fine-Tuning, and Quantization so you can design and build it without confusion.
// TL;DR
The Outcome School AI Engineering Stack Framework is a decision system for mapping any AI product requirement onto the right combination of LLM, RAG, MCP, Agents, Fine-Tuning, and Quantization. Use it whenever you design or evaluate an AI feature, choose between approaches (like RAG vs Fine-Tuning or MCP vs direct API), or need to explain why a technique fits a scenario. It clarifies that AI engineering means USING existing foundation models — not building them — so any developer can design production AI systems without training models from scratch or getting confused by overlapping techniques.
// When should you use the AI Engineering Stack Framework?
Use this skill whenever you need to design or evaluate an AI-powered product feature, choose between AI engineering approaches, or explain why a particular technique (e.g. RAG vs Fine-Tuning, MCP vs direct API) is the right tool for a given scenario.
// What information do you need before designing an AI product with this framework?
- Product or feature descriptionrequired
What the AI-powered product or feature needs to do for end users. - Data sources availablerequired
What external data the system can access: PDFs, databases, live APIs, internal docs, etc. - Deployment constraints
Hardware limits, latency requirements, cost budget, and whether on-device/offline operation is needed. - Domain specificity requirement
Whether the product needs general knowledge or highly specific domain expertise (e.g. company-specific tone, coding language, medical domain).
// What are the core principles behind the AI Engineering Stack Framework?
AI Engineering vs Machine Learning distinction
Machine learning means BUILDING the model — writing algorithms, feeding training data, producing weights. AI engineering means USING the model — making foundation models workable inside real products. Any developer (Android, backend, DevOps) can do AI engineering; it does not require training from scratch.
LLM is a closed room
Think of the LLM as a kid locked in a closed room with no internet: it contains model.py and parameters.bin only. It works COMPLETELY OFFLINE without internet. It cannot make a network call. It can only understand language, predict text, summarise, and frame grounded responses. Any belief that the LLM itself fetches live data is incorrect.
Foundation Base Model
When a company open-sources a model, they release two things: model.py (the architecture code) and parameters.bin (the weights — billions of numbers that are the gold). The parameters.bin is the valuable artifact because replicating it requires $100M+ in GPU compute. This trained artifact is called the Foundation Base Model.
Weights are the gold
Parameters (weights W1, W2, W3 … to billions of W) are numbers produced by the training process and stored in parameters.bin. They encode everything the model learned. If offered data or weights, always prefer the weights — you cannot afford to reproduce them.
RAG — Retrieval Augmented Generation
Because the LLM was trained at a fixed point in time and cannot access live data, you RETRIEVE information from external sources (APIs, databases, PDFs, web), AUGMENT the LLM's input by adding that retrieved information to its context, and let the LLM GENERATE a grounded response. LLM contributes language intelligence; external sources contribute freshness.
Description is the superpower of MCP
The MCP server's metadata — especially the tool DESCRIPTION field — is what allows the LLM to act as the brain of the system. The LLM reads descriptions in plain English and infers which tool to call. A bad or missing description breaks LLM tool selection entirely.
MCP shifts the brain from backend to LLM
Without MCP, a backend developer must write endless if/else logic to route queries to different APIs — the backend becomes the superbrain. MCP eliminates that by exposing tool metadata so the LLM can decide which tool to invoke. The backend's only job becomes: load metadata on startup, loop to/from LLM, execute tool calls the LLM recommends.
Vector DB for relevant chunk retrieval
Rather than dumping an entire PDF into the LLM (slow, hits context window limits), split the document into chunks, convert each chunk to embedding vectors, store in a Vector DB optimised for similarity lookup, then query with the user's request to retrieve only the top-N relevant chunks. Feed only those chunks to the LLM.
Fine-Tuning — further training a Foundation Model
Fine-tuning takes a Foundation Base Model (already trained on general internet data) and does additional training on a smaller, domain-specific dataset. It is faster and cheaper than training from scratch because the model already has general language understanding. Fine-tuning updates the weights slightly so the model performs better on a specific task or domain without losing general intelligence.
Quantization — trading precision for size
Model size = number of parameters × bytes per parameter. At 32-bit (float32), each weight takes 4 bytes; 100B parameters = 400GB. At 16-bit, it halves to 200GB. Quantization reduces bit precision to shrink the model so it fits on machines with limited memory, at a modest cost to accuracy. Use it when you need to run a model on constrained hardware.
// How do you apply the AI Engineering Stack Framework step by step?
- 1
Clarify whether you are BUILDING or USING a model
If the task involves training algorithms, datasets, and producing weights from scratch → that is Machine Learning engineering, not AI engineering. If the task involves taking an existing model and making it useful in a product → proceed with this framework.
- 2
Identify what the LLM can answer from its internal knowledge alone
Ask: is this question answerable from general knowledge the model was trained on (e.g. 'What is reflection in programming?')? If yes, a direct LLM call suffices — no RAG, no tools needed. If the answer requires real-time data, private documents, or post-training events, proceed to step 3.
- 3
Determine which external data sources are needed and apply RAG
For each gap in LLM knowledge, identify the source: live API (crypto price, weather), web search results, internal database, or document/PDF. Plan to RETRIEVE from that source, AUGMENT the LLM prompt with retrieved content, and let the LLM GENERATE the grounded final response. Never ask the LLM to fetch data itself — it cannot make network calls.
- 4
Evaluate whether to use MCP servers to replace backend if/else routing logic
If you have more than one external data source and find yourself writing multiple if/else branches on the backend to route queries, switch to MCP. For each data source, find or build an MCP server that exposes: (1) tool name, (2) description in plain English, (3) input schema. Load all MCP metadata at server startup. Send user query + all metadata to the LLM. Let the LLM recommend which tool to call and with what parameters. Loop backend ↔ LLM until a grounded response is returned.
- 5
Handle document/PDF sources using Vector DB inside the MCP server
Do NOT feed entire documents to the LLM — this causes slow responses and may exceed the context window. Instead: chunk the document into paragraphs, convert chunks to embedding vectors, store in a Vector DB. At query time, embed the user query, retrieve top-N similar chunks from Vector DB, send only those chunks to the LLM. Expose this as an MCP tool with an ingest function and a query function.
- 6
Decide whether Fine-Tuning is needed for domain-specific quality
If the LLM's general responses are good but not specific enough to your domain (e.g. company tone, coding language, medical terminology), plan a fine-tuning run. Use the Foundation Base Model as the starting point. Curate a domain-specific dataset (e.g. company blog posts, code samples). Fine-tune by loading existing parameters.bin, running further training on the domain data, and saving updated parameters.bin. Do NOT fine-tune to inject real-time knowledge — use RAG for that. Fine-tune only to change style, domain depth, or task specialisation.
- 7
Apply Quantization if the model must run on constrained hardware
Calculate model size: parameters × bytes per parameter (32-bit = 4 bytes, 16-bit = 2 bytes). If the result exceeds your machine's memory, reduce bit precision via quantization (32-bit → 16-bit → 8-bit). Accept that lower precision = smaller size = slightly lower accuracy. Choose the lowest precision that still meets your quality bar. This is especially important for on-device or edge deployments.
- 8
Compose the full architecture and validate each component's role
Final checklist: UI/client → Backend (dump routing, loop logic only) → LLM (brain: tool selection + grounded response generation) → MCP Servers (external tools with metadata) → Vector DB (inside document MCP server for chunk retrieval) → Foundation Model optionally fine-tuned → quantized if hardware-constrained. Confirm: LLM never makes network calls. Backend never hard-codes routing logic. Tool descriptions are precise plain English.
// What are real examples of the AI Engineering Stack Framework in action?
A SaaS company wants an AI assistant that answers questions about their internal HR policy PDFs and also shows the current USD/EUR exchange rate.
Step 2: LLM can answer general HR concepts but not company-specific policy or live FX rates. Step 3: RAG needed for both. Step 4: Two MCP servers — one for FX rate API, one for HR PDF. Step 5: HR PDF MCP server chunks the policy documents, stores vectors in Vector DB, retrieves top-3 relevant chunks per query. FX MCP server calls a live rates API. Backend loads both MCP metadata on startup, sends user query + metadata to LLM, LLM recommends which tool to call, backend executes, result fed back to LLM for grounded response. Step 6: No fine-tuning needed — general language quality is sufficient. Step 7: No quantization needed if running on standard cloud GPU.
A coding education platform wants an AI tutor that explains Python concepts in its own simplified teaching style, not the generic internet style.
Step 2: LLM knows Python but explains it in a generalised internet style. Step 6: Fine-tuning required. Curate the platform's existing lesson content and blog posts as the fine-tuning dataset. Load the Foundation Base Model's parameters.bin, run further training on the curated dataset, save updated parameters.bin. The fine-tuned model now explains concepts in the platform's specific style. Step 3/4: If real-time docs or live code execution is also needed, add MCP servers for those on top of the fine-tuned model.
A startup wants to ship an on-device AI feature inside a mobile app on mid-range Android phones with 4GB RAM.
Step 7 is the primary concern. A 7B parameter model at 32-bit = ~28GB — impossible on device. Apply quantization: 16-bit halves to ~14GB, still too large. Move to 8-bit or 4-bit quantization to reach 3-7GB range compatible with device RAM. Accept reduced accuracy. Steps 3-5 still apply if the feature needs external data, but all external calls are made by the app's backend — the on-device LLM only handles language understanding and response generation, never network calls.
// What mistakes should you avoid when using the AI Engineering Stack Framework?
- Believing the LLM can make network calls — it cannot. LLM is model.py + parameters.bin in a closed directory, no internet access whatsoever.
- Retraining the Foundation Base Model from scratch to inject new or real-time information — this costs $100M+ and still won't stay current. Use RAG for live data instead.
- Writing if/else routing logic on the backend for every new data source — this turns the backend into a fragile superbrain. Use MCP to move routing intelligence to the LLM.
- Feeding entire documents (PDFs, long HTML pages) directly into the LLM context — this causes slow responses and hits context window size limits. Use chunking + Vector DB inside an MCP server.
- Using fine-tuning to update the model with recent facts or real-time data — fine-tuning is for style, domain depth, and task specialisation, not for knowledge freshness.
- Writing vague or missing tool descriptions in MCP server metadata — the LLM selects tools based on the description field in plain English. A bad description causes the LLM to pick the wrong tool or none at all.
- Confusing AI Engineering with Machine Learning — AI engineering is USING existing models in products; machine learning is BUILDING models. Most developers should focus on AI engineering.
- Ignoring context window size when designing RAG pipelines — even modern models with 1M token windows will respond slowly and degrade in quality if fed irrelevant bulk content. Always retrieve only the top-N relevant chunks.
// What are the key terms in the AI Engineering Stack Framework?
- Machine Learning
- The discipline of BUILDING models — writing algorithms, feeding training data, and producing weights. Includes deep learning engineers and ML researchers. Distinct from AI engineering.
- AI Engineering
- The discipline of USING models — taking Foundation Base Models and making them workable inside real products via fine-tuning, quantization, RAG, MCP, and agent patterns. Accessible to any developer.
- Foundation Base Model
- A large language model trained from scratch on massive internet-scale data, producing model.py and parameters.bin. Requires $100M+ investment. The starting point for all AI engineering work.
- model.py
- The architecture code file of an LLM — typically a few hundred lines of Python defining how the model processes input and produces output.
- parameters.bin
- The weights file — billions of numbers produced by the training process, stored in binary. Called 'the gold' because it encodes all learned knowledge and is prohibitively expensive to recreate.
- Weights (W1, W2 … Wn)
- Numbers inside parameters.bin that encode what the model learned. Analogous to multipliers in a pricing formula derived from historical data — inferred from the training dataset through the training process.
- Large Language Model (LLM)
- A language model called 'large' because (1) it was trained on a large dataset (the internet) and (2) it outputs a large number of parameters (millions to billions). Works completely offline — model.py + parameters.bin in a directory, no internet required.
- Context Window
- The maximum amount of text (measured in tokens, roughly words) that can be fed into an LLM in a single call. Current state-of-the-art is ~1 million tokens. Feeding irrelevant bulk content wastes this limit and slows responses.
- Tokenization
- The process of converting text into numbers before feeding to the LLM, using a mapping table (e.g. 'he' → 12, 'is' → 7). Required because the model performs mathematical operations and only understands numbers.
- Embeddings
- Converting scalar token numbers into n-dimensional vectors that represent meaning. Used both inside the LLM and separately to convert text chunks into vectors for Vector DB storage and similarity lookup.
- RAG (Retrieval Augmented Generation)
- A pattern to give LLMs access to external or real-time information: RETRIEVE relevant data from an external source, AUGMENT the LLM's input context with that data, and have the LLM GENERATE a grounded response. Solves the LLM's frozen knowledge cutoff problem.
- Grounded Response
- A final answer produced by the LLM that is accurate, well-framed, and directly serveable to the user — combining the LLM's language intelligence with retrieved external data.
- Tool
- An external function, API, or service that the LLM cannot call itself but can recommend the backend to call. The LLM identifies which tool is needed by reading tool descriptions in plain English.
- MCP (Model Context Protocol)
- An open-source standard for connecting AI applications to external systems — the equivalent of HTTP but for the AI context. MCP servers expose metadata (tool name, description, input schema) that the LLM reads to decide which external capability to invoke.
- MCP Server
- A server that exposes one or more tools via standardised metadata (name, description, input schema). The backend fetches this metadata at startup and sends it to the LLM with every user query so the LLM can recommend the right tool.
- MCP Metadata / Tool Description
- The JSON object an MCP server exposes containing tool name, description (plain English explanation of what the tool does), and input schema. The description field is the most critical element — it is how the LLM understands which tool to recommend.
- Vector DB
- A database optimised for similarity-based lookup of embedding vectors. Used in RAG pipelines to store document chunks as vectors and retrieve only the top-N most relevant chunks for a given query, avoiding full-document injection into the LLM.
- Chunking
- Splitting a large document (e.g. 100-page PDF) into smaller pieces (paragraphs or sections) before converting to embeddings and storing in a Vector DB. Enables precise retrieval of only relevant content.
- Fine-Tuning
- The process of further training a Foundation Base Model on a smaller, domain-specific dataset. Cheaper and faster than training from scratch because the model already has general language intelligence. Updates parameters.bin slightly to specialise the model for a specific task, domain, or style.
- Updated parameters.bin
- The output of a fine-tuning run — the same structure as the original parameters.bin but with weights adjusted for the specific domain or task.
- Quantization
- Reducing the bit precision of weights (e.g. from 32-bit float to 16-bit or 8-bit) to shrink model size and memory footprint. Formula: model size = number of parameters × bytes per parameter (32-bit = 4 bytes, 16-bit = 2 bytes). Trades accuracy for deployability on constrained hardware.
- Model Size
- The storage/memory footprint of a model: parameters × bytes per parameter. Example: 100B parameters at 32-bit = 400GB; at 16-bit = 200GB.
- Agent / Looping Code
- The backend pattern that loops between the LLM and tool calls: send query + metadata → LLM recommends tool → backend executes tool → send result back to LLM → repeat until LLM returns a grounded response instead of another tool recommendation.
// FREQUENTLY ASKED QUESTIONS
What is the Outcome School AI Engineering Stack Framework?
It's a decision framework that maps real-world AI product requirements onto the correct combination of LLM, RAG, MCP, Agents, Fine-Tuning, and Quantization. Instead of guessing which technique to use, you walk through steps that clarify whether you need live data (RAG), tool routing (MCP), domain style (fine-tuning), or hardware shrinking (quantization), then compose the full architecture.
What is the difference between AI engineering and machine learning?
AI engineering means USING existing foundation models to make them workable inside products, while machine learning means BUILDING models — writing algorithms, feeding training data, and producing weights. Any developer (Android, backend, DevOps) can do AI engineering because it doesn't require training a model from scratch, which costs $100M+ in GPU compute.
How do I decide between RAG and fine-tuning?
Use RAG when you need live, real-time, or private data the model wasn't trained on, and use fine-tuning when the model's answers are correct but not styled or specialized enough for your domain. RAG solves knowledge freshness; fine-tuning changes tone, domain depth, or task specialization. Never fine-tune to inject recent facts — that's what RAG is for.
How do I apply the AI Engineering Stack Framework step by step?
Start by confirming you're USING a model, not building one. Then check what the LLM already knows, identify data gaps and fill them with RAG, replace backend if/else routing with MCP servers, chunk documents into a Vector DB, add fine-tuning only for domain style, apply quantization if hardware is constrained, and finally validate every component's role in the full architecture.
When should I use MCP instead of direct API calls?
Use MCP when you have multiple external data sources and find yourself writing endless if/else routing logic on the backend. MCP exposes tool metadata (name, description, input schema) so the LLM decides which tool to call, shifting the 'brain' from your backend to the LLM. For a single fixed API with no routing decisions, a direct call is fine.
How does this framework compare to just prompting ChatGPT directly?
Direct prompting only works when the answer lives in the model's frozen training knowledge. This framework handles everything prompting can't: live data via RAG, private documents via Vector DB, tool orchestration via MCP, domain-specific style via fine-tuning, and on-device deployment via quantization. It's the difference between a chatbot and a real production AI product.
Can an LLM fetch live data from the internet by itself?
No — the LLM cannot make any network calls. It's model.py plus parameters.bin in a closed directory, working completely offline. It can only understand language, predict text, and frame grounded responses. To give it live data you must RETRIEVE from an external source and AUGMENT the LLM's input; the LLM never fetches anything itself.
What is quantization and when do I need it?
Quantization reduces the bit precision of model weights (32-bit to 16-bit or 8-bit) to shrink model size and memory footprint. Model size equals parameters × bytes per parameter, so a 100B model at 32-bit needs 400GB. Use quantization when a model must run on constrained hardware like mobile or edge devices, accepting a modest accuracy loss for deployability.
What results can I expect from using this framework?
You'll be able to design and build AI product features without confusion, choosing the correct technique for each requirement instead of over-engineering. Expect a clean architecture where the backend only routes and loops, the LLM acts as the brain for tool selection, RAG grounds responses in fresh data, and fine-tuning or quantization are applied only when genuinely needed.
Why is the tool description the most important part of an MCP server?
Because the LLM reads tool descriptions in plain English to infer which tool to call. The description field is what lets the LLM act as the brain of the system. A vague or missing description breaks tool selection entirely — the LLM will pick the wrong tool or none at all, so precise plain-English descriptions are critical.
Why shouldn't I feed an entire PDF into the LLM?
Feeding whole documents causes slow responses and can exceed the context window limit. Instead, chunk the document into paragraphs, convert each chunk to embedding vectors, store them in a Vector DB, then retrieve only the top-N relevant chunks for each query. Even models with 1M-token windows degrade in quality when fed irrelevant bulk content.