Frequently Asked Questions About Outcome School AI Engineering Stack Framework
21 answers covering everything from basics to advanced usage.
// Basics
What does 'the LLM is a closed room' actually mean?
It means the LLM works completely offline — it's just model.py plus parameters.bin in a directory with no internet access. Think of it as a kid locked in a closed room who can only understand language, predict text, and summarize. It cannot make a single network call, so any belief that the LLM itself fetches live data is incorrect.
What is a Foundation Base Model?
A Foundation Base Model is a large language model trained from scratch on massive internet-scale data, producing two files: model.py (architecture code) and parameters.bin (the weights). It requires $100M+ in GPU compute to create and is the starting point for all AI engineering work. You never rebuild it — you use it or fine-tune it.
Why are the weights called 'the gold'?
Because parameters.bin contains billions of numbers produced by the training process that encode everything the model learned, and replicating them requires $100M+ in GPU compute. The architecture code (model.py) is only a few hundred lines and easily reproduced, but the weights are prohibitively expensive to recreate. If offered data or weights, always prefer the weights.
What does RAG stand for and how does it work?
RAG stands for Retrieval Augmented Generation. Because the LLM was trained at a fixed point and can't access live data, you RETRIEVE information from external sources (APIs, databases, PDFs, web), AUGMENT the LLM's input by adding that content to its context, then let the LLM GENERATE a grounded response. The LLM contributes language intelligence; external sources contribute freshness.
// How To
How do I set up a RAG pipeline for a large document?
Chunk the document into paragraphs or sections, convert each chunk to embedding vectors, and store them in a Vector DB optimized for similarity lookup. At query time, embed the user's query, retrieve the top-N most similar chunks, and send only those chunks to the LLM. Expose this as an MCP tool with an ingest function and a query function.
How do I know if my task is AI engineering or machine learning?
Ask whether you're building or using a model. If the task involves training algorithms, curating datasets, and producing weights from scratch, that's machine learning engineering. If it involves taking an existing model and making it useful inside a product via RAG, MCP, fine-tuning, or quantization, that's AI engineering — proceed with this framework.
How do I build an MCP server for a data source?
For each data source, find or build an MCP server that exposes three things: a tool name, a plain-English description of what it does, and an input schema. Load all MCP metadata at server startup, send the user query plus all metadata to the LLM, and let the LLM recommend which tool to call with what parameters. Loop backend and LLM until a grounded response returns.
How do I calculate whether a model fits my hardware?
Multiply the number of parameters by bytes per parameter: 32-bit is 4 bytes, 16-bit is 2 bytes, 8-bit is 1 byte. A 7B model at 32-bit is ~28GB; at 16-bit ~14GB; at 8-bit ~7GB. Compare that to your device's available memory, then choose the lowest precision quantization that still meets your quality bar.
// Troubleshooting
My LLM keeps picking the wrong MCP tool — what's wrong?
Your tool descriptions are likely vague or missing. The LLM selects tools purely by reading the description field in plain English, so it's the superpower of MCP. Rewrite each description to precisely explain what the tool does and when to use it. A bad description causes the LLM to pick the wrong tool or none at all.
My RAG responses are slow and low quality — how do I fix it?
You're probably feeding too much content into the context window. Even models with 1M-token windows degrade and slow down when fed irrelevant bulk. Make sure you're chunking documents, storing embeddings in a Vector DB, and retrieving only the top-N most relevant chunks per query — never dumping entire documents into the prompt.
I fine-tuned my model but it still doesn't know recent facts — why?
Because fine-tuning is for style, domain depth, and task specialization — not knowledge freshness. Fine-tuning adjusts weights to change how the model responds, not to inject real-time or recent information. For up-to-date facts, use RAG to retrieve live data and augment the prompt. Fine-tuning and RAG solve different problems and are often used together.
My backend is full of if/else routing logic — how do I clean it up?
Switch to MCP. Without it, the backend becomes a fragile superbrain routing queries to different APIs. With MCP, you expose each data source as a tool with metadata, load it at startup, and let the LLM decide which tool to invoke. The backend's only job becomes loading metadata, looping to and from the LLM, and executing tool calls.
// Comparisons
RAG vs fine-tuning: which should I choose for domain knowledge?
It depends on whether you need facts or style. Use RAG when the domain knowledge is factual, changes over time, or lives in documents and databases. Use fine-tuning when the model already has the knowledge but expresses it in a generic way and you need your specific tone, terminology, or teaching style. Many production systems use both together.
MCP vs direct API calls: what's the real difference?
Direct API calls require you to hard-code which endpoint to hit for each scenario, forcing the backend to be the decision-maker. MCP exposes tool metadata so the LLM decides which tool to call based on plain-English descriptions. MCP shines when you have multiple sources and dynamic routing; a single fixed API needs no MCP overhead.
How does this framework compare to using an agent framework like LangChain?
This framework is a decision system for choosing techniques, while agent frameworks are implementation tools. It tells you when to use RAG, MCP, fine-tuning, or quantization and why; a library like LangChain helps you build the looping code once you've decided. The framework prevents over-engineering; the library speeds up building what the framework told you to build.
Is quantization better than using a smaller model?
It depends on your quality bar. Quantizing a large model often preserves more capability than switching to a natively smaller model, because the large model learned richer representations. But quantization does trade precision for size at a modest accuracy cost. Test both against your quality bar and pick the smallest option that still passes for your hardware.
// Advanced
How do the LLM, backend, and MCP servers fit together in a full architecture?
The flow is: UI/client → Backend (routing and loop logic only) → LLM (the brain for tool selection and grounded response generation) → MCP Servers (external tools with metadata) → Vector DB (inside document MCP servers for chunk retrieval) → optionally fine-tuned Foundation Model → quantized if hardware-constrained. The LLM never makes network calls and the backend never hard-codes routing.
Can I combine fine-tuning, RAG, MCP, and quantization in one product?
Yes, and complex products often do. You might fine-tune a Foundation Model for domain style, quantize it to run on constrained hardware, wrap external data sources in MCP servers for tool routing, and use a Vector DB inside a document MCP server for RAG. Each technique solves a distinct problem, so layering them is normal and expected.
What is tokenization and why does it matter for context windows?
Tokenization converts text into numbers using a mapping table before feeding it to the LLM, because the model only understands numbers. The context window is measured in tokens (roughly words), with state-of-the-art around 1 million tokens. Understanding tokenization matters because every chunk you retrieve consumes tokens, so you must budget your context window carefully in RAG pipelines.
What is the agent looping pattern in this framework?
The agent loop is the backend pattern that cycles between the LLM and tool calls: send query plus metadata → LLM recommends a tool → backend executes it → send the result back to the LLM → repeat until the LLM returns a grounded response instead of another tool recommendation. This loop is how the LLM orchestrates multiple tools to solve complex requests.
How do embeddings differ inside the LLM versus in a Vector DB?
Embeddings convert scalar token numbers into n-dimensional vectors representing meaning. Inside the LLM, embeddings are part of how it processes input. In a RAG pipeline, you separately generate embeddings for text chunks and store them in a Vector DB so you can do similarity lookup and retrieve the most relevant chunks for a query. Same concept, two applications.