How Backend Developers Add AI Without Rewriting Routing Logic
For backend developers · Based on Outcome School AI Engineering Stack Framework
// TL;DR
Backend developers can use the AI Engineering Stack Framework to add AI to products without turning their backend into a fragile superbrain of if/else routing. The framework shows you how to expose data sources as MCP servers with clean metadata, let the LLM decide which tool to call, and reduce your backend's job to loading metadata, looping to and from the LLM, and executing tool calls. You'll also learn to handle documents with chunking and a Vector DB, keep the LLM stateless of network calls, and apply RAG correctly so your services stay maintainable as data sources multiply.
Why does my backend keep becoming a fragile superbrain?
Because every time you add a new data source, you write more if/else logic to route queries to the right API. Do this five times and your backend is an unmaintainable decision tree that breaks on every change. The AI Engineering Stack Framework fixes this by shifting the brain from your backend to the LLM. Your backend stops making routing decisions and instead exposes each source's capabilities as metadata the LLM reads to decide what to call.
How do I move routing intelligence to the LLM with MCP?
For each data source, build or find an MCP server exposing three things: a tool name, a plain-English description of what it does, and an input schema. Load all MCP metadata at server startup. On each request, send the user query plus all metadata to the LLM. The LLM reads the descriptions and recommends which tool to call with which parameters. Your backend executes that tool and sends the result back to the LLM. Loop until the LLM returns a grounded response instead of another tool recommendation. That loop — the agent pattern — is now your entire orchestration layer.
The most important detail: the tool description field is the superpower of MCP. The LLM selects tools purely from these plain-English descriptions. Write them vaguely and the LLM picks wrong tools or none. Treat descriptions like API documentation for a very literal reader — precise, specific, and clear about when to use each tool.
How should I handle documents and PDFs in my backend?
Never dump an entire document into the LLM — it's slow and hits context window limits. Instead, inside a document MCP server, implement two functions: an ingest function that chunks the document into paragraphs, converts each chunk to embedding vectors, and stores them in a Vector DB; and a query function that embeds the user's query, retrieves the top-N most similar chunks, and returns only those. The LLM receives just the relevant chunks. Even with million-token context windows, this keeps responses fast and accurate.
What must my backend never do?
Two things. First, never let the LLM make network calls — it can't. The LLM is model.py plus parameters.bin in a closed directory with zero internet access. All fetching happens in your backend or MCP servers. Second, never hard-code routing logic once you have multiple sources — that's what MCP eliminates. Your backend's job narrows to: load metadata on startup, loop to and from the LLM, and execute the tool calls the LLM recommends. Keep it dumb; keep the LLM smart.
What about deployment constraints?
If your feature must run a model on constrained hardware, calculate model size as parameters × bytes per parameter (32-bit = 4 bytes, 16-bit = 2 bytes, 8-bit = 1 byte). If it exceeds available memory, apply quantization step by step down the precision ladder, choosing the lowest precision that still meets your quality bar. For on-device features, all external calls still go through the app's backend — the on-device LLM only does language understanding and generation.
What's my next step?
Audit your current or planned AI service for if/else routing. For each branch, ask whether it should become an MCP tool with a clean description. Refactor toward the agent loop, add a Vector DB behind any document source, and confirm your LLM never touches the network. You'll end up with a backend that scales cleanly as you add data sources instead of collapsing under conditional logic.
// FREQUENTLY ASKED QUESTIONS
Does using MCP mean I stop writing backend code entirely?
No, but your backend's role shrinks dramatically. You still write the loop logic that sends queries and metadata to the LLM, executes the tools the LLM recommends, and returns results. What you stop writing is the branching routing logic that decides which API to call — that decision moves to the LLM based on tool descriptions.
How many chunks should I retrieve from my Vector DB per query?
Retrieve the top-N most relevant chunks, where N is tuned to your context window and quality needs — often 3 to 10. The goal is to give the LLM enough relevant context without wasting the context window on irrelevant content. Start small, test response quality, and increase only if answers lack necessary information.
Can I add MCP on top of a fine-tuned or quantized model?
Yes. MCP, RAG, fine-tuning, and quantization are independent layers. You can fine-tune a Foundation Model for domain style, quantize it for constrained hardware, and still wrap external data sources in MCP servers for tool routing. The agent loop and MCP metadata work the same regardless of whether the underlying model is base, fine-tuned, or quantized.