How to Take a RAG Prototype to Production

For ML engineers migrating a prototype to production · Based on Production RAG Pipeline Engineering Framework

// TL;DR

ML engineers whose RAG prototype worked on 10 documents but collapsed at 10,000 can use this framework to diagnose and fix the migration systematically. The three production rules: confirm the same embedding model at indexing and query time, replace fixed-size chunking with recursive (500–1,000 tokens, 50–200 overlap) or semantic, and test retrieval in isolation before blaming the LLM. Add metadata filtering to reduce noise at scale and use upsert to avoid duplicate documents. Since 90% of failures are retrieval failures, fix retrieval first, then evaluate advanced architectures like late chunking.

Why did my RAG prototype collapse when I added more documents?

Because prototypes hide the failure modes that only surface at scale. On 10 clean documents, a fast embedding model and naive chunking look fine. Add 10,000 and three things break: the chunking strategy shatters meaning across thousands of documents, any embedding mismatch amplifies, and without metadata filtering the retriever drowns in noise. The framework's guiding rule holds: better embeddings on fewer documents beat worse embeddings on more, and 90% of RAG failures are retrieval failures. Diagnose retrieval before you touch the LLM.

What are the three production rules to check first?

Rule one — same embedding model everywhere. Confirm the model used at indexing exactly matches the model at query time, including version. Different models or versions produce incomparable vectors, so similarity search silently fails. If you prototyped with a fast, low-dimension model like BGE-small (384 dims) and want more nuance, you must re-index the whole corpus after switching.

Rule two — real chunking. If you're on fixed-size chunking, migrate immediately; it cuts at arbitrary character counts and destroys meaning before the LLM is involved. Adopt recursive chunking with 500–1,000 token chunks and 50–200 token overlap as your production default. It follows a paragraph→sentence→clause hierarchy and delivers 80% of retrieval quality. Upgrade to semantic chunking only for content where accuracy is critical.

Rule three — test retrieval in isolation. Run the retriever alone on 10 representative queries and inspect the raw chunks. Are they relevant? Complete, or fragments? If retrieval is broken, fix chunking and the embedding model before evaluating generation.

How do I harden the pipeline for scale?

Add metadata filtering to narrow retrieval to relevant subsets — topic, source, or date — which cuts noise dramatically at scale. Set `search_type='similarity'` and a sensible `k` (2–5). Know whether your vector store returns distance scores (closer to 0 = relevant) or similarity scores (closer to 1); Chroma uses distance by default, and you can convert via `similarity = 1 / (1 + distance)`. Critically, populate the store with `get_or_create_collection` and `upsert`, never `add` — repeated indexing runs with `add` duplicate every document and corrupt retrieval.

Ground the LLM with the standard prompt and attach sources with `format_docs_with_source` so answers stay verifiable. Then walk the five failure modes as a checklist: bad chunking, mismatched embeddings, weak grounding, no filtering, no attribution.

When should I reach for advanced RAG architectures?

Only after the baseline is production-stable. Once retrieval is solid, evaluate late chunking (embed the full document, then pool chunk embeddings — ~10–12% accuracy gain, needs models like Jina Embeddings v2), semantic chunking for the last quality margin, Agentic RAG for reasoning-driven retrieval, Graph RAG for multi-hop relationship queries, and Multimodal RAG for images and tables. The creator's verdict: semantic chunking is the practical best today, late chunking is the future. Don't skip the fundamentals to chase these — a stable baseline is the prerequisite.

Next step: Run the three-rule audit on your prototype today — verify the embedding model matches, replace fixed-size chunking, and test retrieval in isolation — then remediate the five failure modes before you call it production-ready.

// FREQUENTLY ASKED QUESTIONS

How do I know if my problem is retrieval or generation?

Run the retriever in isolation and inspect the raw chunks it returns before the LLM sees them. If the chunks are irrelevant or fragmented, it's a retrieval problem — fix chunking and confirm the embedding model matches. If the chunks are complete and relevant but the answer is still wrong, then investigate the grounding prompt and generation. 90% of the time it's retrieval.

Should I use add or upsert to populate my vector store?

Always use upsert with get_or_create_collection. The add method appends documents on every run, so re-indexing duplicates your entire corpus and corrupts retrieval results. Upsert updates existing entries in place, making your indexing pipeline idempotent and safe to re-run during migration and updates.

Do I need to re-index everything if I change chunking strategy?

Yes. Chunking happens before embedding, so changing chunk size, overlap, or strategy produces different chunks and therefore different vectors. You must re-run the indexing phase — load, chunk, embed, upsert — with the same locked embedding model. Test retrieval in isolation afterward to confirm the new chunks actually improve relevance.