Frequently Asked Questions About Production RAG Pipeline Engineering Framework
21 answers covering everything from basics to advanced usage.
// Basics
What does RAG stand for and what problem does it solve?
RAG stands for Retrieval Augmented Generation. It grounds LLM responses in actual retrieved documents, reducing hallucination by augmenting the model's knowledge with relevant external chunks at query time. Instead of relying on what a model memorized during training, RAG fetches current, domain-specific context and passes it into the prompt so answers stay accurate and verifiable.
What are the two phases of a RAG system?
The indexing phase and the querying phase. Indexing is a one-time process: load documents, chunk them, embed each chunk, and store vectors in the vector store. Querying happens per question: embed the user query with the same model, search for similar vectors, retrieve original text chunks, and pass them with the query to the LLM. Consistency between phases is critical.
What are the five main failure modes of production RAG?
One, bad chunking that shatters meaning. Two, wrong or mismatched embedding models between indexing and querying. Three, missing or weak grounding prompts that allow hallucination. Four, no metadata filtering, causing irrelevant retrieval. Five, no source attribution, making answers unverifiable. These five cause roughly 90% of production RAG failures — address each systematically before declaring the system production-ready.
Should I optimize for retrieval accuracy or processing speed?
It depends on your use case constraint. If accuracy is paramount — legal, medical, technical knowledge bases — use semantic chunking and higher-dimension embeddings. If throughput matters more — quick prototypes, large simple corpora — use recursive chunking. Define this quality-versus-speed constraint upfront, because it drives your chunking strategy, embedding model, and cost profile.
// How To
What are the four chunking variables I need to tune?
Chunk size (200–1,000 tokens; too small loses fragments, too large dilutes meaning), overlap (50–200 tokens to preserve context at boundaries), split boundaries (fixed→recursive→semantic in ascending quality), and content type (code keeps functions together, legal needs semantic, markdown splits on headers). Tuning these four determines retrieval quality more than any LLM prompt change.
How do I choose the right document loader in LangChain?
Match the loader to your content type: PyPDFLoader for simple PDFs, PyMuPDFLoader for high-volume PDFs (fastest, best metadata), UnstructuredPDFLoader for complex layouts with tables, TextLoader for plain text, DirectoryLoader for bulk folders (use glob patterns), and WebBaseLoader for URLs. Each outputs LangChain Document objects with page_content and metadata fields like source and page number.
How do I build the RAG chain with a grounding prompt?
Construct it as: {context: retriever | format_docs, question: RunnablePassthrough()} | prompt_template | llm | StrOutputParser(). Context comes from the retriever, RunnablePassthrough ensures the original query reaches the LLM unchanged, and the prompt must explicitly say 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know.'
How do I configure the retriever for similarity search?
Instantiate the retriever from your vector store with search_type='similarity' and set k (documents to retrieve, typically 2–5). Add metadata filtering to narrow retrieval to relevant subsets like topic, source, or date. Understand your score semantics: distance scores mean closer to 0 is more relevant; similarity scores mean closer to 1. Convert if needed via similarity = 1 / (1 + distance).
How do I prevent duplicate documents when re-running indexing?
Use get_or_create_collection and upsert instead of add when populating the vector store. The add method appends documents on every run, so repeated indexing duplicates all your data and corrupts retrieval results. Upsert updates existing entries in place, making re-runs safe and idempotent.
// Troubleshooting
My RAG chatbot returns incomplete answers — how do I diagnose it?
Test retrieval in isolation and inspect the raw chunks. Incomplete answers often mean a concept was split across two chunks — for example an OAuth2 section shattered between chunk 47 and 48, so neither holds the full flow. This usually points to fixed-size chunking. Switch to semantic chunking so the section stays together, re-embed with the same model, and retest.
Why is my similarity search returning irrelevant results?
The most common cause is an embedding model mismatch between indexing and query time — even a version difference makes vectors incomparable and search silently fails. Confirm you're using the identical model everywhere. If that's correct, check for missing metadata filtering adding noise at scale, or fixed-size chunking producing meaningless fragments that embed poorly.
My LLM is hallucinating case citations — what's wrong?
You're likely missing the grounding prompt and source attribution. Add 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know,' and use format_docs_with_source so every answer includes the document and page. Then test retrieval separately to confirm the retrieved chunks actually contain the relevant clauses before blaming generation.
Retrieval quality collapsed after scaling from 10 to 10,000 documents — why?
Check the three production rules. Confirm the same embedding model is used at indexing and query time. Migrate any fixed-size chunking to recursive with 500–1,000 token chunks and 50–200 token overlap. Then run the retriever in isolation on 10 representative queries and inspect output. Add metadata filtering to reduce noise, since remember: better embeddings on fewer docs beat worse embeddings on more.
// Comparisons
How does this framework compare to just throwing documents at an LLM's context window?
Dumping documents into context breaks at scale, wastes tokens, and can't cite sources reliably. This framework indexes documents once, retrieves only the most relevant chunks per query, and grounds the answer with source attribution. It stays fast and cheap at 10,000+ documents where a raw context-stuffing approach becomes impossible or prohibitively expensive.
How does fixed-size chunking compare to recursive chunking?
Fixed-size chunking cuts every n characters regardless of sentence or concept boundaries, destroying meaning and causing retrieval failure before the LLM is involved — never use it in production. Recursive chunking respects a hierarchy of paragraphs, sentences, clauses, then words, always preferring natural boundaries. Recursive is LangChain's reliable default and gets you 80% of retrieval quality.
How does late chunking compare to semantic chunking?
Semantic chunking splits at meaning boundaries by detecting embedding similarity drops and is the practical best today. Late chunking embeds the full document first, then pools embeddings for individual chunks, preserving cross-chunk context and yielding roughly 10–12% accuracy improvement — but it requires models like Jina Embeddings v2. The creator's verdict: semantic is best now, late chunking is the future.
// Advanced
Which embedding model should I choose for production?
The sweet spot is 768–1,536 dimensions. text-embedding-3-small (1,536 dims) is a good balance; text-embedding-3-large (3,072 dims) is highest quality but more storage; Gemini embeddings (768 dims) are free and smaller; BGE-small (384 dims) is fastest but least nuanced. Choose based on your quality/speed/cost trade-off, then lock it before indexing.
What is Agentic RAG and when should I consider it?
Agentic RAG is an architecture where an agent reasoning loop drives retrieval decisions dynamically rather than a fixed single retrieval step. Consider it only after your baseline RAG is production-stable and you need multi-step or iterative retrieval — for example when a query requires reasoning about what to fetch next. Don't reach for it before mastering the fundamentals.
What is Graph RAG and what problem does it solve?
Graph RAG augments retrieval with a knowledge graph to capture entity relationships and enable multi-hop reasoning across documents. It shines when answers require connecting facts scattered across many sources rather than retrieving one relevant chunk. It's an advanced architecture to evaluate after your baseline pipeline is stable and simpler retrieval can't handle relationship-heavy queries.
How do I chunk source code for a RAG system?
Use a dedicated Code Splitter that keeps functions together, never chunk code by character count. Splitting a for-loop or function across two chunks makes both semantically useless for retrieval. The Code Splitter respects language structure so a complete function stays in one chunk, preserving the logic the embedding needs to represent.
What is the difference between distance scores and similarity scores?
Distance scores mean closer to 0 is most relevant (Chroma uses these by default); similarity scores mean closer to 1 is most relevant. Know which your vector store returns before setting thresholds, or you'll filter out your best results. Convert between them with similarity = 1 / (1 + distance).