Production RAG Pipeline Engineering Framework

Transform a prototype RAG system that works on 10 documents into a production-grade, debuggable, and scalable system that works on 10,000+ documents without breaking.

// TL;DR

The Production RAG Pipeline Engineering Framework is a systematic approach for turning a prototype Retrieval Augmented Generation system into a production-grade one that works on 10,000+ documents without breaking. It covers document loading, chunking strategy, embedding model selection, vector store population, retrieval configuration, and source attribution. Use it whenever you're building, auditing, or scaling a RAG system — especially when retrieval quality is poor, the system breaks at scale, or you're moving from a working demo to a real deployment. Its core insight: 90% of RAG failures are retrieval failures, so fix chunking and retrieval before blaming the LLM.

// When should you use the Production RAG Pipeline Engineering Framework?

Use this skill whenever you are building, auditing, or scaling a Retrieval Augmented Generation (RAG) system — especially when retrieval quality is poor, the system breaks at scale, or you need to move from a working prototype to a production deployment.

// What do you need before building a production RAG pipeline?

  • Document corpusrequired
    The raw files to be indexed (PDF, TXT, HTML, DOCX, CSV, etc.) and their approximate size and content type
  • Use case descriptionrequired
    What the RAG system must answer, who the users are, and whether accuracy or speed is the higher priority
  • Embedding model choicerequired
    Which embedding model will be used for indexing (e.g., text-embedding-3-small, text-embedding-3-large, BGE-small). Must be identified upfront and kept consistent.
  • LLM choicerequired
    The chat model used for generation (e.g., GPT-4o, Claude Sonnet). Separate from the embedding model.
  • Vector store choice
    Which vector database will be used (e.g., Chroma, Pinecone, FAISS)
  • Quality vs. speed constraint
    Whether the system must prioritise retrieval accuracy (use semantic chunking) or processing throughput (use recursive chunking)

// What core principles govern production RAG engineering?

Chunking is Architecture, Not Pre-processing

Chunking decisions ripple through the entire RAG pipeline — what gets retrieved, what context the LLM sees, and what answer the user gets. Bad chunking produces bad embeddings, which produces broken RAG. The single biggest lever you have over retrieval quality is how you chunk.

Same Embedding Model Everywhere

The identical embedding model must be used at both the indexing phase and the querying phase. If you switch models between phases, your vectors will not be comparable and the similarity search will fail. Version matters too.

Embedding Quality > Embedding Quantity

Better embeddings on fewer documents beats worse embeddings on more documents. 90% of RAG failures are retrieval failures, not generation failures — so fix retrieval before blaming the LLM.

Ground the LLM with Prompt Engineering

Use the prompt pattern: 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know.' This is low-hanging fruit that prevents hallucination and many people ignore it.

Always Retrieve Sources

Tag every chunk with its source document using format_docs_with_source so users can verify answers. Sources build trust and allow factual verification. This is essential in production, not optional.

The 80/20 Chunking Rule

Recursive chunking with good overlap gets you 80% of retrieval quality. Semantic chunking gets you the last 20% but costs more compute. Start with recursive, measure retrieval quality, upgrade to semantic only if needed.

Test Retrieval Separately from Generation

Before blaming the LLM for bad answers, inspect what documents are being retrieved. 90% of RAG failures are retrieval failures. Retrieval quality determines answer quality — the LLM can only work with what you give it.

// How do you build a production RAG pipeline step by step?

  1. 1

    Load documents using the appropriate Document Loader

    Select loader based on content type: PyPDFLoader for simple PDFs (fast, basic metadata), PyMuPDFLoader for high-volume PDFs (fastest, best metadata), UnstructuredPDFLoader for complex layouts with tables (slower, richest metadata), TextLoader for plain text, DirectoryLoader for bulk folder ingestion (use glob pattern to filter file types), WebBaseLoader for URLs. Output is a list of LangChain Document objects with page_content and metadata fields (source, page number, author, etc.).

  2. 2

    Choose and apply the correct Chunking Strategy

    Apply the Chunking Decision Framework: (1) Is this a quick prototype or simple structured doc? → Use Recursive Chunking. (2) Is quality critical (legal, technical manuals, knowledge bases)? → Use Semantic Chunking. (3) Is it code? → Use the Code Splitter. (4) Is it Markdown? → Use the MD Splitter with header-based chunk size. NEVER use Fixed Size Chunking in production — it destroys meaning by splitting at arbitrary character counts regardless of sentence or concept boundaries. For Recursive Chunking, set chunk_size between 500–1,000 tokens and add overlap (typically 50–200 tokens) to preserve context at boundaries.

  3. 3

    Understand the four Chunking Variables and tune them

    Variable 1 — Chunk Size: too small = fragment loss and retrieval noise; too large = diluted meaning and wasted token budget. Sweet spot: 200–1,000 tokens. Variable 2 — Overlap: preserves context at chunk boundaries; no overlap = context loss between chunks. Variable 3 — Split Boundaries: Fixed (random cut, avoid in production) → Recursive (paragraph/sentence/clause hierarchy, reliable default) → Semantic (meaning boundaries, premium quality). Variable 4 — Content Type: code must keep functions together; legal docs need semantic chunking; markdown splits on headers.

  4. 4

    Select and lock the Embedding Model before indexing

    Choose based on the quality/speed/cost trade-off. Dimensions indicate semantic capacity: text-embedding-3-small = 1,536 dims (good balance); text-embedding-3-large = 3,072 dims (highest quality, more storage); Gemini embeddings = 768 dims (free, smaller); BGE-small = 384 dims (fastest, lowest semantic nuance). Sweet spot for production: 768–1,536 dimensions. Lock this model before indexing and never change it without re-indexing your entire corpus.

  5. 5

    Generate embeddings and populate the Vector Store (Indexing Phase)

    This is a one-time process per corpus update. Each chunk is passed through the embedding model to produce a vector. Vectors, original text chunks, and metadata are stored together in the vector store (e.g., Chroma, Pinecone). Use get_or_create_collection and upsert (not add) so re-runs do not duplicate documents. At the end of indexing, the vector store contains thousands of vectors each representing a chunk of your documents.

  6. 6

    Build the RAG Chain with RunnablePassthrough and a grounding prompt

    Construct the chain as: {context: retriever | format_docs, question: RunnablePassthrough()} | prompt_template | llm | StrOutputParser(). The context input comes from the retriever (relevant document chunks). The question uses RunnablePassthrough so the original query reaches the LLM unchanged. The prompt must explicitly state: 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know.' Pass both context and question into the prompt template.

  7. 7

    Configure the Retriever with similarity search and metadata filtering

    Instantiate the retriever from the vector store with search_type='similarity' and set k (number of documents to retrieve, typically 2–5). Use metadata filtering to narrow retrieval to relevant subsets (e.g., filter by topic, source, date). Understand score semantics: distance scores (closer to 0 = more relevant) vs. similarity scores (closer to 1 = more relevant). Convert if needed: similarity = 1 / (1 + distance).

  8. 8

    Attach source tags to every retrieved chunk

    Use format_docs_with_source to add source tags to each chunk before it enters the prompt. Output format should include the source document name/path alongside the page content. This allows users to verify answers, enables citation, and builds system trust. Sources are essential for production — not optional.

  9. 9

    Test retrieval in isolation before evaluating generation

    Run queries against the retriever alone and inspect the raw documents returned. Ask: Are these chunks actually relevant to the query? Do they contain complete information or are they fragments? If retrieval is broken, fix chunking and/or the embedding model before touching the LLM prompt. 90% of RAG failures are retrieval failures.

  10. 10

    Identify and remediate the Five Main Failure Modes

    The five failure modes that cause 90% of production RAG failures are: (1) Bad chunking → shattered meaning, incomplete embeddings; (2) Wrong or mismatched embedding models between indexing and querying; (3) Missing or weak retrieval prompt grounding → hallucination; (4) No metadata filtering → irrelevant document retrieval; (5) No source attribution → unverifiable answers. Address each systematically before declaring the system production-ready.

  11. 11

    Evaluate advanced architectures for cutting-edge requirements

    Once the baseline RAG is production-stable, consider: Late Chunking (embed full document first, then pool chunk embeddings — 10–12% accuracy improvement, requires models like Jina Embeddings v2); Semantic Chunking (split at meaning boundaries using embedding similarity drops); Agentic RAG (retrieval driven by agent reasoning loops); Graph RAG (knowledge graph-augmented retrieval); Contextual Retrieval; Multimodal RAG (images, tables, mixed media). The verdict: semantic chunking is the practical best for now; late chunking is the future.

// What do real RAG debugging and scaling scenarios look like?

A SaaS company has a 50-page API documentation PDF. Their RAG chatbot returns incomplete answers to authentication questions.

Diagnose with Fixed Size Chunking as the likely culprit — the OAuth2 section is being split across chunk 47 and chunk 48, so neither chunk contains the complete authentication flow. Switch to Semantic Chunking so the OAuth2 section stays together as one chunk with complete context. Re-embed with the same embedding model used at indexing. Add metadata filter for topic='authentication' on the retriever. Re-test retrieval in isolation — the query 'How do I authenticate with OAuth2?' should now return one complete, relevant chunk.

A legal tech firm is building a RAG system over 10,000 legal documents and needs high accuracy but is getting hallucinated case citations.

Apply Semantic Chunking (quality is critical for legal documents — accuracy > speed). Use a high-dimension embedding model (text-embedding-3-large at 3,072 dims) to capture nuanced legal language. Ground the LLM with the prompt pattern: 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know.' Use format_docs_with_source so every answer includes the case document and page number. Test retrieval separately — verify that retrieved chunks actually contain the relevant legal clauses before evaluating LLM output quality.

A developer prototyped a RAG system on 10 documents using a fast embedding model, then added 10,000 documents and retrieval quality collapsed.

Check for the three production rules violations: (1) Same embedding model? Confirm the model used at indexing matches the model used at query time — even version differences break comparability. (2) Chunking strategy? If using Fixed Size Chunking, migrate to Recursive Chunking with 500–1,000 token chunks and 50–200 token overlap as the production default. (3) Retrieval test: run the retriever in isolation on 10 representative queries and inspect raw chunk output before blaming the LLM. Apply metadata filtering to reduce noise at scale.

// What are the most common RAG pipeline mistakes to avoid?

  • Using Fixed Size Chunking in production — it cuts at arbitrary character counts, destroys meaning, and causes retrieval failure before the LLM is even involved. Never use this in production. Ever.
  • Using different embedding models at indexing time vs. query time — vectors become incomparable and similarity search silently fails.
  • Blaming the LLM for bad answers without first testing retrieval in isolation — 90% of RAG failures are retrieval failures, not generation failures.
  • Skipping the grounding prompt pattern — without explicitly instructing the LLM to answer only from context and say 'I don't know' when unsure, hallucination is guaranteed.
  • Omitting source attribution — without sources users cannot verify answers, trust degrades, and the system is not production-ready.
  • Chunking code by character count — splitting a for-loop across two chunks makes both chunks semantically useless. Use the Code Splitter and keep functions together.
  • Treating chunk size as a fixed universal constant — the right chunk size depends on content type, retrieval use case, and token budget. Technical and legal content often needs semantic chunking with auto-determined sizes.
  • Using the add method instead of upsert when populating the vector store — repeated runs will duplicate all documents, corrupting retrieval results.
  • Ignoring overlap between chunks — without overlap, context at chunk boundaries is permanently lost and adjacent concepts become disconnected.
  • Mixing up embedding models and chat models — embedding models output vectors (not text) and are a completely different model type from chat models like GPT-4o or Claude, even from the same provider.

// What key RAG terms should you know?

RAG (Retrieval Augmented Generation)
An architecture that grounds LLM responses in actual retrieved documents, reducing hallucination by augmenting the model's knowledge with relevant external chunks at query time.
Chunking
The process of splitting documents into smaller pieces before embedding. The single biggest lever over RAG quality — described by the creator as architecture, not pre-processing.
Fixed Size Chunking
Splitting every n characters or tokens regardless of sentence or meaning boundaries. Destroys semantic meaning, causes fragment loss, and should never be used in production.
Recursive Chunking
The reliable default chunking strategy used by LangChain. Uses a split hierarchy — paragraphs → sentences → clauses → words → characters — always preferring natural boundaries over arbitrary cuts.
Semantic Chunking
The premium chunking strategy that splits at meaning boundaries by embedding each sentence and splitting where adjacent embedding similarity drops. Best for legal, technical, and high-value RAG where accuracy > speed.
Late Chunking
An advanced approach that embeds the full document first, then pools embeddings for individual chunks — preserving cross-chunk context. Yields ~10–12% accuracy improvement. Requires models like Jina Embeddings v2. Called 'the future of chunking' by the creator.
Embedding Model
A model that takes text in and outputs a vector (list of numbers representing semantic meaning). Completely different from a chat model. Must be the same model at both indexing and querying phases.
Dimensions
The size of the vector output by an embedding model (e.g., 1,536 for text-embedding-3-small). More dimensions = more semantic nuance captured but more storage and slower search. Sweet spot: 768–1,536.
Indexing Phase
The one-time process of loading documents, chunking them, embedding each chunk, and storing vectors in the vector store. The foundation of any RAG system.
Querying Phase
The per-question process of embedding the user query using the same embedding model, searching the vector store for similar vectors, retrieving original text chunks, and passing them with the query to the LLM.
RunnablePassthrough
A LangChain construct that passes the user's question through the RAG chain unchanged — ensuring the original query reaches the LLM without modification.
format_docs_with_source
A function that adds source tags to each retrieved chunk so the final answer includes citation metadata (document name, page, etc.), enabling user verification.
Overlap
A configurable number of shared tokens between adjacent chunks that preserves context at chunk boundaries and prevents meaning loss at split points.
Grounding Prompt Pattern
The prompt engineering pattern: 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know.' The low-hanging fruit for preventing hallucination that most people ignore.
Five Main Failure Modes
The five reasons 90% of RAG systems fail in production: bad chunking, mismatched embedding models, missing grounding prompt, no metadata filtering, and no source attribution.
Distance Score vs. Similarity Score
Distance scores: closer to 0 = most relevant (used by Chroma by default). Similarity scores: closer to 1 = most relevant. Convert via: similarity = 1 / (1 + distance). Know which your vector store returns.
Agentic RAG
An advanced RAG architecture where an agent reasoning loop drives retrieval decisions dynamically, rather than a fixed retrieval step.
Graph RAG
A RAG architecture that augments retrieval with a knowledge graph to capture entity relationships and multi-hop reasoning across documents.
Multimodal RAG
A RAG architecture that handles mixed content types including images, tables, and text in the same retrieval pipeline.

// FREQUENTLY ASKED QUESTIONS

What is a production RAG pipeline?

A production RAG pipeline is a Retrieval Augmented Generation system engineered to work reliably on 10,000+ documents rather than just a demo of 10. It combines correct document loading, deliberate chunking, a locked embedding model, a populated vector store, grounded prompting, metadata filtering, and source attribution — all designed to keep retrieval quality high at scale.

What is chunking in RAG and why does it matter so much?

Chunking is splitting documents into smaller pieces before embedding, and it's the single biggest lever over RAG quality. Bad chunking produces bad embeddings, which produces broken retrieval and wrong answers. The creator calls chunking architecture, not pre-processing — a chunk decision ripples through what gets retrieved, what context the LLM sees, and what the user receives.

How do I improve poor RAG retrieval quality?

Test retrieval in isolation first, since 90% of RAG failures are retrieval failures, not generation. Run queries against the retriever alone and inspect the raw chunks returned. If they're fragments or irrelevant, fix chunking (switch from fixed-size to recursive or semantic) and confirm the same embedding model is used at indexing and query time before touching the LLM prompt.

How do I stop my RAG system from hallucinating?

Use the grounding prompt pattern: 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know.' This low-hanging fruit prevents most hallucination. Also attach source tags to every chunk with format_docs_with_source so answers are verifiable, and ensure retrieval actually returns relevant chunks before blaming the model.

How does recursive chunking compare to semantic chunking?

Recursive chunking gets you 80% of retrieval quality by splitting along a paragraph→sentence→clause hierarchy; semantic chunking gets the last 20% by splitting at meaning boundaries but costs more compute. Start with recursive (500–1,000 tokens, 50–200 overlap), measure retrieval quality, and upgrade to semantic only for high-value content like legal or technical manuals where accuracy beats speed.

When should I use semantic chunking instead of recursive?

Use semantic chunking when accuracy is critical — legal documents, technical manuals, and knowledge bases where a shattered concept produces wrong answers. Semantic chunking keeps meaning-related content together by splitting where embedding similarity drops. For quick prototypes or simple structured documents, recursive chunking is the reliable, cheaper default.

What results can I expect after applying this framework?

You can expect a RAG system that scales from 10 to 10,000+ documents without retrieval collapse, dramatically fewer hallucinations, verifiable answers with source citations, and a debuggable pipeline where you can isolate whether failures are retrieval or generation. Migrating to semantic chunking or late chunking can add roughly 10–12% accuracy on demanding corpora.

Why does my RAG system break when I add more documents?

Scaling failures usually trace to three violations: fixed-size chunking that shatters meaning, a mismatched embedding model between indexing and query time, or no metadata filtering to reduce noise. Better embeddings on fewer documents beat worse embeddings on more. Fix chunking, lock a single embedding model, add metadata filters, and re-test retrieval in isolation.

What is the difference between an embedding model and a chat model in RAG?

An embedding model takes text in and outputs a vector (a list of numbers representing meaning) used for similarity search; a chat model like GPT-4o or Claude generates text answers. They're completely different model types even from the same provider. RAG needs both: the embedding model for indexing and retrieval, the chat model for generation.

Why must I use the same embedding model for indexing and querying?

Because vectors from different models — or even different versions of the same model — aren't comparable, so similarity search silently fails and returns irrelevant chunks. Lock your embedding model before indexing and never change it without re-indexing your entire corpus. This is one of the five main RAG failure modes.

How do I add source citations to RAG answers?

Use format_docs_with_source to tag every retrieved chunk with its source document name and page before it enters the prompt. This lets users verify answers, enables citation, and builds trust. Source attribution is essential in production, not optional — omitting it is one of the five failure modes that makes a system not production-ready.

// GET THIS SKILL — FREE

Use this skill in your AI

Every skill on SkillForge is free. Drop your email and copy this skill straight into Claude, ChatGPT, or any LLM.

We'll email you when new skills drop. Unsubscribe anytime.