Naji RAG from Scratch Architecture
Build a fully functional Retrieval-Augmented Generation (RAG) system using LangChain that answers questions from private or recent documents an LLM was never trained on.
// TL;DR
Naji RAG from Scratch Architecture is a step-by-step method for building a Retrieval-Augmented Generation system with LangChain that lets an LLM answer questions from private or recent documents it was never trained on. Use it when you need to query PDFs, internal wikis, HR policies, or reports too large to fit in a prompt. It has two pipelines: Data Injection (load, chunk, embed, store) and Retrieval (embed query, similarity-search, augment prompt, generate). Both must use the same embedding model. It's the go-to approach for document Q&A over post-training or confidential knowledge.
// When should you use a RAG system?
Use this skill whenever a user needs to query private documents, internal databases, or post-training knowledge that an LLM cannot answer from its weights alone. Also applicable when building a question-answering system over large documents (e.g. PDFs, reports) where passing the entire document as context is impractical.
// What do you need to build a RAG system?
- Source Document(s)required
The raw private or domain-specific documents to build the knowledge base from (e.g. PDF annual reports, internal wikis, policy files). - Embedding Modelrequired
The model used to convert text to numbers (e.g. OpenAI text-embedding-3-small). Must be the same model used in both the injection step and the retrieval step. - LLM Modelrequired
The language model used to generate the final answer from retrieved context (e.g. GPT-4o-mini). - Vector Databaserequired
The vector store where embeddings will be saved and searched (e.g. Qdrant, Chroma, PGVector, or in-memory store). - User Queryrequired
The natural language question the user wants answered from the document. - Chunk Size and Overlap
Parameters controlling how the document is split. Defaults: chunk_size=1000 characters, chunk_overlap=100 characters.
// What core principles make RAG work?
Human Analogy of Document Search
RAG mirrors how a human searches a large document: rather than reading 2,000 pages, a human scans for keywords, retrieves relevant pages, and synthesises an answer. RAG follows the same process programmatically — search, retrieve the specific portion, then answer.
Context is Everything
LLMs have a knowledge cutoff and no access to private data. Without injecting relevant context into the prompt, the LLM will say 'I don't know'. RAG solves this by building a knowledge base and augmenting every query with retrieved context before the LLM generates a response.
Embeddings as the Translator
Humans understand natural language; LLMs understand only numbers. Embedding models are the translators — they convert words and sentences into vectors (numbers). Critically, similar words or similar sentences land in the same region or namespace of the vector space.
Same Namespace = Semantic Similarity
Because similar meanings cluster in the same namespace in vector space, similarity search (e.g. cosine similarity) can find chunks that are semantically related to a query — even if the exact words differ. This is the core mechanism that makes retrieval work.
Chunk Overlap for Context Continuity
When splitting a document into chunks, preserving a chunk_overlap (e.g. 100 characters) between adjacent chunks prevents context from being severed at chunk boundaries. This gives the LLM enough surrounding context to produce coherent answers.
Two-Step RAG Architecture
RAG always has exactly two steps: the Data Injection step (building the knowledge base) and the Retrieval step (querying and generating). These are distinct pipelines and must use the same embedding model to be compatible.
// How do you build a RAG system step by step?
- 1
Load the raw document(s) using a Document Loader
Use an appropriate loader for the file type (e.g. PyMuPDF PDF Loader for PDFs in LangChain). The loader reads the file and splits it into pages or sections. Check the number of pages loaded to confirm success.
- 2
Chunk the document using a Recursive Character Text Splitter
Set chunk_size (e.g. 1000 characters) and chunk_overlap (e.g. 100 characters). The splitter first tries to split on paragraph breaks (\n\n), then newlines (\n), then spaces, then characters — recursively — until each chunk is under the chunk_size limit. This is why it is called a Recursive Character Splitter. Inspect a few chunks and their neighbours to verify overlap is working correctly.
- 3
Instantiate an Embedding Model
Choose a cost-effective embedding model (e.g. OpenAI text-embedding-3-small). This is the translator that converts natural language into numbers. The same embedding model MUST be used in both the injection step and the retrieval step — mixing models will break similarity search.
- 4
Create a Collection in a Vector Database and inject the chunks
Instantiate the vector database client (e.g. local Qdrant via Docker). Create a collection, specifying: (a) the vector size — obtained by embedding a test string and checking its dimension, and (b) the distance metric (cosine similarity recommended). Then call add_documents() with the final chunks. Verify the collection is populated before proceeding. Note: collection name must be consistent across injection and retrieval code.
- 5
Convert the Vector Store into a Retriever
A vector store is a database; a retriever is a LangChain runnable. Convert using .as_retriever(). Runnables are required to build LangChain chains. The retriever will accept a query and return the top-k most similar chunks using cosine similarity search.
- 6
Build a RAG Chain with a prompt, retriever, LLM, and output parser
Construct the prompt template with two variables: {question} and {context}. Include instructions like 'If you don't know the answer, just say you don't know' and 'use three sentences maximum to keep the answer concise'. Build the chain: RunnablePassthrough (passes question unchanged) → retriever → format_documents function (extracts only page_content, dropping metadata to reduce context size) → prompt → LLM → output parser. The format_documents step is important — stripping metadata reduces token usage.
- 7
Invoke the RAG Chain with the user query and return the generated answer
Call chain.invoke(question). The query is automatically embedded, similarity-searched against the vector store, top chunks are retrieved, formatted, injected into the prompt as context, and the LLM generates the answer. Validate the answer against the source document to confirm correctness.
// What are real examples of RAG in action?
A company wants to query a 364-page annual financial report for specific metrics like revenue and net income for a given year, without manually reading the document.
Load the PDF with a PDF Loader → chunk into 1000-character pieces with 100-character overlap using Recursive Character Splitter → embed each chunk with text-embedding-3-small → store in a Qdrant collection with cosine similarity → convert to a retriever → build a RAG chain → ask 'What are the financial highlights of [year]?' The retriever fetches the relevant financial summary chunks; the LLM synthesises the answer with specific figures drawn only from those chunks.
A user needs to query internal HR policy documents that were never publicly available and therefore not in any LLM's training data.
The private HR documents are treated as the knowledge base — the LLM has zero prior knowledge of them. Inject the documents into a vector store. At query time, the user's question is embedded, relevant policy sections are retrieved via similarity search, and those sections are passed as context to the LLM. The LLM answers using only the retrieved context, not its weights.
// What mistakes should you avoid when building RAG?
- Using different embedding models in the Data Injection step and the Retrieval step — this breaks similarity search because vectors are not in the same namespace.
- Passing the entire document to the LLM without chunking — this exceeds context windows, causes hallucination, and is cost-prohibitive.
- Setting chunk_overlap to zero — adjacent chunks lose continuity, cutting off context at boundaries and degrading answer quality.
- Not verifying the vector database collection is populated before running retrieval — a misconfigured collection name will cause the retriever to search an empty store and return no results.
- Including full metadata in the context passed to the LLM — only page_content should be passed; metadata inflates token usage and dilutes the relevant context.
- Forgetting that the LLM only generates answers as good as the retrieved chunks — if the chunk_size is too large or too small, retrieval quality degrades and answers suffer.
// What are the key RAG terms you need to know?
- RAG (Retrieval-Augmented Generation)
- A three-step architecture: Retrieve (get relevant data using exact or similarity search), Augment (package the retrieved data and add it as context to the prompt), Generate (the LLM produces an answer using that context).
- Data Injection Step
- The first of two RAG pipeline steps. Takes raw documents, cleans them, chunks them, converts chunks to embeddings, and stores them in a vector database to build the knowledge base.
- Retrieval Step
- The second of two RAG pipeline steps. Converts the user query into an embedding, runs similarity search against the vector database, retrieves the most relevant chunks, injects them as context into a prompt, and passes the prompt to the LLM to generate an answer.
- Embeddings
- The translation of natural language into numbers (vectors) so LLMs can process it. Similar words or sentences are placed in the same region or namespace of the vector space.
- Namespace (in vector space)
- The region of the vector space where semantically similar words or sentences cluster together. The foundation of similarity search.
- Recursive Character Text Splitter
- A LangChain splitter that chunks documents by recursively trying separators in order (paragraph → newline → space → character) until each chunk is within the specified chunk_size. Named 'recursive' because it repeats the splitting logic at finer granularity.
- Chunk Overlap
- A configurable number of characters shared between adjacent chunks to preserve context continuity at chunk boundaries, preventing the LLM from losing meaning that spans two chunks.
- Vector Database
- A specialised database designed to store and search embeddings (vectors). Examples: Qdrant, Chroma, PGVector, in-memory vector store.
- Collection
- The unit of storage within a vector database (analogous to a table in SQL or a collection in MongoDB). Requires a vector size and a distance metric (e.g. cosine similarity) at creation time.
- Retriever
- A LangChain runnable wrapper around a vector store that accepts a query and returns the top-k most similar chunks. Required for composing LangChain chains.
- Knowledge Base
- The structured store of document chunks (as embeddings in a vector database) that the RAG system searches to answer questions. Analogous to providing a human with contextual reference material before asking them a question.
- Cosine Similarity
- The distance metric used to measure how similar two vectors (embeddings) are. Used by the retriever to rank and return the most relevant chunks from the vector database.
- format_documents
- A helper function in the RAG chain that strips metadata from retrieved chunks, passing only page_content to the LLM prompt — reducing token usage and focusing the LLM on the actual text.
- RunnablePassthrough
- A LangChain component in the RAG chain that passes the user's question to the LLM prompt unchanged, without any transformation.
// FREQUENTLY ASKED QUESTIONS
What is RAG and how does it work?
RAG (Retrieval-Augmented Generation) is a three-step architecture — Retrieve, Augment, Generate — that lets an LLM answer questions from data it wasn't trained on. It searches a knowledge base for relevant document chunks, packages them as context in the prompt, and the LLM generates an answer using only that retrieved context instead of its internal weights.
What is a vector database and why does RAG need one?
A vector database stores and searches embeddings — numerical representations of text. RAG needs one because it enables similarity search: converting a user query into a vector and finding semantically related chunks even when exact words differ. Examples include Qdrant, Chroma, PGVector, and in-memory stores. Each collection requires a vector size and a distance metric like cosine similarity.
How do I build a RAG system with LangChain?
Load your document with a Document Loader, chunk it using a Recursive Character Text Splitter (chunk_size 1000, overlap 100), embed each chunk with an embedding model, and store the vectors in a vector database collection. Then convert the store into a retriever, build a chain with a prompt, LLM, and output parser, and call chain.invoke(question) to get answers grounded in your documents.
How do I query a private PDF that an LLM was never trained on?
Treat the PDF as your knowledge base. Inject it into a vector store by loading, chunking, and embedding it. At query time, your question is embedded, relevant sections are retrieved via similarity search, and those sections are passed as context to the LLM. The model answers using only the retrieved chunks — not its training weights — so it can handle documents it never saw.
How does RAG compare to just pasting the whole document into the prompt?
RAG retrieves only the relevant chunks, while pasting the full document dumps everything into context. Pasting exceeds context windows for large files, inflates cost, dilutes attention, and increases hallucination. RAG scales to thousands of pages by retrieving the small, semantically relevant portion needed for each specific question — mirroring how a human scans for keywords instead of re-reading an entire report.
When should I use RAG instead of fine-tuning?
Use RAG when knowledge changes often, is private, or is post-training — like internal wikis, recent reports, or HR policies. RAG injects fresh context at query time without retraining. Fine-tuning bakes patterns into weights and is better for changing style or behavior, not for injecting large or frequently updated facts. For document Q&A over private data, RAG is faster, cheaper, and easier to update.
What are embeddings in RAG?
Embeddings are the translation of natural language into numbers (vectors) so LLMs and vector databases can process meaning mathematically. An embedding model converts words and sentences into vectors, placing semantically similar text in the same region — or namespace — of vector space. This clustering is what makes similarity search work, letting retrieval find related chunks even when the wording differs.
Why is chunk overlap important in RAG?
Chunk overlap preserves context continuity at chunk boundaries. When a document is split, a concept can be severed across two adjacent chunks. Keeping an overlap (e.g. 100 characters) ensures each chunk carries enough surrounding text so the LLM doesn't lose meaning that spans the split. Setting overlap to zero degrades answer quality by cutting off context mid-thought.
What results can I expect from a RAG system?
You get accurate, source-grounded answers to natural-language questions over documents an LLM couldn't otherwise handle — like extracting revenue from a 364-page annual report or answering HR policy questions. Quality depends on retrieval: if chunks are relevant, answers are precise with specific figures. If chunk size or embedding setup is wrong, answers degrade. Well-tuned RAG reduces hallucination by grounding responses in retrieved text.
Why does RAG use the same embedding model for injection and retrieval?
Both steps must use the same embedding model so all vectors live in the same namespace of vector space. Similarity search compares the query vector against stored chunk vectors; mixing models produces incompatible vectors, breaking the comparison and returning irrelevant results. This is one of the most common RAG bugs — always match the embedding model across your Data Injection and Retrieval pipelines.
What is a retriever in LangChain?
A retriever is a LangChain runnable that wraps a vector store, accepts a query, and returns the top-k most similar chunks using cosine similarity search. You create it by calling .as_retriever() on your vector store. Runnables are required to compose LangChain chains, so converting the store into a retriever is what lets you plug retrieval into your RAG pipeline.
How do I reduce token costs in a RAG chain?
Strip metadata from retrieved chunks before passing them to the LLM — use a format_documents helper that extracts only page_content. Metadata inflates token usage and dilutes relevant context. Also choose a cost-effective embedding model like text-embedding-3-small, tune chunk_size so retrieval returns only what's needed, and instruct the prompt to keep answers concise (e.g. three sentences maximum).