How to Build a RAG System from Scratch with LangChain
For AI engineers at startups · Based on Naji RAG from Scratch Architecture
// TL;DR
This guide shows AI engineers how to build a RAG system from scratch with LangChain, covering the two core pipelines — Data Injection and Retrieval. You'll load documents, chunk them with a Recursive Character Text Splitter, embed with a consistent model, store vectors in Qdrant using cosine similarity, and compose a chain with RunnablePassthrough, a retriever, format_documents, a prompt, an LLM, and an output parser. Use it when you need to ship document Q&A over private or post-training data without fine-tuning, and want a maintainable, cost-controlled architecture you can extend to multiple documents and vector databases.
Why should AI engineers build RAG instead of fine-tuning?
RAG lets you inject private, large, or frequently changing knowledge into an LLM at query time without retraining. For a startup shipping fast, that's decisive: you can update the knowledge base by re-injecting documents instead of running an expensive fine-tune. RAG also keeps answers auditable — every response traces back to retrieved chunks, which matters when stakeholders ask 'where did this number come from?' Fine-tuning changes behavior and style; RAG delivers facts. For document Q&A over internal data, RAG wins on cost, speed to update, and traceability.
How do you structure the two RAG pipelines?
RAG always has exactly two steps, and separating them cleanly is the mark of a maintainable system. The Data Injection pipeline loads raw documents with the right loader (e.g. PyMuPDF for PDFs), chunks them with a Recursive Character Text Splitter at chunk_size 1000 and chunk_overlap 100, embeds each chunk, and stores the vectors in a collection. The Retrieval pipeline embeds the user query, runs cosine similarity search, retrieves top-k chunks, injects them into the prompt, and generates the answer.
The non-negotiable rule: both pipelines must use the same embedding model. Mixing models puts vectors in different namespaces and silently breaks similarity search. Encode this as a shared config constant so injection and retrieval can never drift.
How do you compose the RAG chain in LangChain?
Convert your vector store into a retriever with `.as_retriever()` — chains require runnables, not raw stores. Then compose:
1. `RunnablePassthrough` passes the question through unchanged.
2. The retriever fetches relevant chunks.
3. A `format_documents` function strips metadata, passing only `page_content` — this cuts token cost and sharpens the LLM's focus.
4. A prompt template with `{question}` and `{context}` injects both, plus guardrails like 'if you don't know, say so' and 'use three sentences maximum'.
5. The LLM generates.
6. An output parser returns clean text.
Call `chain.invoke(question)` and the whole flow — embed, search, retrieve, format, prompt, generate — runs automatically.
How do you avoid the common production failures?
The pitfalls that break RAG in production are predictable. Never pass the entire document to the LLM — it blows the context window, invites hallucination, and burns budget. Never set chunk_overlap to zero, or you'll sever context at boundaries. Always verify the collection is populated and the collection name matches across both pipelines — a typo means the retriever searches an empty store and returns nothing. Get the vector size right by embedding a test string and reading its dimension before creating the collection with cosine distance.
Remember: the LLM is only as good as the retrieved chunks. Most quality problems are retrieval problems. If answers are weak, inspect what the retriever returned before touching the generation prompt.
What should you build next?
Start with a single PDF and an in-memory or local Qdrant store to validate the loop end to end. Once retrieval quality is solid, extend to multiple documents in one collection, add source metadata for traceability (strip it before context), and tune top-k and chunk_size for your domain. Then wrap the chain behind an API endpoint.
Next step: Spin up a local Qdrant via Docker, load one document, and build the seven-step pipeline — confirm `chain.invoke()` returns a grounded answer before scaling.
// FREQUENTLY ASKED QUESTIONS
Which vector database should I start with as an engineer?
Start with a local Qdrant via Docker or an in-memory store for prototyping — both let you validate the pipeline without cloud dependencies. Qdrant is a strong default for production because it's fast and self-hostable, keeping private data on your infrastructure. Chroma and PGVector are also solid. The workflow is identical: create a collection with the correct vector size and cosine distance.
How do I keep injection and retrieval embedding models in sync?
Store the embedding model name in a single shared config constant imported by both pipelines, so they can never drift. This prevents the most common RAG bug — vectors landing in different namespaces and breaking similarity search. If you ever change the model, re-inject all documents so every vector uses the new model consistently.
How do I test retrieval quality independently of the LLM?
Call the retriever directly with test queries and inspect the returned chunks before running the full chain. If the relevant text isn't in those chunks, the LLM can't produce a good answer no matter how good the prompt is. This isolates retrieval problems from generation problems and speeds up debugging significantly.