Frequently Asked Questions About Naji RAG from Scratch Architecture

21 answers covering everything from basics to advanced usage.

// Basics

What are the two steps of a RAG pipeline?

RAG always has exactly two distinct pipelines. The Data Injection step takes raw documents, cleans and chunks them, converts chunks to embeddings, and stores them in a vector database to build the knowledge base. The Retrieval step converts the user query into an embedding, runs similarity search, retrieves the most relevant chunks, injects them as context into a prompt, and passes it to the LLM to generate the answer.

What does 'namespace' mean in vector space?

A namespace is the region of vector space where semantically similar words or sentences cluster together. Because an embedding model places similar meanings near each other, a query vector lands close to chunk vectors about the same topic. Similarity search then measures that closeness — usually with cosine similarity — to return the most relevant chunks. This clustering is the foundation of how retrieval works.

What is the human analogy for how RAG works?

RAG mirrors how a human handles a 2,000-page document. Instead of reading every page, a human scans for keywords, retrieves the specific relevant pages, and synthesizes an answer. RAG does the same programmatically: it searches the knowledge base, retrieves the specific relevant portion, augments the prompt with it, and the LLM answers from that context.

Why is it called a Recursive Character Text Splitter?

It's called recursive because it repeatedly applies its splitting logic at finer granularity. It first tries to split on paragraph breaks, then newlines, then spaces, then individual characters — recursively — until each resulting chunk is under the chunk_size limit. This ordering keeps chunks as semantically intact as possible before falling back to cruder separators.

// How To

How do I set the correct vector size when creating a collection?

Get the vector size by embedding a test string with your chosen embedding model and checking the dimension of the resulting vector. Use that number as the collection's vector size, and set the distance metric to cosine similarity. If the vector size doesn't match your embedding model's output dimension, insertion will fail or search will misbehave.

How do I build the RAG chain in LangChain?

Build a prompt template with {question} and {context} variables, then compose the chain: RunnablePassthrough passes the question unchanged, the retriever fetches relevant chunks, a format_documents function strips metadata to page_content only, the prompt injects both variables, the LLM generates, and an output parser returns clean text. Include instructions like 'if you don't know, say so' and 'use three sentences maximum'.

How do I verify my vector database was populated correctly?

After calling add_documents(), inspect the collection to confirm it holds the expected number of vectors before running any retrieval. A common bug is a mismatched collection name between injection and retrieval code, which makes the retriever search an empty store and return nothing. Always confirm the collection is populated and the name is consistent across both pipelines.

How do I choose a chunk size and overlap?

Start with chunk_size 1000 characters and chunk_overlap 100 characters. Chunks too large dilute retrieval relevance and inflate tokens; chunks too small sever context and lose meaning. Inspect a few chunks and their neighbors to confirm overlap is working. Tune based on your document structure — dense reports may benefit from smaller chunks, narrative text from larger ones.

// Troubleshooting

Why does my RAG system return 'I don't know' or irrelevant answers?

Usually retrieval failed. Check that injection and retrieval use the same embedding model, that the collection name matches across both pipelines, and that the collection is actually populated. Also verify chunk_size isn't so large or small that relevant text isn't retrieved. The LLM only answers as well as the chunks it receives — bad retrieval means bad answers.

Why is my RAG system giving inconsistent or cut-off answers?

This often stems from chunk_overlap set to zero, which severs context at chunk boundaries so meaning spanning two chunks is lost. Add overlap (e.g. 100 characters) to preserve continuity. It can also happen if metadata is bloating the context — strip it with format_documents so only page_content reaches the LLM, keeping attention focused on the relevant text.

Why did my similarity search suddenly break after I changed something?

You likely changed the embedding model on only one side. If the Data Injection step used one model and the Retrieval step uses another, the vectors sit in different namespaces and comparisons return garbage. Re-inject all documents with the new model, or revert to the original. The embedding model must be identical across both pipelines.

My RAG costs are too high — what's causing it?

Common causes are passing full metadata into context, using oversized chunks, retrieving too many chunks (high top-k), or an expensive embedding or LLM model. Fix by stripping metadata with format_documents, tuning chunk_size, lowering top-k to only what's needed, using text-embedding-3-small, and instructing concise answers. Never dump the whole document into the prompt — that defeats RAG's purpose.

// Comparisons

How does RAG compare to fine-tuning an LLM?

RAG injects external knowledge at query time without touching model weights, making it ideal for private, large, or frequently changing data. Fine-tuning bakes patterns into weights and suits changing tone, format, or behavior — not injecting facts. For document Q&A, RAG is cheaper, faster to update, and easier to audit since answers trace back to retrieved chunks. Many production systems combine both.

How does RAG compare to a long context window model?

Long-context models can hold more text in a single prompt but still cost more as input grows, can suffer 'lost in the middle' attention issues, and can't scale to millions of pages. RAG retrieves only relevant chunks regardless of total corpus size, keeping cost and attention focused. For very large or continually growing knowledge bases, RAG remains more economical and scalable.

How does cosine similarity compare to other distance metrics?

Cosine similarity measures the angle between vectors, focusing on direction (meaning) rather than magnitude, which makes it robust for text embeddings — that's why it's the recommended default in this workflow. Alternatives like Euclidean distance factor in magnitude and can be sensitive to vector length. For semantic text retrieval, cosine similarity is generally the reliable choice when creating your collection.

What's the difference between a vector store and a retriever?

A vector store is the database that holds and searches embeddings. A retriever is a LangChain runnable wrapper around that store, created with .as_retriever(), which accepts a query and returns top-k similar chunks. The distinction matters because LangChain chains require runnables — you can't compose a chain directly from a raw vector store, so you convert it into a retriever first.

// Advanced

How do I handle multiple documents in one RAG knowledge base?

Load each document with the appropriate loader, chunk them all with the same splitter and settings, embed with one consistent model, and add them to the same collection. You can attach source metadata to each chunk for traceability, but strip it before passing context to the LLM. The retriever will search across all documents and return the most relevant chunks regardless of source.

Can I use a local vector database instead of a cloud service?

Yes. You can run Qdrant locally via Docker, or use Chroma, PGVector, or an in-memory store for prototyping. Local databases keep private data on your infrastructure — valuable for confidential HR or financial documents. The workflow is identical: create a collection with the correct vector size and cosine distance, then add your chunks and convert to a retriever.

How can I improve retrieval quality beyond default settings?

Tune chunk_size and overlap for your document type, experiment with top-k to balance recall and noise, and ensure metadata is stripped from context. Consider a better embedding model for nuanced domains, and validate answers against the source document. Because the LLM is only as good as the retrieved chunks, most quality gains come from improving retrieval, not the generation prompt.

What does the format_documents step do and why is it important?

format_documents is a helper in the RAG chain that strips metadata from retrieved chunks, passing only page_content into the prompt. It matters because metadata inflates token usage and dilutes the LLM's focus on the actual relevant text. This single step reduces cost and often improves answer quality by keeping the context clean and concentrated on meaningful content.

Can RAG answer questions requiring information spread across many chunks?

Yes, if retrieval surfaces all the necessary chunks. Increase top-k so more relevant sections are returned, and use chunk overlap to preserve continuity. For complex multi-part questions, the retriever fetches several chunks and the LLM synthesizes across them. If answers miss information, the cause is usually retrieval not returning enough relevant chunks rather than the LLM's reasoning.