How to Build Accurate RAG for Legal Documents
For Legal tech engineers · Based on Production RAG Pipeline Engineering Framework
// TL;DR
Legal tech engineers building RAG over case law, contracts, or regulations should prioritize accuracy over speed using this framework. That means semantic chunking so clauses stay intact, a high-dimension embedding model like text-embedding-3-large to capture nuanced legal language, strict grounding prompts, and mandatory source attribution with document and page. Hallucinated case citations are almost always retrieval or grounding failures — test retrieval separately to confirm chunks actually contain the relevant clauses before blaming the model. In legal, unverifiable answers are unusable, so citations aren't optional.
Why is my legal RAG system hallucinating case citations?
Usually because the LLM isn't grounded and the retrieved chunks don't actually contain the clause. In high-stakes legal work, an ungrounded model will confidently invent citations. First, apply the grounding prompt pattern: 'Answer based only on the following context. If the context doesn't contain the answer, say I don't know.' Then test retrieval in isolation — run the query against the retriever alone and verify the returned chunks contain the relevant legal clauses. If they don't, the problem is retrieval, not generation, and 90% of the time that's exactly where the failure lives.
Why does chunking strategy matter more in legal tech?
Because a split clause is a wrong answer. Legal language is dense and interdependent — a definition, a condition, and its exception can span several sentences that must stay together. Fixed-size chunking cuts at arbitrary character counts and destroys this structure; never use it in production. For legal documents, quality is critical, so use semantic chunking, which splits at meaning boundaries by detecting where embedding similarity drops. This keeps a complete clause or holding in a single chunk. Semantic chunking costs more compute than recursive, but for legal accuracy the last 20% of retrieval quality is worth it.
Which embedding model should legal RAG use?
Reach for a high-dimension model to capture nuanced legal language — text-embedding-3-large at 3,072 dimensions is a strong choice, since more dimensions mean more semantic nuance. Whatever you pick, lock it before indexing 10,000 documents and never swap it without a full re-index; a model or even version mismatch between indexing and querying makes vectors incomparable and silently returns the wrong precedents. Remember the distinction: the embedding model produces vectors for retrieval, while your chat model (GPT-4o, Claude) generates the answer — two completely different model types.
How do I make legal answers verifiable and defensible?
Attach sources to everything. Use `format_docs_with_source` so every answer includes the case document name and page number. In legal work, an answer a lawyer can't trace to a source is worthless and potentially dangerous. Source attribution enables verification, supports citation in briefs, and builds the trust the domain demands. Combine this with metadata filtering — narrow retrieval by jurisdiction, date, or document type — so the system surfaces the right body of law and reduces noise across a large corpus.
What's the end-to-end build for legal RAG?
Load documents with UnstructuredPDFLoader for complex layouts with tables and rich metadata. Apply semantic chunking. Embed with text-embedding-3-large and populate the vector store with `upsert`. Build the chain with `RunnablePassthrough` and a strict grounding prompt. Configure the retriever with `search_type='similarity'`, a modest `k`, and metadata filters for jurisdiction and date. Attach source tags. Then systematically walk the five failure modes — bad chunking, mismatched embeddings, weak grounding, no filtering, no attribution — and remediate each before going live. Once stable, evaluate late chunking (~10–12% accuracy gain) for demanding matters.
Next step: Run 10 representative legal queries against your retriever in isolation and confirm every returned chunk contains a complete, relevant clause before you trust a single generated answer.
// FREQUENTLY ASKED QUESTIONS
Is semantic chunking worth the extra cost for legal documents?
Yes. Legal accuracy demands that clauses, conditions, and exceptions stay together, and semantic chunking splits at meaning boundaries to preserve that structure. Recursive chunking gets 80% of quality, but for legal work the last 20% — the difference between a complete holding and a fragment — is decisive. Prioritize accuracy over speed here.
How do I stop my legal RAG from inventing citations?
Combine a strict grounding prompt ('Answer based only on the following context... say I don't know') with source attribution via format_docs_with_source, then verify retrieval in isolation. Most invented citations come from an ungrounded model or from chunks that never contained the clause. Confirm retrieval is correct before evaluating generation.
Can I switch embedding models to save cost after indexing legal docs?
No — not without re-indexing your entire corpus. Vectors from a different model or version aren't comparable, so similarity search silently returns wrong precedents. Lock text-embedding-3-large (or your chosen model) before indexing and keep it consistent at query time. This is one of the five main RAG failure modes.