How to Query Annual Reports with RAG

For Financial and business analysts · Based on Naji RAG from Scratch Architecture

// TL;DR

This guide shows financial and business analysts how to use RAG to query long reports — like a 364-page annual report — in plain English and get source-grounded figures. Instead of manually scanning for revenue or net income, you inject the PDF into a vector database and ask 'What are the financial highlights of 2023?' The retriever fetches the relevant summary chunks and the LLM synthesizes an answer using only those figures. Use it when documents are too large to read manually or paste into a prompt, and when you need answers traceable back to the source.

Why is RAG better than reading a 300-page report manually?

RAG mirrors exactly how you already work: instead of reading every page, you scan for the relevant section, extract the figures, and summarize. RAG does this programmatically — it searches the report, retrieves only the chunks about your question, and the LLM synthesizes an answer from those chunks. For a 364-page annual report, this turns hours of scanning into a single natural-language question. And because the answer comes from retrieved text, you can verify every figure against the source.

How do you set up RAG for a financial report?

The process is a two-step pipeline. First, Data Injection: load the PDF with a document loader, split it into 1000-character chunks with 100-character overlap using a Recursive Character Text Splitter, embed each chunk with a model like text-embedding-3-small, and store the vectors in a Qdrant collection using cosine similarity. Second, Retrieval: your question is embedded, similarity search finds the relevant financial summary chunks, and the LLM answers using those figures.

The chunk overlap matters here — financial context often spans sentences (a figure and its qualifying footnote), and overlap keeps them together so you don't get a number without its context.

What kinds of questions can you actually ask?

You can ask targeted questions like 'What was net income in 2023?', 'What are the financial highlights of the year?', or 'How did revenue change year over year?' The retriever pulls the chunks containing those specific metrics, and the LLM returns the figures drawn only from the report — not from its training data. If the answer isn't in the document, a well-configured prompt instructs the model to say it doesn't know rather than guess, which protects you from fabricated numbers.

How do you make sure the numbers are trustworthy?

Trust comes from grounding and verification. Because RAG answers from retrieved chunks, always validate the figure against the source document — the retrieved text tells you exactly where it came from. Configure the prompt with guardrails: 'if you don't know, say you don't know' and 'keep answers concise.' Avoid the pitfall of oversized chunks, which can dilute retrieval so the right financial table isn't returned. If a figure looks off, inspect which chunks the retriever surfaced — the problem is almost always retrieval, not the model's math.

One more safeguard: never paste the whole report into a chatbot prompt. It exceeds context limits, drives up cost, and increases the chance of a hallucinated figure. RAG's selective retrieval is precisely what keeps large-document analysis both accurate and affordable.

What's the fastest way to get started?

You don't need to be a developer to benefit, but this workflow does involve setup. Partner with an engineer or use a no-code RAG tool built on this architecture. Start with one report, inject it, and ask a question you already know the answer to — verify the system returns the correct figure before trusting it on new questions.

Next step: Pick one annual report, run it through the inject-and-retrieve pipeline, and test with a known figure like prior-year revenue to confirm accuracy before rolling it out across your document library.

// FREQUENTLY ASKED QUESTIONS

Can RAG pull exact figures like revenue and net income?

Yes. When you ask for a specific metric, the retriever fetches the chunks containing that figure and the LLM reports it drawn only from those chunks. Always verify against the source, since accuracy depends on retrieval surfacing the right section. Keeping chunk overlap ensures figures stay attached to their qualifying context, like footnotes or period labels.

What happens if the report doesn't contain the answer?

With a properly configured prompt instructing the model to say 'I don't know' when unsure, RAG won't fabricate a figure it can't find in the retrieved chunks. This is a key advantage over asking a raw chatbot, which might guess. If you get 'I don't know' but believe the data exists, check whether retrieval returned the right chunks.

Is my confidential financial data safe with RAG?

It can be. Run a local vector database like Qdrant via Docker so document embeddings stay on your own infrastructure. The private report is treated purely as a knowledge base — the LLM only sees the chunks retrieved for each question. For sensitive filings, self-hosting the vector store keeps your data off third-party services.