Imagine asking an AI assistant, “What is our company’s current reimbursement limit for overseas hotels?”
The answer may live in an internal policy document that was updated last week. The base language model may never have seen it.
One approach is Retrieval-Augmented Generation (RAG): retrieve relevant information first, then give that information to the model while it generates the answer.
Think of an open-book exam
A closed-book student answers from memory.
An open-book student can first find the relevant pages, read them and then write an answer.
RAG follows a similar pattern. It does not necessarily change the model’s weights. It changes the context available at answer time.
A typical RAG pipeline
A simple document RAG system often does this:
- collect documents,
- split them into smaller chunks,
- create embeddings for those chunks,
- store them in a vector database,
- embed the user’s question,
- retrieve relevant chunks,
- place those chunks into the model context,
- ask the model to answer from that material.
The embeddings from Lesson 021 and vector database from Lesson 022 are common building blocks.
Why split documents into chunks?
A 300-page manual is usually too broad to retrieve as one item.
Smaller chunks make retrieval more precise, but chunks that are too small can lose context.
Chunk size, overlap and document structure are therefore design decisions, not meaningless preprocessing details.
A heading-aware chunker may work better for a policy manual than blindly cutting every 500 characters.
Retrieval quality limits answer quality
If the retrieval stage returns the wrong passages, the language model receives poor evidence.
A beautifully written answer can still be wrong because the right document was never placed into context.
This is why RAG evaluation should separate at least two questions:
- Did retrieval find the right evidence?
- Did generation use that evidence correctly?
RAG can reduce hallucination, not eliminate it
Providing evidence helps ground an answer, but the model can still misread, ignore or combine passages incorrectly.
Good systems may:
- instruct the model to say when evidence is insufficient,
- show citations or source links,
- limit answers to retrieved documents,
- rerank retrieved passages,
- use metadata filters and access controls.
RAG versus fine-tuning
RAG is useful when facts change often or must come from a private knowledge base.
Fine-tuning is more useful when you want to change recurring behavior, style, task performance or output format by updating the model parameters.
They can also be combined.
Lesson 024 explains that distinction next.
One thing to remember
RAG retrieves relevant external information at request time and gives it to the model as context before generation.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.