← Back home
LESSON 023AI Development10 min

What is RAG? Let the model look up relevant material before it answers

Retrieval-Augmented Generation gives an LLM relevant external context at answer time. Lesson 023 explains chunking, retrieval, grounding, citations and common failure modes.

Today’s analogyan open-book exam where the assistant first finds the right pages, then answers using them

Imagine asking an AI assistant, “What is our company’s current reimbursement limit for overseas hotels?”

The answer may live in an internal policy document that was updated last week. The base language model may never have seen it.

One approach is Retrieval-Augmented Generation (RAG): retrieve relevant information first, then give that information to the model while it generates the answer.

Think of an open-book exam

A closed-book student answers from memory.

An open-book student can first find the relevant pages, read them and then write an answer.

RAG follows a similar pattern. It does not necessarily change the model’s weights. It changes the context available at answer time.

A typical RAG pipeline

A simple document RAG system often does this:

  1. collect documents,
  2. split them into smaller chunks,
  3. create embeddings for those chunks,
  4. store them in a vector database,
  5. embed the user’s question,
  6. retrieve relevant chunks,
  7. place those chunks into the model context,
  8. ask the model to answer from that material.

The embeddings from Lesson 021 and vector database from Lesson 022 are common building blocks.

Why split documents into chunks?

A 300-page manual is usually too broad to retrieve as one item.

Smaller chunks make retrieval more precise, but chunks that are too small can lose context.

Chunk size, overlap and document structure are therefore design decisions, not meaningless preprocessing details.

A heading-aware chunker may work better for a policy manual than blindly cutting every 500 characters.

Retrieval quality limits answer quality

If the retrieval stage returns the wrong passages, the language model receives poor evidence.

A beautifully written answer can still be wrong because the right document was never placed into context.

This is why RAG evaluation should separate at least two questions:

RAG can reduce hallucination, not eliminate it

Providing evidence helps ground an answer, but the model can still misread, ignore or combine passages incorrectly.

Good systems may:

RAG versus fine-tuning

RAG is useful when facts change often or must come from a private knowledge base.

Fine-tuning is more useful when you want to change recurring behavior, style, task performance or output format by updating the model parameters.

They can also be combined.

Lesson 024 explains that distinction next.

One thing to remember

RAG retrieves relevant external information at request time and gives it to the model as context before generation.

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. OpenAI — Retrieval ↗
  2. Pinecone — Retrieval-Augmented Generation ↗
← Previous022What is a vector database? A search system built for finding nearby embeddings
Next →024What is fine-tuning? Continue training a model so a behavior becomes part of the model itself
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200