Phase 5 · Grounding & KnowledgeModule 19~36 min read

Retrieval-Augmented Generation (RAG)

Give the model an open book. RAG retrieves relevant documents and drops them into the prompt so answers are grounded in your own, up-to-date data.

What you'll learn

Retrieval-Augmented Generation (RAG) is grounding at scale. Instead of hoping a fact lives in the model's memory, you look it up in your own documents and paste the relevant passages into the prompt. It's the dominant pattern for building AI over private, current, or large knowledge bases.

By the end of this module you'll be able to:

  • Explain RAG as giving the model an open book
  • Describe how embeddings power semantic search
  • Walk through the retrieve → augment → generate pipeline
  • Know why RAG often beats fine-tuning for knowledge tasks

The open-book idea

A closed-book exam tests memory; an open-book exam lets you look things up. RAG turns every query into an open-book question: it finds the passages most likely to contain the answer and puts them right in front of the model. The model then does what it's great at — reading and synthesizing — instead of what it's shaky at: recalling exact facts.

The RAG pipeline

1 · Question

The user's query comes in.

2 · Retrieve

Search your documents for the most relevant chunks.

3 · Augment

Paste those chunks into the prompt as context.

4 · Generate

The model answers using the retrieved facts.

Retrieve the relevant chunks, augment the prompt with them, then generate a grounded answer.

Embeddings & semantic search

How does step 2 find "relevant" text? With embeddings — numerical vectors that capture meaning, so that passages about similar topics sit close together in vector space. You embed the user's question, then search a vector database for the nearest chunks. Unlike keyword search, this matches on meaning: a query about "getting my money back" can retrieve a passage titled "Refund policy."

Note

Embeddings are the same "meaning has geometry" idea behind word vectors. If you've taken the AI course, this is the applied payoff; if not, just picture related ideas clustering together in space.

Retrieve, then generate

Put together, the pipeline is short — most of RAG's quality comes from doing each step well, not from complex code:

rag.py
# Retrieval-Augmented Generation, in miniature.
query = "What is our refund window?"

# 1. embed the query and find the closest document chunks
q_vec  = embed(query)
chunks = vector_db.search(q_vec, top_k=3)     # semantic search

# 2. build a grounded prompt from the retrieved text
context = "\n\n".join(chunks)
prompt = f'''Answer using ONLY the context. If unsupported, say "Not found."

Context:
{context}

Question: {query}'''

# 3. generate
answer = call_model(prompt)
Embed the query, retrieve the top chunks, drop them into a grounded prompt, and generate. The prompt is exactly the grounding pattern from Module 18.

Key idea

RAG is prompt engineering plus search. The retrieval finds the facts; the grounded prompt — "use only this context, cite it, say if it's missing" — is what makes the answer trustworthy. Weak retrieval or a weak prompt each sink the result.

Chunking your documents

You don't embed whole documents — you split them into chunks (a few paragraphs each) so retrieval can return just the relevant part. Chunking is a real lever on quality:

  • Too large and each chunk is mostly irrelevant filler that wastes context and dilutes the answer.
  • Too small and you slice sentences apart, losing the context that makes them meaningful.
  • Overlap a little between chunks so an answer that straddles a boundary isn't cut in half.

RAG vs. fine-tuning

A common question: why not just fine-tune the model on your data? For knowledge that changes or must be cited, RAG usually wins:

ConcernRAGFine-tuning
Updating knowledgeEdit the documents — instantRetrain the model
Citing sourcesNatural — you have the chunksHard — knowledge is baked in
Up-front costLow (build an index)Higher (training run)
Best forFacts, docs, changing infoStyle, format, specialized behavior
Rule of thumb: RAG for knowledge, fine-tuning for behavior. They can also be combined.

Recap & quick check

Key takeaways

  • RAG is grounding at scale: retrieve relevant passages from your documents and paste them into the prompt.
  • Embeddings turn text into meaning-vectors, so semantic search finds relevant chunks even without keyword overlap.
  • The pipeline is retrieve → augment → generate; the grounded prompt is what makes the answer trustworthy.
  • Chunk documents thoughtfully (not too big or small, with a little overlap) — it strongly affects quality.
  • RAG beats fine-tuning for changing, citable knowledge; fine-tuning is for style and specialized behavior.

Quick check

1. What does RAG do?

2. Why use embeddings for retrieval instead of keyword search?

3. Why does chunk size matter?

4. When is RAG usually preferable to fine-tuning?

RAG keeps the prompt focused by retrieving only what matters. But modern models can also take in enormous prompts directly — so when should you? Next up: Module 20 — Working with Long Context & Documents.