What you'll learn
Retrieval-Augmented Generation (RAG) is grounding at scale. Instead of hoping a fact lives in the model's memory, you look it up in your own documents and paste the relevant passages into the prompt. It's the dominant pattern for building AI over private, current, or large knowledge bases.
By the end of this module you'll be able to:
- Explain RAG as giving the model an open book
- Describe how embeddings power semantic search
- Walk through the retrieve → augment → generate pipeline
- Know why RAG often beats fine-tuning for knowledge tasks
The open-book idea
A closed-book exam tests memory; an open-book exam lets you look things up. RAG turns every query into an open-book question: it finds the passages most likely to contain the answer and puts them right in front of the model. The model then does what it's great at — reading and synthesizing — instead of what it's shaky at: recalling exact facts.
1 · Question
The user's query comes in.
2 · Retrieve
Search your documents for the most relevant chunks.
3 · Augment
Paste those chunks into the prompt as context.
4 · Generate
The model answers using the retrieved facts.
Embeddings & semantic search
How does step 2 find "relevant" text? With embeddings — numerical vectors that capture meaning, so that passages about similar topics sit close together in vector space. You embed the user's question, then search a vector database for the nearest chunks. Unlike keyword search, this matches on meaning: a query about "getting my money back" can retrieve a passage titled "Refund policy."
Note
Retrieve, then generate
Put together, the pipeline is short — most of RAG's quality comes from doing each step well, not from complex code:
# Retrieval-Augmented Generation, in miniature.
query = "What is our refund window?"
# 1. embed the query and find the closest document chunks
q_vec = embed(query)
chunks = vector_db.search(q_vec, top_k=3) # semantic search
# 2. build a grounded prompt from the retrieved text
context = "\n\n".join(chunks)
prompt = f'''Answer using ONLY the context. If unsupported, say "Not found."
Context:
{context}
Question: {query}'''
# 3. generate
answer = call_model(prompt)Key idea
Chunking your documents
You don't embed whole documents — you split them into chunks (a few paragraphs each) so retrieval can return just the relevant part. Chunking is a real lever on quality:
- Too large and each chunk is mostly irrelevant filler that wastes context and dilutes the answer.
- Too small and you slice sentences apart, losing the context that makes them meaningful.
- Overlap a little between chunks so an answer that straddles a boundary isn't cut in half.
RAG vs. fine-tuning
A common question: why not just fine-tune the model on your data? For knowledge that changes or must be cited, RAG usually wins:
| Concern | RAG | Fine-tuning |
|---|---|---|
| Updating knowledge | Edit the documents — instant | Retrain the model |
| Citing sources | Natural — you have the chunks | Hard — knowledge is baked in |
| Up-front cost | Low (build an index) | Higher (training run) |
| Best for | Facts, docs, changing info | Style, format, specialized behavior |
Recap & quick check
Key takeaways
- RAG is grounding at scale: retrieve relevant passages from your documents and paste them into the prompt.
- Embeddings turn text into meaning-vectors, so semantic search finds relevant chunks even without keyword overlap.
- The pipeline is retrieve → augment → generate; the grounded prompt is what makes the answer trustworthy.
- Chunk documents thoughtfully (not too big or small, with a little overlap) — it strongly affects quality.
- RAG beats fine-tuning for changing, citable knowledge; fine-tuning is for style and specialized behavior.
Quick check
1. What does RAG do?
2. Why use embeddings for retrieval instead of keyword search?
3. Why does chunk size matter?
4. When is RAG usually preferable to fine-tuning?
RAG keeps the prompt focused by retrieving only what matters. But modern models can also take in enormous prompts directly — so when should you? Next up: Module 20 — Working with Long Context & Documents.