What you'll learn
You don't need to train a model to build with one — you need to use it well. This practical module covers prompting, why models confidently make things up, and retrieval-augmented generation (RAG), the technique that grounds LLMs in real facts.
By the end of this module you'll be able to:
- Write clearer, more effective prompts
- Use few-shot and chain-of-thought prompting
- Explain why LLMs hallucinate
- Describe how RAG and agents extend an LLM
Prompting fundamentals
The prompt is your whole interface to the model, and small changes matter. The reliable habits: be specific about the task, give any needed context, state the desired format, and assign a role ("You are a careful editor…"). Vague in, vague out.
Few-shot & chain-of-thought
Two techniques punch above their weight. Few-shot prompting includes a couple of worked examples so the model copies the pattern. Chain-of-thought asks the model to "think step by step," which dramatically improves reasoning on math and logic by letting it work through intermediate steps instead of blurting an answer.
Why models hallucinate
LLMs sometimes state false things with total confidence — a hallucination. It's not lying: the model is a next-token predictor optimised for plausible-sounding text, with no built-in notion of truth and no live access to facts. If a fluent-but-wrong continuation is likely, it may produce it.
Watch out
Retrieval-augmented generation
The main cure for hallucination and stale knowledge is RAG: before answering, fetch relevant documents and put them in the prompt, so the model reasons over real retrieved text instead of its fuzzy memory. This is how "chat with your PDFs" and most company knowledge assistants work:
Question
“What is our refund policy?”
Retrieve
Search your documents (vector DB) for relevant passages
Augment
Paste those passages into the prompt as context
Generate
LLM answers using the retrieved facts — grounded & citable
Retrieval uses the embeddings from Module 22: your documents are chunked and embedded into a vector database, and the question's embedding finds the most similar chunks.
# Retrieval-Augmented Generation, in essence:
question = "What is our refund policy?"
# 1. embed the question and find similar chunks in your docs
chunks = vector_db.search(embed(question), top_k=3)
# 2. build a prompt that includes them
prompt = f"""Answer using ONLY the context below.
Context:
{chunks}
Question: {question}"""
# 3. let the LLM answer — grounded in your data
answer = llm.generate(prompt)Tools & agents
LLMs can also call tools: instead of guessing, the model can decide to run a calculator, search the web, query a database, or execute code, then use the result. Chain several such steps — plan, act, observe, repeat — and you have an agent: an LLM that pursues a goal by taking actions in the world. This is one of the most active frontiers in AI (Module 32).
Recap & quick check
Key takeaways
- The prompt is your interface: be specific, give context, state the format, set a role.
- Few-shot (examples) and chain-of-thought ('think step by step') sharply improve results.
- Hallucinations happen because LLMs predict plausible text, not verified truth.
- RAG retrieves relevant documents into the prompt to ground answers in real data.
- Tool use plus a plan-act-observe loop turns an LLM into an agent.
Quick check
1. Why do LLMs hallucinate?
2. What does RAG do?
3. Chain-of-thought prompting improves reasoning by…
4. An LLM that can call tools and loop plan-act-observe toward a goal is called…
Text is one frontier of generation; images are another, and they work in a beautifully different way. Next up: Module 29 — Generating Images: GANs & Diffusion.