What you'll learn
A single chain of thought reasons in a straight line, using only what it already knows. The frontier techniques go further: ReAct lets a model act on the world mid-reasoning, reflection lets it critique and fix its own work, and tree-of-thought lets it explore several paths and keep the best. These ideas power modern AI agents.
By the end of this module you'll be able to:
- Explain why a lone chain of thought hits a ceiling
- Describe the ReAct loop of thought, action, and observation
- Use reflection to have a model catch its own mistakes
- Pick the right reasoning technique for a given problem
The limits of a single chain
Chain-of-thought is powerful, but it's a one-shot straight line: the model commits to a path and can't look anything up, and if it goes wrong early, the whole chain inherits the error. Real problems often need information the model doesn't have, or a chance to reconsider, or a way to compare alternatives. Each advanced technique adds one of those.
ReAct: reason + act
ReAct (Reason + Act) interleaves thinking with actions that use tools — a search, a calculation, a database lookup. The model writes a Thought, takes an Action, reads the Observation that comes back, and loops — grounding each step in real results instead of guessing.
Thought
Reason about what to do next.
Action
Use a tool (search, calculate, look up).
Observation
Read the tool's result.
↻ repeat until confident, then answer
Prompt
Who is the current CEO of the company that makes the Kindle, and what year was that company founded? Use search.
AI response
Action: search("current CEO of Amazon")
Observation: Andy Jassy (CEO since 2021).
Action: search("year Amazon founded")
Observation: 1994.
Answer: Amazon makes the Kindle; its CEO is Andy Jassy, and it was founded in 1994.
Key idea
Reflection & self-critique
Models are often better critics than first-draft authors. Reflection exploits this: after producing an answer, the model is prompted to critique it against the requirements and then revise. A simple "draft → critique → improve" loop can noticeably raise quality.
Tip
Tree-of-thought exploration
Where chain-of-thought follows one path, tree-of-thought branches into several candidate approaches, evaluates which look promising, and expands the best — pruning dead ends, a bit like how you'd tackle a puzzle. It's the most expensive technique here, so save it for genuinely hard problems with many possible strategies.
promising ✓
dead end ✗
dead end ✗
Choosing the right technique
Match the technique to what the problem actually lacks — and don't over-engineer:
| If the problem needs… | Reach for | Cost |
|---|---|---|
| Step-by-step reasoning it already knows | Chain-of-thought (Module 10) | Low |
| Reliability on a hard, checkable answer | Self-consistency (Module 12) | Medium |
| External facts or live data | ReAct + tools | Medium |
| Higher quality via revision | Reflection / self-critique | Medium |
| Exploring many possible strategies | Tree-of-thought | High |
Watch out
Recap & quick check
Key takeaways
- A single chain of thought can't look things up, reconsider, or compare alternatives — its main limits.
- ReAct interleaves Thought → Action (tool) → Observation, grounding reasoning in real results; it's the core of agents.
- Reflection has the model critique its own draft against the requirements, then revise — models make good critics.
- Tree-of-thought explores multiple approaches and expands the best; it's powerful but the most expensive.
- Pick the cheapest technique that fits the problem's actual gap; don't over-engineer.
Quick check
1. What does the ReAct loop consist of?
2. Why is ReAct able to answer questions about very recent events?
3. What is the core idea of reflection?
4. When is tree-of-thought the right tool?
That completes the reasoning phase. Next we make the model's output not just smart but structured and dependable — starting with clean, machine-readable JSON. Next up: Phase 4, Module 14 — Structured Output: JSON & Schemas.