Phase 3 · Making Models ReasonModule 13~36 min read

Advanced Reasoning: ReAct, Reflection & Tree-of-Thought

Beyond a single chain of thought: interleave reasoning with actions (ReAct), let a model critique and revise itself, and explore multiple reasoning paths.

What you'll learn

A single chain of thought reasons in a straight line, using only what it already knows. The frontier techniques go further: ReAct lets a model act on the world mid-reasoning, reflection lets it critique and fix its own work, and tree-of-thought lets it explore several paths and keep the best. These ideas power modern AI agents.

By the end of this module you'll be able to:

  • Explain why a lone chain of thought hits a ceiling
  • Describe the ReAct loop of thought, action, and observation
  • Use reflection to have a model catch its own mistakes
  • Pick the right reasoning technique for a given problem

The limits of a single chain

Chain-of-thought is powerful, but it's a one-shot straight line: the model commits to a path and can't look anything up, and if it goes wrong early, the whole chain inherits the error. Real problems often need information the model doesn't have, or a chance to reconsider, or a way to compare alternatives. Each advanced technique adds one of those.

ReAct: reason + act

ReAct (Reason + Act) interleaves thinking with actions that use tools — a search, a calculation, a database lookup. The model writes a Thought, takes an Action, reads the Observation that comes back, and loops — grounding each step in real results instead of guessing.

The ReAct loop

Thought

Reason about what to do next.

Action

Use a tool (search, calculate, look up).

Observation

Read the tool's result.

↻ repeat until confident, then answer

Reason, act with a tool, observe the result, repeat. This loop is the beating heart of an agent (Module 24).

Prompt

Who is the current CEO of the company that makes the Kindle, and what year was that company founded? Use search.

AI response

Thought: The Kindle is made by Amazon. I need Amazon's current CEO and founding year.
Action: search("current CEO of Amazon")
Observation: Andy Jassy (CEO since 2021).
Action: search("year Amazon founded")
Observation: 1994.
Answer: Amazon makes the Kindle; its CEO is Andy Jassy, and it was founded in 1994.
Each Action fetches a fact; each Observation grounds the next Thought. The model reasons with real data, not memory alone.

Key idea

ReAct is how a model overcomes its training cutoff and its shaky memory for specifics: it doesn't recall the fact, it looks it up. You'll build exactly this in Module 23 (Tools) and Module 24 (Agents).

Reflection & self-critique

Models are often better critics than first-draft authors. Reflection exploits this: after producing an answer, the model is prompted to critique it against the requirements and then revise. A simple "draft → critique → improve" loop can noticeably raise quality.

Tip

Make the critique concrete: "List every requirement the draft fails to meet, then rewrite to fix each one." A vague "make it better" produces a vague edit; a checklist produces real fixes. Add a verifier prompt (a separate check that the answer is correct) for high-stakes tasks.

Tree-of-thought exploration

Where chain-of-thought follows one path, tree-of-thought branches into several candidate approaches, evaluates which look promising, and expands the best — pruning dead ends, a bit like how you'd tackle a puzzle. It's the most expensive technique here, so save it for genuinely hard problems with many possible strategies.

Tree-of-thought
Problem
Path A
promising ✓
Path B
dead end ✗
Path C
dead end ✗
Solution
Explore several approaches, prune the dead ends, and expand the most promising path to the solution.

Choosing the right technique

Match the technique to what the problem actually lacks — and don't over-engineer:

If the problem needs…Reach forCost
Step-by-step reasoning it already knowsChain-of-thought (Module 10)Low
Reliability on a hard, checkable answerSelf-consistency (Module 12)Medium
External facts or live dataReAct + toolsMedium
Higher quality via revisionReflection / self-critiqueMedium
Exploring many possible strategiesTree-of-thoughtHigh
Start with the cheapest technique that fits. Reserve tree-of-thought for problems that truly branch.

Watch out

More machinery is not always better. Each technique adds latency, cost, and failure points. Solve the problem with the simplest approach that works, and escalate only when results demand it.

Recap & quick check

Key takeaways

  • A single chain of thought can't look things up, reconsider, or compare alternatives — its main limits.
  • ReAct interleaves Thought → Action (tool) → Observation, grounding reasoning in real results; it's the core of agents.
  • Reflection has the model critique its own draft against the requirements, then revise — models make good critics.
  • Tree-of-thought explores multiple approaches and expands the best; it's powerful but the most expensive.
  • Pick the cheapest technique that fits the problem's actual gap; don't over-engineer.

Quick check

1. What does the ReAct loop consist of?

2. Why is ReAct able to answer questions about very recent events?

3. What is the core idea of reflection?

4. When is tree-of-thought the right tool?

That completes the reasoning phase. Next we make the model's output not just smart but structured and dependable — starting with clean, machine-readable JSON. Next up: Phase 4, Module 14 — Structured Output: JSON & Schemas.