Phase 3 · Making Models ReasonModule 11~34 min read

Task Decomposition & Prompt Chaining

Big tasks fail as one giant prompt. Break them into a chain of small, reliable steps where each prompt's output feeds the next.

What you'll learn

The instinct to cram an entire job into one giant prompt is natural — and it's where most complex prompts break. The professional move is decomposition: split the task into a chain of small, focused prompts, each doing one thing well and handing its result to the next.

By the end of this module you'll be able to:

  • Recognize when a task is too big for a single prompt
  • Break a job into a sequence of single-purpose steps
  • Pass one step's output cleanly into the next
  • Debug a chain by isolating the step that failed

Why one huge prompt fails

A prompt that says "read this email thread, figure out the priority, draft a reply, translate it, and log it" asks the model to juggle five jobs at once. It tends to drop steps, blur them together, or do a mediocre job of each. Smaller prompts are more reliable for the same reason numbered steps helped in Module 5 — focus. They're also far easier to test and reuse.

Key idea

A single prompt's quality drops as you pile on responsibilities. Two focused prompts that each do one thing well almost always beat one prompt trying to do both.

Decomposing a task

Ask yourself: what are the natural stages a careful human would go through? Usually there's an extract stage (get the raw material into a clean shape), a reason stage (make decisions over that clean material), and a produce stage (write the final output). Each becomes its own prompt.

A three-step prompt chain

Step 1 · Extract

Pull the key facts / structured data from the raw input.

output ↓ becomes next input

Step 2 · Analyze

Reason over just those facts — classify, score, decide.

output ↓ becomes next input

Step 3 · Draft

Write the final output from the analysis.

Each stage has one job. Its output is the next stage's input — smaller, cleaner, and testable in isolation.

Chaining the steps

A chain is just running those prompts in order, feeding each result forward — typically through the API (Phase 6). Because step 2 sees only the clean output of step 1, it isn't distracted by the raw, messy input, and it works on a much smaller, clearer context.

chain.py
# A prompt chain: each call's output feeds the next.
raw_email = load_email()

# Step 1 - extract structured facts
facts = ask(f"Extract sender, request, and deadline as JSON:\n{raw_email}")

# Step 2 - decide priority using ONLY those facts
priority = ask(f"Given these facts, rate priority low/med/high:\n{facts}")

# Step 3 - draft a reply appropriate to the priority
reply = ask(f"Write a reply for a {priority}-priority request:\n{facts}")

send(reply)
A minimal chain: extract → prioritize → draft. Each ask() is a separate, focused prompt; outputs flow forward.

Passing output between steps

The seam between steps is where chains succeed or fail. Make the handoff clean:

  • Use a structured format (like JSON) for intermediate results so the next step can rely on the shape — more in Module 14.
  • Pass only what the next step needs. Forwarding the entire history reintroduces the clutter you decomposed away.
  • Validate between steps. Catch a bad extraction before it poisons everything downstream.

Tip

Treat each step like a function with a typed output. If step 1 promises JSON with three fields, step 2 can be written and tested against exactly that — just like normal software.

Branching & debugging

Chains can branch: a classification step can route to different follow-up prompts (a bug report goes one way, a refund request another). And when the final output is wrong, a chain is far easier to debug than a monolith — you can inspect each step's output and find exactly where things went sideways.

Note

This is the bridge to agents (Module 24). An agent is essentially a chain the model builds itself at runtime — deciding the next step instead of you hard-coding it. Master fixed chains first; they're more predictable and cover most needs.

Recap & quick check

Key takeaways

  • One giant prompt juggling many jobs drops steps and does each poorly — decompose instead.
  • A typical chain is extract → reason → produce, each stage a single-purpose prompt.
  • Feed each step's output forward, ideally in a structured format, passing only what's needed.
  • Validate between steps so a bad early result doesn't poison everything downstream.
  • Chains are easy to test, reuse, branch, and debug — and they're the stepping stone to agents.

Quick check

1. Why does splitting a big task into a prompt chain usually beat one huge prompt?

2. What are the typical three stages of a decomposition?

3. What's the best format for an intermediate result passed to the next step?

4. How does a prompt chain relate to an agent?

Chaining makes each step reliable. But even a single step can get unlucky on a hard problem — so next we ask the model more than once and take a vote. Next up: Module 12 — Self-Consistency & Ensembling.