Phase 7 · Production PromptingModule 27~30 min read

Iterating & Debugging Prompts

A repeatable loop for turning a mediocre prompt into a great one: read failures closely, change one thing at a time, and let the model help debug itself.

What you'll learn

Great prompts are debugged into existence, not written in one shot. When a prompt underperforms, resist the urge to rewrite everything in frustration. This module gives you a systematic loop — read the failure, form one hypothesis, change one thing, re-measure — that turns fixing prompts from guesswork into a craft.

By the end of this module you'll be able to:

  • Run a disciplined draft → run → inspect → refine loop
  • Read a failure closely instead of guessing
  • Change one variable at a time so you learn what actually helped
  • Enlist the model to help debug itself, and know when to stop

The iteration loop

This is the loop from Module 1, now backed by the evals from Module 26. The addition that makes it rigorous: every change is measured, not just glanced at.

The prompt debugging loop
Step 1

Observe

Look closely at a specific failure.

Step 2

Hypothesize

Guess the one reason it failed.

Step 3

Change one thing

Edit a single variable in the prompt.

Step 4

Re-run evals

Did it fix the case without breaking others?

Observe a real failure, hypothesize one cause, change a single variable, then re-run your evals to confirm.

Reading failures like a detective

The fix almost always hides in the failure. Before changing anything, study the bad output and ask what, specifically, went wrong — and why. Precise diagnosis beats a blind rewrite every time.

The failure looks like…Likely cause → fix
Ignored part of the requestBuried or bundled instruction → put it up front, number the steps (Module 5)
Wrong format / unparseableOutput spec unclear → specify format, show an example (Modules 9, 14)
Made-up factsNo grounding → provide context, allow 'I don't know' (Module 18)
Inconsistent across runsTemperature too high, or ambiguous prompt (Module 16)
Right idea, wrong depth/toneMissing role or audience → set them (Module 6)
Most failures map to a technique you already know. Diagnose first, then apply the matching fix.

Change one thing at a time

The cardinal rule of debugging anything: change a single variable, then re-measure. If you rewrite the role, add examples, and change the format all at once and quality improves, you've learned nothing about which change helped — and you may have added two useless edits and one real fix. One change, one measurement, every time.

Key idea

Isolating variables is what separates debugging from flailing. It's slower per step and far faster overall, because you build real knowledge of what moves the needle for your task.

Ask the model why

A uniquely useful trick: the model can help debug itself. Feed it the failed output and ask it to explain its reasoning, or to point out what in the prompt was ambiguous. The answer often reveals the exact misreading — and sometimes suggests the fix.

  • "Why did you answer this way?" — surfaces the misinterpretation.
  • "Which part of my instructions was unclear?" — finds the ambiguity to fix.
  • "What information would have helped you answer correctly?" — reveals a missing context or constraint.

Tip

Also build a minimal reproduction: strip the prompt down to the smallest version that still fails. Removing everything irrelevant usually makes the real cause obvious — the same instinct as a minimal bug report in software.

Knowing when to stop

Iteration has diminishing returns. Stop when the prompt clears your eval bar, when further gains cost more than they're worth, or when the ceiling is the approach, not the wording. If no phrasing gets you there, the answer may be a different technique — few-shot, grounding, decomposition — or a different model, not another rewrite.

Recap & quick check

Key takeaways

  • Great prompts are debugged, not written once — run a measured observe → hypothesize → change → re-run loop.
  • Diagnose the specific failure first; most map to a known technique (grounding, format spec, roles, temperature).
  • Change exactly one variable at a time and re-measure, so you learn which change actually helped.
  • Ask the model why it failed or what was unclear, and build a minimal reproduction to expose the cause.
  • Stop when you clear the eval bar or hit diminishing returns — sometimes the fix is a new technique, not new wording.

Quick check

1. What's the first step when a prompt underperforms?

2. Why change only one variable at a time?

3. How can the model help you debug your own prompt?

4. When should you stop iterating on wording?

A reliable prompt still faces adversaries. When your prompt meets untrusted text, someone may try to hijack it — so next we cover security. Next up: Module 28 — Prompt Injection & Security.