What you'll learn
Great prompts are debugged into existence, not written in one shot. When a prompt underperforms, resist the urge to rewrite everything in frustration. This module gives you a systematic loop — read the failure, form one hypothesis, change one thing, re-measure — that turns fixing prompts from guesswork into a craft.
By the end of this module you'll be able to:
- Run a disciplined draft → run → inspect → refine loop
- Read a failure closely instead of guessing
- Change one variable at a time so you learn what actually helped
- Enlist the model to help debug itself, and know when to stop
The iteration loop
This is the loop from Module 1, now backed by the evals from Module 26. The addition that makes it rigorous: every change is measured, not just glanced at.
Observe
Look closely at a specific failure.
Hypothesize
Guess the one reason it failed.
Change one thing
Edit a single variable in the prompt.
Re-run evals
Did it fix the case without breaking others?
Reading failures like a detective
The fix almost always hides in the failure. Before changing anything, study the bad output and ask what, specifically, went wrong — and why. Precise diagnosis beats a blind rewrite every time.
| The failure looks like… | Likely cause → fix |
|---|---|
| Ignored part of the request | Buried or bundled instruction → put it up front, number the steps (Module 5) |
| Wrong format / unparseable | Output spec unclear → specify format, show an example (Modules 9, 14) |
| Made-up facts | No grounding → provide context, allow 'I don't know' (Module 18) |
| Inconsistent across runs | Temperature too high, or ambiguous prompt (Module 16) |
| Right idea, wrong depth/tone | Missing role or audience → set them (Module 6) |
Change one thing at a time
The cardinal rule of debugging anything: change a single variable, then re-measure. If you rewrite the role, add examples, and change the format all at once and quality improves, you've learned nothing about which change helped — and you may have added two useless edits and one real fix. One change, one measurement, every time.
Key idea
Ask the model why
A uniquely useful trick: the model can help debug itself. Feed it the failed output and ask it to explain its reasoning, or to point out what in the prompt was ambiguous. The answer often reveals the exact misreading — and sometimes suggests the fix.
- "Why did you answer this way?" — surfaces the misinterpretation.
- "Which part of my instructions was unclear?" — finds the ambiguity to fix.
- "What information would have helped you answer correctly?" — reveals a missing context or constraint.
Tip
Knowing when to stop
Iteration has diminishing returns. Stop when the prompt clears your eval bar, when further gains cost more than they're worth, or when the ceiling is the approach, not the wording. If no phrasing gets you there, the answer may be a different technique — few-shot, grounding, decomposition — or a different model, not another rewrite.
Recap & quick check
Key takeaways
- Great prompts are debugged, not written once — run a measured observe → hypothesize → change → re-run loop.
- Diagnose the specific failure first; most map to a known technique (grounding, format spec, roles, temperature).
- Change exactly one variable at a time and re-measure, so you learn which change actually helped.
- Ask the model why it failed or what was unclear, and build a minimal reproduction to expose the cause.
- Stop when you clear the eval bar or hit diminishing returns — sometimes the fix is a new technique, not new wording.
Quick check
1. What's the first step when a prompt underperforms?
2. Why change only one variable at a time?
3. How can the model help you debug your own prompt?
4. When should you stop iterating on wording?
A reliable prompt still faces adversaries. When your prompt meets untrusted text, someone may try to hijack it — so next we cover security. Next up: Module 28 — Prompt Injection & Security.