What you'll learn
"This prompt feels better" is where amateurs stop and professionals begin. To improve a prompt reliably — and to know a change didn't secretly break something — you need to measure it. This module turns prompting from art into engineering with evaluation sets, metrics, and LLM-as-a-judge.
By the end of this module you'll be able to:
- Explain why "vibes" fail as a quality signal
- Build an evaluation set that represents real usage
- Pick metrics for open-ended output, including LLM-as-a-judge
- A/B test prompts and guard against regressions over time
Why vibes aren't enough
Eyeballing a few outputs is fine for a first draft and dangerous for anything real. You'll remember the one great answer and forget the three mediocre ones; a change that fixes your test case may quietly break five others you didn't check. At scale, the only honest question is: across many representative inputs, how often does this prompt produce a good result?
Building an eval set
An evaluation set ("evals") is a collection of representative inputs, ideally paired with the expected output or clear success criteria. It's your prompt's test suite. A good set:
- Mirrors real usage — draw from actual inputs, not just tidy examples.
- Includes the hard cases — edge cases, empty inputs, and past failures (Module 17).
- Is big enough to trust — a handful of cases is noisy; dozens start to be meaningful.
- Grows over time — every bug you find in production becomes a new eval case.
Key idea
Metrics for open-ended output
Some tasks are easy to score: a classification is right or wrong; extracted JSON matches or doesn't. Others — summaries, answers, drafts — have no single correct string. For those, score against criteria rather than an exact match:
| Output type | How to score it |
|---|---|
| Classification / extraction | Exact match or accuracy vs. a labeled answer |
| Short factual answer | Contains the key fact; matches a reference |
| Summary / long answer | Rubric scoring: accurate? complete? on-format? (often via an LLM judge) |
| Format compliance | Programmatic check — does it parse / match the schema? (Module 14) |
LLM-as-a-judge
For open-ended output at scale, human grading doesn't keep up — so use another model as the grader. An LLM-as-a-judge prompt scores each output against a rubric. It's not perfect, but it's consistent and fast, which makes large evals practical.
System
You are a strict grader. Score the answer 1-5 for factual accuracy against the reference. Reply as JSON: {score, reason}.
Prompt
Question: What is the refund window? Reference: "14 days." Answer: "You can get a refund within two weeks of purchase."
AI response
Tip
A/B testing & regression testing
With an eval set and a metric, improving a prompt becomes experimentation: change the prompt, run both versions over the set, and keep the winner. And crucially, re-run the whole set after every change — that's regression testing, and it's how you catch a "fix" that quietly broke other cases.
# Turn "it seems better" into a number.
cases = load_eval_set() # inputs + expected/reference outputs
THRESHOLD = 4 # judge score out of 5
def evaluate(prompt_version):
scores = []
for c in cases:
out = run(prompt_version, c.input)
scores.append(judge(c.input, c.expected, out)) # metric or LLM-judge
return mean(s >= THRESHOLD for s in scores) # pass rate
print("v1:", evaluate(PROMPT_V1))
print("v2:", evaluate(PROMPT_V2)) # compare, don't guessRecap & quick check
Key takeaways
- 'It feels better' doesn't scale — measure prompt quality across many representative inputs.
- An eval set is your prompt's test suite: representative, includes hard cases, big enough to trust, and grows over time.
- Match the metric to the output — programmatic checks where possible, rubric scoring for open-ended text.
- LLM-as-a-judge scores open-ended output consistently at scale; give it a reference and a concrete rubric.
- A/B test prompt versions over the same set and re-run it after every change to catch regressions.
Quick check
1. Why are 'vibes' an unreliable way to judge a prompt?
2. What is an evaluation set?
3. What is 'LLM-as-a-judge'?
4. Why re-run the entire eval set after every prompt change?
Evaluation tells you whether a prompt improved. Next: a disciplined loop for making it improve when it's falling short. Next up: Module 27 — Iterating & Debugging Prompts.