Phase 7 · Production PromptingModule 26~34 min read

Evaluating & Testing Prompts

Stop guessing whether a prompt is better. Build a test set, define metrics, and use LLM-as-judge to measure prompt quality objectively.

What you'll learn

"This prompt feels better" is where amateurs stop and professionals begin. To improve a prompt reliably — and to know a change didn't secretly break something — you need to measure it. This module turns prompting from art into engineering with evaluation sets, metrics, and LLM-as-a-judge.

By the end of this module you'll be able to:

  • Explain why "vibes" fail as a quality signal
  • Build an evaluation set that represents real usage
  • Pick metrics for open-ended output, including LLM-as-a-judge
  • A/B test prompts and guard against regressions over time

Why vibes aren't enough

Eyeballing a few outputs is fine for a first draft and dangerous for anything real. You'll remember the one great answer and forget the three mediocre ones; a change that fixes your test case may quietly break five others you didn't check. At scale, the only honest question is: across many representative inputs, how often does this prompt produce a good result?

Building an eval set

An evaluation set ("evals") is a collection of representative inputs, ideally paired with the expected output or clear success criteria. It's your prompt's test suite. A good set:

  • Mirrors real usage — draw from actual inputs, not just tidy examples.
  • Includes the hard cases — edge cases, empty inputs, and past failures (Module 17).
  • Is big enough to trust — a handful of cases is noisy; dozens start to be meaningful.
  • Grows over time — every bug you find in production becomes a new eval case.

Key idea

Start an eval set on day one, even with ten cases. The moment you have one, "is this prompt better?" becomes a number you can check instead of an argument you can have.

Metrics for open-ended output

Some tasks are easy to score: a classification is right or wrong; extracted JSON matches or doesn't. Others — summaries, answers, drafts — have no single correct string. For those, score against criteria rather than an exact match:

Output typeHow to score it
Classification / extractionExact match or accuracy vs. a labeled answer
Short factual answerContains the key fact; matches a reference
Summary / long answerRubric scoring: accurate? complete? on-format? (often via an LLM judge)
Format complianceProgrammatic check — does it parse / match the schema? (Module 14)
Match the metric to the output. Programmatic checks where you can; rubric scoring where you can't.

LLM-as-a-judge

For open-ended output at scale, human grading doesn't keep up — so use another model as the grader. An LLM-as-a-judge prompt scores each output against a rubric. It's not perfect, but it's consistent and fast, which makes large evals practical.

System

You are a strict grader. Score the answer 1-5 for factual accuracy against the reference. Reply as JSON: {score, reason}.

Prompt

Question: What is the refund window? Reference: "14 days." Answer: "You can get a refund within two weeks of purchase."

AI response

{"score": 5, "reason": "Two weeks equals 14 days, so the answer matches the reference."}
A judge prompt applies the same rubric to every output — consistent scoring that scales past human review.

Tip

Make the judge's job narrow and concrete: give it a reference answer and a specific rubric, and ask for a reason with the score. A vague "is this good?" judge is nearly as unreliable as vibes.

A/B testing & regression testing

With an eval set and a metric, improving a prompt becomes experimentation: change the prompt, run both versions over the set, and keep the winner. And crucially, re-run the whole set after every change — that's regression testing, and it's how you catch a "fix" that quietly broke other cases.

evaluate.py
# Turn "it seems better" into a number.
cases = load_eval_set()          # inputs + expected/reference outputs
THRESHOLD = 4                    # judge score out of 5

def evaluate(prompt_version):
    scores = []
    for c in cases:
        out = run(prompt_version, c.input)
        scores.append(judge(c.input, c.expected, out))   # metric or LLM-judge
    return mean(s >= THRESHOLD for s in scores)          # pass rate

print("v1:", evaluate(PROMPT_V1))
print("v2:", evaluate(PROMPT_V2))   # compare, don't guess
Score each prompt version over the same eval set and compare the pass rates — decide with data, not memory.

Recap & quick check

Key takeaways

  • 'It feels better' doesn't scale — measure prompt quality across many representative inputs.
  • An eval set is your prompt's test suite: representative, includes hard cases, big enough to trust, and grows over time.
  • Match the metric to the output — programmatic checks where possible, rubric scoring for open-ended text.
  • LLM-as-a-judge scores open-ended output consistently at scale; give it a reference and a concrete rubric.
  • A/B test prompt versions over the same set and re-run it after every change to catch regressions.

Quick check

1. Why are 'vibes' an unreliable way to judge a prompt?

2. What is an evaluation set?

3. What is 'LLM-as-a-judge'?

4. Why re-run the entire eval set after every prompt change?

Evaluation tells you whether a prompt improved. Next: a disciplined loop for making it improve when it's falling short. Next up: Module 27 — Iterating & Debugging Prompts.