Phase 2 · The Prompting ToolkitModule 7~34 min read

Zero-shot & Few-shot Prompting

Sometimes the fastest way to explain a task is to show it. Zero-shot vs. few-shot, how to choose and format examples, and the traps that make examples backfire.

What you'll learn

Sometimes describing a task is harder than showing it. Few-shot prompting — including a couple of worked examples right in the prompt — is one of the most reliable ways to lock in a format, a style, or a tricky judgment call. This module covers when to just ask (zero-shot) and when to demonstrate.

By the end of this module you'll be able to:

  • Tell zero-shot and few-shot apart and know when to reach for each
  • Write clean, consistent examples that teach the exact behavior you want
  • Decide how many examples are worth the tokens
  • Avoid the subtle ways examples can backfire

Zero-shot: just ask

Zero-shot means no examples — you simply describe the task. Modern models are strong enough that for common, well-understood tasks this is all you need. Always start here: it's the cheapest prompt, and often it just works.

Prompt

Classify the sentiment of this review as positive, negative, or neutral: "The battery lasts forever, but the camera is mediocre."

AI response

Neutral — it praises battery life but criticizes the camera, so the overall sentiment is mixed.
No examples needed: sentiment classification is a task the model understands from the description alone.

Few-shot: teach by example

When zero-shot output drifts from the format or judgment you want, add examples. A few input → output pairs show the model the exact pattern to imitate — far more precisely than adjectives can. Here the examples pin down both the label set and the terse output format:

few-shot-intent.txt
Classify each support message's intent as one of:
billing, bug, feature_request, praise.

Message: "You've charged me twice this month."
Intent: billing

Message: "The export button does nothing on Safari."
Intent: bug

Message: "Please add a way to sort by due date."
Intent: feature_request

Message: "Honestly the new dashboard is gorgeous."
Intent:
Three labeled examples establish the pattern; the model completes the fourth in the same shape.

Key idea

Few-shot works because the model is a pattern-continuer (Module 2). Your examples are the pattern — so the answer inherits their format, tone, and labeling exactly. Make the examples look like the output you want.

How many examples?

More isn't always better — each example costs tokens and context. Start small and add only if quality needs it:

ExamplesWhen it's the right call
0 (zero-shot)Common tasks the model already understands; start here
1-3Locking a specific output format or a consistent style
4-8Nuanced judgment, tricky edge cases, or an unusual label set
Many (dozens)Rarely worth it — consider fine-tuning or RAG instead
Add examples until quality plateaus, then stop. Past a point you're just spending tokens.

Choosing & formatting examples

Good examples share a few qualities. Aim for these:

  • Representative. Cover the real variety of inputs, including a hard case or two.
  • Consistent. Identical structure every time — same labels, same separators, same casing.
  • Correct. The model copies mistakes faithfully, so a wrong example teaches wrong behavior.
  • Balanced. If you show three "positive" examples and one "negative," the model leans positive.

Tip

Use a clear, repeated delimiter between examples (like Message: / Intent: above). It signals where each example starts and ends, and tells the model exactly where to write its answer.

When few-shot hurts

Examples are powerful, which means they can also mislead:

  • Accidental patterns. If all your examples happen to be short, the model may think short is the rule and truncate real answers.
  • Label bias. An unbalanced set skews predictions toward the over-represented label.
  • Overfitting the format. Examples that are too narrow make the model brittle on inputs that look different.
  • Token cost. Long examples repeated on every call add up fast at scale.

Watch out

If few-shot output looks oddly uniform, inspect your examples for a pattern you didn't intend — order, length, or an over-represented label. The model is often imitating something you didn't mean to teach.

Recap & quick check

Key takeaways

  • Zero-shot (no examples) is the cheapest prompt and often enough for common tasks — start there.
  • Few-shot adds input → output examples to lock in a format, style, or tricky judgment.
  • It works because the model continues patterns: your examples ARE the pattern it imitates.
  • Keep examples representative, consistent, correct, and balanced, with a clear repeated delimiter.
  • Examples can backfire via accidental patterns, label bias, or token cost — watch for uniform output.

Quick check

1. What distinguishes few-shot from zero-shot prompting?

2. Why does few-shot prompting work so well for fixing output format?

3. You show 4 examples, 3 labeled 'positive' and 1 'negative'. What's the risk?

4. Your few-shot answers all come out oddly short. What's the most likely cause?

Examples teach patterns — but only if the model can tell your instructions apart from your data. That's a job for structure. Next up: Module 8 — Structure, Delimiters & Formatting Input.