Phase 5 · Performance, Transactions & SecurityModule 34~76 min read

EXPLAIN, ANALYZE & Planner Statistics

Read PostgreSQL execution plans, compare estimates with actuals, identify scans and join algorithms, and improve statistics.

What you'll learn

Read PostgreSQL execution plans, compare estimates with actuals, identify scans and join algorithms, and improve statistics. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply EXPLAIN to a realistic data question
  • Apply EXPLAIN ANALYZE to a realistic data question
  • Apply Cost and rows to a realistic data question
  • Apply Scan types to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
EstimatePlanner prediction of rows and costCompare estimated and actual rows to find statistics or correlation problems
ScanA method for locating table rowsJudge sequential and index scans by work and selectivity, not labels alone
Join algorithmNested loop, hash, or merge strategyEvaluate it in context of input sizes, order, indexes, and memory

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the execution-plan diagnosis question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Measure a real query safely

BUFFERS and timing expose work, but ANALYZE executes the statement and should be used carefully for data changes.

explain_orders.sql
EXPLAIN (ANALYZE, BUFFERS, VERBOSE, SETTINGS)
SELECT id, status, total
FROM sales.orders
WHERE tenant_id = 42
  AND archived_at IS NULL
ORDER BY created_at DESC, id DESC
LIMIT 50;

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

EXPLAIN ANALYZE executes the query. Running it on UPDATE or DELETE without a protective transaction changes data.

Independent workshop

Build a review-ready execution-plan diagnosis lab against the course commerce dataset.

Your finished workshop must include:

  • EXPLAIN
  • EXPLAIN ANALYZE
  • Cost and rows
  • Scan types
  • Join algorithms
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • Estimate: Compare estimated and actual rows to find statistics or correlation problems
  • Scan: Judge sequential and index scans by work and selectivity, not labels alone
  • Join algorithm: Evaluate it in context of input sizes, order, indexes, and memory

Quick check

1. Which rule best applies to Estimate?

2. Which rule best applies to Scan?

3. Which rule best applies to Join algorithm?

Next: Query Optimization & Sargability