Phase 5 · Performance, Transactions & SecurityModule 33~70 min read

Index Fundamentals & B-Tree Design

Understand index structure and cost, then design single, composite, covering, expression, and partial indexes from workload.

What you'll learn

Understand index structure and cost, then design single, composite, covering, expression, and partial indexes from workload. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply B-tree mental model to a realistic data question
  • Apply Selectivity to a realistic data question
  • Apply Composite order to a realistic data question
  • Apply Covering indexes to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
SelectivityHow strongly a predicate narrows rowsPrioritize indexes that remove substantial work on important queries
Composite indexAn ordered index over several expressionsMatch leading equality, range, and ordering needs from the workload
Partial indexAn index containing rows meeting a predicateUse for stable hot subsets such as active records

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the workload-driven index question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Support an active-customer feed

The partial composite index matches tenant equality and newest-first ordering while excluding archived rows.

customer_feed_index.sql
CREATE INDEX CONCURRENTLY orders_tenant_active_created_idx
ON sales.orders (tenant_id, created_at DESC, id DESC)
INCLUDE (status, total)
WHERE archived_at IS NULL;

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

Indexing every column slows writes, consumes cache, and does not guarantee a useful plan. Index complete access patterns and verify them.

Independent workshop

Build a review-ready workload-driven index lab against the course commerce dataset.

Your finished workshop must include:

  • B-tree mental model
  • Selectivity
  • Composite order
  • Covering indexes
  • Expression indexes
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • Selectivity: Prioritize indexes that remove substantial work on important queries
  • Composite index: Match leading equality, range, and ordering needs from the workload
  • Partial index: Use for stable hot subsets such as active records

Quick check

1. Which rule best applies to Selectivity?

2. Which rule best applies to Composite index?

3. Which rule best applies to Partial index?

Next: EXPLAIN, ANALYZE & Planner Statistics