Phase 4 · Advanced Querying & AnalyticsModule 25~64 min read

Common Table Expressions & Query Decomposition

Use WITH queries to name intermediate results, stage transformations, and improve reasoning while understanding planner behavior.

What you'll learn

Use WITH queries to name intermediate results, stage transformations, and improve reasoning while understanding planner behavior. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply WITH syntax to a realistic data question
  • Apply Multiple CTEs to a realistic data question
  • Apply Data-modifying CTEs to a realistic data question
  • Apply Materialization to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
CTEA named query scoped to one statementUse names that describe business meaning, not step numbers
MaterializationComputing and storing an intermediate resultUnderstand version-specific planner choices before relying on it
Data-modifying CTEA WITH term that changes and returns rowsUse cautiously when one statement must connect changes

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the CTE query pipeline question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Separate cohort and measure

Named stages make the eligible customer population distinct from its aggregated outcome.

cohort_revenue.sql
WITH eligible_customers AS (
  SELECT id FROM sales.customers
  WHERE created_at < current_date - INTERVAL '1 year'
), revenue AS (
  SELECT o.customer_id, sum(o.total) AS lifetime_revenue
  FROM sales.orders AS o
  WHERE o.status = 'paid'
  GROUP BY o.customer_id
)
SELECT c.id, COALESCE(r.lifetime_revenue, 0) AS lifetime_revenue
FROM eligible_customers AS c
LEFT JOIN revenue AS r ON r.customer_id = c.id;

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

Breaking a query into many CTEs can obscure grain transitions and, on some plans, force unnecessary work. Names do not guarantee performance.

Independent workshop

Build a review-ready CTE query pipeline lab against the course commerce dataset.

Your finished workshop must include:

  • WITH syntax
  • Multiple CTEs
  • Data-modifying CTEs
  • Materialization
  • Query decomposition
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • CTE: Use names that describe business meaning, not step numbers
  • Materialization: Understand version-specific planner choices before relying on it
  • Data-modifying CTE: Use cautiously when one statement must connect changes

Quick check

1. Which rule best applies to CTE?

2. Which rule best applies to Materialization?

3. Which rule best applies to Data-modifying CTE?

Next: Recursive CTEs & Hierarchical Data