Phase 4 · Advanced Querying & AnalyticsModule 29~62 min read

GROUPING SETS, ROLLUP & CUBE

Generate multiple aggregation levels in one query and distinguish subtotal NULLs from stored NULL values.

What you'll learn

Generate multiple aggregation levels in one query and distinguish subtotal NULLs from stored NULL values. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply GROUPING SETS to a realistic data question
  • Apply ROLLUP to a realistic data question
  • Apply CUBE to a realistic data question
  • Apply GROUPING() to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
GROUPING SETSAn explicit list of grouping grainsUse when a report needs selected subtotal levels
ROLLUPHierarchical prefixes of grouping columnsOrder columns from broadest to most detailed hierarchy
GROUPING()Distinguishes subtotal placeholders from stored NULLUse it to label generated totals correctly

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the multi-level aggregation question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Produce detail and subtotals

ROLLUP emits region/category detail, region totals, and a grand total; GROUPING identifies their levels.

sales_rollup.sql
SELECT
  CASE WHEN GROUPING(region) = 1 THEN 'All regions' ELSE region END AS region,
  CASE WHEN GROUPING(category) = 1 THEN 'All categories' ELSE category END AS category,
  sum(revenue) AS revenue
FROM analytics.sales
GROUP BY ROLLUP (region, category)
ORDER BY GROUPING(region), region, GROUPING(category), category;

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

Displaying every NULL as Total incorrectly relabels genuine missing values. Use GROUPING() to distinguish generated subtotal rows.

Independent workshop

Build a review-ready multi-level aggregation lab against the course commerce dataset.

Your finished workshop must include:

  • GROUPING SETS
  • ROLLUP
  • CUBE
  • GROUPING()
  • Subtotals
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • GROUPING SETS: Use when a report needs selected subtotal levels
  • ROLLUP: Order columns from broadest to most detailed hierarchy
  • GROUPING(): Use it to label generated totals correctly

Quick check

1. Which rule best applies to GROUPING SETS?

2. Which rule best applies to ROLLUP?

3. Which rule best applies to GROUPING()?

Next: Views, Materialized Views & Stable Interfaces