Phase 3 · Relationships & ReportingModule 21~62 min read

Aggregate Functions, GROUP BY & HAVING

Calculate counts, totals, averages, and conditional measures at a clearly defined grouping grain.

What you'll learn

Calculate counts, totals, averages, and conditional measures at a clearly defined grouping grain. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply COUNT SUM AVG to a realistic data question
  • Apply MIN and MAX to a realistic data question
  • Apply GROUP BY to a realistic data question
  • Apply HAVING to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
Group grainOne output row per unique grouping keyName the grain before selecting columns
HAVINGA predicate applied after groups are formedUse WHERE for input rows and HAVING for aggregate results
FILTERA condition attached to one aggregateUse for multiple conditional measures in one grouping pass

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the grouped measure question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Calculate customer KPIs

FILTER produces open and paid measures without changing the shared group input.

customer_kpis.sql
SELECT
  customer_id,
  count(*) AS orders,
  count(*) FILTER (WHERE status = 'open') AS open_orders,
  sum(total) FILTER (WHERE status = 'paid') AS paid_revenue
FROM sales.orders
WHERE created_at >= current_date - INTERVAL '1 year'
GROUP BY customer_id
HAVING count(*) >= 3;

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

Selecting a non-grouped, non-aggregated column either fails or—on permissive systems—returns an arbitrary value unrelated to the group.

Independent workshop

Build a review-ready grouped measure lab against the course commerce dataset.

Your finished workshop must include:

  • COUNT SUM AVG
  • MIN and MAX
  • GROUP BY
  • HAVING
  • FILTER
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • Group grain: Name the grain before selecting columns
  • HAVING: Use WHERE for input rows and HAVING for aggregate results
  • FILTER: Use for multiple conditional measures in one grouping pass

Quick check

1. Which rule best applies to Group grain?

2. Which rule best applies to HAVING?

3. Which rule best applies to FILTER?

Next: Subqueries, EXISTS & Correlated Logic