Phase 3 · Relationships & ReportingModule 22~68 min read

Subqueries, EXISTS & Correlated Logic

Use scalar, table, and correlated subqueries while choosing EXISTS, IN, joins, or aggregation by semantics.

What you'll learn

Use scalar, table, and correlated subqueries while choosing EXISTS, IN, joins, or aggregation by semantics. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply Scalar subqueries to a realistic data question
  • Apply IN and NOT IN to a realistic data question
  • Apply EXISTS and NOT EXISTS to a realistic data question
  • Apply Correlated subqueries to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
EXISTSTrue when a subquery returns at least one rowUse for presence without multiplying outer rows
Correlated subqueryA subquery referencing the current outer rowUse when semantics are clear, then inspect the plan for scale
Scalar subqueryA subquery required to return at most one row and columnGuarantee uniqueness or use an aggregate

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the subquery composition question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Find customers with no recent payment

NOT EXISTS expresses the anti-condition directly and avoids NULL traps from NOT IN.

inactive_customers.sql
SELECT c.id, c.email
FROM sales.customers AS c
WHERE NOT EXISTS (
  SELECT 1
  FROM sales.payments AS p
  WHERE p.customer_id = c.id
    AND p.paid_at >= current_date - INTERVAL '90 days'
);

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

NOT IN against a subquery containing NULL can make every comparison UNKNOWN and return no rows.

Independent workshop

Build a review-ready subquery composition lab against the course commerce dataset.

Your finished workshop must include:

  • Scalar subqueries
  • IN and NOT IN
  • EXISTS and NOT EXISTS
  • Correlated subqueries
  • Derived tables
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • EXISTS: Use for presence without multiplying outer rows
  • Correlated subquery: Use when semantics are clear, then inspect the plan for scale
  • Scalar subquery: Guarantee uniqueness or use an aggregate

Quick check

1. Which rule best applies to EXISTS?

2. Which rule best applies to Correlated subquery?

3. Which rule best applies to Scalar subquery?

Next: UNION, INTERSECT & EXCEPT