Phase 6 · Professional Database EngineeringModule 43~78 min read

Table Partitioning & Data Lifecycle

Partition large tables by range, list, or hash, enable pruning, manage local indexes, and automate retention safely.

What you'll learn

Partition large tables by range, list, or hash, enable pruning, manage local indexes, and automate retention safely. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply Declarative partitioning to a realistic data question
  • Apply Range list hash to a realistic data question
  • Apply Partition pruning to a realistic data question
  • Apply Partition keys to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
Partition keyThe value routing rows to child tablesChoose from high-volume pruning and lifecycle operations
PruningPlanner or executor excludes impossible partitionsWrite predicates that constrain the partition key
DetachRemoving a partition from the parent without row-by-row deleteUse for archiving or retention with safeguards

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the partitioned-table lifecycle question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Partition events by month

Half-open range bounds prevent overlap and support predictable monthly creation and retirement.

event_partitions.sql
CREATE TABLE telemetry.events (
  occurred_at timestamptz NOT NULL,
  tenant_id bigint NOT NULL,
  payload jsonb NOT NULL
) PARTITION BY RANGE (occurred_at);

CREATE TABLE telemetry.events_2026_08
PARTITION OF telemetry.events
FOR VALUES FROM ('2026-08-01') TO ('2026-09-01');

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

Partitioning is not a universal speed switch; poor pruning, too many partitions, and missing local indexes can make operations worse.

Independent workshop

Build a review-ready partitioned-table lifecycle lab against the course commerce dataset.

Your finished workshop must include:

  • Declarative partitioning
  • Range list hash
  • Partition pruning
  • Partition keys
  • Index strategy
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • Partition key: Choose from high-volume pruning and lifecycle operations
  • Pruning: Write predicates that constrain the partition key
  • Detach: Use for archiving or retention with safeguards

Quick check

1. Which rule best applies to Partition key?

2. Which rule best applies to Pruning?

3. Which rule best applies to Detach?

Next: Import, Export, ETL & Data Quality