Phase 6 · Professional Database EngineeringModule 45~84 min read

Backup, Restore, Monitoring & Maintenance

Plan logical and physical recovery, test restores, monitor workload health, and understand vacuum, analyze, bloat, and routine operations.

What you'll learn

Plan logical and physical recovery, test restores, monitor workload health, and understand vacuum, analyze, bloat, and routine operations. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply pg_dump and restore to a realistic data question
  • Apply Physical backup and WAL to a realistic data question
  • Apply RPO and RTO to a realistic data question
  • Apply Monitoring to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
RPOMaximum acceptable data loss measured in timeChoose backup and WAL retention to meet it
RTOMaximum acceptable restoration timeMeasure full restore and validation, not only backup duration
VACUUMReclaims reusable space and supports transaction-ID healthMonitor autovacuum and long transactions rather than running blindly

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the recoverable database operation question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Record restore evidence

Recovery becomes a tested capability when each drill records artifact identity, timing, validation, and gaps.

restore_drills.sql
CREATE TABLE operations.restore_drills (
  id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
  backup_reference text NOT NULL,
  started_at timestamptz NOT NULL,
  completed_at timestamptz,
  recovered_to timestamptz,
  validation_status text NOT NULL,
  findings text
);

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

A backup that has never been restored is an assumption. Corruption, missing WAL, credentials, extensions, or slow transfer may appear only during recovery.

Independent workshop

Build a review-ready recoverable database operation lab against the course commerce dataset.

Your finished workshop must include:

  • pg_dump and restore
  • Physical backup and WAL
  • RPO and RTO
  • Monitoring
  • VACUUM and ANALYZE
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • RPO: Choose backup and WAL retention to meet it
  • RTO: Measure full restore and validation, not only backup duration
  • VACUUM: Monitor autovacuum and long transactions rather than running blindly

Quick check

1. Which rule best applies to RPO?

2. Which rule best applies to RTO?

3. Which rule best applies to VACUUM?

Next: SQL from Applications & Prepared Statements