Phase 6 · Professional Database EngineeringModule 44~74 min read

Import, Export, ETL & Data Quality

Move data through staging tables, validate and transform it set-wise, reconcile results, and make pipelines idempotent.

What you'll learn

Move data through staging tables, validate and transform it set-wise, reconcile results, and make pipelines idempotent. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.

By the end of this lesson, you'll be able to:

  • Apply COPY to a realistic data question
  • Apply CSV formats to a realistic data question
  • Apply Staging tables to a realistic data question
  • Apply Set-based transforms to a realistic data question

Core mental model

SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.

ConceptWhat it meansDecision rule
Staging tableA controlled landing area matching source shapeKeep raw imports separate from trusted domain tables
ReconciliationCounts and totals proving source-to-target completenessRecord controls before and after every load
IdempotencyRepeating a batch produces the same final stateUse stable source keys and batch identity

Professional workflow

Work from a defined question and result grain, then verify correctness before performance.

  1. State the idempotent data pipeline question and the exact grain of the expected result.
  2. Inspect table definitions, keys, constraints, representative values, and row counts.
  3. Write the smallest correct query with explicit columns, aliases, and predicates.
  4. Test missing, duplicate, boundary, and NULL cases before trusting the result.
  5. Inspect the execution plan or affected rows when cost or data change matters.
  6. Save the query with its assumptions, parameters, verification, and recovery notes.

Make results explainable

Keep each query in a saved SQL file with a short statement of its purpose, expected grain, assumptions, and verification query.

Guided SQL lab

Validate staging before merge

Rejected rows remain queryable while valid rows are transformed and upserted using the source key.

load_customers.sql
INSERT INTO core.customers (source_id, email, joined_at)
SELECT source_id, lower(trim(email)), joined_at_text::timestamptz
FROM staging.customers_20260825
WHERE email IS NOT NULL
  AND email ~* '^[^@]+@[^@]+$'
ON CONFLICT (source_id) DO UPDATE
SET email = EXCLUDED.email,
    joined_at = EXCLUDED.joined_at;

Production practice

Contract

Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.

Verification

Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.

Operations

Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.

Common failure mode

Loading directly into production tables mixes parsing, quality rejection, deduplication, and domain changes with no auditable boundary.

Independent workshop

Build a review-ready idempotent data pipeline lab against the course commerce dataset.

Your finished workshop must include:

  • COPY
  • CSV formats
  • Staging tables
  • Set-based transforms
  • Data-quality rules
  • Verification notes and edge-case evidence

Definition of done

Run the expected case and at least two edge cases, verify row counts and grain, and add comments explaining any vendor-specific behavior.

Recap & quick check

Key takeaways

  • Staging table: Keep raw imports separate from trusted domain tables
  • Reconciliation: Record controls before and after every load
  • Idempotency: Use stable source keys and batch identity

Quick check

1. Which rule best applies to Staging table?

2. Which rule best applies to Reconciliation?

3. Which rule best applies to Idempotency?

Next: Backup, Restore, Monitoring & Maintenance