What you'll learn
Design facts, dimensions, grain, surrogate keys, slowly changing dimensions, and incremental analytical loads. The lab uses PostgreSQL while identifying the semantics that transfer to other relational systems.
By the end of this lesson, you'll be able to:
- Apply OLTP vs OLAP to a realistic data question
- Apply Fact grain to a realistic data question
- Apply Dimensions to a realistic data question
- Apply Star schemas to a realistic data question
Core mental model
SQL is declarative: describe the result or invariant you need, then let the database choose a physical execution strategy. Use this table to connect syntax to design decisions.
| Concept | What it means | Decision rule |
|---|---|---|
| Fact grain | Exactly what one fact row measures | Declare it before dimensions or measures |
| Surrogate key | Warehouse-controlled dimension identity | Use to preserve history across changing source keys |
| Slowly changing dimension | A method for tracking attribute history | Select Type 1 or Type 2 from the analytical question |
Professional workflow
Work from a defined question and result grain, then verify correctness before performance.
- State the dimensional analytical model question and the exact grain of the expected result.
- Inspect table definitions, keys, constraints, representative values, and row counts.
- Write the smallest correct query with explicit columns, aliases, and predicates.
- Test missing, duplicate, boundary, and NULL cases before trusting the result.
- Inspect the execution plan or affected rows when cost or data change matters.
- Save the query with its assumptions, parameters, verification, and recovery notes.
Make results explainable
Guided SQL lab
Model an order-line fact
The fact records additive measures at one order-line event while dimensions provide analysis context.
CREATE TABLE warehouse.fact_order_line (
order_line_key bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
order_id bigint NOT NULL,
line_number integer NOT NULL,
date_key integer NOT NULL REFERENCES warehouse.dim_date(date_key),
customer_key bigint NOT NULL REFERENCES warehouse.dim_customer(customer_key),
product_key bigint NOT NULL REFERENCES warehouse.dim_product(product_key),
quantity integer NOT NULL,
gross_amount numeric(14,2) NOT NULL,
discount_amount numeric(14,2) NOT NULL,
UNIQUE (order_id, line_number)
);Production practice
Contract
Define the expected row grain, inputs, output columns, invariants, and failure or empty-result behavior before writing SQL.
Verification
Use representative fixtures and independent row-count, uniqueness, NULL, and boundary checks; compare plans when cost matters.
Operations
Save reviewed SQL with explicit schema names where appropriate, bounded scope, least privilege, observability, and a recovery path for changes.
Common failure mode
Independent workshop
Build a review-ready dimensional analytical model lab against the course commerce dataset.
Your finished workshop must include:
- OLTP vs OLAP
- Fact grain
- Dimensions
- Star schemas
- Slowly changing dimensions
- Verification notes and edge-case evidence
Definition of done
Recap & quick check
Key takeaways
- Fact grain: Declare it before dimensions or measures
- Surrogate key: Use to preserve history across changing source keys
- Slowly changing dimension: Select Type 1 or Type 2 from the analytical question
Quick check
1. Which rule best applies to Fact grain?
2. Which rule best applies to Surrogate key?
3. Which rule best applies to Slowly changing dimension?
Next: Capstone: Commerce Database & Analytics Portfolio