Phase 3 · Databases & Persistent DataModule 21~62 min read

MongoDB & Document Modeling

Model aggregates as documents, use the official MongoDB driver, design indexes, and choose embedding or references deliberately.

What you'll learn

MongoDB stores aggregates as documents, changing where joins, consistency, and duplication live. Shape documents from access patterns, manage one MongoClient, and select indexes and consistency options intentionally.

By the end of this lesson, you'll be able to:

  • Choose embedding or references
  • Use one managed MongoClient
  • Build safe CRUD and projections
  • Design indexes from access patterns

Core mental model

Node.js becomes easier when you separate the JavaScript language from the runtime and the operating-system capabilities it exposes. Use this table as a decision guide.

ConceptWhat it meansDecision rule
AggregateData changed and read as one consistency boundaryEmbed bounded child data sharing the parent lifecycle
ReferenceAn identifier connecting independent documentsUse for unbounded data or separate lifecycles
Compound indexAn ordered index over multiple fieldsMatch equality, sort, and range query shape

Professional workflow

Build and verify Node.js programs from the terminal in small, observable steps.

  1. Define the document model boundary: inputs, outputs, invariants, ownership, and expected failures.
  2. Design the data or message contract before choosing implementation details.
  3. Implement the smallest correct path with dependencies passed explicitly.
  4. Add validation, failure translation, cleanup, and concurrency behavior.
  5. Verify the boundary with realistic data and at least one adversarial case.
  6. Measure or observe the behavior before optimizing or extracting abstractions.

Keep the feedback loop short

Run the smallest useful command after every meaningful change. Read the complete error message before editing again, and keep inputs and outputs visible while you learn.

Guided code lab

Connect once and project

A process-level client owns pooling; projections avoid returning heavy fields.

tasks.js
const client = new MongoClient(process.env.MONGODB_URI);
await client.connect();
const tasks = client.db('course').collection('tasks');

const page = await tasks.find(
  { ownerId, completedAt: null },
  { projection: { title: 1, createdAt: 1 } },
).sort({ createdAt: -1 }).limit(20).toArray();

Support the query shape

Equality fields precede the newest-first sort field.

indexes.js
await tasks.createIndex(
  { ownerId: 1, completedAt: 1, createdAt: -1 },
  { name: 'owner_open_created' },
);

Production practice

Contract

A document is an aggregate boundary, not permission to store an unbounded object graph.

Verification

Test schema edges, duplicate keys, query plans, pagination stability, and realistic document sizes.

Operations

Reuse the client, cap arrays, monitor slow operations and pool pressure, and choose read/write concerns deliberately.

Common failure mode

Embedding an ever-growing event array approaches document limits and makes every update more expensive.

Independent workshop

Model a collaborative board and implement its two highest-volume reads.

Your finished workshop must include:

  • Aggregate boundaries
  • Embedding rationale
  • Validated CRUD
  • Two compound indexes
  • Explain notes
  • Client shutdown

Definition of done

Run the happy path and at least two edge cases, keep responsibilities separated, and add a short README explaining how to run the program.

Recap & quick check

Key takeaways

  • Documents model aggregates
  • Embedding favors locality
  • References preserve lifecycles
  • Indexes follow queries
  • Consistency is explicit

Quick check

1. When is embedding strongest?

2. How many MongoClient instances should a process usually use?

3. What drives compound index order?

Next: Mongoose Schemas & Validation