Phase 6 · Testing, Delivery & ProductionModule 47~72 min read

Deployment, Scaling & Reliability

Operate stateless Node.js services behind proxies with replicas, health probes, timeouts, load shedding, graceful shutdown, and reliability targets.

What you'll learn

Operate stateless Node.js replicas behind a proxy with distinct liveness and readiness, bounded timeouts, retry budgets, load shedding, graceful draining, and service-level objectives that guide tradeoffs.

By the end of this lesson, you'll be able to:

  • Scale stateless services
  • Design health and graceful shutdown
  • Budget timeouts and retries
  • Define SLOs and error budgets

Core mental model

Node.js becomes easier when you separate the JavaScript language from the runtime and the operating-system capabilities it exposes. Use this table as a decision guide.

ConceptWhat it meansDecision rule
ReadinessWhether a replica should receive trafficFail during startup, draining, or inability to serve
Retry budgetBound on repeated workRetry transient, idempotent operations with jitter within deadline
SLOTarget reliability over a windowTie availability or latency to user outcomes

Professional workflow

Build and verify Node.js programs from the terminal in small, observable steps.

  1. Define the reliable service boundary: inputs, outputs, invariants, ownership, and expected failures.
  2. Design the data or message contract before choosing implementation details.
  3. Implement the smallest correct path with dependencies passed explicitly.
  4. Add validation, failure translation, cleanup, and concurrency behavior.
  5. Verify the boundary with realistic data and at least one adversarial case.
  6. Measure or observe the behavior before optimizing or extracting abstractions.

Keep the feedback loop short

Run the smallest useful command after every meaningful change. Read the complete error message before editing again, and keep inputs and outputs visible while you learn.

Guided code lab

Drain before stopping

Readiness changes first, the server stops accepting work, and infrastructure closes inside a hard deadline.

shutdown.js
async function shutdown(signal) {
  readiness.set(false);
  server.close();
  const deadline = AbortSignal.timeout(20_000);
  await Promise.race([
    Promise.all([httpFinished(server), workers.drain(), pool.end(), redis.quit()]),
    new Promise((_, reject) => deadline.addEventListener('abort', () => reject(deadline.reason))),
  ]);
}
for (const signal of ['SIGTERM', 'SIGINT']) process.once(signal, () => shutdown(signal));

Production practice

Contract

Each dependency call has a deadline, cancellation path, retry policy, idempotency assumption, fallback, and observable failure.

Verification

Terminate replicas under load, remove dependencies, saturate pools, exceed capacity, deploy mixed versions, and measure SLO impact.

Operations

Use autoscaling signals beyond CPU, capacity headroom, load shedding, circuit breaking where justified, runbooks, and post-incident learning.

Common failure mode

Retries at every layer multiply traffic during an outage and can prevent the dependency from recovering.

Independent workshop

Write and exercise the reliability plan for the production API.

Your finished workshop must include:

  • Statelessness review
  • Health contracts
  • Graceful drain
  • Timeout/retry budget
  • Load-shed policy
  • SLO/error-budget dashboard

Definition of done

Run the happy path and at least two edge cases, keep responsibilities separated, and add a short README explaining how to run the program.

Recap & quick check

Key takeaways

  • Replicas need shared state elsewhere
  • Readiness differs from liveness
  • Deadlines bound work
  • Retries consume capacity
  • SLOs express user reliability

Quick check

1. What should happen before process exit?

2. Why add jitter to retries?

3. What does an error budget guide?

Next: Phase Project: Production Delivery