What you'll learn
Operate stateless Node.js replicas behind a proxy with distinct liveness and readiness, bounded timeouts, retry budgets, load shedding, graceful draining, and service-level objectives that guide tradeoffs.
By the end of this lesson, you'll be able to:
- Scale stateless services
- Design health and graceful shutdown
- Budget timeouts and retries
- Define SLOs and error budgets
Core mental model
Node.js becomes easier when you separate the JavaScript language from the runtime and the operating-system capabilities it exposes. Use this table as a decision guide.
| Concept | What it means | Decision rule |
|---|---|---|
| Readiness | Whether a replica should receive traffic | Fail during startup, draining, or inability to serve |
| Retry budget | Bound on repeated work | Retry transient, idempotent operations with jitter within deadline |
| SLO | Target reliability over a window | Tie availability or latency to user outcomes |
Professional workflow
Build and verify Node.js programs from the terminal in small, observable steps.
- Define the reliable service boundary: inputs, outputs, invariants, ownership, and expected failures.
- Design the data or message contract before choosing implementation details.
- Implement the smallest correct path with dependencies passed explicitly.
- Add validation, failure translation, cleanup, and concurrency behavior.
- Verify the boundary with realistic data and at least one adversarial case.
- Measure or observe the behavior before optimizing or extracting abstractions.
Keep the feedback loop short
Guided code lab
Drain before stopping
Readiness changes first, the server stops accepting work, and infrastructure closes inside a hard deadline.
async function shutdown(signal) {
readiness.set(false);
server.close();
const deadline = AbortSignal.timeout(20_000);
await Promise.race([
Promise.all([httpFinished(server), workers.drain(), pool.end(), redis.quit()]),
new Promise((_, reject) => deadline.addEventListener('abort', () => reject(deadline.reason))),
]);
}
for (const signal of ['SIGTERM', 'SIGINT']) process.once(signal, () => shutdown(signal));Production practice
Contract
Each dependency call has a deadline, cancellation path, retry policy, idempotency assumption, fallback, and observable failure.
Verification
Terminate replicas under load, remove dependencies, saturate pools, exceed capacity, deploy mixed versions, and measure SLO impact.
Operations
Use autoscaling signals beyond CPU, capacity headroom, load shedding, circuit breaking where justified, runbooks, and post-incident learning.
Common failure mode
Independent workshop
Write and exercise the reliability plan for the production API.
Your finished workshop must include:
- Statelessness review
- Health contracts
- Graceful drain
- Timeout/retry budget
- Load-shed policy
- SLO/error-budget dashboard
Definition of done
Recap & quick check
Key takeaways
- Replicas need shared state elsewhere
- Readiness differs from liveness
- Deadlines bound work
- Retries consume capacity
- SLOs express user reliability
Quick check
1. What should happen before process exit?
2. Why add jitter to retries?
3. What does an error budget guide?
Next: Phase Project: Production Delivery