What you'll learn
Observability connects a user request across services and lets operators ask new questions. Combine structured application logs, RED metrics, distributed traces, context propagation, sampling, dashboards, and actionable alerts.
By the end of this lesson, you'll be able to:
- Emit structured contextual logs
- Design counters, gauges, and histograms
- Trace distributed work with OpenTelemetry
- Build SLO-oriented dashboards and alerts
Core mental model
Node.js becomes easier when you separate the JavaScript language from the runtime and the operating-system capabilities it exposes. Use this table as a decision guide.
| Concept | What it means | Decision rule |
|---|---|---|
| Log | Discrete contextual event | Use structured fields and redact sensitive data |
| Metric | Aggregated numeric time series | Use low-cardinality dimensions for trends and alerts |
| Trace | Causal path across operations | Propagate context through HTTP, queues, and workers |
Professional workflow
Build and verify Node.js programs from the terminal in small, observable steps.
- Define the observability signal boundary: inputs, outputs, invariants, ownership, and expected failures.
- Design the data or message contract before choosing implementation details.
- Implement the smallest correct path with dependencies passed explicitly.
- Add validation, failure translation, cleanup, and concurrency behavior.
- Verify the boundary with realistic data and at least one adversarial case.
- Measure or observe the behavior before optimizing or extracting abstractions.
Keep the feedback loop short
Guided code lab
Create one application span
Semantic attributes describe the use case without high-cardinality secrets, and exceptions affect span status.
return tracer.startActiveSpan('task.create', async (span) => {
try {
span.setAttribute('app.tenant_tier', subject.tenantTier);
const task = await createTask(subject, input);
span.setAttribute('app.task.created', true);
return task;
} catch (error) {
span.recordException(error);
span.setStatus({ code: SpanStatusCode.ERROR });
throw error;
} finally {
span.end();
}
});Production practice
Contract
Every critical operation carries correlation context and emits enough safe evidence to identify rate, errors, duration, saturation, and dependency cause.
Verification
Trace HTTP-to-queue-to-worker flows, force errors and timeouts, verify log correlation, metric units/buckets, cardinality, sampling, and redaction.
Operations
Set retention and sampling, protect telemetry credentials, alert on symptoms tied to runbooks, and test collector outage behavior.
Common failure mode
Independent workshop
Instrument the production API and background worker end to end.
Your finished workshop must include:
- Log schema/redaction
- RED metrics
- OpenTelemetry setup
- Queue context propagation
- Service dashboard
- Two actionable alerts
Definition of done
Recap & quick check
Key takeaways
- Signals answer different questions
- Context connects work
- Metrics require bounded labels
- Traces reveal causality
- Alerts map to user impact
Quick check
1. Which metric labels are dangerous?
2. What does a trace show?
3. What makes an alert actionable?
Next: Deployment, Scaling & Reliability