Phase 6 · Testing, Delivery & ProductionModule 46~68 min read

Observability: Logs, Metrics & Traces

Instrument services with structured logs, RED metrics, distributed traces, context propagation, sampling, dashboards, and alerts.

What you'll learn

Observability connects a user request across services and lets operators ask new questions. Combine structured application logs, RED metrics, distributed traces, context propagation, sampling, dashboards, and actionable alerts.

By the end of this lesson, you'll be able to:

  • Emit structured contextual logs
  • Design counters, gauges, and histograms
  • Trace distributed work with OpenTelemetry
  • Build SLO-oriented dashboards and alerts

Core mental model

Node.js becomes easier when you separate the JavaScript language from the runtime and the operating-system capabilities it exposes. Use this table as a decision guide.

ConceptWhat it meansDecision rule
LogDiscrete contextual eventUse structured fields and redact sensitive data
MetricAggregated numeric time seriesUse low-cardinality dimensions for trends and alerts
TraceCausal path across operationsPropagate context through HTTP, queues, and workers

Professional workflow

Build and verify Node.js programs from the terminal in small, observable steps.

  1. Define the observability signal boundary: inputs, outputs, invariants, ownership, and expected failures.
  2. Design the data or message contract before choosing implementation details.
  3. Implement the smallest correct path with dependencies passed explicitly.
  4. Add validation, failure translation, cleanup, and concurrency behavior.
  5. Verify the boundary with realistic data and at least one adversarial case.
  6. Measure or observe the behavior before optimizing or extracting abstractions.

Keep the feedback loop short

Run the smallest useful command after every meaningful change. Read the complete error message before editing again, and keep inputs and outputs visible while you learn.

Guided code lab

Create one application span

Semantic attributes describe the use case without high-cardinality secrets, and exceptions affect span status.

instrumented-task.js
return tracer.startActiveSpan('task.create', async (span) => {
  try {
    span.setAttribute('app.tenant_tier', subject.tenantTier);
    const task = await createTask(subject, input);
    span.setAttribute('app.task.created', true);
    return task;
  } catch (error) {
    span.recordException(error);
    span.setStatus({ code: SpanStatusCode.ERROR });
    throw error;
  } finally {
    span.end();
  }
});

Production practice

Contract

Every critical operation carries correlation context and emits enough safe evidence to identify rate, errors, duration, saturation, and dependency cause.

Verification

Trace HTTP-to-queue-to-worker flows, force errors and timeouts, verify log correlation, metric units/buckets, cardinality, sampling, and redaction.

Operations

Set retention and sampling, protect telemetry credentials, alert on symptoms tied to runbooks, and test collector outage behavior.

Common failure mode

Putting user IDs, URLs, or error messages into metric labels creates unbounded cardinality and can overwhelm the monitoring system.

Independent workshop

Instrument the production API and background worker end to end.

Your finished workshop must include:

  • Log schema/redaction
  • RED metrics
  • OpenTelemetry setup
  • Queue context propagation
  • Service dashboard
  • Two actionable alerts

Definition of done

Run the happy path and at least two edge cases, keep responsibilities separated, and add a short README explaining how to run the program.

Recap & quick check

Key takeaways

  • Signals answer different questions
  • Context connects work
  • Metrics require bounded labels
  • Traces reveal causality
  • Alerts map to user impact

Quick check

1. Which metric labels are dangerous?

2. What does a trace show?

3. What makes an alert actionable?

Next: Deployment, Scaling & Reliability