Monitoring is not observability

Two services are down. The first was designed with a dashboard: request count, error rate, p95 latency, CPU, memory. Green graphs everywhere — then at 2:14 a.m. the errors spike, and the dashboard tells you exactly what you already knew: errors are up. You can see that something is wrong. You cannot see why.

The second service was built to be observed. Its logs carry trace IDs and structured fields. Its spans know which database call ate three seconds. Its metrics are sliced by tenant and endpoint. When it degrades, you don't ask "is something wrong?" — you ask "which request, which component, which dependency?" and get an answer.

This is the difference between monitoring and observability. Monitoring asks is the system healthy? Observability asks why is it not? Monitoring is dashboards and thresholds — a set of questions you decided in advance to ask. Observability is the property of a system that lets you answer questions you didn't know you'd need to ask, using data you already emitted. You cannot retrofit observability onto an unobserved system by adding a better dashboard. The data has to be there.

The practical consequence: the goal is not "more dashboards." The goal is that every request, every error, every state change leaves evidence you can reconstruct later. That's what the three pillars are for.

The three pillars, and what each one cannot see

The industry converged on three kinds of telemetry: metrics, logs, and traces. Each is cheap to produce and individually blind. Together they cover most production questions.

Metrics are numbers aggregated over time: request rate, error count, latency distribution, queue depth. Strengths: cheap to store, easy to aggregate, great for dashboards and alerts. Blind spot: they tell you a 503 happened, not which user hit it or what that user was doing. A metric is an average or a distribution — it has already thrown away the individual request.

Logs are discrete records of events: "order created", "connection refused", "payment succeeded". Strengths: high detail, a natural place to record state, and the tool you reach for when you need to know what exactly happened. Blind spot: a log line is a snapshot from one process. It says nothing about the request's context across services, and correlation is manual — unless you put the trace ID in every line.

Traces describe the path of a single request through your system: a root span for the request, child spans for each component, annotated with timing and attributes. Strengths: they connect the dots. A trace shows the whole causal chain — the HTTP call, the two database queries, the cache miss — and exactly where the time went. Blind spot: they only help if you actually sample and store them. A single trace is a story about one request, not about the fleet.

The three pillars overlap. Metrics can be derived from spans (OpenTelemetry can produce RED metrics from trace data). Logs can carry trace context. The system works when you can move between them: start from an alert on a metric, drill into a correlated trace, jump from the trace to the relevant log lines via the trace ID. That cross-reference is the actual payoff, not any single pillar.

The golden signals, and two ways to score them

Google's SRE book defines five golden signals for a user-facing service:

  1. Latency — the time it takes to serve a request. Measure distributions (p50, p95, p99), not averages.
  2. Traffic — how much demand is on the system: requests per second, active users, bytes per second.
  3. Errors — the rate of requests that fail. Define "fail" carefully: 500s obviously, but also 404s that should have been 200s, or responses that come back fast but are still wrong.
  4. Saturation — how full the system is: CPU, memory, queue depth, connection-pool usage. This is your early-warning signal; it degrades before latency does.
  5. Utilization — what fraction of a resource is busy. Saturation is "how much is left"; utilization is "how much is used".

Two acronyms tell you which metrics to track depending on the shape of the service. RED — Rate, Errors, Duration — is for request-driven services: an API, a web service, a job consumer. It asks: how many requests, how many failed, how long did they take?

USE — Utilization, Saturation, Errors — is for resources: a database, a cache, a queue, a worker pool. It asks per resource: is it busy, is it full, is it failing?

RED USE
Question Is the request path healthy? Is the resource healthy?
Applies to Services that handle requests Databases, queues, pools, workers
Metrics Rate, Errors, Duration Utilization, Saturation, Errors
Example RPS, error %, p99 latency CPU %, queue depth, pool exhaustion

A real service uses both: RED on the HTTP layer, USE on the Postgres it sits in front of. The common mistake is applying one framework everywhere and then being surprised that it doesn't describe your dependencies.

Structured logging: shape before volume

Here is a log line from a legacy system:

code
2026-08-17T02:14:31Z ERROR payment failed: connection to db timed out after 5000ms

It reads fine to a human. To a machine it is one string: "payment failed: connection to db timed out after 5000ms". You cannot filter by user ID, by order, by tenant, by database host. You cannot alert on "connection timeouts" without a regex over free text. You cannot join it to a trace. The moment you have more than one service and more than a handful of requests per second, human-parsable log lines stop scaling — they become an archive you grep and hope.

Structured logging emits fields, not sentences. In Node.js/TypeScript, pino is the standard choice — it is fast, JSON by default, and supports child loggers:

typescript
import pino from "pino";

const logger = pino({
  level: process.env.LOG_LEVEL ?? "info",
  base: { service: "payments-api", env: "prod" },
});

// A child logger carries context without you having to repeat it
const orderLogger = logger.child({ orderId: "ord_19f3", tenantId: "acme" });

orderLogger.info({ event: "payment_attempt", amountCents: 2999 }, "payment started");
orderLogger.error(
  { event: "payment_failed", reason: "db_timeout", dbHost: "pg-primary-01", retries: 2 },
  "payment failed after retries"
);

The fields do the work. event is a stable, machine-readable name; reason is a bounded enum, not prose; dbHost lets you correlate with infrastructure. The human-readable message becomes a summary, not the payload. And because the structure is in the line, your aggregation backend (Loki, Elasticsearch, Cloud Logging) becomes useful for something more than a bigger search box — you can filter, group, and alert on the fields.

The threshold to hold: every log line must be independently queryable and joinable. If you cannot filter the logs by the fields you would search for during an incident, the logging investment is mostly wasted.

Distributed tracing: one request, many services

A request to a checkout endpoint touches: an API gateway, an auth service, a cart service, an external payment provider, the order database, and the email worker. When checkout is slow, is it your code, the provider, or a lock in Postgres? Without a trace, the answer is "read twelve different logs and guess."

A trace models a request as a tree of spans. The root span is the whole operation; child spans are the work each component did. Each span records a start time, a duration, attributes, and a status. The spans are linked by a shared traceId, and each span carries its own spanId plus the ID of its parent — that parent link is what builds the tree.

Trace context is propagated, not invented. The W3C traceparent header is the standard: it carries version-traceId-spanId-flags, for example 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01. Each service reads the incoming header, creates a child span under it, and writes its own traceparent on the way out. That is how the whole path stitches together across service and even vendor boundaries.

Correlation IDs are the low-tech cousin of the same idea: put a request ID on every log line and every HTTP header so you can grep a request across systems that are not yet traced. A trace is strictly more useful, but a correlation ID costs one field and works on the systems you have not instrumented. If you can only do one thing today, ship the correlation ID.

OpenTelemetry is the de facto standard for this. Auto-instrumentation (such as @opentelemetry/instrumentation-http) captures inbound and outbound HTTP calls with no code changes and propagates trace context automatically. For the parts it cannot see, you create spans yourself. A minimal Node.js setup:

typescript
import { NodeTracerProvider } from "@opentelemetry/sdk-trace-node";
import { Resource } from "@opentelemetry/resources";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { SimpleSpanProcessor } from "@opentelemetry/sdk-trace-base";
import { trace, SpanKind, SpanStatusCode, context } from "@opentelemetry/api";

const provider = new NodeTracerProvider({
  resource: new Resource({
    "service.name": "checkout-api",
  }),
});

provider.addSpanProcessor(new SimpleSpanProcessor(new OTLPTraceExporter()));
provider.register();

const tracer = trace.getTracer("checkout-api");

Then instrument a piece of work that auto-instrumentation cannot see — say, a call to a payment provider:

typescript
async function chargeCard(provider: PaymentProvider, amountCents: number) {
  const span = tracer.startSpan("charge.credit_card", {
    kind: SpanKind.CLIENT,
    attributes: {
      "payments.provider": "stripe",
      "payments.amount_cents": amountCents,
      "http.method": "POST",
    },
  });

  try {
    const result = await provider.charge(amountCents);
    span.setAttribute("payments.status", result.status);
    span.end();
    return result;
  } catch (err) {
    span.recordException(err);
    span.setStatus({ code: SpanStatusCode.ERROR, message: "provider charge failed" });
    span.end();
    throw err;
  }
}

Note the discipline in that snippet: the span gets a status on every path (never leave a failed call looking healthy), the exception is recorded, and the attributes describe the work rather than the implementation. The HTTP instrumentation picks up the traceparent the caller sent, so this span attaches to the correct trace even though it was created by hand.

Sampling is the part people get wrong. Traces are relatively cheap but not free, so high-traffic systems store only a fraction. Tail-based sampling (decide after the trace completes) is ideal because you can keep every trace that ended in error and drop the boring successful ones; head-based sampling (decide at the start) rolls the dice before you know the outcome and tends to lose exactly the rare failures you care about. Whatever you choose, never let sampling throw away the errors. If your sample rate discards erroring traces, the most important debugging data you have is being silently deleted.

Metric types: know what each one is for

Metrics come in a small set of shapes, and each shape answers a different question. Mixing them up produces dashboards that lie.

A counter only ever increases over a process's lifetime. It measures cumulative totals: requests handled, bytes written, errors raised. To get a rate you compute the change over a time window — a counter that went from 0 to 90 in a minute is 1.5 req/s. You never use a counter for a value that can go down.

typescript
import { metrics } from "@opentelemetry/api";

const meter = metrics.getMeter("checkout-api");

const requestCounter = meter.createCounter("http.requests.total", {
  description: "Total HTTP requests handled",
});

// in the request handler:
requestCounter.add(1, { route: req.route ?? "unknown", method: req.method });

A gauge is a snapshot of a value that goes up and down: current queue depth, connections in the pool, CPU percent. In OpenTelemetry a gauge is observable — you provide a callback that reads the current value when it is collected:

typescript
const poolDepth = meter.createObservableGauge("db.connection_pool.depth", {
  description: "Connections currently in use from the pool",
});
poolDepth.addCallback((observable) => {
  observable.observe(getPoolConnectionsInUse());
});

A histogram records a distribution of observations — latency is the canonical case. You do not store every value; the client buckets them into ranges and the backend computes percentiles from the buckets. A histogram answers "what does p99 look like" without storing every duration:

typescript
const latency = meter.createHistogram("http.request.duration", {
  unit: "ms",
  description: "Request latency",
});

// in the request handler:
latency.record(durationMs, { route: req.route ?? "unknown" });

A counter, a gauge, and a histogram are roughly all you need to describe a service. Add an up-down counter when you need a counter that can also decrement — active in-flight requests, or a queue size that changes in both directions.

The recurring error with metrics is high-cardinality labels. Cardinality is the number of distinct values a label can take. route with 50 routes is fine. user_id, order_id, or request_id as a label is a catastrophe: every new user creates a new time series, storage and query cost blow up, and the metric backend slows to a crawl. Rule of thumb: labels should be drawn from a small, bounded set — endpoint, service, status class, tenant, error code. If you need per-user numbers, put them in logs or traces, not in a metric label.

Connecting the pillars

The pillars are not three separate projects. They are most valuable stitched together:

A pragmatic rollout plan for an existing service

You cannot instrument everything at once. A sane order:

  1. Structured logging first. It is the cheapest change and it improves every future step. Convert your app to a structured logger and put service, environment, and a request correlation_id on every line. This alone fixes the worst incidents.
  2. Add the correlation ID everywhere. Put it in the HTTP response headers (X-Request-Id), propagate it in your clients, and include it in logs. This works on systems you will not get to trace this quarter.
  3. Turn on OpenTelemetry auto-instrumentation. Mostly configuration: the HTTP instrumentation captures inbound and outbound calls and generates a span per request with a working trace. Point the OTLP exporter at your backend and you have end-to-end traces with no manual spans.
  4. Instrument the critical path manually. Add spans to the calls that matter and are invisible to auto-instrumentation: external APIs, background jobs, the payment flow, the cache. Add one histogram for request latency and one counter for errors on the two or three endpoints that carry your business.
  5. Ship the golden-signal dashboards. RED for the HTTP layer, USE for your dependencies. Now the SRE framework has a home.
  6. Add tail-based sampling, then set storage limits. Decide what you keep, set a budget, and never sample away errors.
  7. Only then, build alerts — from the metrics you trust, with thresholds derived from real data, not vibes.

The principle behind the order: every step produces something useful on its own and sets up the next. Structured logging fixes today's incidents; tracing fixes next quarter's.

Common pitfalls

High-cardinality tags. This is the single most common way to kill a metrics backend. Keep label sets bounded.

Unbounded logging. Logging every request body and every internal debug value at info level. The cost is not storage — it is that the signal gets buried. Log at debug what you might want during an incident, flip the level at runtime, and keep info for events with a decision attached. Nothing should default to "log everything always."

Sampling that loses errors. If your trace sampling drops 99% of traces at the head, the 0.1% incident is likely sampled away too. Sample to keep errors and unusual latencies, and drop the boring successful requests. Your sampling configuration is a reliability decision, not a storage decision.

Alert fatigue. Alerting on everything, or on low-signal conditions like "CPU above 50%", guarantees that real pages get ignored. Every alert should encode a user-visible consequence or an imminent one: a latency breach for a named SLO, a rising error rate, saturation near exhaustion. Everything else is a dashboard, not a page. And when an alert fires, it should point to a runbook or at least to the right dashboard and trace search.

Treating the three pillars as three projects. Teams that stand up logs, metrics, and tracing separately end up with three backends, three taxonomies, and no way to cross the joins. Ship one pipeline that carries context — trace ID in logs, metrics from spans — even if that means starting smaller.

Monitoring without action. A dashboard nobody looks at and an alert nobody pages on is decoration. Observability earns its keep only when an incident that would have taken an hour takes five minutes. Keep a small feedback loop: every postmortem should ask "did our telemetry actually contain the answer, and was it findable?" If the answer is no, that is the highest-priority improvement.

Where to start tomorrow

The shortest useful path: put a structured logger on every service with a correlation ID on every line, turn on OpenTelemetry's HTTP auto-instrumentation pointed at any OTLP-capable backend, add one histogram and one error counter to your main endpoint, and build a RED dashboard plus alerts from those two metrics. That is roughly a day of work, and it converts your system from "monitored" to "observable" for the questions that actually come up. Everything else — manual spans, tail sampling, derived metrics, per-tenant slicing — is additive on top of that spine.

The measure of success is not the number of dashboards. It is whether, at 2 a.m. with a degraded service, you can answer why — and whether the system you built can answer questions you have not asked it yet.