myibrahim.cloud

Backend · Operations

Observability: traces beat logs

Logs find what you knew to look for. Traces find what you didn't. The pragmatic case for OpenTelemetry as the foundation of production debugging.

For years my debugging workflow on production incidents was: SSH to a box, tail -f, grep until I found the line that explained what was happening. It worked, until services multiplied. By the time you have eight microservices and one user request fans across all of them, log-grep stops scaling.

Traces fix this. This is the playbook.

The three pillars (and why you need all three)#

The "three pillars of observability" — logs, metrics, traces — get cited so often the phrase is meaningless. Here's what each is actually for:

  • Metrics answer "how much is happening?" — request rate, error rate, p99 latency, queue depth. Aggregate. Cheap to retain. The signal you watch on dashboards.
  • Logs answer "what specifically happened in this one event?" — context-dense, structured (ideally), expensive at scale. The signal you read when investigating a known failure.
  • Traces answer "how did this one request flow through the system?" — span tree from frontend → service-A → service-B → DB. Newest of the three; most powerful for distributed debugging.

You need all three because they answer different questions. Stop arguing about which "wins."

OpenTelemetry: a 5-min intro

Why traces are the unlock#

Imagine: customer reports "my dashboard is slow." Without traces, you check service-A latency (fine), service-B latency (fine), DB latency (fine). Everything looks fine because each service's average is fine. The customer's specific request was the one that took 4 seconds, and you can't find it.

With traces, you find the request by customer_id and see the span tree. service-A → service-B → service-C (slow! 3.8s in service-C). Click into service-C's span — its child span shows a DB query that took 3.5s. Click into the DB span — query text + parameters. Done.

Logs would have found the same info if you knew exactly which service to look at and which time window. Traces find it without that prior knowledge.

OpenTelemetry: the standard#

OpenTelemetry (OTel) is the open standard for instrumentation. It's a vendor-neutral format for emitting traces, metrics, and logs. Honeycomb, Datadog, New Relic, Tempo, Jaeger — all consume OTel-format data. This means: instrument once, swap backends.

Adopt OTel. It's the closest thing to a sure bet in this space.

Minimum viable trace instrumentation#

In Python with auto-instrumentation:

pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap --action=install
opentelemetry-instrument \
  --traces_exporter otlp \
  --metrics_exporter otlp \
  --service_name my-service \
  python app.py

That's it. With auto-instrumentation, every Flask/Django request, every requests/httpx call, every SQLAlchemy query, every Redis op gets a span automatically. You haven't written any instrumentation code.

For Node:

const { NodeSDK } = require("@opentelemetry/sdk-node");
const { getNodeAutoInstrumentations } = require("@opentelemetry/auto-instrumentations-node");

new NodeSDK({
  serviceName: "my-service",
  instrumentations: [getNodeAutoInstrumentations()],
}).start();

Same idea. Auto-instrumentation covers express, http, fetch, pg, redis, etc.

Custom spans for business logic#

Auto-instrumentation gives you the framework layer. For business logic that matters, add custom spans:

from opentelemetry import trace
tracer = trace.get_tracer(__name__)

def process_payment(order_id: int):
    with tracer.start_as_current_span("process_payment") as span:
        span.set_attribute("order.id", order_id)
        span.set_attribute("user.id", current_user.id)

        try:
            charge_card(order_id)
            span.set_attribute("payment.outcome", "success")
        except Exception as e:
            span.record_exception(e)
            span.set_attribute("payment.outcome", "failed")
            raise

The spans you add show up in the trace tree alongside the auto-generated ones. order.id and user.id become searchable attributes — you can pull up "all traces for user 42" or "all failed payments in the last hour" instantly.

Log enrichment with trace context#

Logs aren't replaced by traces; they're correlated with traces. Inject the trace_id into every log line:

import logging
from opentelemetry import trace

class TraceFormatter(logging.Formatter):
    def format(self, record):
        span = trace.get_current_span()
        ctx = span.get_span_context() if span else None
        record.trace_id = format(ctx.trace_id, "032x") if ctx else "-"
        record.span_id = format(ctx.span_id, "016x") if ctx else "-"
        return super().format(record)

handler = logging.StreamHandler()
handler.setFormatter(TraceFormatter("%(asctime)s [%(trace_id)s] %(message)s"))

Now your logs have trace_id on every line. Find a slow trace; pivot to its logs; pivot back. The combined view is what wins on incidents.

Metrics that matter#

The "RED" method (Rate, Errors, Duration) covers most service health:

  • Rate: requests per second. Catch traffic shifts.
  • Errors: % of requests that errored. Catch regressions.
  • Duration: p50, p95, p99 latency. p99 is what users feel.

For systems with queues, add the "USE" method (Utilization, Saturation, Errors): how busy is the resource, how queued is it, how often is it failing.

Don't track averages. Always p50/p95/p99. Averages hide tail behavior.

Sampling: keep traces affordable#

Tracing every request gets expensive at scale. Sampling strategies:

  • Head-based: decide at the start of the request whether to keep its trace. Cheap. Stupid (you might drop the slow ones).
  • Tail-based: keep all spans in memory briefly, decide at the end. Smart (always keeps slow/error traces). Expensive.
  • Adaptive: sample 100% of errors, 10% of normal requests, 100% of requests > p95 latency.

OTel Collector supports tail sampling via the tail_sampling processor. For most teams: head-sample at 1-10% normal, with overrides for error/slow traces.

What to display on the dashboard#

The dashboard should answer "is anything wrong?" in under 10 seconds. My standard layout:

  1. Top row: rate, error %, p95 latency for each service.
  2. Middle: top 5 slowest endpoints (by p95).
  3. Bottom: recent error count + link to traces.

Anything else (CPU, memory, DB connections) goes on a secondary "system health" dashboard. The primary dashboard shows what users feel.

The bar to clear#

When an incident hits at 2am, can your on-call engineer:

  1. See within 30 seconds which service is degraded? (metrics)
  2. Find a representative slow/failed request within 2 minutes? (traces, by error or duration)
  3. See the request's full path through your system? (trace tree)
  4. Pull the relevant logs for that request? (log correlation by trace_id)
  5. Identify the slow span / error point within 5 minutes total? (good span attributes)

If yes, your observability is good. If any of those takes longer, fix that gap before adding more dashboards.

Further reading#

  • observability
  • tracing
  • opentelemetry
  • logs
  • metrics
  • production
Need this built? I build full-stack web app or saas projects for clients worldwide. Tell me about yours.