What is Observability? Monitoring vs Observability Explained
The term observability comes from control theory. A system is observable if you can determine its internal state from its external outputs. For software systems, the practical translation is: can you understand what is happening in production without guessing, and can you answer questions you did not think to ask before the incident started?
That second part is what separates observability from monitoring.
Monitoring vs Observability
Monitoring is checking known things. You define a set of metrics and thresholds, alert when values cross those thresholds, and build dashboards to watch the ones you care about. This works well for failure modes you have seen before. It fails when something new goes wrong -- a query pattern you did not expect, a dependency you did not model, a resource contention that only appears at a specific combination of load and concurrency.
Observability is the property that lets you answer novel questions. An observable system emits enough data, in enough detail, that you can reconstruct what happened after the fact -- even for failures nobody predicted.
This is not an either/or. Production systems need both. Monitoring with SLOs and alerts gives you the pager going off. Observability gives you the tools to figure out what happened once it does.
The Three Pillars
The observability community settled on three fundamental signal types:
Logs
Logs are timestamped records of discrete events. They are the most familiar signal type -- almost every application has been logging since before "observability" was a word. The problem is that unstructured logs are expensive to search at scale and hard to correlate across services.
Structured logging -- emitting JSON instead of free-form strings -- dramatically improves the situation. Every log entry carries consistent fields (service name, trace ID, user ID, severity) that make filtering fast and cross-service correlation possible.
{
"timestamp": "2026-09-01T14:23:11Z",
"level": "error",
"service": "payments-api",
"trace_id": "7c3b9f2a1e4d8c6b",
"user_id": "usr_8823",
"message": "payment processor timeout",
"duration_ms": 5002,
"processor": "stripe"
}
The trace_id field is what links this log entry to the distributed trace for the same request.
Metrics
Metrics are numerical measurements aggregated over time. They are compact, cheap to store at scale, and ideal for alerting. Common metric types are counters (requests total), gauges (current memory usage), and histograms (request duration distribution).
Metrics answer questions like "what fraction of requests are failing?" and "is this service's latency getting worse over time?" They are poor at answering "why?" -- for that, you need logs and traces.
Prometheus is the dominant metrics system in cloud-native environments. Applications expose a /metrics endpoint, Prometheus scrapes it on an interval, and Grafana visualises the results.
Traces
Distributed traces record the path of a single request as it travels through multiple services. Each service adds a span -- a timed segment with metadata -- to a trace that was started when the request entered the system. Together, the spans form a tree showing where time was spent and where errors occurred.
Traces answer questions that metrics and logs cannot: "which service in this 12-hop request chain is adding 400ms of latency?" and "where exactly does this cascade of errors begin?"
OpenTelemetry: The Instrumentation Standard
For years, the observability space was fragmented. Every vendor had its own instrumentation SDK, agent, and wire format. Adding Datadog meant one SDK; switching to Honeycomb meant rewriting your instrumentation.
OpenTelemetry (OTel) is the answer. It is a CNCF project providing vendor-neutral APIs, SDKs, and a collector for generating and exporting all three signal types. You instrument your code once against the OTel API and configure where the data goes separately -- Prometheus, Jaeger, Datadog, Grafana Tempo, or your own backend.
The practical upside: vendor lock-in for instrumentation is largely gone. You can change backends without touching application code. See the devops tools guide for a comparison of backend options.
Connecting the Signals
The real power of observability comes from correlation across signal types. Modern tooling makes this possible:
- An alert fires on high error rate (metrics)
- You jump to the Grafana dashboard and see the error spike started 12 minutes ago
- You navigate to the logs for the affected service and filter by
level=error - Each log entry carries a
trace_id-- you click through to the trace - The trace shows that 90% of latency is in a downstream database call
- You check the database metrics and see connection pool saturation
That sequence took minutes, not hours, because the signals are correlated by trace ID and the tooling lets you navigate between them. This is what practitioners mean when they say a system is observable.
Without proper instrumentation, CI/CD pipelines can deploy broken code that only manifests under production load. Observability is what closes that loop. For the full monitoring stack picture, the infrastructure as code article covers how monitoring infrastructure is provisioned alongside the services it watches.
