What is Prometheus? Monitoring and Alerting Explained
You cannot operate what you cannot measure. When a production service slows down at 2 AM, the first question is always: what do the metrics look like? Prometheus gives you the instrumentation, collection, storage, and alerting stack to answer that question -- and because it is open-source and has become the industry standard for Kubernetes monitoring, understanding it is a core DevOps skill.
Prometheus was developed at SoundCloud in 2012 by ex-Google engineers who missed the Borgmon monitoring system. It was open-sourced in 2015 and became the second project to graduate from the CNCF (after Kubernetes itself) in 2018.
How Prometheus Collects Metrics
Prometheus uses a pull model. Rather than having applications push metrics to a central server, Prometheus scrapes HTTP endpoints that applications expose. By convention, this endpoint is GET /metrics.
An example /metrics response (Prometheus text format):
# HELP http_requests_total Total number of HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",status="200"} 1234
http_requests_total{method="POST",status="200"} 89
http_requests_total{method="GET",status="500"} 12
# HELP http_request_duration_seconds Request latency
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{le="0.05"} 980
http_request_duration_seconds_bucket{le="0.1"} 1100
http_request_duration_seconds_bucket{le="0.5"} 1230
http_request_duration_seconds_bucket{le="+Inf"} 1234
http_request_duration_seconds_sum 45.2
http_request_duration_seconds_count 1234
Prometheus client libraries for Go, Python, Node.js, Java, and others make exposing these endpoints straightforward. For infrastructure metrics (CPU, memory, disk, network), the Prometheus ecosystem provides exporters: node_exporter for Linux hosts, kube-state-metrics for Kubernetes object state, blackbox_exporter for endpoint probing.
The Four Metric Types
Counter: monotonically increasing value. Total requests, total errors, bytes transferred. Counters only go up (reset to 0 on restart). You query the rate of increase, not the raw value.
Gauge: value that can go up and down. Current memory usage, active connections, queue depth.
Histogram: samples observations and counts them in configurable buckets. Used for request latency and response size. Enables percentile calculations with histogram_quantile().
Summary: similar to histogram but calculates configurable quantiles on the client side. Less flexible for aggregation across instances -- prefer histograms in most cases.
PromQL: Querying Your Metrics
PromQL is the language you use to query metrics. It takes time to get fluent, but a handful of patterns cover most use cases.
# Per-second HTTP request rate over the last 5 minutes, by status code
rate(http_requests_total[5m])
# Error rate as a fraction of total traffic
rate(http_requests_total{status=~"5.."}[5m])
/
rate(http_requests_total[5m])
# 99th percentile request latency, aggregated across all instances
histogram_quantile(
0.99,
sum by (le) (rate(http_request_duration_seconds_bucket[5m]))
)
# Memory usage above 80% of the container limit
container_memory_usage_bytes
/
container_spec_memory_limit_bytes
> 0.8
The rate() function is fundamental. Never graph a counter directly -- it only ever goes up and resets to zero on restart. rate() converts it into a per-second rate, which is what you actually want.
Prometheus Configuration
Prometheus is configured with a YAML file that defines scrape targets, alerting rules, and recording rules:
global:
scrape_interval: 15s
evaluation_interval: 15s
rule_files:
- "alerts/*.yaml"
alerting:
alertmanagers:
- static_configs:
- targets: ["alertmanager:9093"]
scrape_configs:
- job_name: "kubernetes-pods"
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: "true"
In Kubernetes, the kube-prometheus-stack Helm chart (from the prometheus-community org) is the standard way to deploy the full observability stack: Prometheus, Alertmanager, Grafana, node_exporter, kube-state-metrics, and a set of pre-built dashboards and alerting rules. A 5-minute install with Helm:
helm repo add prometheus-community \
https://prometheus-community.github.io/helm-charts
helm repo update
helm install kube-prometheus-stack \
prometheus-community/kube-prometheus-stack \
--namespace monitoring \
--create-namespace \
--set grafana.adminPassword=changeme
Alerting with Alertmanager
Prometheus evaluates alerting rules on a fixed interval. When a rule fires, it sends the alert to Alertmanager, which handles routing, grouping, deduplication, and silencing. Alertmanager routes alerts to PagerDuty, Slack, email, or webhook endpoints.
A practical example alert:
groups:
- name: api-server
rules:
- alert: HighErrorRate
expr: |
rate(http_requests_total{status=~"5.."}[5m])
/
rate(http_requests_total[5m])
> 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "High error rate on {{ $labels.instance }}"
description: "Error rate is {{ $value | humanizePercentage }}"
The for: 2m clause requires the condition to be true for 2 consecutive minutes before the alert fires -- this prevents spurious alerts from brief spikes.
Prometheus in the Observability Stack
Prometheus handles the metrics pillar of observability. For logs, most teams use Loki (also from Grafana Labs) or the ELK stack. For distributed tracing, Jaeger or Tempo. Grafana ties all three together in dashboards. The combination of Prometheus, Loki, and Tempo with Grafana is called the PLG stack and is increasingly common in Kubernetes environments.
For the broader context of where monitoring fits in a DevOps practice, see the DevOps tools guide.
