DevOps Prometheus and Grafana
Prometheus collects numbers about your systems, and Grafana turns those numbers into charts. Teams use the pair to watch servers, containers, and applications. Both tools are free and open source, and both appear in most Kubernetes setups.
The Two Roles in Plain Words
Think of a hospital ward. A nurse checks each patient every few minutes and writes the readings in a chart. A large screen at the nurse station shows those readings as easy graphs. Prometheus plays the nurse who collects and stores the readings. Grafana plays the screen that displays them.
How the Pieces Connect
[ App + /metrics ] \
[ Server exporter ] ---> Prometheus ---> Grafana (dashboards)
[ Database exporter]/ |
v
Alertmanager ---> Email / Chat
Prometheus uses a pull model. The server visits each target at a fixed interval and reads a page called /metrics. Each visit is called a scrape.
What Is a Metric?
A metric is a number that changes over time, such as CPU use, memory, or the count of web requests. Prometheus stores every metric with a timestamp and a set of labels. Labels describe the source, such as the server name or the page address.
Metric Types
| Type | Behavior | Example |
|---|---|---|
| Counter | Only goes up | Total requests served |
| Gauge | Goes up and down | Current memory use |
| Histogram | Counts values in ranges | Response time buckets |
Exporters
Many tools do not publish metrics by themselves. An exporter is a small helper program that reads data from a tool and exposes it in the Prometheus format. The Node Exporter reports server CPU, memory, and disk. Other exporters cover databases, web servers, and message queues.
Basic Configuration
global:
scrape_interval: 15s
scrape_configs:
- job_name: "web-servers"
static_configs:
- targets: ["server1:9100", "server2:9100"]The file tells Prometheus to scrape two servers every 15 seconds. Each target runs a Node Exporter on port 9100.
PromQL Queries
PromQL is the query language of Prometheus. A few short queries cover most daily needs.
# Requests per second over the last 5 minutes
rate(http_requests_total[5m])
# Memory used by each server
node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes
# Percentage of requests that failed
sum(rate(http_requests_total{status="500"}[5m]))
/ sum(rate(http_requests_total[5m])) * 100The rate function converts a growing counter into a speed. The curly braces filter by label.
Alerting
An alert rule watches a query and fires when the result stays bad for a set time.
groups:
- name: basic-alerts
rules:
- alert: HighErrorRate
expr: rate(http_requests_total{status="500"}[5m]) > 5
for: 10m
labels:
severity: critical
annotations:
summary: "Web errors are too high"Alertmanager receives fired alerts, groups similar ones, and sends them to email, chat, or paging tools.
Grafana Dashboards
Grafana connects to Prometheus as a data source. A dashboard holds panels, and each panel shows the result of one query as a line chart, gauge, table, or number. Good dashboards follow a few rules:
- Place the most important numbers at the top.
- Show traffic, errors, and response time for every service.
- Limit each dashboard to one clear purpose.
- Name panels with plain words and units.
Service Discovery
Static target lists work for a few servers. Containers and cloud machines appear and disappear all day, so a fixed list goes stale. Service discovery lets Prometheus ask Kubernetes, AWS, or Consul for the current list of targets and scrape each new one automatically.
scrape_configs:
- job_name: "kubernetes-pods"
kubernetes_sd_configs:
- role: podRecording Rules
Some queries read huge amounts of data and slow down dashboards. A recording rule runs the query on a schedule and saves the result as a new metric. Dashboards read the small saved metric instead of recalculating it.
groups:
- name: speed
rules:
- record: job:http_requests:rate5m
expr: sum(rate(http_requests_total[5m])) by (job)What to Measure: Golden Signals
Google's site reliability teams popularized four signals that describe the health of almost any service.
| Signal | Question It Answers |
|---|---|
| Latency | How long do requests take? |
| Traffic | How many requests arrive? |
| Errors | How many requests fail? |
| Saturation | How full are the resources? |
The RED method suits services: Rate, Errors, and Duration. The USE method suits machines: Utilization, Saturation, and Errors.
Retention and Long-Term Storage
Prometheus stores data on local disk and keeps it for 15 days by default. Longer history needs extra tools. Thanos, Mimir, Cortex, and VictoriaMetrics receive a copy of the data through the remote_write setting and store it in cheap object storage. Teams use that history to study trends over months and plan capacity.
Prometheus --remote_write--> Long-term store --> Grafana (monthly charts) (15 days local) (months or years)
Pushing Short Jobs
Batch jobs may finish before Prometheus visits them. The Pushgateway accepts metrics from such jobs and holds them until the next scrape. Use it only for short-lived jobs and not for regular services.
Key Points
- Prometheus scrapes, stores, and queries metrics.
- Grafana draws dashboards from those metrics.
- Exporters bring metrics from tools that lack built-in support.
- Alertmanager delivers alerts to the right people.
