DevOps Prometheus and Grafana

Prometheus collects numbers about your systems, and Grafana turns those numbers into charts. Teams use the pair to watch servers, containers, and applications. Both tools are free and open source, and both appear in most Kubernetes setups.

The Two Roles in Plain Words

Think of a hospital ward. A nurse checks each patient every few minutes and writes the readings in a chart. A large screen at the nurse station shows those readings as easy graphs. Prometheus plays the nurse who collects and stores the readings. Grafana plays the screen that displays them.

How the Pieces Connect

 [ App + /metrics ]  \
 [ Server exporter ]  ---> Prometheus ---> Grafana (dashboards)
 [ Database exporter]/        |
                              v
                        Alertmanager ---> Email / Chat

Prometheus uses a pull model. The server visits each target at a fixed interval and reads a page called /metrics. Each visit is called a scrape.

What Is a Metric?

A metric is a number that changes over time, such as CPU use, memory, or the count of web requests. Prometheus stores every metric with a timestamp and a set of labels. Labels describe the source, such as the server name or the page address.

Metric Types

TypeBehaviorExample
CounterOnly goes upTotal requests served
GaugeGoes up and downCurrent memory use
HistogramCounts values in rangesResponse time buckets

Exporters

Many tools do not publish metrics by themselves. An exporter is a small helper program that reads data from a tool and exposes it in the Prometheus format. The Node Exporter reports server CPU, memory, and disk. Other exporters cover databases, web servers, and message queues.

Basic Configuration

global:
  scrape_interval: 15s

scrape_configs:
  - job_name: "web-servers"
    static_configs:
      - targets: ["server1:9100", "server2:9100"]

The file tells Prometheus to scrape two servers every 15 seconds. Each target runs a Node Exporter on port 9100.

PromQL Queries

PromQL is the query language of Prometheus. A few short queries cover most daily needs.

# Requests per second over the last 5 minutes
rate(http_requests_total[5m])

# Memory used by each server
node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes

# Percentage of requests that failed
sum(rate(http_requests_total{status="500"}[5m]))
  / sum(rate(http_requests_total[5m])) * 100

The rate function converts a growing counter into a speed. The curly braces filter by label.

Alerting

An alert rule watches a query and fires when the result stays bad for a set time.

groups:
  - name: basic-alerts
    rules:
      - alert: HighErrorRate
        expr: rate(http_requests_total{status="500"}[5m]) > 5
        for: 10m
        labels:
          severity: critical
        annotations:
          summary: "Web errors are too high"

Alertmanager receives fired alerts, groups similar ones, and sends them to email, chat, or paging tools.

Grafana Dashboards

Grafana connects to Prometheus as a data source. A dashboard holds panels, and each panel shows the result of one query as a line chart, gauge, table, or number. Good dashboards follow a few rules:

  • Place the most important numbers at the top.
  • Show traffic, errors, and response time for every service.
  • Limit each dashboard to one clear purpose.
  • Name panels with plain words and units.

Service Discovery

Static target lists work for a few servers. Containers and cloud machines appear and disappear all day, so a fixed list goes stale. Service discovery lets Prometheus ask Kubernetes, AWS, or Consul for the current list of targets and scrape each new one automatically.

scrape_configs:
  - job_name: "kubernetes-pods"
    kubernetes_sd_configs:
      - role: pod

Recording Rules

Some queries read huge amounts of data and slow down dashboards. A recording rule runs the query on a schedule and saves the result as a new metric. Dashboards read the small saved metric instead of recalculating it.

groups:
  - name: speed
    rules:
      - record: job:http_requests:rate5m
        expr: sum(rate(http_requests_total[5m])) by (job)

What to Measure: Golden Signals

Google's site reliability teams popularized four signals that describe the health of almost any service.

SignalQuestion It Answers
LatencyHow long do requests take?
TrafficHow many requests arrive?
ErrorsHow many requests fail?
SaturationHow full are the resources?

The RED method suits services: Rate, Errors, and Duration. The USE method suits machines: Utilization, Saturation, and Errors.

Retention and Long-Term Storage

Prometheus stores data on local disk and keeps it for 15 days by default. Longer history needs extra tools. Thanos, Mimir, Cortex, and VictoriaMetrics receive a copy of the data through the remote_write setting and store it in cheap object storage. Teams use that history to study trends over months and plan capacity.

 Prometheus --remote_write--> Long-term store --> Grafana (monthly charts)
   (15 days local)              (months or years)

Pushing Short Jobs

Batch jobs may finish before Prometheus visits them. The Pushgateway accepts metrics from such jobs and holds them until the next scrape. Use it only for short-lived jobs and not for regular services.

Key Points

  • Prometheus scrapes, stores, and queries metrics.
  • Grafana draws dashboards from those metrics.
  • Exporters bring metrics from tools that lack built-in support.
  • Alertmanager delivers alerts to the right people.

Leave a Comment

Your email address will not be published. Required fields are marked *