Fractal Techware

Prometheus alert runbooks

Mid-incident and staring at an unfamiliar alert name? Each page below explains what a common Prometheus alert means in plain words, what usually causes it, and the first commands to run (kubectl, PromQL, shell) to find out what is actually wrong.

The alerts come from the Prometheus Alert Rules & Runbook Pack: 179 alerts across 20 domains, every one covered by promtool unit tests. 12 of them, with their rules, tests and full runbooks, are free under MIT in prometheus-alert-rules on GitHub; those pages are marked free rule.

Kubernetes workloads

Kubernetes nodes & capacity

Hosts (node_exporter)

Kubernetes persistent volumes

Kubernetes Jobs & CronJobs

Kubernetes autoscaling (HPA)

Kubernetes quotas, limits & disruption budgets

Kubernetes control plane & kubelet

Kubernetes certificates (kubelet & cert-manager)

Prometheus self-monitoring

Alertmanager self-monitoring

Endpoint probes (blackbox_exporter)

PostgreSQL (postgres_exporter)

Redis (redis_exporter)

NGINX Ingress Controller

Apache Kafka (kafka_exporter)

Grafana Loki

CoreDNS

etcd

HTTP availability SLO (multi-window, multi-burn-rate)