All guides

Filter by category, difficulty, or free text to find the right material for your team.

Metrics~40 min

Monitor Gateway API TCPRoute and UDPRoute without cardinality sprawl

Created: August 23, 2026 · Published: August 23, 2026

A hands-on guide for Prometheus/Grafana dashboards and alerts on TCPRoute and UDPRoute: inventory, PromQL, metric relabeling, canary validation, and proof that L4 traffic is neither blind nor expensive.

LinuxDocker
Intermediate
Read guide
Metrics~45 min

Keep target churn and scrape gaps from blinding your burn-rate alerts

Created: August 17, 2026 · Published: August 17, 2026

Harden SLO burn-rate alerts against pod churn, scrape gaps, unstable endpoints, and no-data in Prometheus with PromQL, promtool, and measurable canaries.

LinuxDocker
Advanced
Read guide
Metrics~40 min

Cut kube-state-metrics CustomResourceState noise without breaking alerts

Created: August 13, 2026 · Published: August 13, 2026

A practical guide to reducing kube-state-metrics and CustomResourceState cardinality with allowlists, metricRelabelings, impact PromQL, and rollback gates.

LinuxDocker
Intermediate
Read guide
Metrics~40 min

Roll out Prometheus native histograms without blinding your SLOs

Created: August 11, 2026 · Published: August 11, 2026

Learn how to enable native histograms with dual publishing, parity PromQL, cardinality limits, and remote_write guardrails so p95, p99, and burn alerts remain trustworthy.

LinuxDocker
Advanced
Read guide
Metrics~35 min

Manage golden-signal dashboards as code without trusting drifted panels

Created: August 5, 2026 · Published: August 5, 2026

Turn critical dashboards into reviewable artifacts: Terraform, promtool tests, bounded Prometheus probes, and a visual canary before production.

Linux
Intermediate
Read guide
Metrics~45 min

Control HTTP route cardinality in Prometheus before Kubernetes API golden signals lie

Created: August 1, 2026 · Published: August 1, 2026

Learn how to find explosive HTTP labels, normalize routes, protect remote_write, and validate dashboards and alerts before publishing Kubernetes API golden signals.

Advanced
Read guide
Metrics~40 min

Controller-runtime golden signals: catch cache lag before blaming the API server

Created: July 30, 2026 · Published: July 30, 2026

Learn how to instrument and diagnose Kubernetes controllers with controller-runtime, workqueue, cache, and REST-client metrics before tuning concurrency or restarting pods blindly.

Linux
Advanced
Read guide
Reliability~35 min

Triage SLO burn-rate alerts from Prometheus to traces without inventing culprits

Created: July 30, 2026 · Published: July 30, 2026

Learn how to design multi-window alerts, validate telemetry health, and pivot from PromQL to traces when investigating burn rate without chasing ghosts.

Linux
Intermediate
Read guide
Metrics~40 min

Clean kube-state-metrics noise before your dashboards drift

Created: July 21, 2026 · Published: July 21, 2026

Learn how to audit kube-state-metrics cardinality, design safe allowlists, validate alert parity, and roll out changes without breaking Kubernetes dashboards.

Linux
Intermediate
Read guide
Reliability~45 min

Design multi-window SLO burn-rate alerts without paging people for noise

Created: July 11, 2026 · Published: July 11, 2026

Learn how to define SLIs, error budget, PromQL, and promtool tests for burn-rate alerts that catch fast incidents without turning every short spike into a broken on-call shift.

Linux
Intermediate
Read guide
Metrics~40 min

Build golden signals for Kubernetes HTTP APIs without cardinality sprawl

Created: July 7, 2026 · Published: July 7, 2026

Define a label contract, create RED/USE recording rules, validate cardinality with promtool, and roll out Kubernetes API dashboards without turning every route into a series factory.

LinuxDocker
Intermediate
Read guide
Metrics~35 min

Clean up kube-state-metrics noise before device metrics trigger Prometheus cardinality spikes

Created: June 26, 2026 · Published: June 26, 2026

Find which kube-state-metrics families inflate Prometheus, prune volatile labels with allowlists, and validate alerts and dashboards before rolling the change out.

Linux
Intermediate
Read guide
Reliability~35 min

Design burn-rate alerts with Kubernetes PSI before latency pages anyone

Created: June 24, 2026 · Published: June 24, 2026

A practical guide to turning burn-rate alerts into actionable signals: clear SLOs, multiple windows, PSI as context, promtool tests, and gradual rollout.

Linux
Intermediate
Read guide
Metrics~35 min

Diagnose Kubernetes API WATCH/LIST pressure before blaming etcd

Created: June 23, 2026 · Published: June 23, 2026

An actionable guide for detecting Kubernetes API WATCH/LIST pressure by combining API server, etcd, Prometheus, and client validation before timeouts spread.

Linux
Advanced
Read guide
Metrics~35 min

Use Kubernetes PSI metrics to detect real API saturation without noisy paging

Created: June 12, 2026 · Published: June 12, 2026

A practical guide to adding PSI to Kubernetes API dashboards, correlating CPU/memory/IO pressure with SLIs, and designing alerts that page only on real impact.

Linux
Intermediate
Read guide
Reliability~35 min

Design SLO burn-rate alerts that page on real impact, not noise

Created: June 11, 2026 · Published: June 11, 2026

Learn how to define a useful SLI, combine multi-window burn alerts, control cardinality, and validate that an SLO page fires only on real impact.

Linux
Intermediate
Read guide
Metrics~35 min

Clean up kube-state-metrics noise before your dashboards and alerts start lying

Created: June 10, 2026 · Published: June 10, 2026

kube-state-metrics can turn Kubernetes state into a storm of series, labels, and dashboards that look precise but do not help. This guide shows how to measure the noise, remove unstable labels, keep critical signals, and validate Prometheus before touching production.

LinuxDocker
Intermediate
Read guide
Metrics~40 min

Understand Kubernetes memory metrics without firing false OOM alerts

Created: April 26, 2026 · Published: April 26, 2026

A practical guide to diagnosing Kubernetes container memory with Prometheus and Grafana without confusing usage, working set, RSS, or reclaimable page cache.

LinuxDocker
Intermediate
Read guide
Metrics~35 min

Build useful golden signals for Kubernetes APIs without triggering Prometheus cardinality traps

Created: April 26, 2026 · Published: April 26, 2026

A practical guide to traffic, latency, errors, and saturation for Kubernetes APIs without filling Prometheus with useless series or breaking alerts.

LinuxDocker
Intermediate
Read guide
Reliability~36 min

Design burn rate alerts that do not wake people up for sport

Created: April 22, 2026 · Published: April 22, 2026

Recent Prometheus, OpenTelemetry Collector, Loki, and Alloy releases all point to the same uncomfortable truth: alerts wired straight to raw metrics and fragile labels become noisy or broken very easily. This guide shows how to anchor burn-rate alerts on stable recording rules, validate them with promtool, and roll them out without turning every short spike into a fake emergency.

LinuxDocker
Intermediate
Read guide
Metrics~38 min

Clean up kube-state-metrics noise so your dashboards mean something again

Created: April 20, 2026 · Published: April 20, 2026

kube-state-metrics is still valuable, but in 2026 it exposes more surface area, more stable metrics, and recent defaults such as EndpointSlices. If your dashboards filled up with irrelevant series, fragile joins, or duplicated states, this guide shows how to reduce noise at the source, fix your queries, and validate that the cleanup does not break alerts or troubleshooting.

LinuxDocker
Intermediate
Read guide
Metrics~45 min

Practical golden signals for APIs on Kubernetes without inflating the stack

Created: April 15, 2026 · Published: April 15, 2026

A technical, actionable guide to map golden signals to Prometheus metrics and PromQL, build Grafana panels, create alerts and follow a reproducible troubleshooting flow for Kubernetes APIs without adding unnecessary agents.

LinuxDocker
Intermediate
Read guide
Dashboards~20 min

Grafana to unify metrics, logs, and traces (cross-platform)

Created: April 7, 2026 · Published: April 7, 2026

Deploy Grafana, provision datasources, and keep one place to explore metrics, logs, and trace correlation.

Docker
Intermediate
Read guide