Grafana Heatmaps for Job Duration
A p95 line tells you that some jobs are slow; it does not tell you whether all jobs got a bit slower, a separate group of jobs appeared, or a few outliers dragged the percentile up. A heatmap shows the whole distribution of durations over time, and those three situations look completely different on it. This guide builds one from a Prometheus histogram and explains how to read it, as part of Grafana Dashboards for Queues in Observability & Monitoring for Job Queues.
Problem Statement
A report-generation queue has a p95 duration alert at 30 seconds. It started firing intermittently after a release, but the p95 line looks noisy rather than clearly worse, and averages barely changed. Engineers argued for a day about whether the release caused it. When someone finally plotted a heatmap, it showed two distinct bands: most reports still took 2–4 seconds, and a new band appeared at 40–60 seconds — reports for large accounts that now included a year of data instead of a month. You want heatmaps on the dashboard from the start, built correctly, so the shape of the problem is visible immediately.
Prerequisites
- A Prometheus histogram for job duration with sensible buckets — for example
bullmq_job_duration_seconds,celery_task_runtime, or your own. See choosing histogram buckets for job duration. - Grafana 10 or newer (the heatmap panel was rebuilt in version 9 and handles Prometheus histograms natively).
- Labels for queue and job name so heatmaps can be filtered.
Step 1 — Understand What a Heatmap Shows
A heatmap has time on the x-axis, duration buckets on the y-axis, and colour for how many jobs fell into each bucket in each time slice. Each column is a histogram turned on its side; reading left to right shows how the distribution changes.
Three patterns cover most of what you will see: the whole band moving up (everything got slower — a dependency or resource problem), a new band appearing (a subset of jobs behaves differently — data size, a code path, a tenant), and scattered dots far above the band (occasional outliers — timeouts, retries, lock waits).
Step 2 — Write the Query Correctly
Grafana needs the bucket counts per time slice, not cumulative totals. Query the rate of the _bucket series, summed by le, and set the query format to Heatmap:
sum by (le) (
increase(celery_task_runtime_bucket{name=~"$task", queue_name=~"$queue"}[$__rate_interval])
)
In the query options, set Format to Heatmap and Legend to {{le}}. Grafana then converts Prometheus's cumulative buckets (le="1" includes everything below one second) into per-bucket counts. If you forget the Heatmap format, the panel shows cumulative counts and the whole chart is dominated by the top bucket.
increase shows the number of jobs per slice; rate shows jobs per second. Either works — increase is easier to read as "how many jobs took this long."
Step 3 — Configure the Panel
In the heatmap panel settings:
- Calculate from data: No — the data is already bucketed by Prometheus.
- Y axis: unit seconds; scale logarithmic if your buckets are exponential (1, 2, 5, 10, 30…), which they usually are.
- Y bucket layout: "Upper" (Prometheus
leis an upper bound). - Colour: mode "Scheme", scale "Exponential" so a few slow jobs are still visible next to thousands of fast ones. A single-hue scheme (blues) reads better than a rainbow.
- Cell gap: 1 pixel, so individual cells are distinguishable.
- Tooltip: show histogram, so hovering over a column shows its full distribution.
The exponential colour scale is the setting that makes heatmaps useful for job queues, where the interesting jobs are rare. Without it, the panel shows only the busiest band.
Step 4 — Put Heatmaps Where They Help
One heatmap per job type is more useful than one for the whole queue, because different job types have different normal durations and a queue-wide heatmap blurs them together. Use a repeating panel over a $task variable limited to the top job types by volume, or a single heatmap with the task selector next to it.
Place the heatmap next to the p95 line rather than instead of it: the line is what alerts use and what trends show over weeks; the heatmap is what you open when the line moves. Link the p95 alert to the dashboard with the task variable set, so the person investigating lands on the matching heatmap.
A third companion panel is often more useful than either: the share of jobs slower than a threshold, computed straight from the same buckets. It turns "a band appeared at 40–60 seconds" into a number you can track and alert on:
1 - (
sum(rate(celery_task_runtime_bucket{name=~"$task", le="30"}[10m]))
/ sum(rate(celery_task_runtime_count{name=~"$task"}[10m]))
)
Unlike a percentile, this fraction does not jump between bands when the mix of jobs changes; it rises smoothly as more jobs cross the threshold, and it maps directly onto an objective such as "99% of reports finish within 30 seconds." Choose the threshold to match an existing bucket boundary so the result is exact rather than interpolated.
Step 5 — Link Slow Cells to Traces with Exemplars
A heatmap shows that slow jobs exist; exemplars let you click through to one. If your instrumentation attaches trace IDs as exemplars when observing the histogram, Grafana can overlay them as dots on the heatmap and open the trace in Tempo or Jaeger.
from prometheus_client import Histogram
from opentelemetry import trace
DURATION = Histogram("job_duration_seconds", "Job duration", ["task"])
def observe(task_name: str, seconds: float) -> None:
ctx = trace.get_current_span().get_span_context()
exemplar = {"trace_id": format(ctx.trace_id, "032x")} if ctx.is_valid else None
DURATION.labels(task_name).observe(seconds, exemplar=exemplar)
Exemplars require Prometheus's exemplar storage (--enable-feature=exemplar-storage) and the OpenMetrics exposition format. In the panel, enable the exemplars toggle on the query and configure the data source's internal link to your tracing backend. Clicking a dot in the slow band then shows exactly which job it was, with its spans — the fastest way to answer "what are these slow ones?" See OpenTelemetry tracing for BullMQ for producing the traces.
Step 6 — Recognise the Common Patterns
With a heatmap in front of you, most investigations start by naming the pattern:
For a shifted band, compare with dependency latency and worker CPU throttling. For a second band, split the histogram by a candidate label temporarily (account size tier, job variant) or use exemplars to inspect examples. For scattered outliers, check retry counts and timeout settings; outliers at exactly the timeout value are jobs being killed. The report queue from the problem statement was a textbook second band, and the fix — paginating large reports — became obvious once the band was visible.
Verification
- The heatmap shows a clear band at the typical duration for each job type, matching what the p95 panel suggests.
- A deliberately slowed test job appears as cells above the band within a minute or two.
- Switching the query format away from Heatmap visibly breaks the panel (confirming the format is what makes it correct).
- With exemplars enabled, clicking a dot opens the matching trace.
Gotchas & Edge Cases
Bucket boundaries define the resolution. A heatmap cannot show detail finer than your histogram buckets. If all jobs fall into one bucket, the heatmap is a single stripe; add buckets around the typical duration.
Changing buckets breaks history. Adding or removing bucket boundaries creates new le series; old and new data do not line up for a while. Change buckets rarely and annotate when you do.
Too many series. A heatmap sums over all labels except le. Filter by task and queue in the query, or a single noisy job type will dominate.
Empty slices. Slices with no jobs show as gaps. For low-volume queues, use a longer $__rate_interval or a larger minimum interval so each column has enough jobs to be meaningful.
FAQ
Can I build a heatmap from a summary instead of a histogram? No. Summaries expose precomputed quantiles, not bucket counts, and cannot be aggregated across processes. Use histograms for anything you want to visualise as a distribution.
Do native histograms change this? Prometheus native histograms have automatic, high-resolution buckets, and Grafana's heatmap supports them directly. They make heatmaps sharper without choosing boundaries by hand; the reading and panel settings are the same.
Should I alert on heatmap patterns? Alert on percentiles and on the fraction of jobs above a threshold (from the bucket counts). Use the heatmap to investigate, not to page.
How long a time range works best? Six to twenty-four hours for investigations, so the normal band is established before the change. Over weeks, columns become too narrow to read; use the p95 and slow-fraction panels for long-term trends.
Related
- Grafana Dashboards for Queues — shared panel patterns.
- Choosing Histogram Buckets for Job Duration — the buckets that make heatmaps readable.
- Building a Celery Grafana Dashboard — where the heatmap fits in a full dashboard.
- OpenTelemetry Tracing for BullMQ — traces behind the exemplars.