Forecasting Backlog Drain Time
During a queue incident the first question from everyone outside the on-call rotation is "when will it be back to normal?", and this guide turns that into a number on a dashboard, as part of Capacity Planning for Job Queues in Observability & Monitoring for Job Queues. The forecast is simple arithmetic on metrics you already have — backlog, arrival rate, completion rate — but getting it right needs care with rates that change, recovery that is not linear, and backlogs that never drain at all.
Problem Statement
A reporting pipeline's database had a 40-minute slowdown, and the report-generation queue grew to 380,000 jobs. After the database recovered, the incident channel filled with questions: customer support wanted an ETA for delayed reports, the account team wanted to know whether to warn a large customer, and the on-call engineer wanted to know whether adding workers would help. The answer given — "a couple of hours, probably" — turned out to be off by a factor of three, because arrivals kept climbing into the evening peak. You want a live ETA that accounts for arrivals, a clear signal when the backlog is not draining at all, and an alert that fires on "will take too long to drain" rather than on raw depth.
Prerequisites
- Metrics for backlog depth, enqueue rate, and completion rate per queue (see Prometheus Metrics for Workers).
- Grafana or another dashboard tool that can plot PromQL expressions.
- Optionally, historical arrival patterns (same weekday last week) for forecasting arrivals during the drain.
Step 1 — Get the Drain Equation Right
The backlog shrinks at the net rate: completions minus arrivals. Dividing the backlog by the completion rate alone — the most common mistake — ignores that new work keeps arriving.
net_drain_rate = completion_rate - arrival_rate (jobs/s)
eta_seconds = backlog / net_drain_rate (only if net_drain_rate > 0)
if net_drain_rate <= 0: the backlog is not draining; ETA is infinite
With 380,000 jobs queued, workers completing 520 jobs per second, and 380 jobs per second still arriving, the net rate is 140 per second and the ETA is 45 minutes. Dividing by completions alone gives 12 minutes — the kind of answer that gets repeated to customers and then missed.
Step 2 — Build the ETA in PromQL
Use a smoothing window long enough to ignore second-to-second noise but short enough to react to scaling changes — 10 minutes works for most queues.
# Net drain rate (positive = shrinking)
sum by (queue) (rate(jobs_completed_total[10m]))
- sum by (queue) (rate(jobs_enqueued_total[10m]))
# ETA in minutes, only where the backlog is actually shrinking
(
sum by (queue) (queue_depth)
/
(sum by (queue) (rate(jobs_completed_total[10m])) - sum by (queue) (rate(jobs_enqueued_total[10m])))
) / 60
and on (queue)
(sum by (queue) (rate(jobs_completed_total[10m])) - sum by (queue) (rate(jobs_enqueued_total[10m]))) > 0
The and ... > 0 clause drops the series when the queue is not draining, so the panel shows "no data" instead of a negative or enormous ETA. Pair it with a separate stat panel showing the net rate, coloured red when it is zero or negative: "not draining" is a different message from "draining slowly".
If you lack a depth metric, derive the backlog from its rate of change: -deriv(queue_depth[10m]) is the measured drain rate and cross-checks the counter arithmetic — they should agree within a few percent.
Step 3 — Account for Arrivals That Will Change
A flat-rate ETA assumes arrivals stay where they are. During a recovery that crosses a daily peak, they will not. Use last week's arrivals at the same time of day as the forecast for the next few hours.
# eta.py — step through time with forecast arrivals instead of a constant
def eta_with_forecast(backlog: float, capacity_per_s: float,
forecast_arrivals: list[float], step_s: int = 300) -> float | None:
"""forecast_arrivals: expected arrivals/s for each upcoming step (e.g., last week's)."""
t = 0
for arrivals in forecast_arrivals:
backlog -= (capacity_per_s - arrivals) * step_s
t += step_s
if backlog <= 0:
return t - (-backlog) / max(capacity_per_s - arrivals, 1e-9)
return None # not drained within the forecast horizon
last_week = prom_range('sum(rate(jobs_enqueued_total{queue="reports"}[5m]) offset 7d)',
start="now", end="now+6h", step="5m")
print(eta_with_forecast(380_000, 520, last_week) / 60) # minutes
In the scenario, arrivals rose from 380 to 480 per second as the evening peak began, shrinking the net rate from 140 to 40 per second. The stepped forecast gave 2 hours 20 minutes — close to what actually happened — while the flat estimate said 45 minutes.
Step 4 — Decide Whether More Workers Will Help
Before scaling out to speed up the drain, check whether completion rate is limited by workers or by something they depend on. If workers are saturated and dependencies have headroom, more workers raise capacity; if a dependency is saturated, more workers only increase contention.
# Are workers the limit? Busy ratio near 1 means yes, if dependencies are healthy
sum(worker_busy_slots{queue="reports"}) / sum(worker_total_slots{queue="reports"})
# Is job duration inflating? Rising p50 during the drain points at a dependency
histogram_quantile(0.5, sum by (le) (rate(job_duration_seconds_bucket{queue="reports"}[5m])))
With workers at 98% busy and job duration flat, doubling workers roughly doubles completion rate: the net rate would go from 140 to 660 per second and the ETA from 45 to about 10 minutes.
If duration were climbing, the right move would be to protect the dependency — see circuit breakers for worker dependencies — and accept the slower drain.
Step 5 — Communicate the ETA with a Range
A single number sounds precise and will be quoted. Give a range from two assumptions — current arrivals and forecast arrivals — and update it on a cadence.
Incident update 16:40 — report queue
Backlog: 380k (peak), now 352k. Draining at 140/s net.
ETA to normal: 45 min at current arrivals; up to 2h20m if evening traffic follows last week.
Action: scaled report workers 20 -> 40 at 16:35; expect net rate ~600/s by 16:45.
Next update: 17:00.
Oldest-job age matters more to customers than total backlog. Report both: "the oldest pending report is 52 minutes old" tells support what to say.
Step 6 — Alert on Drain Time, Not Depth
A depth threshold fires on harmless spikes that clear in seconds and misses slow backlogs that never clear. Alert when the forecast crosses what users can tolerate, or when the queue stops draining.
groups:
- name: queue-drain
rules:
- alert: QueueNotDraining
expr: |
sum by (queue) (queue_depth) > 1000
and on (queue)
(sum by (queue) (rate(jobs_completed_total[10m])) - sum by (queue) (rate(jobs_enqueued_total[10m]))) <= 0
for: 15m
labels: { severity: page }
annotations:
summary: "{{ $labels.queue }}: backlog growing or flat for 15 minutes"
- alert: QueueDrainTooSlow
expr: |
(sum by (queue) (queue_depth)
/ clamp_min(sum by (queue) (rate(jobs_completed_total[10m])) - sum by (queue) (rate(jobs_enqueued_total[10m])), 0.001))
> 3600
for: 10m
labels: { severity: page }
annotations:
summary: "{{ $labels.queue }}: forecast drain time over 1 hour"
These complement the burn-rate alerts in burn-rate alerts for queue backlogs: drain-time alerts answer "will this fix itself soon?", burn-rate alerts answer "how much of our reliability budget is it costing?".
Verification
Replay a past incident's metrics (or run a burst in staging, as in load testing queue throughput) and compare the forecast at several points with the actual time-to-empty:
t+0 forecast 45m (flat) / 140m (weekly) | actual 146m
t+30 forecast 71m / 118m | actual 116m
t+60 forecast 48m / 84m | actual 86m
If the weekly-arrival forecast is consistently within 10–15%, it is good enough to communicate. If it is not, check whether retries from the incident are inflating arrivals — they should be included in arrival counters but may not appear in last week's pattern.
Gotchas & Edge Cases
Counters that reset. Worker restarts reset in-process counters; rate() handles resets, but a fleet-wide restart during the incident creates a short gap. Use the broker's own depth metric as the anchor.
Delayed and scheduled jobs. Jobs scheduled for the future count in some depth metrics but cannot drain until due. Exclude them from the backlog used for ETAs.
Priority queues. When workers prefer a high-priority queue, low-priority backlog drains only with leftover capacity. Compute ETAs per queue with the capacity actually available to it.
Dead-lettered jobs. Jobs that fail permanently leave the backlog without completing. Count them in the drain rate or the ETA will look pessimistic.
FAQ
Why not use predict_linear on the depth?
predict_linear(queue_depth[30m], 3600) extrapolates the recent trend and is a decent quick check, but it ignores changing arrivals and scaling actions. The net-rate formula responds immediately when capacity changes.
What smoothing window should the ETA use? 10 minutes for most queues; 30 minutes for very bursty arrivals. Shorter windows make the ETA jump around; longer ones lag scaling changes.
How do I show the ETA to customers? Show oldest-job age and a range, not a precise time. Update on a fixed cadence and stop quoting ETAs once the queue is back within its normal wait objective.
Related
- Capacity Planning for Job Queues — the planning side of drain arithmetic.
- Alerting on Queue Backlog with Prometheus — backlog alerts this refines.
- Writing Runbooks for Queue Incidents — where the ETA procedure belongs.
- Load Testing Queue Throughput — measuring drain rates before incidents.