Alerting on Stuck and Stalled Jobs
Most queue alerts watch the backlog: too many jobs waiting, or waiting too long. A different failure slips past them — a job that has started and then stops making progress. It holds a worker slot, sometimes a lock, and never finishes or fails. A few of these can quietly consume a whole worker pool. This guide defines the signals that detect them and turns them into alerts that fire for real problems and stay quiet for jobs that are simply long, as part of SLOs & Alerting for Queues in Observability & Monitoring for Job Queues.
Problem Statement
A payment-reconciliation worker calls a bank's SFTP server. One morning the server accepted connections but never sent data, and the client had no read timeout. Over two hours, each of the 16 worker slots picked up a reconciliation job and hung. The queue's backlog grew slowly, because reconciliations are infrequent, and never crossed the backlog alert threshold. Separately, a BullMQ image service logs "job stalled" warnings several times a day, which nobody reads; some of those jobs run twice and produce duplicate thumbnails. The team wants alerts for jobs that stop progressing, visibility into stalls, and no pages for the monthly export that legitimately takes 90 minutes.
Prerequisites
- Prometheus (or a similar metrics system) and Alertmanager.
- Per-job start timestamps available from the queue (BullMQ
processedOn, Sidekiq's busy set, Celery'sactiveinspection) or from worker instrumentation. - Known expected durations per job type, ideally from a duration histogram.
- Worker processes you can add a heartbeat or progress metric to.
Step 1 — Separate Waiting, Stuck, and Stalled
Three different conditions need three different signals:
Backlog and wait-time alerts cover waiting jobs — see measuring queue wait time with enqueue timestamps. The reconciliation incident was a stuck-job problem; the duplicate thumbnails were a stalled-job problem. Neither shows up in wait time until the entire pool is blocked.
Step 2 — Export the Age of the Oldest Running Job
The most direct signal for stuck jobs is how long the oldest currently running job has been running, per job type. In Celery, record start times in each worker process and report the oldest at scrape time; max across processes in PromQL then gives the fleet-wide value:
# Celery: track running tasks per worker process
import time, threading
from celery.signals import task_prerun, task_postrun
from prometheus_client import Gauge
_running: dict[str, tuple[str, float]] = {} # task_id -> (task name, start)
_lock = threading.Lock()
@task_prerun.connect
def _start(task_id=None, task=None, **_):
with _lock:
_running[task_id] = (task.name, time.time())
@task_postrun.connect
def _end(task_id=None, **_):
with _lock:
_running.pop(task_id, None)
OLDEST_RUNNING = Gauge("job_oldest_running_seconds", "Age of oldest running job", ["task"])
def _collect():
now, oldest = time.time(), {}
with _lock:
for name, started in _running.values():
oldest[name] = max(oldest.get(name, 0.0), now - started)
OLDEST_RUNNING.clear()
for name, age in oldest.items():
OLDEST_RUNNING.labels(name).set(age)
Call _collect() from the metrics endpoint handler before serving, and query max by (task) (job_oldest_running_seconds). Because the value comes from inside the worker process, it keeps updating even while the task itself is blocked, as long as the metrics server runs in its own thread. For BullMQ, read processedOn from queue.getActive() in a single collector; for Sidekiq, Sidekiq::WorkSet lists every running job with its run_at. The gauge shape is the same for all three.
Step 3 — Compare Against Each Job Type's Normal Duration
A single threshold fails: 10 minutes is alarming for a thumbnail and normal for an export. Compare each job type's oldest running job to a multiple of its own p99 duration, computed from the duration histogram:
groups:
- name: stuck-jobs
rules:
- record: job:duration_seconds:p99_7d
expr: |
histogram_quantile(0.99,
sum by (task, le) (rate(job_duration_seconds_bucket[7d])))
- alert: JobStuck
expr: |
job_oldest_running_seconds
> on (task) group_left 3 * job:duration_seconds:p99_7d
and job_oldest_running_seconds > 300
for: 5m
labels: { severity: ticket }
annotations:
summary: "{{ $labels.task }} has a job running {{ $value | humanizeDuration }}"
The absolute floor stops short jobs from alerting on trivial delays; the 7-day p99 adapts as job durations change. For job types with too little history, fall back to an explicit per-task threshold in a static recording rule.
Step 4 — Alert When Many Slots Are Stuck at Once
One stuck job is a ticket. Many stuck jobs mean the pool is being consumed and the queue will stop — that is a page. Combine running-job counts with worker capacity:
- alert: WorkerPoolMostlyStuck
expr: |
sum by (queue) (job_running_over_threshold)
/ sum by (queue) (worker_concurrency) > 0.5
for: 5m
labels: { severity: page }
job_running_over_threshold is a gauge from the same collector counting running jobs past their threshold (not just the oldest). With 16 slots and 8 stuck reconciliations, this fires well before the whole pool is gone. Link the alert to a runbook entry that says how to find and kill the stuck jobs — see writing runbooks for queue incidents.
Step 5 — Count and Alert on Stalled Jobs
A stalled job is one whose worker stopped renewing its lock (BullMQ) or visibility timeout (SQS, Celery with Redis), so the queue hands it to another worker. Stalls mean duplicate execution risk and usually point to event-loop blocking, CPU starvation, or crashes. Count them explicitly:
import { QueueEvents } from "bullmq";
const events = new QueueEvents("thumbnails", { connection });
events.on("stalled", ({ jobId }) => stalledCounter.labels("thumbnails").inc());
For SQS, ApproximateReceiveCount above 1 on messages that did not fail indicates a visibility-timeout expiry; count it in the consumer. For Celery with acks_late, redelivered tasks carry delivery_info.redelivered = True.
- alert: JobsStalling
expr: sum by (queue) (increase(bullmq_jobs_stalled_total[30m])) > 5
for: 10m
labels: { severity: ticket }
A steady trickle of stalls is a bug to fix, not noise to ignore: in the image service, stalls came from a synchronous image resize blocking the event loop longer than the lock duration. Moving resizing to a sandboxed processor removed them. See BullMQ lock duration and stalled jobs for the mechanics.
Step 6 — Add Progress Heartbeats for Long Jobs
For long jobs, "running for a long time" is expected; what matters is whether they are still making progress. Have long jobs report progress, and alert when progress stops:
PROGRESS = Gauge("job_last_progress_timestamp", "Last progress report", ["task", "job_id"])
def reconcile(batch_id):
for i, chunk in enumerate(fetch_chunks(batch_id)):
process(chunk)
PROGRESS.labels("reconcile", batch_id).set_to_current_time()
PROGRESS.remove("reconcile", batch_id)
- alert: LongJobNoProgress
expr: time() - job_last_progress_timestamp > 900
for: 5m
Keep the job_id label only for long, infrequent jobs — for high-volume jobs it would create too many series. BullMQ's job.updateProgress() gives the same information inside the queue, where a collector can read it.
Verification
- A test job that sleeps forever triggers
JobStuckafter its threshold plus theforduration, and appears in the oldest-running gauge. - Running the 90-minute export does not trigger
JobStuckorLongJobNoProgress. - Blocking the event loop in a BullMQ test worker for longer than the lock duration increments the stalled counter and fires
JobsStallingafter repeated stalls. - Starting eight stuck test jobs on a 16-slot pool fires
WorkerPoolMostlyStuck.
Gotchas & Edge Cases
The real fix is a timeout. Alerts find stuck jobs; timeouts prevent them. Every network call in a job needs a connect and read timeout, and every job type needs a hard time limit — see Celery task time limits.
Inspection can hang too. celery inspect active waits for replies from workers; a worker whose main thread is blocked may not reply. Treat "no reply from a worker that should exist" as its own signal.
Clock differences. Start timestamps come from worker clocks. Keep NTP running on every host.
Stuck jobs hide in averages. Duration histograms record a job only when it finishes, so a job that never finishes never appears in them. That is why the oldest-running gauge exists: p99 duration can look perfectly healthy while half the pool is blocked.
Collector gaps. If the collector that reads running jobs fails, the gauge disappears and every stuck-job alert goes quiet. Add an absent() alert on the gauge so a broken collector is noticed.
Deploys interrupt long jobs. A job killed by a deploy may show as stalled and be picked up again. Expect a small bump in stalls after deploys, and exclude that window if it causes noise.
FAQ
Should stuck jobs page someone? One stuck job usually does not need a page; a pool that is mostly stuck does. Route single stuck jobs to a ticket queue and page on pool-level impact.
Can I automatically kill stuck jobs? Time limits do that safely. Killing jobs from an alert handler is riskier, because the alert may be wrong about a legitimately long job. Prefer limits in the job configuration.
What if job durations vary enormously? Split the job type by input size (small, large) with separate names or labels, so each has a meaningful p99, or rely on progress heartbeats rather than duration thresholds.
How does this relate to SLOs? Stuck jobs eventually show up as latency SLO burn once users notice the missing work, but by then the damage is done. These alerts are early warnings that sit alongside the SLO alerts rather than replacing them.
Related
- SLOs & Alerting for Queues — the alerting framework this fits into.
- Burn-Rate Alerts for Queue Backlogs — alerts for waiting work.
- BullMQ Lock Duration and Stalled Jobs — why jobs stall.
- Writing Runbooks for Queue Incidents — what to do when these alerts fire.