Alerting on Stuck and Stalled Jobs

Most queue alerts watch the backlog: too many jobs waiting, or waiting too long. A different failure slips past them — a job that has started and then stops making progress. It holds a worker slot, sometimes a lock, and never finishes or fails. A few of these can quietly consume a whole worker pool. This guide defines the signals that detect them and turns them into alerts that fire for real problems and stay quiet for jobs that are simply long, as part of SLOs & Alerting for Queues in Observability & Monitoring for Job Queues.

Problem Statement

A payment-reconciliation worker calls a bank's SFTP server. One morning the server accepted connections but never sent data, and the client had no read timeout. Over two hours, each of the 16 worker slots picked up a reconciliation job and hung. The queue's backlog grew slowly, because reconciliations are infrequent, and never crossed the backlog alert threshold. Separately, a BullMQ image service logs "job stalled" warnings several times a day, which nobody reads; some of those jobs run twice and produce duplicate thumbnails. The team wants alerts for jobs that stop progressing, visibility into stalls, and no pages for the monthly export that legitimately takes 90 minutes.

Prerequisites

  • Prometheus (or a similar metrics system) and Alertmanager.
  • Per-job start timestamps available from the queue (BullMQ processedOn, Sidekiq's busy set, Celery's active inspection) or from worker instrumentation.
  • Known expected durations per job type, ideally from a duration histogram.
  • Worker processes you can add a heartbeat or progress metric to.

Step 1 — Separate Waiting, Stuck, and Stalled

Three different conditions need three different signals:

Waiting, stuck and stalled jobs A waiting job is ready but no worker has started it; the signal is the age of the oldest waiting job. A stuck job has started and holds its worker, but makes no progress, for example blocked on a network call without a timeout; the signal is the age of the oldest running job compared with its expected duration, or a progress heartbeat that stops. A stalled job is one whose worker stopped renewing its lock, so the queue assumes the worker died and hands the job to another worker; the signal is stall events. Three conditions, three signals waiting ready, not started signal: oldest waiting age stuck started, no progress signal: oldest running age or silent heartbeat stalled lock lost, job handed on signal: stall event count

Backlog and wait-time alerts cover waiting jobs — see measuring queue wait time with enqueue timestamps. The reconciliation incident was a stuck-job problem; the duplicate thumbnails were a stalled-job problem. Neither shows up in wait time until the entire pool is blocked.

Step 2 — Export the Age of the Oldest Running Job

The most direct signal for stuck jobs is how long the oldest currently running job has been running, per job type. In Celery, record start times in each worker process and report the oldest at scrape time; max across processes in PromQL then gives the fleet-wide value:

# Celery: track running tasks per worker process
import time, threading
from celery.signals import task_prerun, task_postrun
from prometheus_client import Gauge

_running: dict[str, tuple[str, float]] = {}      # task_id -> (task name, start)
_lock = threading.Lock()

@task_prerun.connect
def _start(task_id=None, task=None, **_):
    with _lock:
        _running[task_id] = (task.name, time.time())

@task_postrun.connect
def _end(task_id=None, **_):
    with _lock:
        _running.pop(task_id, None)

OLDEST_RUNNING = Gauge("job_oldest_running_seconds", "Age of oldest running job", ["task"])

def _collect():
    now, oldest = time.time(), {}
    with _lock:
        for name, started in _running.values():
            oldest[name] = max(oldest.get(name, 0.0), now - started)
    OLDEST_RUNNING.clear()
    for name, age in oldest.items():
        OLDEST_RUNNING.labels(name).set(age)

Call _collect() from the metrics endpoint handler before serving, and query max by (task) (job_oldest_running_seconds). Because the value comes from inside the worker process, it keeps updating even while the task itself is blocked, as long as the metrics server runs in its own thread. For BullMQ, read processedOn from queue.getActive() in a single collector; for Sidekiq, Sidekiq::WorkSet lists every running job with its run_at. The gauge shape is the same for all three.

Step 3 — Compare Against Each Job Type's Normal Duration

A single threshold fails: 10 minutes is alarming for a thumbnail and normal for an export. Compare each job type's oldest running job to a multiple of its own p99 duration, computed from the duration histogram:

groups:
  - name: stuck-jobs
    rules:
      - record: job:duration_seconds:p99_7d
        expr: |
          histogram_quantile(0.99,
            sum by (task, le) (rate(job_duration_seconds_bucket[7d])))

      - alert: JobStuck
        expr: |
          job_oldest_running_seconds
            > on (task) group_left 3 * job:duration_seconds:p99_7d
          and job_oldest_running_seconds > 300
        for: 5m
        labels: { severity: ticket }
        annotations:
          summary: "{{ $labels.task }} has a job running {{ $value | humanizeDuration }}"
Thresholds relative to each job type For thumbnails, p99 duration is 4 seconds, so three times p99 is 12 seconds; the absolute floor of 5 minutes applies to avoid noise. For monthly exports, p99 is 95 minutes, so the threshold is about 4.75 hours and a 90-minute export never alerts. For reconciliation, p99 is 40 seconds; a job running for 2 hours is far beyond the 5-minute threshold and alerts. Stuck means "much longer than usual for this job" thumbnail p99 4 s → threshold 5 min (floor) running 3 s: fine export p99 95 min → threshold ≈ 4.75 h running 90 min: fine reconciliation p99 40 s → threshold 5 min running 2 h: stuck threshold = max(3 × p99 over 7 days, 5 minutes)

The absolute floor stops short jobs from alerting on trivial delays; the 7-day p99 adapts as job durations change. For job types with too little history, fall back to an explicit per-task threshold in a static recording rule.

Step 4 — Alert When Many Slots Are Stuck at Once

One stuck job is a ticket. Many stuck jobs mean the pool is being consumed and the queue will stop — that is a page. Combine running-job counts with worker capacity:

- alert: WorkerPoolMostlyStuck
  expr: |
    sum by (queue) (job_running_over_threshold)
      / sum by (queue) (worker_concurrency) > 0.5
  for: 5m
  labels: { severity: page }

job_running_over_threshold is a gauge from the same collector counting running jobs past their threshold (not just the oldest). With 16 slots and 8 stuck reconciliations, this fires well before the whole pool is gone. Link the alert to a runbook entry that says how to find and kill the stuck jobs — see writing runbooks for queue incidents.

Step 5 — Count and Alert on Stalled Jobs

A stalled job is one whose worker stopped renewing its lock (BullMQ) or visibility timeout (SQS, Celery with Redis), so the queue hands it to another worker. Stalls mean duplicate execution risk and usually point to event-loop blocking, CPU starvation, or crashes. Count them explicitly:

import { QueueEvents } from "bullmq";
const events = new QueueEvents("thumbnails", { connection });
events.on("stalled", ({ jobId }) => stalledCounter.labels("thumbnails").inc());

For SQS, ApproximateReceiveCount above 1 on messages that did not fail indicates a visibility-timeout expiry; count it in the consumer. For Celery with acks_late, redelivered tasks carry delivery_info.redelivered = True.

- alert: JobsStalling
  expr: sum by (queue) (increase(bullmq_jobs_stalled_total[30m])) > 5
  for: 10m
  labels: { severity: ticket }

A steady trickle of stalls is a bug to fix, not noise to ignore: in the image service, stalls came from a synchronous image resize blocking the event loop longer than the lock duration. Moving resizing to a sandboxed processor removed them. See BullMQ lock duration and stalled jobs for the mechanics.

Step 6 — Add Progress Heartbeats for Long Jobs

For long jobs, "running for a long time" is expected; what matters is whether they are still making progress. Have long jobs report progress, and alert when progress stops:

PROGRESS = Gauge("job_last_progress_timestamp", "Last progress report", ["task", "job_id"])

def reconcile(batch_id):
    for i, chunk in enumerate(fetch_chunks(batch_id)):
        process(chunk)
        PROGRESS.labels("reconcile", batch_id).set_to_current_time()
    PROGRESS.remove("reconcile", batch_id)
- alert: LongJobNoProgress
  expr: time() - job_last_progress_timestamp > 900
  for: 5m

Keep the job_id label only for long, infrequent jobs — for high-volume jobs it would create too many series. BullMQ's job.updateProgress() gives the same information inside the queue, where a collector can read it.

Progress heartbeats distinguish long from stuck A 90-minute export reports progress every few minutes; the time since last progress never exceeds the 15-minute limit, so no alert fires despite the long runtime. A reconciliation job reports progress twice, then blocks on the SFTP read; after 15 minutes with no progress report, the alert fires, long before the pool is exhausted. Long is fine; silent is not export (90 min) progress every few minutes: no alert reconciliation 15 min silent: alert Dots are progress reports.

Verification

  • A test job that sleeps forever triggers JobStuck after its threshold plus the for duration, and appears in the oldest-running gauge.
  • Running the 90-minute export does not trigger JobStuck or LongJobNoProgress.
  • Blocking the event loop in a BullMQ test worker for longer than the lock duration increments the stalled counter and fires JobsStalling after repeated stalls.
  • Starting eight stuck test jobs on a 16-slot pool fires WorkerPoolMostlyStuck.

Gotchas & Edge Cases

The real fix is a timeout. Alerts find stuck jobs; timeouts prevent them. Every network call in a job needs a connect and read timeout, and every job type needs a hard time limit — see Celery task time limits.

Inspection can hang too. celery inspect active waits for replies from workers; a worker whose main thread is blocked may not reply. Treat "no reply from a worker that should exist" as its own signal.

Clock differences. Start timestamps come from worker clocks. Keep NTP running on every host.

Stuck jobs hide in averages. Duration histograms record a job only when it finishes, so a job that never finishes never appears in them. That is why the oldest-running gauge exists: p99 duration can look perfectly healthy while half the pool is blocked.

Collector gaps. If the collector that reads running jobs fails, the gauge disappears and every stuck-job alert goes quiet. Add an absent() alert on the gauge so a broken collector is noticed.

Deploys interrupt long jobs. A job killed by a deploy may show as stalled and be picked up again. Expect a small bump in stalls after deploys, and exclude that window if it causes noise.

FAQ

Should stuck jobs page someone? One stuck job usually does not need a page; a pool that is mostly stuck does. Route single stuck jobs to a ticket queue and page on pool-level impact.

Can I automatically kill stuck jobs? Time limits do that safely. Killing jobs from an alert handler is riskier, because the alert may be wrong about a legitimately long job. Prefer limits in the job configuration.

What if job durations vary enormously? Split the job type by input size (small, large) with separate names or labels, so each has a meaningful p99, or rely on progress heartbeats rather than duration thresholds.

How does this relate to SLOs? Stuck jobs eventually show up as latency SLO burn once users notice the missing work, but by then the damage is done. These alerts are early warnings that sit alongside the SLO alerts rather than replacing them.

Related