BullMQ Lock Duration and Stalled Jobs
BullMQ has no setting called "visibility timeout", but its job lock is exactly that: a lease that a worker must keep renewing, and that another worker may reclaim if it lapses. This guide explains the mechanism and tunes it, as part of the Visibility Timeout Deep Dive in Queue Fundamentals & Architecture.
Problem Statement
A PDF-generation service built on BullMQ logs dozens of "job stalled more than allowable limit" failures per day, and some PDFs are generated twice. The jobs are not slow — p99 is 12 seconds against a default 30-second lock — but PDF rendering runs synchronously in the worker process for up to several seconds at a time. Investigation shows that during large renders, the lock-renewal timer cannot fire, the lock expires, the stalled-job checker in another worker moves the job back to waiting, and a second worker starts the same job. You want to understand exactly when BullMQ considers a job stalled, stop false stalls, keep real crash recovery working, and alert on stalls that do happen.
Prerequisites
- BullMQ 4.x or 5.x with a Redis 6.2+ instance.
- Access to worker options and the ability to run processors in sandboxed child processes.
QueueEventsor metrics to observestalledevents.- Idempotent processors, since a stalled job is by definition run again.
Step 1 — Understand the Lock Lifecycle
When a worker takes a job it acquires a lock key in Redis with a TTL of lockDuration (30 seconds by default). While the job runs, the worker renews the lock every lockRenewTime (by default half the lock duration). Separately, every worker runs a stalled-job check every stalledInterval (30 seconds): any active job whose lock no longer exists is considered stalled and moved back to waiting — or failed, if it has already stalled maxStalledCount times (default 1).
const worker = new Worker("pdf", processor, {
connection,
concurrency: 8,
lockDuration: 30_000, // lease length: how long a job may go without renewal
lockRenewTime: 15_000, // renew halfway through (default lockDuration / 2)
stalledInterval: 30_000, // how often any worker scans for expired locks
maxStalledCount: 1, // stalls tolerated before the job is failed
});
A job stalls only when renewal stops — because the process crashed, lost its Redis connection, or could not run the renewal timer. How long the job takes does not matter as long as renewals keep happening.
Step 2 — Find What Blocks Renewal
Renewal is a setTimeout in the worker's event loop. Anything that blocks the event loop for longer than lockDuration − lockRenewTime (15 seconds by default) can make a renewal late enough for the lock to expire. The usual culprits are synchronous CPU work — PDF rendering, image processing, large JSON.parse, regex on huge strings, synchronous crypto or compression.
// Measure event-loop blocking in the worker process
import { monitorEventLoopDelay } from "node:perf_hooks";
const h = monitorEventLoopDelay({ resolution: 20 });
h.enable();
setInterval(() => {
const maxMs = h.max / 1e6;
eventLoopMaxDelay.set(maxMs); // Prometheus gauge
if (maxMs > 5000) logger.warn({ maxMs }, "event loop blocked");
h.reset();
}, 10_000);
In the scenario, maximum event-loop delay reached 22 seconds during large PDF renders — well past the 15-second renewal window. Every job running concurrently in that process (concurrency 8) lost its lock at the same time, which is why stalls came in bursts.
Step 3 — Move CPU-Bound Work into Sandboxed Processors
The fix for event-loop blocking is not a longer lock; it is running the blocking work where it cannot block the lock timer. BullMQ's sandboxed processors run each job in a child process (or worker thread) while the parent process — whose event loop stays free — handles locks and renewals.
// worker.ts — parent process: locks, renewals, events
import { Worker } from "bullmq";
import path from "node:path";
const worker = new Worker("pdf", path.join(__dirname, "render.sandbox.js"), {
connection,
concurrency: 4, // ≈ CPU cores: each job gets its own process
useWorkerThreads: true, // threads instead of forked processes (lighter)
lockDuration: 30_000,
});
// render.sandbox.ts — runs in a worker thread; blocking here cannot starve renewals
export default async function (job) {
const pdf = renderPdfSync(job.data.template, job.data.fields); // CPU-heavy, synchronous
await storePdf(job.data.documentId, pdf);
return { bytes: pdf.length };
}
With the render off the main thread, event-loop delay in the parent stays in milliseconds and stalls from blocking disappear. Concurrency for CPU-bound sandboxed work should be close to the core count; more just time-slices. The general sandboxing guide is BullMQ sandboxed processors.
Step 4 — Tune Lock and Stall Settings for Real Crash Recovery
With blocking removed, the lock settings trade off how quickly a crashed worker's jobs are recovered against tolerance for transient hiccups (Redis failover, GC pauses, network blips).
new Worker("pdf", sandboxFile, {
connection,
lockDuration: 60_000, // tolerate up to ~30 s of missed renewals (Redis failover)
lockRenewTime: 20_000, // renew well before expiry
stalledInterval: 30_000, // recovery latency ≈ lockDuration + up to stalledInterval
maxStalledCount: 2, // a genuine crash-looping job fails after its 3rd stall
});
Worst-case recovery time for a crashed worker's job is roughly lockDuration + stalledInterval — 90 seconds here. Shorter locks recover faster but turn any pause longer than the renewal margin into a stall and a duplicate run.
maxStalledCount protects against a job that crashes its worker every time (for example, an out-of-memory input): after that many stalls, the job fails instead of being retried forever and taking workers down with it.
Step 5 — Handle Long Jobs Explicitly
For jobs that legitimately run for many minutes, renewal still works automatically in the parent; nothing extra is needed beyond keeping the event loop free. If a processor must run in the main thread and has a natural loop, yield regularly so timers can fire:
async function processLargeExport(job: Job) {
for (const [i, chunk] of chunks(job.data.rows, 500).entries()) {
writeChunk(chunk); // synchronous but short
if (i % 10 === 0) {
await job.updateProgress(Math.round((i / total) * 100));
await new Promise((r) => setImmediate(r)); // let lock renewal run
}
}
}
Checkpointing progress (writing which chunk was last completed) also makes a genuinely stalled long job resume rather than restart — the same idea as in configuring visibility timeouts for long-running workers.
Step 6 — Alert on Stalls and Failed-Due-to-Stall
A stall is not an error the processor sees; it surfaces as a stalled event and, after maxStalledCount, a failure with reason "job stalled more than allowable limit". Count both.
const events = new QueueEvents("pdf", { connection });
events.on("stalled", ({ jobId }) => {
stalledTotal.inc({ queue: "pdf" });
logger.warn({ job_id: jobId, queue: "pdf" }, "job stalled");
});
events.on("failed", ({ jobId, failedReason }) => {
if (failedReason?.includes("stalled more than allowable limit")) {
stalledFailedTotal.inc({ queue: "pdf" });
}
});
# Any stalls outside deploy windows deserve a look; bursts point at event-loop blocking
sum(increase(bullmq_jobs_stalled_total{queue="pdf"}[10m])) > 5
Correlate stall bursts with event-loop delay from Step 2 and with worker restarts. Stalls during deploys are expected if workers are killed mid-job; stalls at other times are almost always blocking or Redis connectivity. Worker-level alerting patterns are in alerting on stuck and stalled jobs.
Verification
// Reproduce a stall deterministically in an integration test (see the Testcontainers guide)
it("recovers a job when its worker stops renewing", async () => {
const worker = new Worker("t", async () => { busyWait(1500); }, { connection, lockDuration: 500, stalledInterval: 300 });
const job = await queue.add("x", {});
await expect(waitForEvent(events, "stalled", job.id)).resolves.toBeTruthy();
});
In production, after moving rendering into a sandbox, confirm that eventLoopMaxDelay in the parent stays under 100 ms and that stall counts drop to near zero outside deploys. The harness for such tests is in integration testing BullMQ workers with Testcontainers.
Gotchas & Edge Cases
Redis failover. During a failover, renewals fail for a few seconds. A lockDuration shorter than typical failover time produces a burst of stalls on every failover.
Clock assumptions. Lock TTLs are enforced by Redis, so worker clock skew does not matter — but a paused VM or container freeze does, because renewals stop.
Concurrency and blocking multiply. With concurrency 8 on the main thread, one blocking job stalls all 8. Sandboxing isolates them.
skipStalledCheck. Disabling stall checks removes crash recovery entirely; jobs held by dead workers stay active forever. Do not use it to silence false stalls.
FAQ
Is a longer lockDuration the fix for stalls? Only if the cause is short pauses longer than your current margin (failovers, GC). For event-loop blocking of unpredictable length, move the work off the main thread instead.
Can I extend the lock manually for one job?
Yes — job.extendLock(token, duration) pushes a specific job's lock out, which is useful when a processor knows it is about to do one unusually long synchronous step. It does not help if the event loop is already blocked, because the call itself needs the loop to run; call it before the blocking step, not during.
Do stalled jobs count as attempts?
Stalls are tracked separately (stalledCounter) and governed by maxStalledCount; they do not consume attempts in the same way failures do. A job can stall and later fail normally with its retry budget intact.
How is this different from an SQS visibility timeout?
It is the same concept implemented client-side with renewal: SQS requires explicit ChangeMessageVisibility calls, BullMQ renews automatically — as long as the event loop lets it.
Related
- Visibility Timeout Deep Dive — leases across brokers.
- RabbitMQ Consumer Timeout for Unacked Messages — the RabbitMQ equivalent.
- Configuring BullMQ Concurrency Limits for High Throughput — concurrency and its interaction with blocking.
- BullMQ Sandboxed Processors — running CPU work safely.