BullMQ Retries and Backoff Strategies

BullMQ retries a job when its processor throws, according to the job's attempts and backoff options. The defaults are easy to get almost right, and "almost" is where retry storms and silent data loss come from. This guide configures retries for real workloads as part of BullMQ for Node.js Ecosystems in Backend Frameworks & Worker Scaling.

Problem Statement

A Node.js integration service syncs CRM records through a partner API. Jobs are added with attempts: 3 and backoff: { type: "exponential", delay: 1000 }. During a 10-minute partner outage, every job exhausted its three attempts in about 7 seconds and landed in the failed set; 12,000 syncs had to be retried by hand. Validation errors from bad CRM data were retried three times each for no reason, and 429 responses with a Retry-After of 60 seconds were retried after 1, 2, and 4 seconds, earning more 429s. You want retries that outlast realistic outages, different behaviour per error class, respect for Retry-After, and visibility into retry volume.

Prerequisites

  • BullMQ 5.x workers and producers.
  • A classification of the errors your processor can throw (transient, rate-limited, permanent).
  • Default job options set in one place (queue defaultJobOptions or a shared enqueue helper).
  • Metrics on job outcomes, or QueueEvents access to build them.

Step 1 — Size Attempts and Delay to Cover Real Outages

Total retry time is what matters, not the attempt count. With exponential backoff, BullMQ waits delay Ɨ 2^(attempt āˆ’ 1) before each retry, so the total time across n attempts is roughly delay Ɨ (2^(nāˆ’1) āˆ’ 1).

function totalRetrySpan(delayMs: number, attempts: number) {
  let t = 0;
  for (let a = 1; a < attempts; a++) t += delayMs * 2 ** (a - 1);
  return t / 1000;
}
console.log(totalRetrySpan(1000, 3));     // 3 s     -> the scenario: gone before any outage ends
console.log(totalRetrySpan(5000, 10));    // 2,555 s -> ~43 minutes of retrying
console.log(totalRetrySpan(10000, 12));   // 20,470 s -> ~5.7 hours

Choose the span from the longest outage you want to ride out without human action — typically tens of minutes to a few hours for third-party APIs — then set attempts and delay to reach it.

const queue = new Queue("crm-sync", {
  connection,
  defaultJobOptions: {
    attempts: 10,
    backoff: { type: "exponential", delay: 5_000, jitter: 0.5 },   // ~43 min span, jittered
    removeOnFail: { age: 14 * 24 * 3600, count: 50_000 },
  },
});

jitter (BullMQ 5.x) randomises each delay by up to the given fraction, spreading retries of jobs that failed at the same moment — the defence against synchronised retry waves described in preventing retry storms after an outage.

Retry span, not attempt count With exponential backoff, three attempts starting at a one-second delay give up after about three seconds, far shorter than a ten-minute outage. Ten attempts starting at five seconds span about forty-three minutes. Twelve attempts starting at ten seconds span nearly six hours. How long a job keeps retrying 3 x 1 s 3 s: gone before the outage ends 10 x 5 s ~43 min 12 x 10 s ~5.7 hours

Step 2 — Fail Fast on Permanent Errors

Retrying a job that can never succeed wastes attempts and delays the alert. Throw UnrecoverableError for permanent failures; BullMQ moves the job straight to failed regardless of remaining attempts.

import { Worker, UnrecoverableError } from "bullmq";

const worker = new Worker("crm-sync", async (job) => {
  const record = await crm.get(job.data.recordId);
  if (!record) throw new UnrecoverableError(`record ${job.data.recordId} no longer exists`);
  const res = await partner.upsert(mapRecord(record));
  if (res.status === 400 || res.status === 422) {
    throw new UnrecoverableError(`partner rejected: ${res.body?.error ?? res.status}`);
  }
  if (res.status >= 500) throw new TransientError(`partner ${res.status}`);
  return { partnerId: res.body.id };
}, { connection, concurrency: 25 });

The classification mirrors retrying only transient errors by exception type: data problems fail immediately with a clear reason; infrastructure problems retry.

Step 3 — Use a Custom Backoff Strategy per Error

A single backoff curve cannot fit both "server error, try again soon" and "rate limited, the server told you exactly when". Custom backoff strategies receive the attempt number and the error, and return a delay (or -1 to stop retrying).

class RateLimitedError extends Error {
  constructor(public retryAfterMs: number) { super(`rate limited for ${retryAfterMs} ms`); }
}

const worker = new Worker("crm-sync", processor, {
  connection,
  settings: {
    backoffStrategy: (attemptsMade: number, type: string, err: Error) => {
      if (err instanceof RateLimitedError) return err.retryAfterMs + Math.random() * 2000;
      if (type === "crm") {
        const base = Math.min(5_000 * 2 ** (attemptsMade - 1), 30 * 60_000);   // cap at 30 min
        return Math.random() * base;                                            // full jitter
      }
      return -1;                                                                // unknown type: no retry
    },
  },
});

// jobs opt into the custom strategy by type
await queue.add("sync", { recordId }, { attempts: 12, backoff: { type: "crm" } });

Honouring Retry-After exactly fixes the 429 loop in the problem statement. When the whole queue should slow down on a 429 — not just the one job — use the worker-level rate limit instead, as in configuring the BullMQ rate limiter.

One strategy, three behaviours The processor throws. If the error is a RateLimitedError, the strategy returns the server's Retry-After plus a little jitter. If the job's backoff type is crm and the error is transient, it returns a fully jittered exponential delay capped at thirty minutes. For anything else it returns minus one, which fails the job instead of retrying. UnrecoverableError bypasses the strategy entirely. backoffStrategy(attemptsMade, type, err) processor throws RateLimitedError: Retry-After + up to 2 s jitter type crm, transient: random(0, min(5 s x 2^n, 30 min)) unknown type: -1, fail now

Step 4 — Delay a Job Without Counting an Attempt

Some situations are not failures: a record is locked by another sync, a dependency's circuit breaker is open, or data is not ready yet. Throwing would consume an attempt. Instead, move the job back to delayed and signal BullMQ not to treat it as a failure.

import { DelayedError } from "bullmq";

const worker = new Worker("crm-sync", async (job, token) => {
  if (await breaker.isOpen("partner")) {
    await job.moveToDelayed(Date.now() + 30_000 + Math.random() * 15_000, token);
    throw new DelayedError();                       // tells BullMQ: not a failure, no attempt used
  }
  // ... normal processing ...
}, { connection });
Postpone without spending an attempt A job on attempt 2 of 10 finds the partner breaker open. Throwing an error would move it to attempt 3 and eventually exhaust its budget during a long outage. Calling moveToDelayed and throwing DelayedError returns it to the delayed set for about thirty seconds and leaves it on attempt 2. Job on attempt 2 of 10, breaker open throw new Error() now attempt 3 of 10 long outage exhausts budget moveToDelayed + DelayedError still attempt 2 of 10 waits out the outage Use it for "not now", never for "this failed": real failures should count.

This is the building block for dependency-aware pausing; the breaker itself is described in circuit breakers for worker dependencies.

Step 5 — Observe Retries, Not Just Failures

A queue that eventually succeeds can still be unhealthy if most jobs need several attempts. Track attempts at completion and retry events.

const events = new QueueEvents("crm-sync", { connection });

worker.on("completed", (job) =>
  attemptsAtSuccess.observe({ queue: job.queueName }, job.attemptsMade + 1));
worker.on("failed", (job, err) => {
  if (!job) return;
  const final = job.attemptsMade >= (job.opts.attempts ?? 1) || err instanceof UnrecoverableError;
  jobFailures.inc({ queue: job.queueName, final: String(final), error: err.name });
});
# Share of successful jobs that needed retries: rising values mean a degrading dependency
sum(rate(job_attempts_at_success_bucket{le="1"}[15m])) / sum(rate(job_attempts_at_success_count[15m]))

# Final failures by error type
sum by (error) (rate(job_failures_total{final="true"}[15m]))

A drop in first-attempt success is often the earliest sign of a dependency problem, well before anything reaches the failed set.

Step 6 — Write Retry Settings Down per Job Type

Retry settings encode business decisions — how long a sync may lag, how much a stale email is worth — so they deserve a table that product and engineering agree on, rather than values scattered across enqueue calls.

Job type Attempts Backoff Span Why
CRM sync 12 custom crm, 5 s base, full jitter, 30 min cap ~5 h partner outages last hours; data must arrive
Password reset email 4 exponential, 2 s ~15 s useless after a few minutes; user can request again
Invoice PDF 8 exponential, 10 s ~20 min user waits in-app; after that, alert and fix
Nightly export 3 fixed, 10 min ~20 min next night's run supersedes it

Encode the table once, in the enqueue helper, keyed by job name, and have code review treat changes to it like any other contract change. The broader reasoning about budgets is in setting retry budgets and max attempts.

const RETRY: Record<string, Partial<JobsOptions>> = {
  "crm-sync":        { attempts: 12, backoff: { type: "crm" } },
  "password-reset":  { attempts: 4,  backoff: { type: "exponential", delay: 2_000 } },
  "invoice-pdf":     { attempts: 8,  backoff: { type: "exponential", delay: 10_000, jitter: 0.5 } },
  "nightly-export":  { attempts: 3,  backoff: { type: "fixed", delay: 600_000 } },
};
export const enqueue = (q: Queue, name: string, data: unknown, opts: JobsOptions = {}) =>
  q.add(name, data, { ...RETRY[name], ...opts });

Verification

it("retries transient errors and fails permanent ones immediately", async () => {
  partnerFake.failTimes(2, 503);
  const ok = await queue.add("sync", { recordId: "r1" }, { attempts: 5, backoff: { type: "fixed", delay: 10 } });
  await ok.waitUntilFinished(events, 5000);
  expect((await queue.getJob(ok.id!))!.attemptsMade).toBe(3);

  partnerFake.respond(422);
  const bad = await queue.add("sync", { recordId: "r2" }, { attempts: 5 });
  await expect(bad.waitUntilFinished(events, 5000)).rejects.toThrow(/partner rejected/);
  expect((await queue.getJob(bad.id!))!.attemptsMade).toBe(1);
});

The integration harness is described in integration testing BullMQ workers with Testcontainers.

Gotchas & Edge Cases

Default attempts is 1. Without attempts, a job never retries. Set it in defaultJobOptions, not per call site.

Backoff in seconds by mistake. delay is milliseconds. delay: 5 retries after 5 ms.

Retries and ordering. A retried job goes back behind newer jobs; code that assumes jobs for one entity run in order breaks under retries. See Message Ordering Guarantees.

Stalls are separate. Jobs that stall (lost lock) follow maxStalledCount, not attempts.

FAQ

Should every queue retry for hours? No — match the span to the job's value over time. A login code is useless after five minutes; an invoice sync matters for days.

Can I change retry options for jobs already queued? Options are stored per job at enqueue time. New defaults apply to new jobs; queued jobs keep their options.

How do I retry a job manually after it failed? Call job.retry() on the failed job, or queue.retryJobs({ state: "failed" }) in bulk once the cause is fixed. The attempt count resets, so the job gets its full retry budget again. Retry by failure reason rather than blindly, so jobs that failed for a still-unfixed reason do not go straight back to the failed set.

Should the processor catch errors and retry internally? Only for very short, local retries (a single immediate retry on a connection reset). Anything longer belongs to BullMQ's backoff, so the job releases its worker slot while it waits and the retry is visible in metrics and the UI.

What's the difference between exponential and fixed backoff? Fixed waits the same delay each time — fine for short, predictable hiccups. Exponential grows the delay, which suits outages of unknown length.

Related