BullMQ Failed Jobs as a Dead-Letter Queue

BullMQ does not have a component called a dead-letter queue, but every queue already has one: the failed set, where jobs land once they exhaust their attempts. This guide shows how to run that set deliberately — retention, inspection, alerting, and replay — as part of Dead-Letter Queues & Poison Messages in Queue Fundamentals & Architecture.

Problem Statement

An order-sync service processes BullMQ jobs that push orders to an ERP system. Workers are configured with removeOnFail: true to "keep Redis clean", so when an ERP outage caused 8,000 jobs to exhaust their retries, they vanished without a trace — the orders never reached the ERP, and finance discovered the gap during month-end reconciliation. On another queue the opposite happened: removeOnFail was unset, and 2 million failed jobs accumulated over a year, consuming 6 GB of Redis memory. You want failed jobs kept long enough to triage and replay, bounded so they cannot fill Redis, alerting when they appear, and a safe bulk-retry procedure after a fix.

Prerequisites

  • BullMQ 5.x workers and producers, a Redis instance with noeviction.
  • The ability to change job options (attempts, backoff, removeOnFail) at enqueue or as queue defaults.
  • A place to run operational scripts (a small admin CLI or a one-off job).
  • Metrics collection for queue counts (BullMQ's getJobCounts, a Prometheus exporter, or Bull Board).

Step 1 — Make Jobs Reach the Failed Set for the Right Reasons

A job enters the failed set when it throws on its final attempt, or immediately if it throws UnrecoverableError. Configure attempts and backoff for transient failures, and fail fast on permanent ones so the failed set contains only what needs attention.

import { Queue, Worker, UnrecoverableError } from "bullmq";

const queue = new Queue("erp-sync", {
  connection,
  defaultJobOptions: {
    attempts: 8,
    backoff: { type: "exponential", delay: 30_000 },   // 30 s, 60 s, 2 min ... ~1 h total
    removeOnComplete: { age: 3600, count: 1000 },
    removeOnFail: { age: 14 * 24 * 3600, count: 50_000 },   // Step 2
  },
});

const worker = new Worker("erp-sync", async (job) => {
  const order = await loadOrder(job.data.orderId);
  if (!order) throw new UnrecoverableError(`order ${job.data.orderId} deleted`);  // straight to failed
  const res = await erp.upsertOrder(order, { idempotencyKey: `order-${order.id}` });
  if (res.status === 422) throw new UnrecoverableError(`ERP rejected: ${res.body.error}`);
  if (res.status >= 500 || res.status === 429) throw new Error(`ERP ${res.status}`);  // retry
}, { connection, concurrency: 20 });

Eight attempts with exponential backoff span about an hour; an ERP outage shorter than that recovers without anyone touching the failed set. A longer outage fills it — which is exactly when you want the jobs preserved rather than deleted.

How a job reaches the failed set A job moves from waiting to active. If it succeeds it becomes completed. If it throws and attempts remain, it moves to delayed for its backoff and then back to waiting. If it throws on its final attempt, or throws UnrecoverableError at any attempt, it moves to failed, which acts as the queue's dead-letter set. The failed set is the dead-letter queue waiting active completed delayed (backoff) failed then back to waiting last attempt or Unrecoverable

Step 2 — Bound Retention by Age and Count

removeOnFail accepts true (delete immediately — no dead letters at all), false (keep forever — unbounded memory), a number (keep the last N), or { age, count } (keep up to N jobs no older than age seconds). The last form is the one to use.

removeOnFail: {
  age: 14 * 24 * 3600,   // 14 days: covers triage, weekends, and a fix-and-deploy cycle
  count: 50_000,         // hard cap: bounds memory even during a mass-failure event
}

Size the count from memory: a failed job stores its data, options, stack trace, and attempt history. Measure with MEMORY USAGE bull:erp-sync:<jobId> on a few samples — typically 2–10 KB each — so 50,000 failed jobs cost 100–500 MB. If a mass failure could exceed the cap, the oldest failures are removed first; Step 4 moves them somewhere durable before that happens. Memory planning for queues is covered in sizing Redis memory for queue backlogs.

Step 3 — Inspect Failures by Reason

Before retrying anything, group failures by reason to see whether you are looking at one cause or many.

// scripts/failed-report.ts
const failed = await queue.getFailed(0, 4999);              // newest first
const byReason = new Map<string, { count: number; sample: string }>();
for (const job of failed) {
  const key = (job.failedReason ?? "unknown").replace(/\d+/g, "N").slice(0, 120);
  const e = byReason.get(key) ?? { count: 0, sample: job.id! };
  e.count++; byReason.set(key, e);
}
console.table([...byReason].sort((a, b) => b[1].count - a[1].count)
  .map(([reason, v]) => ({ reason, count: v.count, sampleJob: v.sample })));
// ERP 503                          7_912   job 88121
// ERP rejected: invalid tax code      61   job 90112
// order N deleted                     27   job 90555

Normalising digits collapses messages that differ only in ids. The report above says: retry the 503s once the ERP is healthy; fix the tax-code mapping before retrying those 61; discard the deleted-order jobs. Bull Board or Taskforce provide the same view interactively.

One failed set, three different actions Grouping 8,000 failed jobs by normalised reason shows 7,912 ERP 503 responses from the outage, which can be retried once the ERP is healthy; 61 validation rejections for an invalid tax code, which need a code fix before retrying; and 27 jobs for deleted orders, which should be removed. Failed jobs by reason ERP 503: 7,912 retry when healthy invalid tax code: 61 — fix mapping, then retry order deleted: 27 — remove

Step 4 — Retry in Bulk, Safely

Retry by reason, in batches, with a rate limit, and only after confirming the cause is fixed. job.retry() moves a failed job back to waiting with its attempts reset.

// scripts/retry-failed.ts --reason "ERP 503" --batch 500 --pause-ms 2000
async function retryByReason(match: RegExp, batch = 500, pauseMs = 2000, dryRun = true) {
  let retried = 0;
  for (let start = 0; ; start += batch) {
    const jobs = await queue.getFailed(start, start + batch - 1);
    if (jobs.length === 0) break;
    const selected = jobs.filter((j) => match.test(j.failedReason ?? ""));
    if (!dryRun) await Promise.all(selected.map((j) => j.retry("failed")));
    retried += selected.length;
    await new Promise((r) => setTimeout(r, pauseMs));      // don't flood a recovering ERP
  }
  console.log(`${dryRun ? "would retry" : "retried"} ${retried} jobs`);
}

Paging with getFailed while retrying shifts indices (retried jobs leave the set), so production scripts should either collect ids first and then retry them, or loop until no matching jobs remain. For "retry everything", BullMQ also offers queue.retryJobs({ state: "failed", count: 1000 }), which does the move server-side in batches. Replay discipline — fix first, validate on a few, then rate-limited bulk — is the same as in replaying dead-letter messages in RabbitMQ.

Step 5 — Optionally Forward Failures to a Durable Store

The failed set lives in Redis with bounded retention. For jobs that must never be lost (financial syncs), forward each final failure to a durable store as it happens, so retention limits and Redis incidents cannot erase evidence.

const events = new QueueEvents("erp-sync", { connection });
events.on("failed", async ({ jobId, failedReason, prev }) => {
  const job = await Job.fromId(queue, jobId);
  if (!job || job.attemptsMade < (job.opts.attempts ?? 1)) return;   // not final yet
  await db.query(
    `INSERT INTO dead_letters (queue, job_id, name, data, reason, attempts, failed_at)
     VALUES ($1,$2,$3,$4,$5,$6, now()) ON CONFLICT (queue, job_id) DO NOTHING`,
    ["erp-sync", jobId, job.name, job.data, failedReason, job.attemptsMade]);
});
A durable copy of every final failure When a job fails on its final attempt, a QueueEvents listener writes its queue, id, name, data, reason, and attempt count to a dead_letters table in Postgres. The failed set in Redis keeps up to fourteen days or fifty thousand jobs; the database copy is kept as long as reconciliation needs, and replays can be driven from it. Redis for operations, Postgres for the record failed set (Redis) 14 days / 50k cap QueueEvents "failed" final attempt only dead_letters (Postgres) durable, queryable Month-end reconciliation queries the table, not a Redis set that may have been trimmed.

The database record survives Redis eviction, flushes, and the removeOnFail limits, and it is easy to query in reconciliation. Replaying from it means enqueueing a new job with the stored data.

Step 6 — Alert on Failed-Set Growth

A failed set that nobody watches is how 8,000 orders disappeared. Alert on growth rate and on absolute size, per queue.

# New failures in the last 10 minutes (from a counter incremented on final failure)
sum by (queue) (increase(bullmq_jobs_failed_final_total[10m])) > 20

# Failed-set size approaching its retention cap (from getJobCounts exported as a gauge)
bullmq_queue_jobs{state="failed"} / 50000 > 0.5

The first alert catches incidents in progress; the second catches slow accumulation and warns before retention starts deleting evidence. Routing and thresholds follow alerting on dead-letter queue growth.

Verification

it("keeps failed jobs with reasons and retries them after a fix", async () => {
  erpFake.failWith(503);
  const job = await queue.add("sync", { orderId: "o-1" }, { attempts: 2, backoff: { type: "fixed", delay: 50 } });
  await expect(job.waitUntilFinished(events, 5000)).rejects.toThrow(/ERP 503/);
  expect((await queue.getJobCounts("failed")).failed).toBe(1);
  erpFake.recover();
  await (await queue.getJob(job.id!))!.retry("failed");
  await expect((await queue.getJob(job.id!))!.waitUntilFinished(events, 5000)).resolves.toBeUndefined();
});

Gotchas & Edge Cases

removeOnFail: true in examples. Many tutorials set it to keep Redis tidy. In production it deletes your dead-letter queue.

Retrying stale jobs. A job failed two weeks ago may no longer be valid (the order was edited or cancelled). Handlers should re-read current state, not trust job data blindly.

Flows and parent jobs. A failed child can leave its parent waiting; decide per flow whether children use failParentOnFailure or ignoreDependencyOnFailure.

Retention cap during mass failure. If failures exceed count, the oldest are removed. Forward to durable storage (Step 5) for anything critical.

FAQ

Should I use a separate "dlq" queue instead of the failed set? Only if you need different tooling or retention for dead letters. The failed set already stores reasons, stack traces, and attempts, and retry() works in place; a separate queue adds a move step without much benefit.

How do I stop an outage from filling the failed set in the first place? Pause the queue when a dependency is known to be down (queue.pause(), or a circuit breaker in the processor that throws a retryable error with a long delay), so jobs wait instead of burning through their attempts. A failed set full of identical 503s is usually a sign that retries ran out before the outage did — lengthen the backoff horizon for dependencies with a history of long outages.

Do retried jobs keep their history? retry() resets attemptsMade and moves the job back to waiting; the failed reason and stack trace from previous attempts remain in job.stacktrace.

How do I discard jobs that should not be retried? job.remove() for individual jobs, or queue.clean(grace, limit, "failed") for bulk removal by age. Record what you discard and why.

Related