Configuring the BullMQ Rate Limiter

BullMQ ships a queue-wide rate limiter enforced in Redis, so the limit holds across every worker process no matter how many you run. This guide configures it for a rate-limited third-party API, including reacting to 429 responses at runtime, as part of Rate Limiting & Throttling Jobs in Queue Fundamentals & Architecture.

Problem Statement

A marketing service sends SMS through a provider that allows 50 requests per second per account and returns 429 with a Retry-After header when exceeded. The service runs 6 worker pods with concurrency 20 each. During campaigns, 120 concurrent jobs fire at once, the provider returns hundreds of 429s, BullMQ retries them with exponential backoff, and some messages arrive 40 minutes late while others exhaust their attempts. Setting concurrency to 50 total did not help, because concurrency limits jobs in flight, not requests per second. You want at most 50 SMS per second across the whole fleet, automatic slowdown when the provider signals a limit, and no failed attempts spent on throttling.

Prerequisites

  • BullMQ 5.x (the rateLimit / RateLimitError API described here is from BullMQ 3+).
  • All workers for the queue using the same queue name and Redis instance.
  • The provider's documented limits and its throttling response format.
  • Metrics for completed jobs per second and 429 responses.

Step 1 — Understand Concurrency vs Rate

Concurrency caps how many jobs a worker runs at once; the rate limiter caps how many jobs the queue starts per time window, across all workers. They solve different problems and are usually both needed.

concurrency = 20 per worker x 6 workers = 120 jobs in flight
job duration ≈ 200 ms  -> up to 600 jobs/s could start
provider limit = 50 req/s
=> concurrency alone does not bound the rate; a limiter of 50 per 1000 ms does

A job that takes 200 ms with 120 slots can start 600 times a second. Only a limiter measured in jobs per duration matches a provider limit measured in requests per second.

In flight vs per second With 120 concurrency slots and 200-millisecond jobs, the fleet could start 600 jobs per second, twelve times the provider limit. The BullMQ limiter, stored in Redis and shared by all six workers, allows only 50 job starts per 1000 milliseconds, so the fleet never exceeds the provider's 50 requests per second regardless of how many slots are free. Job starts per second concurrency only up to 600/s: 429 storm limiter 50 / 1 s 50/s, shared by all 6 workers via Redis Keep concurrency too: it bounds memory and connections while the limiter bounds rate.

Step 2 — Configure the Queue-Wide Limiter

The limiter is set on the worker options, but its state lives in Redis per queue, so every worker enforces the same shared budget. Give every worker for the queue the same limiter configuration.

import { Worker } from "bullmq";

const worker = new Worker("sms", sendSms, {
  connection,
  concurrency: 20,                          // memory/connection bound per process
  limiter: {
    max: 45,                                // stay a little under the provider's 50/s
    duration: 1000,                         // per 1000 ms window
  },
});

Leaving about 10% headroom absorbs clock differences between BullMQ's window and the provider's, plus any other clients sharing the same provider account. When the limit is reached, workers stop fetching until the window resets; waiting jobs stay in the queue and are not charged an attempt.

Step 3 — React to 429 Responses with Manual Rate Limiting

Even with a limiter, providers throttle for their own reasons (account-wide limits, burst rules, incidents). When the provider returns 429, tell BullMQ to pause the whole queue for the Retry-After period, and hand the job back without counting a failed attempt.

import { Worker, RateLimitError } from "bullmq";

async function sendSms(job: Job<SmsData>, token?: string) {
  const res = await provider.send(job.data.to, job.data.body, { idempotencyKey: job.id });
  if (res.status === 429) {
    const retryAfterMs = Number(res.headers["retry-after"] ?? 1) * 1000;
    await worker.rateLimit(retryAfterMs);          // pause fetching queue-wide
    throw new RateLimitError();                    // job returns to waiting; no attempt used
  }
  if (res.status >= 500) throw new Error(`provider ${res.status}`);   // normal retry path
  if (res.status >= 400) throw new UnrecoverableError(`rejected: ${res.body.code}`);
  return { messageId: res.body.id };
}

worker.rateLimit(ms) sets a queue-wide pause visible to every worker, so one 429 slows the whole fleet rather than each worker discovering the limit separately. Throwing RateLimitError moves the job back to waiting with its attempt count unchanged — the fix for jobs exhausting attempts on throttling in the problem statement.

One 429 slows the whole fleet Worker 3 receives a 429 with Retry-After of 2 seconds. It calls rateLimit for 2000 milliseconds, which sets a queue-wide pause in Redis, and throws RateLimitError, which returns the job to waiting with its attempts unchanged. All six workers stop fetching for two seconds, then resume under the normal limiter. Handling a 429 from the provider 429 Retry-After: 2 rateLimit(2000) queue-wide, in Redis all 6 workers pause fetching 2 s job back to waiting, attempts unchanged Throttling is not a failure: it should never consume retry budget.

Step 4 — Separate Limits per Account or Tenant

The limiter applies to the whole queue. When limits are per provider account (each customer has its own SMS sender with its own quota), a single queue would let one busy account consume everyone's budget. Two common approaches:

// A. One queue per account, each with its own limiter (good for tens of accounts)
function workerFor(accountId: string, max: number) {
  return new Worker(`sms:${accountId}`, sendSms, { connection, concurrency: 5,
    limiter: { max, duration: 1000 } });
}

// B. BullMQ Pro groups: one queue, per-group rate limits and fair rotation (many accounts)
await queue.add("send", data, { group: { id: accountId } });
// worker: new WorkerPro("sms", sendSms, { connection, group: { limit: { max: 10, duration: 1000 } } })
Shared budget vs per-account budgets With one queue and one limiter of 45 per second, a campaign from account A fills the queue and consumes almost the whole budget, so account B's transactional messages wait behind it. With a queue or group per account, A is limited to its own quota and B's messages flow at B's rate independently. Account A's campaign vs account B's receipts one queue, 45/s shared A: 43/s B squeezed to 2/s, receipts late queue or group per account A: own 25/s B: own 20/s B unaffected by A's campaign Match the limiter's scope to the provider's: per account limits need per account budgets.

Per-queue limiters are available in open-source BullMQ and work well for a modest number of accounts. For thousands of tenants, BullMQ Pro's groups (or a custom token bucket per tenant, as in sliding window rate limiting with Redis Lua) avoid one queue per tenant. The fairness side of the problem is covered in preventing tenant starvation with weighted queues.

Step 5 — Monitor the Limiter Instead of Guessing

A rate-limited queue always has a backlog during bursts; the questions are whether it drains in acceptable time and whether 429s are rare.

const events = new QueueEvents("sms", { connection });
let completed = 0;
events.on("completed", () => completed++);
setInterval(async () => {
  const counts = await queue.getJobCounts("waiting", "delayed", "active");
  const ttl = await queue.getRateLimitTtl();          // ms until the limiter window/pause ends
  smsCompletedPerSec.set(completed / 10); completed = 0;
  smsWaiting.set(counts.waiting);
  smsRateLimitTtl.set(ttl);
}, 10_000);
# Throughput should sit at the limit during campaigns, not below it
avg_over_time(sms_completed_per_sec[5m])

# 429s should be rare after the limiter is tuned
sum(rate(sms_provider_responses_total{status="429"}[5m])) / sum(rate(sms_provider_responses_total[5m])) > 0.01

# Backlog drain estimate: waiting / limit
sms_waiting / 45

If completions sit well below the limit while jobs wait, workers are the bottleneck (too little concurrency for the job duration); if 429s stay high, lower max or check for other clients sharing the account.

Verification

it("never exceeds the configured rate across two workers", async () => {
  const stamps: number[] = [];
  const handler = async () => { stamps.push(Date.now()); };
  const w1 = new Worker(q.name, handler, { connection, concurrency: 50, limiter: { max: 10, duration: 1000 } });
  const w2 = new Worker(q.name, handler, { connection, concurrency: 50, limiter: { max: 10, duration: 1000 } });
  await q.addBulk(Array.from({ length: 50 }, () => ({ name: "x", data: {} })));
  await waitUntil(() => stamps.length === 50, 10_000);
  for (const t of stamps) {
    const inWindow = stamps.filter((s) => s >= t && s < t + 1000).length;
    expect(inWindow).toBeLessThanOrEqual(10);
  }
});

Gotchas & Edge Cases

Different limiter settings per worker. The limiter state is shared, but each worker applies its own max. Deploy the same configuration everywhere, or the effective limit depends on which worker fetches.

Delayed jobs and the limiter. Jobs delayed by backoff re-enter waiting and compete for the same budget; a large retry wave can crowd out fresh jobs. Keep retry volume low by fixing the causes of failure.

Priority under rate limits. Priorities decide which job starts next within the limit; they do not raise the limit. Urgent messages still wait for their turn in the window.

Rate limiting and graceful shutdown. A worker closing while the queue is rate-limited waits for in-flight jobs only; jobs waiting for the window are untouched and will be picked up by the remaining workers. There is no need to drain the limiter before deploys.

Idempotency on 429 retries. Some providers return 429 after partially accepting a request. Always send an idempotency key (the job id works) so a retried send after a throttle cannot produce a duplicate SMS.

Provider windows differ. A provider enforcing a rolling window may still throttle a fixed-window limiter at window boundaries. Headroom and the 429 handler cover it.

FAQ

Does the limiter count failed jobs? Every job start counts, including ones that fail. Jobs returned with RateLimitError count against the window they were started in.

How long will a campaign take to send under the limit? Divide the number of messages by the effective rate: 90,000 SMS at 45 per second is about 33 minutes. Put that number in front of whoever schedules campaigns, because it is a property of the provider contract, not of the worker fleet — adding workers will not shorten it. If the business needs faster delivery, the lever is a higher provider limit or splitting traffic across accounts that each have their own quota.

Should the limit live in code or configuration? Configuration, read at worker start, so it can be adjusted when the provider changes a quota without a code deploy. Log the active limit on startup and export it as a metric, so dashboards show throughput against the limit actually in force.

Can I limit by job name within one queue? Not with the built-in limiter — it is per queue. Use separate queues per limit, or implement a token bucket keyed by name.

What's the difference between worker.rateLimit and pausing the queue? rateLimit(ms) is a timed pause that expires on its own and is meant for throttling; queue.pause() stops processing until someone resumes it and is meant for incidents or maintenance.

Related