BullMQ Job Priorities

BullMQ can order jobs within a queue by a numeric priority, which is often the first tool teams reach for when urgent work waits behind bulk work. This guide shows how priorities actually behave, where they cause starvation, and when separate queues are the better design, as part of Priority Queues & Job Fairness in Queue Fundamentals & Architecture.

Problem Statement

A document service processes three kinds of jobs on one BullMQ queue: interactive previews a user is waiting for, uploads that should finish within minutes, and a nightly re-index of 2 million documents. During the re-index, previews waited up to 40 minutes. The team added priority: 1 to previews, priority: 5 to uploads, and priority: 10 to re-index jobs. Previews became fast — and then the re-index never finished, because during business hours there was always a higher-priority job waiting. A week later the re-index backlog was 9 million jobs. You want previews fast, uploads timely, the re-index guaranteed to make progress, and a way to see waits per class.

Prerequisites

  • BullMQ 5.x with Redis 6.2+.
  • A classification of job types into latency classes with rough volume per class.
  • Metrics for queue wait time (enqueue to start), ideally per job name or class.
  • Workers you can deploy with different queue subscriptions if you split queues.

Step 1 — Understand How BullMQ Orders Prioritized Jobs

A job added with priority goes into a prioritized set instead of the plain wait list. Lower numbers run first (1 is highest; the maximum is 2,097,152). Jobs with the same priority run in insertion order. Jobs without a priority are treated as the highest class and, in current versions, are taken before prioritized ones.

await queue.add("preview",  { docId }, { priority: 1 });    // highest
await queue.add("upload",   { docId }, { priority: 5 });
await queue.add("reindex",  { docId }, { priority: 10 });   // lowest
await queue.add("legacy",   { docId });                     // no priority: taken before all of the above

Two consequences follow. First, mixing prioritized and non-prioritized jobs on one queue is confusing — give every job a priority or none. Second, priority is strict: a worker always takes the lowest-numbered waiting job. As long as priority-1 or priority-5 jobs keep arriving, a priority-10 job never starts.

Strict priority starves the bottom class Workers take the waiting job with the lowest priority number. Previews at priority 1 and uploads at priority 5 arrive continuously during business hours, so there is always a higher-priority job waiting. Re-index jobs at priority 10 start only when both higher classes are empty, which during the day is never, and their backlog grows. One queue, strict priorities, business hours priority 1: previews priority 5: uploads priority 10: re-index workers take lowest number re-index backlog grows to 9 million

Step 2 — Know the Cost of Prioritized Jobs

Prioritized jobs are stored in a sorted set, so adding and fetching them is O(log n) rather than O(1) for the plain wait list. With millions of prioritized jobs waiting, that cost is measurable on the Redis main thread. Priorities also make the queue's order harder to reason about in dashboards.

// Inspect the prioritized set size vs waiting list
const counts = await queue.getJobCounts("waiting", "prioritized", "active", "delayed");
console.log(counts);   // { waiting: 0, prioritized: 9_214_332, active: 40, delayed: 12 }

A 9-million-entry sorted set is not a problem for correctness, but it is a signal that priority is being used for a volume class (bulk re-index) rather than an urgency class. Bulk work belongs in its own queue.

Step 3 — Split Latency Classes into Separate Queues

The robust design for classes with very different volumes and latency needs is one queue per class, each with its own workers or a weighted share of workers. Separate queues give guaranteed capacity to each class, which priorities cannot.

// Producers pick the queue by class
const previews = new Queue("docs-interactive", { connection });
const uploads  = new Queue("docs-standard",    { connection });
const reindex  = new Queue("docs-bulk",        { connection });

// Workers: capacity is allocated, not contested
new Worker("docs-interactive", render, { connection, concurrency: 20 });   // always available for users
new Worker("docs-standard",    process, { connection, concurrency: 20 });
new Worker("docs-bulk",        reindexDoc, { connection, concurrency: 10,
                                             limiter: { max: 200, duration: 1000 } });  // bounded impact

The re-index now always has 10 slots, so it finishes in a predictable time; previews always have 20, so a re-index burst cannot delay them. Scale each class independently on its own backlog. The general approach is in routing high-priority jobs in Celery, which applies equally here.

Allocated capacity instead of contested order Previews go to docs-interactive with twenty dedicated slots, uploads to docs-standard with twenty, and re-index jobs to docs-bulk with ten slots and a rate limit of two hundred per second. Each class progresses at its own rate: previews stay fast during the re-index, and the re-index always makes progress during business hours. One queue per latency class docs-interactive (previews) docs-standard (uploads) docs-bulk (re-index) 20 slots 20 slots 10 slots, 200/s cap p95 wait under 2 s p95 wait under 3 min finishes in ~3 h

Step 4 — Use Priorities Within a Class

Priorities remain useful inside a queue where jobs share a latency class but some should jump ahead: a paying customer's upload before a free-tier one, or the first page of a document before the rest.

// Within docs-standard: plan tiers as priorities, with a small range
const TIER_PRIORITY = { enterprise: 1, pro: 2, free: 3 } as const;
await uploads.add("upload", { docId }, { priority: TIER_PRIORITY[account.tier] });

This works because the volumes are the right shape: enterprise and pro uploads are a minority of traffic, so they get ahead of the line without ever saturating the queue's 20 slots, and free-tier uploads still start within their class's objective. The rule of thumb is that the sum of all classes above the lowest should stay comfortably below the queue's capacity — if it does not, priority stops being a tie-breaker and becomes starvation.

Priority as a tie-breaker inside one class In the docs-standard queue, enterprise uploads at priority 1 are about 10 percent of load and pro uploads at priority 2 about 20 percent. Together they use under a third of the queue's capacity, so they start first whenever they arrive, and free-tier uploads at priority 3 still receive the remaining two thirds and meet their wait objective. docs-standard: share of capacity used by tier ent 10% pro 20% free 50% headroom 20% Higher tiers jump ahead but cannot fill the queue, so the lowest tier never starves. If the top tiers could reach 100% of capacity, split them into their own queue instead.

Keep the range small and the volumes of the top classes modest relative to capacity. If the higher tiers alone can saturate the workers, lower tiers starve here too — the fairness techniques in preventing tenant starvation with weighted queues apply.

Step 5 — Change Priority of Waiting Jobs

Sometimes a waiting job becomes urgent: a user opens a document whose preview is queued as a background job. BullMQ lets you change the priority of a job that has not started.

async function expedite(jobId: string) {
  const job = await Job.fromId(uploads, jobId);
  if (job && (await job.isWaiting() || await job.getState() === "prioritized")) {
    await job.changePriority({ priority: 1 });          // moves it ahead in the prioritized set
  }
}

For cross-class promotion (a bulk job that should become interactive), add a new job to the interactive queue with a deduplication id and let the bulk copy become a no-op when it eventually runs — simpler than moving jobs between queues.

Step 6 — Monitor Wait Time per Class

Priorities and class queues are only as good as the waits they produce. Measure queue wait per class and alert on the classes that have objectives.

const worker = new Worker("docs-interactive", async (job) => {
  jobWaitSeconds.observe({ queue: job.queueName, name: job.name },
                         (Date.now() - job.timestamp) / 1000);      // enqueue -> start
  return render(job);
}, { connection, concurrency: 20 });
histogram_quantile(0.95, sum by (le, queue) (rate(job_wait_seconds_bucket[5m])))

Watch the lowest class as closely as the highest: starvation shows up as a steadily growing wait in the bottom class while the top looks perfect. Measuring waits from enqueue timestamps is covered in measuring queue wait time with enqueue timestamps.

Verification

After splitting queues, run the nightly re-index during a busy period in staging and confirm: interactive p95 wait stays under its objective, the bulk queue's completion rate matches its allocated capacity, and the prioritized set on the standard queue stays small.

Gotchas & Edge Cases

Default jobs jump the line. Jobs added without a priority to a queue that uses priorities may run first. Make priority mandatory in your enqueue helper.

Priority is not a rate limit. Low-priority jobs still hit dependencies as hard as high ones once they start. Use a limiter for bulk classes.

Large prioritized sets. Millions of prioritized jobs slow adds and fetches. Move volume classes to their own queue.

Retries keep their priority. A failed priority-10 job retries at priority 10 and waits behind new higher-priority work, so its retry delay in practice can be much longer than its backoff setting. Account for that when a low-priority job has a deadline.

Workers on shared queues. If the same workers consume all class queues, give interactive queues their own worker deployment so a burst in another class cannot occupy its processes.

FAQ

How many priority levels should I use? Two or three within a queue. More levels are hard to reason about and rarely change outcomes.

Does BullMQ support weighted fair scheduling? Not across priorities in open-source BullMQ; strict order only. Use separate queues with allocated concurrency, or BullMQ Pro groups for per-group fairness.

How do I migrate from priorities on one queue to separate queues? Create the new queues and their workers first, then switch producers per job type so new jobs land on the right queue, and let the old queue drain with its existing workers. Because the old queue still holds the prioritized backlog, keep its workers until it is empty — for a 9-million-job bulk backlog, move those workers' capacity gradually to the new bulk queue as the old one shrinks.

Should the interactive path use a queue at all? If a user is waiting synchronously and the work is short, doing it inline may be better. Use the queue when the work is long, bursty, or needs retries.

Related