BullMQ Job Schedulers for Repeatable Jobs

BullMQ's job schedulers produce jobs on a cron pattern or fixed interval without a separate scheduler process, and this guide sets them up for production, as part of Scheduled & Delayed Jobs in Queue Fundamentals & Architecture. Job schedulers (upsertJobScheduler, BullMQ 5.16+) replace the older "repeatable jobs" API, fixing its main operational pain: duplicate schedules created every time an app booted with slightly different options.

Problem Statement

A Node.js service registers recurring jobs at startup with queue.add(name, data, { repeat: { pattern } }). Over a year of deploys, small changes to the options — a new time zone, an edited cron pattern — created new repeatable entries without removing the old ones. The queue now has 14 repeat configurations for 5 logical schedules; the nightly invoice job runs three times, at 01:00 UTC, 02:00 UTC, and 02:00 Berlin. Nobody is sure which entries are live. You want exactly one schedule per logical job, defined in code and updated in place on deploy, correct time zones, visibility into next run times, and a clean migration away from the duplicates.

Prerequisites

  • BullMQ 5.16 or later (for upsertJobScheduler, getJobSchedulers, removeJobScheduler).
  • Workers for the queue running with the matching processor for each scheduled job name.
  • A decision on each schedule's time zone (UTC for machine tasks; a named zone for human-facing ones).
  • Idempotent processors — a scheduled job can occasionally run late or be retried.

Step 1 — Define Schedulers with Stable Ids

A job scheduler is identified by an id you choose. upsertJobScheduler creates it if missing and updates it in place if it exists, so calling it on every boot is safe and idempotent.

// schedules.ts — the single source of truth for recurring jobs
import { Queue } from "bullmq";

const queue = new Queue("maintenance", { connection });

export async function registerSchedules() {
  await queue.upsertJobScheduler(
    "nightly-invoices",                                   // stable id: one per logical schedule
    { pattern: "0 2 * * *", tz: "Europe/Berlin" },        // 02:00 Berlin, DST-aware
    { name: "generate-invoices", data: { region: "eu" },
      opts: { attempts: 5, backoff: { type: "exponential", delay: 60_000 },
              removeOnComplete: { count: 100 }, removeOnFail: { age: 7 * 86400 } } },
  );
  await queue.upsertJobScheduler(
    "refresh-rates",
    { every: 15 * 60_000 },                              // every 15 minutes, elapsed time
    { name: "refresh-exchange-rates", data: {} },
  );
}

The id is the key difference from the legacy API, where the identity was derived from the name and options — so changing an option created a new schedule. With stable ids, changing the pattern or data updates the one existing scheduler.

Stable ids stop schedule duplication With legacy repeatable jobs, the schedule key is derived from the job name and repeat options, so each deploy that changed the cron pattern or time zone added another schedule, leaving three live copies of the nightly invoice job. With upsertJobScheduler, the id nightly-invoices identifies the schedule, and each deploy updates that single entry. Three deploys that changed the schedule legacy repeat options 0 1 * * * UTC (still live) 0 2 * * * UTC (still live) 0 2 * * * Berlin (live) upsertJobScheduler id nightly-invoices 0 2 * * * Berlin (updated in place) One logical schedule, one entry, whatever the deploy history.

Step 2 — Register Schedulers Once per Deploy, Not per Worker

upsertJobScheduler is idempotent, so calling it from every worker on startup is harmless, but it is cleaner to run registration from one place — a deploy step or a single "scheduler" entry point — so the definition of what is scheduled is not scattered.

// scripts/register-schedules.ts — run in the deploy pipeline after migrations
await registerSchedules();
await pruneUnknownSchedulers(["nightly-invoices", "refresh-rates"]);   // Step 5
await queue.close();

The scheduler state lives in Redis; workers only need to process the jobs it produces. Unlike a separate cron process, there is no single scheduler instance to keep alive — BullMQ adds the next job when the previous one is taken, so the schedule continues as long as any worker is running.

Step 3 — Choose Pattern vs Every, and Set the Time Zone

A pattern is a cron expression evaluated in tz (default UTC); every is a fixed interval in milliseconds, unaffected by time zones. Match the kind of schedule to the requirement.

// Human-facing: follows the local clock, including DST changes
{ pattern: "0 9 * * 1-5", tz: "America/New_York" }     // 09:00 New York, weekdays

// Machine-facing: elapsed time
{ every: 60_000 }                                       // every minute

// Start later, stop after N runs or a date
{ pattern: "*/30 * * * *", startDate: new Date("2026-10-01"), endDate: new Date("2026-12-31") }
{ every: 3_600_000, limit: 24 }                         // 24 hourly runs, then stop

Cron patterns in a named zone inherit that zone's DST behaviour for skipped and repeated hours; if exact semantics matter (a business close that must run once per local day), keep the job idempotent by business date as described in timezone-safe job scheduling across DST.

Step 4 — Understand How Runs Are Produced

A job scheduler keeps exactly one delayed job queued: the next occurrence. When a worker picks it up (it becomes active), BullMQ creates the following occurrence as a new delayed job. That design has two consequences worth knowing.

const schedulers = await queue.getJobSchedulers(0, -1, true);
for (const s of schedulers) {
  console.log(s.key, s.pattern ?? `every ${s.every}ms`, s.tz ?? "UTC", new Date(s.next!).toISOString());
}
// nightly-invoices   0 2 * * *        Europe/Berlin  2026-09-19T00:00:00.000Z
// refresh-rates      every 900000ms   UTC            2026-09-18T12:15:00.000Z

First, runs do not overlap by construction only in the sense of production — if a run takes longer than the interval, the next occurrence is already scheduled and another worker can start it. Guard long jobs with a lock or make them idempotent. Second, if no worker is running when an occurrence is due, it simply waits in the delayed set and runs when workers return — once, not once per missed interval.

One occurrence queued at a time The scheduler for refresh-rates holds one delayed job for 12:15. At 12:15 the job moves to waiting, a worker takes it, and when it becomes active BullMQ adds the next delayed job for 12:30. If workers are down from 12:20 to 13:10, the 12:30 job waits and runs once at 13:10, then 13:15 is scheduled; missed intervals are not replayed. refresh-rates, every 15 minutes delayed 12:15 active 12:15 delayed 12:30 added workers down 12:20-13:10: 12:30 job waits runs once at 13:10, then 13:15 next Missed intervals are not replayed; jobs needing catch-up must compute what they missed.

Step 5 — Remove Orphaned and Legacy Schedules

Clean up by comparing what exists in Redis with the ids defined in code. Remove anything unknown, and remove legacy repeatable entries created with the old API.

async function pruneUnknownSchedulers(knownIds: string[]) {
  const existing = await queue.getJobSchedulers(0, -1);
  for (const s of existing) {
    if (!knownIds.includes(s.key)) {
      console.warn(`removing unknown scheduler ${s.key}`);
      await queue.removeJobScheduler(s.key);
    }
  }
  // Legacy repeatable jobs (pre-5.16 API)
  const legacy = await queue.getRepeatableJobs();
  for (const r of legacy) {
    console.warn(`removing legacy repeatable ${r.name} ${r.pattern ?? r.every} ${r.tz ?? ""}`);
    await queue.removeRepeatableByKey(r.key);
  }
}

Run the prune in a dry-run mode first and review the list: in the scenario it found nine stale entries, including two nobody could explain. After pruning, the next-run listing from Step 4 is the definitive answer to "what is scheduled?". Removing a scheduler removes its pending delayed job too; an already-active run finishes normally.

Step 6 — Protect Long Runs and Alert on Missed Ones

A scheduled job that overruns its interval can overlap with the next occurrence. Guard it with a short-lived lock, and alert when a schedule's job has not completed recently.

const worker = new Worker("maintenance", async (job) => {
  if (job.name === "generate-invoices") {
    const lock = await redis.set(`lock:${job.name}`, job.id!, "PX", 3_600_000, "NX");
    if (!lock) return { skipped: "previous run still active" };
    try { await generateInvoices(job.data.region); }
    finally { await redis.eval(RELEASE_IF_OWNER, 1, `lock:${job.name}`, job.id!); }
  }
}, { connection });

worker.on("completed", (job) => lastSuccess.set({ job: job.name }, Date.now() / 1000));
# Nightly invoices should have succeeded within the last 26 hours
time() - last_success_timestamp_seconds{job="generate-invoices"} > 26 * 3600
One freshness alert, three failure modes Three different problems stop the nightly invoice job from completing: the scheduler was removed by a bad prune, no workers are running for the queue, or the job itself fails every night. In all three cases the last successful completion timestamp stops advancing, so an alert on time since last success above 26 hours fires regardless of cause. Alert on the outcome, not on each cause scheduler removed no workers running job fails every night last success stops advancing alert age > 26 h

A "last success" freshness alert catches every failure mode at once — scheduler removed, workers down, job failing — which is why it is more useful than alerting on each cause separately. See alerting on stuck and stalled jobs.

Verification

it("upserting twice keeps one scheduler and updates it", async () => {
  await queue.upsertJobScheduler("s1", { every: 1000 }, { name: "tick" });
  await queue.upsertJobScheduler("s1", { every: 2000 }, { name: "tick" });
  const all = await queue.getJobSchedulers();
  expect(all).toHaveLength(1);
  expect(all[0].every).toBe(2000);
});

In production, after the migration, confirm that getRepeatableJobs() returns nothing, that getJobSchedulers() lists exactly the ids in code, and that the nightly invoice job ran once on the next night.

Gotchas & Edge Cases

Mixed API during migration. Deploy the prune and the new registration together; running old code that still calls add(..., { repeat }) after pruning recreates the legacy entries.

Changing every alignment. Interval schedulers align to the time of creation or update, not to clock boundaries. Use a cron pattern if runs must happen on the quarter hour.

Data changes and in-flight jobs. Updating a scheduler's data affects future occurrences; the already-queued next job may still carry the old data depending on version. Make processors tolerant of both.

Large numbers of schedulers. Thousands of per-customer schedulers work but add Redis keys and delayed jobs; for per-customer local times, one scheduler per zone that fans out is lighter.

FAQ

Do I still need a separate cron process? No — BullMQ produces scheduled jobs from within Redis operations performed by workers and producers. Keep at least one worker running for the queue.

What happens during Redis failover? Scheduler state and the pending delayed job are in Redis; with replication, a failover preserves them subject to the replication lag. A missed run after failover shows up in the freshness alert.

How do I trigger a scheduled job manually, outside its schedule? Add a normal job with the same name and data (queue.add("generate-invoices", { region: "eu" })). It runs through the same processor and does not affect the scheduler's next occurrence. Give manual runs a distinguishing field or job id prefix so logs and metrics show which runs were scheduled and which were operator-initiated.

Can one scheduler produce jobs on another queue? A scheduler belongs to its queue. Have the scheduled job enqueue work onto other queues if needed.

Related