BullMQ Sandboxed Processors

Node.js runs job handlers on one event loop, so a single CPU-heavy or leaky job can starve every other job in the process — and the timers that keep job locks alive. BullMQ's sandboxed processors run each job in a separate child process or worker thread, and this guide shows when and how to use them, as part of BullMQ for Node.js Ecosystems in Backend Frameworks & Worker Scaling.

Problem Statement

A media service processes image jobs with sharp and a few pure-JavaScript transforms, and PDF jobs with a JavaScript PDF library, all in one BullMQ worker at concurrency 10. Three problems keep recurring: jobs stall because the event loop is blocked for many seconds during large PDF renders; a memory leak in the PDF library grows the worker to 3 GB over a day until it is OOM-killed, taking all 10 in-flight jobs with it; and a native crash in one image codec terminates the whole process. You want CPU-heavy work off the main event loop, a leak or crash in one job contained to that job, parallelism matched to CPU cores, and progress and logs still flowing from sandboxed jobs.

Prerequisites

  • BullMQ 5.x and Node 20+.
  • Processor code that can live in its own module file (it will be loaded separately in each sandbox).
  • Knowledge of which job types are CPU-bound, memory-hungry, or depend on native code with a crash history.
  • Container CPU and memory limits known for the worker pods.

Step 1 — Know What the Sandbox Changes

With an in-process processor, the worker's main thread runs your handler. With a sandboxed processor, the main thread only coordinates — fetching jobs, renewing locks, emitting events — while each job runs in a child process (default) or worker thread (useWorkerThreads: true). The child receives the job data, runs the module's default export, and sends back the result or error.

in-process:   main thread = fetch + locks + YOUR CODE          (blocking code blocks everything)
sandboxed:    main thread = fetch + locks + events
              child/thread = YOUR CODE, one job at a time per child

Isolation is the point: CPU work in a child cannot delay lock renewals in the parent, and a crash or leak in a child ends that child, not the worker. The costs are startup time for children, serialization of job data and results across the process boundary, and more memory per concurrent job.

Coordinator plus isolated children In the in-process model, ten jobs share the main event loop with lock renewal, so one long synchronous render blocks all of them and a crash ends all of them. In the sandboxed model, the main thread only fetches jobs and renews locks, and each job runs in its own child process, so blocking, leaks, and crashes stay inside one child. Where job code runs in-process one event loop: locks + 10 jobs together one crash = 10 jobs lost sandboxed main: locks, fetch, events child 1: job child 2: job child 3: job Sandboxing trades some startup and serialization cost for isolation.

Step 2 — Move the Processor into Its Own File

A sandboxed processor is a module whose default export takes a SandboxedJob. Pass the file path to the Worker instead of a function.

// src/processors/pdf.sandbox.ts  (compiled to dist/processors/pdf.sandbox.js)
import type { SandboxedJob } from "bullmq";
import { renderPdf } from "../pdf/render";
import { uploadToStorage } from "../storage";

export default async function (job: SandboxedJob<{ docId: string; template: string }>) {
  const pdf = renderPdf(job.data.template, await loadFields(job.data.docId));  // CPU-heavy, sync
  await job.updateProgress(80);
  const key = await uploadToStorage(`pdf/${job.data.docId}.pdf`, pdf);
  return { key, bytes: pdf.length };
}
// src/worker.ts
import { Worker } from "bullmq";
import path from "node:path";

const pdfWorker = new Worker("pdf", path.join(__dirname, "processors/pdf.sandbox.js"), {
  connection,
  concurrency: 4,                    // ≈ cores available to the pod for CPU-bound work
  useWorkerThreads: false,           // child processes: full memory/crash isolation
});

The sandbox module is loaded fresh in each child, so it must import everything it needs itself — database clients, storage clients, configuration. Keep those imports lightweight; heavy initialisation happens once per child, not per job, as children are reused.

Step 3 — Choose Child Processes or Worker Threads

useWorkerThreads: true runs processors in worker_threads instead of forked processes. Threads start faster and use less memory; processes isolate more strongly.

Concern Child processes (default) Worker threads
Blocking the parent's event loop Isolated Isolated
Memory leak in job code Contained; child can be recycled Shared heap limits per thread, but same process
Native crash (segfault) Kills the child only Kills the whole process
Startup and memory overhead Higher (a Node process per slot) Lower
Data transfer Serialized over IPC Structured clone, faster

Use child processes for native libraries with a crash history and for leaky code; use worker threads for pure-JavaScript CPU work where overhead matters. The media service in the scenario uses processes for the PDF and codec jobs and threads for a lightweight transform queue.

Processes for isolation, threads for speed Jobs using native code with a history of segfaults, or libraries that leak memory, run in child processes so a crash or leak ends only that child. Pure-JavaScript CPU-bound jobs with no crash history run in worker threads, which start faster and use less memory. Both keep the parent's event loop free for lock renewal. Which sandbox for which job child processes native codecs, PDF library leaks a crash ends one child worker threads pure JS transforms, hashing lower overhead, faster start I/O-bound jobs need neither: keep them in-process at high concurrency.

Step 4 — Match Concurrency to Cores and Memory

For CPU-bound sandboxed work, concurrency above the pod's CPU count just time-slices, raising every job's duration. Memory sets the other bound: each child holds its own heap.

# Kubernetes resources for the pdf worker
resources:
  requests: { cpu: "4", memory: "4Gi" }
  limits:   { memory: "6Gi" }          # 4 children x ~1.2 GB peak + parent
import os from "node:os";
const cores = Number(process.env.CPU_REQUEST ?? os.availableParallelism());
new Worker("pdf", pdfFile, { connection, concurrency: cores });

Scale throughput by adding pods, not by raising concurrency past the core count. The sizing method is in right-sizing worker concurrency per CPU.

Step 5 — Contain Leaks by Recycling Children

A leaking library in a long-lived child still grows without bound — just inside the child. Recycle children periodically, or cap their heap so a leaking child dies and is replaced instead of growing the pod until the OOM killer takes everything.

new Worker("pdf", pdfFile, {
  connection,
  concurrency: 4,
  workerForkOptions: { execArgv: ["--max-old-space-size=1024"] },   // child heap cap: 1 GB
});

A child that exceeds its heap cap crashes with an out-of-memory error; the job fails (and retries per its options) and BullMQ starts a fresh child. One job pays, instead of ten.

A heap cap turns a pod crash into a job retry Without sandboxing, a leaking PDF library grows the single worker process steadily over a day until it reaches the container limit and the pod is OOM-killed with ten jobs in flight. With child processes capped at one gigabyte of heap, each child grows until it hits the cap, crashes with one job in flight, and is replaced by a fresh child, keeping total memory bounded. Worker memory over one day pod limit in-process: OOM kill, 10 jobs lost child hits 1 GB cap, replaced 00:00 24:00

If crashes recur on specific inputs, those inputs are the bug — see the leak diagnosis approach in fixing Celery worker memory leaks, which translates directly to Node heap snapshots.

Step 6 — Report Progress and Logs from the Sandbox

job.updateProgress and job.log work inside sandboxes; calls are forwarded to the parent over IPC. Structured logs from children go to the same stdout, so give them job context explicitly.

export default async function (job: SandboxedJob) {
  const log = baseLogger.child({ job_id: job.id, job_name: job.name, pid: process.pid });
  log.info("render start");
  await job.updateProgress(10);
  await job.log("fields loaded");                   // visible in Bull Board / Taskforce
  // ...
}

Logging with the child's pid makes it easy to connect a crash to the last lines that child printed. The logging setup is covered in structured JSON logging for BullMQ workers.

Verification

Measure the two symptoms from the problem statement before and after moving to sandboxes: event-loop delay in the parent and stalled-job count.

import { monitorEventLoopDelay } from "node:perf_hooks";
const h = monitorEventLoopDelay({ resolution: 20 }); h.enable();
setInterval(() => { parentLoopMaxMs.set(h.max / 1e6); h.reset(); }, 10_000);

Expect parent event-loop delay to drop from seconds to milliseconds and stalls to disappear. Then kill a child process manually (kill -9 on a child pid) and confirm only its job fails and retries while the other three continue. The stall mechanics are in BullMQ lock duration and stalled jobs.

Gotchas & Edge Cases

Large job data or results. Everything crosses the process boundary by serialization. Pass ids and storage keys; return small results.

TypeScript paths. The worker loads the compiled .js file by path. Point it at the build output, and make sure the file is included in the container image.

Shared connections. Each child opens its own database and Redis connections. With concurrency 4 and 10 pods, that is 40 extra connection sets; size pools accordingly.

Child startup on the first jobs. Children are created lazily as jobs arrive, so the first job per slot after a deploy pays the startup cost of loading the module and its dependencies. For latency-sensitive queues, warm the workers with a no-op job after deploy, or accept a slower first minute.

Graceful shutdown. worker.close() waits for sandboxed jobs to finish, then terminates the children. Keep the pod's grace period above the longest sandboxed job, or cancelled jobs will be retried after the lock expires — the drain pattern is in graceful shutdown & worker deployments.

Debugging. Stack traces from children show the child's frames; attach inspectors to children with workerForkOptions.execArgv when debugging locally.

FAQ

Should every BullMQ worker use sandboxes? No. I/O-bound handlers (HTTP calls, database writes) do not block the event loop and are cheaper in-process at high concurrency. Sandbox CPU-heavy, leaky, or crash-prone code.

Do sandboxes make jobs slower? For CPU-bound jobs, total throughput usually improves because work runs in parallel across cores. For tiny jobs, IPC overhead can dominate.

Can one worker mix sandboxed and in-process job types? A worker has one processor. Run separate workers (or separate queues) for sandboxed and in-process job types.

Is a separate service better than a sandbox? For very heavy work (video transcoding, ML inference), a dedicated service or container per job may be cleaner. Sandboxes suit moderate CPU work that belongs with the rest of the worker code.

Related