Benchmarking Redis Broker Throughput

Redis is fast enough that teams rarely measure it as a broker โ€” until the day it becomes the bottleneck. This guide benchmarks Redis the way a job framework actually uses it, as part of Capacity Planning for Job Queues in Observability & Monitoring for Job Queues. The headline number from redis-benchmark (hundreds of thousands of operations per second) says little about how many jobs per second a Sidekiq, BullMQ, Celery, or Asynq deployment can push through one Redis instance.

Problem Statement

A BullMQ deployment handles 3,000 jobs per second on a single Redis instance. Growth projections put next year's peak at 12,000. The platform team's back-of-envelope check โ€” redis-benchmark reports 180,000 SET operations per second on the same instance type โ€” suggests plenty of headroom. A load test disagrees: at about 7,500 jobs per second, the Redis main thread sits at 100% CPU, command latency jumps from 0.2 ms to 8 ms, and workers spend most of their time waiting. You need a benchmark that predicts that ceiling, explains it in commands per job and CPU per command, and tells you whether tuning, a bigger instance, or sharding across instances is the fix.

Prerequisites

  • A Redis instance matching production: version, instance type, persistence settings (AOF/RDB), TLS, and network placement relative to workers.
  • The job framework and version used in production, plus a no-op job type for benchmarking.
  • redis-cli access for INFO commandstats and LATENCY commands, and Prometheus redis_exporter metrics.
  • Enough load-generator and worker capacity to exceed the broker's limit without being the bottleneck themselves.

Step 1 โ€” Count the Redis Commands Behind One Job

A job is not one Redis operation. Enqueue, fetch, lock renewal, completion, event publishing, and cleanup each issue commands, often as Lua scripts that run several commands atomically. Measure it: reset stats, run a known number of no-op jobs, and divide.

redis-cli CONFIG RESETSTAT
node bench/run-noop-jobs.js --count 100000          # enqueue + process 100k no-op jobs
redis-cli INFO commandstats | awk -F'[:=,]' '/^cmdstat_/ {print $1, $3, $5}' | sort -k2 -n -r | head
# cmdstat_evalsha   calls=500213  usec=3910221     <- Lua scripts: add, moveToActive, moveToFinished...
# cmdstat_bzpopmin  calls=100522  usec=180114
# cmdstat_xadd      calls=300080  usec=210300      <- events stream
# cmdstat_hset      calls=100000  usec=51211
# ...

Summing calls and dividing by 100,000 gives commands per job (for BullMQ with default options, typically 8โ€“12, several of them Lua scripts). Summing usec and dividing gives Redis CPU microseconds per job โ€” the number that actually predicts the ceiling. In this scenario it comes to about 130 ยตs of Redis CPU per job, so a single core saturates at roughly 1,000,000 / 130 โ‰ˆ 7,700 jobs per second โ€” matching the load test.

One job, many Redis commands A single no-op BullMQ job costs about 130 microseconds of Redis CPU: the add-job Lua script about 30, fetching and moving to active about 35, lock handling about 10, the move-to-finished script about 40, and event stream writes about 15. At 130 microseconds per job, one Redis core saturates near 7,700 jobs per second. Redis CPU per job, microseconds add script 30 fetch + active 35 lock finish script 40 events 15 total โ‰ˆ 130 ยตs per job 1,000,000 ยตs / 130 ยตs โ‰ˆ 7,700 jobs/s per Redis core redis-benchmark's 180k SET/s measures a 5 ยตs command, not a 130 ยตs job.

Step 2 โ€” Benchmark End to End with the Real Framework

Drive the full path โ€” producers enqueueing, workers processing no-op jobs โ€” with enough of both that neither is the limit. Measure jobs per second completed and Redis CPU, stepping up the load.

// bench/throughput.ts โ€” BullMQ: N producers, M workers, no-op jobs
import { Queue, Worker } from "bullmq";

const connection = { host: process.env.REDIS_HOST, port: 6379 };
const queue = new Queue("bench", { connection });
let done = 0;

for (let i = 0; i < Number(process.env.WORKERS ?? 8); i++) {
  new Worker("bench", async () => {}, {
    connection, concurrency: 100,
    removeOnComplete: { count: 1000 },      // production-like cleanup, avoids unbounded growth
  }).on("completed", () => done++);
}

async function produce(rate: number, seconds: number) {
  const batch = 100, interval = (batch / rate) * 1000;
  const end = Date.now() + seconds * 1000;
  while (Date.now() < end) {
    await queue.addBulk(Array.from({ length: batch }, () => ({ name: "noop", data: {} })));
    await new Promise((r) => setTimeout(r, interval));
  }
}

for (const rate of [2000, 4000, 6000, 8000, 10000]) {
  const start = done;
  await produce(rate, 120);
  console.log(rate, "offered ->", ((done - start) / 120).toFixed(0), "completed/s");
}

Run producers and workers on separate hosts from Redis, in the same network zone as production. Record redis_cpu_sys_seconds_total + redis_cpu_user_seconds_total rates and redis_commands_duration_seconds_total during each step.

Step 3 โ€” Identify What Saturates

Three things can cap a Redis broker, and the metrics distinguish them:

# 1. Main-thread CPU: ~1.0 means the single command thread is saturated
rate(redis_cpu_user_seconds_total[1m]) + rate(redis_cpu_sys_seconds_total[1m])

# 2. Network: bytes in/out approaching the instance's bandwidth allowance
rate(redis_net_input_bytes_total[1m]) + rate(redis_net_output_bytes_total[1m])

# 3. Persistence: AOF fsync delays and fork time for RDB/rewrite
rate(redis_aof_delayed_fsync_total[1m])
redis_latest_fork_usec

In the scenario, main-thread CPU reached 0.98 at 7,500 jobs per second while network sat at 15% of the instance's bandwidth and AOF fsync was clean โ€” a classic CPU-bound broker. A network-bound broker shows the opposite: modest CPU, bandwidth near the limit, usually because payloads are large (see the claim-check pattern). A persistence-bound broker shows latency spikes aligned with fsyncs or forks.

CPU, network, or persistence? CPU-bound: main thread CPU near one core while bandwidth and fsync are fine; the fix is fewer or cheaper commands per job or sharding. Network-bound: bandwidth near the instance limit with modest CPU; the fix is smaller payloads. Persistence-bound: latency spikes aligned with AOF fsync delays or forks; the fix is persistence tuning or faster disks. Three ceilings, three fixes CPU-bound main thread ~1 core fix: fewer commands/job, or shard queues network-bound bandwidth near limit fix: smaller payloads, claim-check pattern persistence-bound spikes at fsync / fork fix: AOF everysec, faster disk, more RAM

Step 4 โ€” Reduce Redis Work per Job Before Scaling Hardware

Framework options change commands per job significantly. Measure each change with the Step 1 method.

// Options that cut Redis work per job in BullMQ
new Worker("bench", handler, {
  connection,
  concurrency: 100,
  removeOnComplete: { age: 3600, count: 1000 },  // bounded retention: less memory and cleanup
  removeOnFail: { age: 86400 },
  skipStalledCheck: false,                        // keep correctness; tune stalledInterval instead
  stalledInterval: 60_000,                        // fewer stalled-check scripts than the 30 s default
});

// Producers: add in bulk (one round trip, one script call per batch in recent versions)
await queue.addBulk(jobs);

// QueueEvents consumers read the events stream; cap it
new Queue("bench", { connection, streams: { events: { maxLen: 10_000 } } });

Other levers: avoid per-job progress updates for short jobs, keep job data small, and prefer removeOnComplete limits over keeping every completed job (which also bloats memory โ€” see sizing Redis memory for queue backlogs). In the scenario, these changes cut Redis CPU per job from 130 ยตs to about 85 ยตs, raising the single-instance ceiling to roughly 11,500 jobs per second.

Tuning moves the ceiling With default options, each job costs about 130 microseconds of Redis CPU and one instance tops out near 7,700 jobs per second. After bulk adds, bounded retention, a capped events stream, and a longer stalled-check interval, each job costs about 85 microseconds and the same instance reaches about 11,500 jobs per second, close to next year's 12,000 target. Jobs per second on one Redis instance defaults, 130 ยตs ~7,700 tuned, 85 ยตs ~11,500 Still short of 12,000 with headroom: sharding is needed next year, but not yet.

Step 5 โ€” Check Persistence and TLS Overhead

Benchmark with the persistence and TLS settings production uses; both cost CPU on the main thread or I/O threads.

# Compare the same benchmark under different settings
redis-cli CONFIG SET appendonly yes
redis-cli CONFIG SET appendfsync everysec     # production default: fsync in background thread
# run benchmark step 8000/s -> record CPU, p99 command latency

redis-cli CONFIG SET appendfsync always       # every write fsynced: much lower ceiling
# run again -> expect higher latency and lower max throughput

With TLS, Redis 6+ can offload TLS work to I/O threads (io-threads 4 with io-threads-do-reads yes), recovering much of the cost; without it, TLS can take a large share of the main thread. Persistence trade-offs for queues are covered in Redis persistence: AOF vs RDB for queues.

Step 6 โ€” Plan Sharding When One Core Is Not Enough

Redis executes commands on one thread, so a bigger instance helps only through faster cores and I/O threads. Beyond that, split queues across instances. Job frameworks shard naturally by queue: put different queues on different Redis instances, or use Redis Cluster with hash tags so each queue's keys stay on one slot.

// Shard by queue across Redis instances: each queue lives entirely on one instance
const shards = [
  { host: "redis-jobs-0", port: 6379 },
  { host: "redis-jobs-1", port: 6379 },
  { host: "redis-jobs-2", port: 6379 },
];
const shardFor = (queueName: string) => shards[hash(queueName) % shards.length];

// One logical high-volume queue split into N physical queues, each on its own shard
const physical = (logical: string, key: string) => `${logical}-${hash(key) % 6}`;
await new Queue(physical("events", tenantId), { connection: shardFor(physical("events", tenantId)) })
  .add("ingest", data);

Splitting one logical queue into several physical ones gives up global FIFO across the split, which rarely matters for independent jobs but does for ordered workloads โ€” see Message Ordering Guarantees. BullMQ on Redis Cluster requires queue names with a hash-tag prefix, covered in running BullMQ on Redis Cluster.

Verification

A benchmark report is trustworthy when three measurements agree:

predicted ceiling (1e6 / ยตs per job)         โ‰ˆ 11,500 jobs/s
measured end-to-end plateau (Step 2)         โ‰ˆ 11,100 jobs/s
Redis main-thread CPU at the plateau          โ‰ˆ 0.95-1.0

If the measured plateau is well below the prediction while Redis CPU is low, the bottleneck is elsewhere โ€” producers, workers, or network โ€” and the benchmark is not yet measuring the broker.

Gotchas & Edge Cases

redis-benchmark results. It measures a single simple command with pipelining. Use it to compare instance types, never to predict job throughput.

Benchmarking with an empty queue. Workers waiting on blocking pops use little CPU; a backlog changes the command mix. Run steps with a standing backlog too.

Slow Lua scripts under large sets. Some cleanup scripts scale with set size (completed or failed sets). Keep retention bounded, or throughput decays as the sets grow.

Managed Redis limits. Cloud offerings may cap connections, bandwidth, or command rate per node below the raw hardware. Benchmark the managed service itself.

FAQ

Is Redis Cluster faster than several standalone instances? For queues, not inherently โ€” both spread load across cores. Standalone instances per queue group are simpler to reason about; Cluster helps when you need automatic slot rebalancing and failover across many shards.

Would a bigger instance fix a CPU-bound broker? Only through a faster clock and I/O threads for networking. More cores do not speed up the command thread. Reduce commands per job or shard.

What ceiling should I plan against? Plan peak load at 50โ€“60% of the measured broker ceiling. Broker saturation hurts every queue on the instance at once, so it deserves more headroom than workers.

Related