Choosing a Redis maxmemory Policy for Queues

Redis decides what to do when it reaches its memory limit from a single setting, maxmemory-policy. For a cache, evicting old keys is exactly right. For a job queue, eviction means jobs disappear without an error anywhere. This guide sets the policy correctly and handles the consequences, as part of In-Memory vs Persistent Queue Storage in Backend Frameworks & Worker Scaling.

Problem Statement

A team uses one managed Redis instance for both the application cache and BullMQ. The provider's default policy is volatile-lru, and the team later switched it to allkeys-lru to improve cache hit rates. During a large import, the backlog grew to 3 GB and Redis reached its 4 GB limit. Nothing crashed. Weeks later, customers reported missing invoices: Redis had evicted job hashes and, in some cases, the internal keys BullMQ uses to track job state, leaving IDs in lists that pointed at nothing. Workers logged "missing key for job" warnings that nobody watched. You want jobs never to be evicted, a clear error when memory is exhausted, and enough warning to act before that happens.

Prerequisites

  • Access to the Redis configuration (CONFIG SET or the provider's parameter group).
  • Knowledge of which applications share the instance.
  • Metrics for used_memory, maxmemory, and evicted_keys (from INFO or an exporter).
  • Producers that can handle an enqueue error (retry, fail the request, or spill to another store).

Step 1 — Understand What Each Policy Does to a Queue

Redis offers eight policies. Every one except noeviction deletes keys when memory is full:

Policy What gets evicted Effect on a queue
noeviction nothing; writes fail with OOM enqueue errors, no data loss
allkeys-lru / allkeys-lfu any key, least recently / frequently used oldest waiting jobs vanish first
allkeys-random any key at random random jobs and internal keys vanish
volatile-lru / volatile-lfu / volatile-random only keys with a TTL completed jobs with TTLs go; if none, behaves like noeviction
volatile-ttl keys closest to expiry same as above, shortest TTL first

The allkeys-* policies are the dangerous ones: a job that has been waiting a long time is, by definition, the least recently used key. The volatile-* policies are safer only as long as no queue key has a TTL — and some libraries set TTLs on completed-job data or locks. BullMQ checks this on startup and logs IMPORTANT! Eviction policy is allkeys-lru. It should be "noeviction"; that warning is worth turning into an alert.

Eviction vs noeviction at the memory limit Two outcomes when Redis reaches maxmemory. With allkeys-lru, a new enqueue succeeds but Redis silently deletes the least recently used key, which is the oldest waiting job, so the job is lost with no error. With noeviction, the new enqueue fails with an OOM error, nothing is deleted, and the producer can retry, reject, or divert the job. Redis at maxmemory: what happens to the next enqueue producer allkeys-lru enqueue returns OK oldest job evicted lost, no error anywhere producer noeviction enqueue fails: OOM producer retries or rejects existing jobs intact A visible failure you can handle beats a silent one you discover weeks later.

Step 2 — Separate the Cache from the Queue

A cache wants eviction; a queue must never have it. One instance cannot have both policies, so the clean fix is two instances (or two databases on a managed service that supports per-instance parameters):

# queue instance
redis-cli -h queue-redis CONFIG SET maxmemory-policy noeviction
redis-cli -h queue-redis CONFIG SET maxmemory 6gb

# cache instance keeps its eviction policy
redis-cli -h cache-redis CONFIG GET maxmemory-policy
# 1) "maxmemory-policy"  2) "allkeys-lfu"
Separate cache and queue instances The application talks to two Redis instances. Cache entries go to a cache instance configured with allkeys-lfu, where eviction is harmless because a miss is re-computed. Jobs go to a queue instance configured with noeviction and persistence, which the workers consume. A backlog in the queue never pushes out cache entries, and cache growth never evicts jobs. Two instances, two policies application cache Redis allkeys-lfu, no persistence queue Redis noeviction, AOF everysec workers eviction = a cache miss

Point the queue library at the queue instance only. If a separate instance is not possible, set noeviction on the shared one and make sure every cache key has a TTL, so the cache still expires entries on its own schedule — but the cache can then fill memory and cause enqueue failures, which is why separation is preferred.

On managed services the setting lives in a parameter group (ElastiCache), a configuration (Azure Cache for Redis, Memorystore), or the provider's console. Changing it on ElastiCache applies without a restart.

Step 3 — Persist the Setting and Check It at Startup

CONFIG SET is lost on restart unless it is also in redis.conf or the parameter group. Write it there, then make workers refuse to start on a wrong value:

import redis, sys

r = redis.Redis.from_url(QUEUE_REDIS_URL)
policy = r.config_get("maxmemory-policy")["maxmemory-policy"]
if policy != "noeviction":
    print(f"refusing to start: maxmemory-policy is {policy}", file=sys.stderr)
    sys.exit(1)

Some managed services disallow CONFIG GET; in that case check through the provider's API in your deployment pipeline instead.

Step 4 — Handle OOM Errors in Producers

With noeviction, a full Redis returns OOM command not allowed when used memory > 'maxmemory' on writes. Reads and deletes still work, so workers can keep draining the queue — which frees memory. Producers need a plan for the error:

try {
  await queue.add("invoice", data, { jobId: invoiceId });
} catch (err) {
  if (String(err).includes("OOM")) {
    metrics.enqueueOom.inc();
    await outbox.save({ queue: "invoice", data, jobId: invoiceId }); // retried later
    return;
  }
  throw err;
}

For a web request, failing with a 503 and a retry hint is acceptable when the job is not critical. For critical jobs, write them to a durable store (a transactional outbox in Postgres) and relay them when memory recovers. Note that the worker also writes — moving a job to active, storing results — so a truly full instance can stall workers too. Keep headroom (Step 5) so that never happens.

Step 5 — Leave Headroom and Alert Early

Set maxmemory below the machine's memory to leave room for fork-based persistence (copy-on-write during BGSAVE or AOF rewrite can need up to double the changed pages), replication buffers, and fragmentation. A common rule is maxmemory at 60–75% of RAM on a persistent instance.

Memory budget for a queue Redis An 8 gigabyte node sets maxmemory to 6 gigabytes and keeps 2 gigabytes free for copy-on-write during persistence, client and replication buffers, and fragmentation. A warning alert fires at 70 percent of maxmemory and a critical alert at 85 percent, well before enqueues fail. 8 GB node: where the memory goes normal backlog fork / buffers / frag warn 70% page 85% maxmemory 6 GB Percentages are of maxmemory; the last 2 GB of RAM is never given to data.

Alert on the ratio, and on any eviction at all:

- alert: QueueRedisMemoryHigh
  expr: redis_memory_used_bytes{instance="queue-redis"} / redis_memory_max_bytes > 0.85
  for: 5m
- alert: QueueRedisEvicting
  expr: increase(redis_evicted_keys_total{instance="queue-redis"}[5m]) > 0

The eviction alert should never fire with noeviction; if it does, someone has changed the policy. Size the instance from your expected peak backlog using sizing Redis memory for queue backlogs.

Step 6 — Trim What the Queue Keeps

Much of a queue's memory is not waiting jobs but finished ones. Configure retention so completed and failed jobs do not accumulate:

await queue.add("invoice", data, {
  removeOnComplete: { age: 3600, count: 1000 },
  removeOnFail: { age: 7 * 24 * 3600 },
});

In Celery, set result_expires; in RQ, result_ttl and failure_ttl. Keep payloads small by storing large blobs in object storage and passing a reference — see message size limits & serialization.

Verification

  • CONFIG GET maxmemory-policy returns noeviction on the queue instance, and the value survives a restart or failover.
  • INFO stats shows evicted_keys:0, and the eviction alert has never fired.
  • In staging, fill Redis to maxmemory: producers log OOM errors and use their fallback path; no existing job disappears; workers continue draining and enqueues resume once memory is freed.
  • The memory-ratio alert fires at the configured threshold during the fill test.

Gotchas & Edge Cases

Replicas and failover. A replica promoted by Sentinel or the provider uses its own configuration. Set the policy on every node, not only the primary.

Lua scripts under OOM. BullMQ and similar libraries run multi-key Lua scripts. With noeviction, a script that would allocate memory fails as a whole, so state stays consistent — but the worker sees errors and retries. Headroom avoids this.

Keyspace shared with sessions. Session stores often rely on eviction. Moving sessions to the queue instance with noeviction can make logins fail when the queue backs up. Keep them on the cache instance.

Provider defaults change between engines. ElastiCache defaults to volatile-lru, and some Valkey or hosted offerings default to other eviction policies. Never assume the default; check it on every new instance, including those created by infrastructure templates for staging.

maxmemory 0. On 64-bit builds, 0 means no limit; Redis grows until the OS kills it. Always set an explicit limit.

FAQ

Is volatile-lru safe if my queue keys have no TTL? It behaves like noeviction for those keys, but it depends on no library ever setting a TTL on queue data. BullMQ, for example, uses expiring lock keys; evicting a lock early causes a job to be treated as stalled. Use noeviction explicitly.

What happens to workers when Redis is full? Commands that only read or delete still succeed, so workers can fetch and remove jobs. Commands that add data (moving a job to completed with its result) can fail with OOM. Workers retry, and memory frees as jobs are removed.

Should I just give Redis more memory instead? Do both. More memory raises the ceiling; noeviction makes hitting the ceiling safe and visible.

Does Celery with a Redis broker need the same setting? Yes. Celery keeps waiting messages in Redis lists and unacknowledged messages in a hash; evicting either loses tasks. The same applies to RQ, Sidekiq, and any other Redis-backed queue.

How do I spot past evictions? Check evicted_keys in INFO stats (it resets on restart) and look for jobs whose IDs appear in a list without a matching job hash. BullMQ workers log a missing-key error for these; searching logs for it shows when evictions began.

Related