Choosing a Redis maxmemory Policy for Queues
Redis decides what to do when it reaches its memory limit from a single setting, maxmemory-policy. For a cache, evicting old keys is exactly right. For a job queue, eviction means jobs disappear without an error anywhere. This guide sets the policy correctly and handles the consequences, as part of In-Memory vs Persistent Queue Storage in Backend Frameworks & Worker Scaling.
Problem Statement
A team uses one managed Redis instance for both the application cache and BullMQ. The provider's default policy is volatile-lru, and the team later switched it to allkeys-lru to improve cache hit rates. During a large import, the backlog grew to 3 GB and Redis reached its 4 GB limit. Nothing crashed. Weeks later, customers reported missing invoices: Redis had evicted job hashes and, in some cases, the internal keys BullMQ uses to track job state, leaving IDs in lists that pointed at nothing. Workers logged "missing key for job" warnings that nobody watched. You want jobs never to be evicted, a clear error when memory is exhausted, and enough warning to act before that happens.
Prerequisites
- Access to the Redis configuration (
CONFIG SETor the provider's parameter group). - Knowledge of which applications share the instance.
- Metrics for
used_memory,maxmemory, andevicted_keys(fromINFOor an exporter). - Producers that can handle an enqueue error (retry, fail the request, or spill to another store).
Step 1 — Understand What Each Policy Does to a Queue
Redis offers eight policies. Every one except noeviction deletes keys when memory is full:
| Policy | What gets evicted | Effect on a queue |
|---|---|---|
noeviction |
nothing; writes fail with OOM | enqueue errors, no data loss |
allkeys-lru / allkeys-lfu |
any key, least recently / frequently used | oldest waiting jobs vanish first |
allkeys-random |
any key at random | random jobs and internal keys vanish |
volatile-lru / volatile-lfu / volatile-random |
only keys with a TTL | completed jobs with TTLs go; if none, behaves like noeviction |
volatile-ttl |
keys closest to expiry | same as above, shortest TTL first |
The allkeys-* policies are the dangerous ones: a job that has been waiting a long time is, by definition, the least recently used key. The volatile-* policies are safer only as long as no queue key has a TTL — and some libraries set TTLs on completed-job data or locks. BullMQ checks this on startup and logs IMPORTANT! Eviction policy is allkeys-lru. It should be "noeviction"; that warning is worth turning into an alert.
Step 2 — Separate the Cache from the Queue
A cache wants eviction; a queue must never have it. One instance cannot have both policies, so the clean fix is two instances (or two databases on a managed service that supports per-instance parameters):
# queue instance
redis-cli -h queue-redis CONFIG SET maxmemory-policy noeviction
redis-cli -h queue-redis CONFIG SET maxmemory 6gb
# cache instance keeps its eviction policy
redis-cli -h cache-redis CONFIG GET maxmemory-policy
# 1) "maxmemory-policy" 2) "allkeys-lfu"
Point the queue library at the queue instance only. If a separate instance is not possible, set noeviction on the shared one and make sure every cache key has a TTL, so the cache still expires entries on its own schedule — but the cache can then fill memory and cause enqueue failures, which is why separation is preferred.
On managed services the setting lives in a parameter group (ElastiCache), a configuration (Azure Cache for Redis, Memorystore), or the provider's console. Changing it on ElastiCache applies without a restart.
Step 3 — Persist the Setting and Check It at Startup
CONFIG SET is lost on restart unless it is also in redis.conf or the parameter group. Write it there, then make workers refuse to start on a wrong value:
import redis, sys
r = redis.Redis.from_url(QUEUE_REDIS_URL)
policy = r.config_get("maxmemory-policy")["maxmemory-policy"]
if policy != "noeviction":
print(f"refusing to start: maxmemory-policy is {policy}", file=sys.stderr)
sys.exit(1)
Some managed services disallow CONFIG GET; in that case check through the provider's API in your deployment pipeline instead.
Step 4 — Handle OOM Errors in Producers
With noeviction, a full Redis returns OOM command not allowed when used memory > 'maxmemory' on writes. Reads and deletes still work, so workers can keep draining the queue — which frees memory. Producers need a plan for the error:
try {
await queue.add("invoice", data, { jobId: invoiceId });
} catch (err) {
if (String(err).includes("OOM")) {
metrics.enqueueOom.inc();
await outbox.save({ queue: "invoice", data, jobId: invoiceId }); // retried later
return;
}
throw err;
}
For a web request, failing with a 503 and a retry hint is acceptable when the job is not critical. For critical jobs, write them to a durable store (a transactional outbox in Postgres) and relay them when memory recovers. Note that the worker also writes — moving a job to active, storing results — so a truly full instance can stall workers too. Keep headroom (Step 5) so that never happens.
Step 5 — Leave Headroom and Alert Early
Set maxmemory below the machine's memory to leave room for fork-based persistence (copy-on-write during BGSAVE or AOF rewrite can need up to double the changed pages), replication buffers, and fragmentation. A common rule is maxmemory at 60–75% of RAM on a persistent instance.
Alert on the ratio, and on any eviction at all:
- alert: QueueRedisMemoryHigh
expr: redis_memory_used_bytes{instance="queue-redis"} / redis_memory_max_bytes > 0.85
for: 5m
- alert: QueueRedisEvicting
expr: increase(redis_evicted_keys_total{instance="queue-redis"}[5m]) > 0
The eviction alert should never fire with noeviction; if it does, someone has changed the policy. Size the instance from your expected peak backlog using sizing Redis memory for queue backlogs.
Step 6 — Trim What the Queue Keeps
Much of a queue's memory is not waiting jobs but finished ones. Configure retention so completed and failed jobs do not accumulate:
await queue.add("invoice", data, {
removeOnComplete: { age: 3600, count: 1000 },
removeOnFail: { age: 7 * 24 * 3600 },
});
In Celery, set result_expires; in RQ, result_ttl and failure_ttl. Keep payloads small by storing large blobs in object storage and passing a reference — see message size limits & serialization.
Verification
CONFIG GET maxmemory-policyreturnsnoevictionon the queue instance, and the value survives a restart or failover.INFO statsshowsevicted_keys:0, and the eviction alert has never fired.- In staging, fill Redis to
maxmemory: producers log OOM errors and use their fallback path; no existing job disappears; workers continue draining and enqueues resume once memory is freed. - The memory-ratio alert fires at the configured threshold during the fill test.
Gotchas & Edge Cases
Replicas and failover. A replica promoted by Sentinel or the provider uses its own configuration. Set the policy on every node, not only the primary.
Lua scripts under OOM. BullMQ and similar libraries run multi-key Lua scripts. With noeviction, a script that would allocate memory fails as a whole, so state stays consistent — but the worker sees errors and retries. Headroom avoids this.
Keyspace shared with sessions. Session stores often rely on eviction. Moving sessions to the queue instance with noeviction can make logins fail when the queue backs up. Keep them on the cache instance.
Provider defaults change between engines. ElastiCache defaults to volatile-lru, and some Valkey or hosted offerings default to other eviction policies. Never assume the default; check it on every new instance, including those created by infrastructure templates for staging.
maxmemory 0. On 64-bit builds, 0 means no limit; Redis grows until the OS kills it. Always set an explicit limit.
FAQ
Is volatile-lru safe if my queue keys have no TTL?
It behaves like noeviction for those keys, but it depends on no library ever setting a TTL on queue data. BullMQ, for example, uses expiring lock keys; evicting a lock early causes a job to be treated as stalled. Use noeviction explicitly.
What happens to workers when Redis is full? Commands that only read or delete still succeed, so workers can fetch and remove jobs. Commands that add data (moving a job to completed with its result) can fail with OOM. Workers retry, and memory frees as jobs are removed.
Should I just give Redis more memory instead?
Do both. More memory raises the ceiling; noeviction makes hitting the ceiling safe and visible.
Does Celery with a Redis broker need the same setting? Yes. Celery keeps waiting messages in Redis lists and unacknowledged messages in a hash; evicting either loses tasks. The same applies to RQ, Sidekiq, and any other Redis-backed queue.
How do I spot past evictions?
Check evicted_keys in INFO stats (it resets on restart) and look for jobs whose IDs appear in a list without a matching job hash. BullMQ workers log a missing-key error for these; searching logs for it shows when evictions began.
Related
- In-Memory vs Persistent Queue Storage — durability trade-offs.
- Redis Persistence: AOF vs RDB for Queues — surviving restarts.
- Sizing Redis Memory for Queue Backlogs — how much memory to provision.
- Redis Sentinel for Queue High Availability — surviving node failures.