Tuning the Celery Result Backend
The result backend stores each task's state and return value so callers can fetch it later. Most tasks never need that, yet many deployments store every result by default and pay for it in memory, write load, and occasional outages. This guide tunes the backend deliberately, as part of Celery Architecture & Configuration in Backend Frameworks & Worker Scaling.
Problem Statement
A Celery deployment uses one Redis instance as both broker and result backend. It runs 3 million tasks a day, all storing results with the default one-day expiry. Redis memory hovers at 11 GB, most of it celery-task-meta-* keys nobody reads. During a traffic spike, Redis hit its memory limit; with noeviction set (correctly, for the broker), writes started failing — including enqueues, which stopped the whole job system. Chords, which genuinely need the backend, were caught in the same outage. You want results stored only where they are used, bounded memory, chords working reliably, and the broker isolated from result-backend load.
Prerequisites
- Celery 5.3+; a Redis or database result backend.
- An inventory of which tasks' results are actually read (
AsyncResult.get(), chords, workflows). - The ability to run a second Redis instance or use a database for results.
Step 1 — Find Out Which Results Anyone Reads
Most tasks are fire-and-forget: the caller enqueues and moves on. Their stored results are never read. Before changing settings, identify the tasks whose results are consumed.
# Rough census of result keys by task name (requires result_extended=True to include names)
redis-cli --scan --pattern 'celery-task-meta-*' | head -20000 \
| xargs -n 100 redis-cli mget | grep -o '"name": "[^"]*"' | sort | uniq -c | sort -rn | head
# In code: grep for consumers of results
# .get( AsyncResult( chord( group(...)() .join( result.status
In the scenario, 94% of result keys belonged to three notification and cache-refresh tasks whose callers never looked at them. Only reporting tasks (polled by a UI) and chord members used results.
Step 2 — Ignore Results by Default, Opt In Where Needed
Flip the default: store nothing unless a task declares it needs a result.
# celeryconfig.py
task_ignore_result = True # default: no result written
result_expires = 3600 # results that ARE stored live 1 h (default 1 day)
result_extended = True # store name/args for stored results (debuggability)
# tasks.py — opt in for tasks whose results are read
@app.task(ignore_result=False)
def build_report(report_id): ...
@app.task(ignore_result=False) # chord members MUST store results
def fetch_usage(region, customer_id): ...
Two cautions. Chord header tasks must store results, or the chord counter never completes and the body never runs — the failure mode described in Celery chains, groups, and chords. And task_track_started=True writes an extra STARTED state per task; enable it only where a UI shows running state.
Step 3 — Separate the Result Backend from the Broker
A broker needs noeviction and priority on writes; a result backend is closer to a cache. Sharing one Redis couples them: result growth can fill the broker's memory and stop enqueues. Give results their own instance (or database).
broker_url = "redis://redis-broker:6379/0" # noeviction, AOF, small, critical
result_backend = "redis://redis-results:6379/0" # sized for results, can use volatile-lru
redis_backend_health_check_interval = 30
result_backend_transport_options = {"retry_policy": {"timeout": 5.0}}
On the results instance, every key has a TTL (result_expires), so volatile-lru eviction is acceptable as a last resort: under pressure, the oldest results go first instead of all writes failing. Chord counters also have TTLs, so extreme eviction can still break an in-flight chord — size the instance to avoid relying on eviction. Memory sizing is covered in sizing Redis memory for queue backlogs.
Step 4 — Choose Redis or a Database Backend
| Concern | Redis backend | Database backend (SQLAlchemy / Django) |
|---|---|---|
| Write latency | Sub-millisecond | Milliseconds; adds DB write load |
| Chord performance | Fast atomic counters | Slower; counters via polling or row updates |
| Durability | Per Redis persistence | Full DB durability |
| Querying results | Key lookups only | SQL (django-celery-results admin) |
| Expiry | Automatic TTL | Needs celery.backend_cleanup periodic task |
Use Redis for chords and high volume; use the database when results are business records someone will query later — though in that case, writing the result to your own table from the task is usually cleaner than depending on backend internals. With the database backend, schedule celery.backend_cleanup (Celery beat adds it automatically when result_expires is set) or the table grows forever.
Step 5 — Keep Result Payloads Small
The backend stores whatever the task returns. A task that returns a 5 MB list writes 5 MB per call and holds it for result_expires.
# Anti-pattern
@app.task(ignore_result=False)
def export_rows(query_id):
return list(run_query(query_id)) # megabytes into Redis
# Better: store the artifact, return a reference
@app.task(ignore_result=False)
def export_rows(query_id):
key = storage.put_json(f"exports/{query_id}.json", list(run_query(query_id)))
return {"key": key, "rows": count}
The same principle applies to arguments; see the claim-check pattern for large payloads.
Step 6 — Avoid Polling Results in Tight Loops
AsyncResult.get() polls the backend (or subscribes, with Redis). Web requests that wait on task results tie up request workers and hammer the backend. For UI progress, have the task write status to your own table and let the UI poll that, or push updates — the approach in tracking progress of multi-step jobs.
# If you must wait in a script, bound it and use a reasonable interval
result = build_report.delay(report_id)
value = result.get(timeout=300, interval=1.0) # default interval is 0.5 s
Never call .get() inside another task; Celery warns because it can deadlock the pool when all processes are waiting on tasks queued behind them.
Step 7 — Roll Out Without Breaking Result Readers
Flipping task_ignore_result globally breaks any caller that reads a result you did not know about. Roll out in an order that surfaces those callers safely:
- Add
ignore_result=Falseexplicitly to every task found in Step 1 that has a reader, plus every chord member. - Deploy, then lower
result_expiresfrom a day to an hour — callers that read results within minutes are unaffected; any that read them much later will show up as missing-result errors in logs. - Enable
task_ignore_result = Trueglobally. Any remaining reader now getsNoneor a pending state instead of a value; watch for errors and add opt-ins where needed. - Move the result backend to its own Redis instance, point
result_backendat it, and let old keys on the broker instance expire.
Each step is independently reversible, and the memory drop after step 3 is usually visible within one expiry period.
Verification
# After the change: result keys and memory on each instance
redis-cli -h redis-results info keyspace
redis-cli -h redis-results info memory | grep used_memory_human
redis-cli -h redis-broker info memory | grep used_memory_human
In the scenario, result memory fell from 11 GB to about 400 MB, and the broker instance shrank to what the queued messages need. Confirm chords still complete by running the chord integration test from testing Celery tasks with pytest.
Gotchas & Edge Cases
ignore_result on chord members. The most common way to break chords silently. Test chords end to end after any change to result settings.
Result expiry vs long workflows. Results that feed later steps must outlive the workflow. For multi-day workflows, store intermediate outputs in durable storage, not the backend.
task_track_started load. Adds one write per task. Keep it off unless you display running state.
Flower and results. Flower reads events, not the backend, for its task list; ignoring results does not break Flower.
FAQ
Do I need a result backend at all?
Only for chords, for callers that read results, and for UIs that show task state from Celery. Many deployments can run with task_ignore_result=True everywhere and no backend except for chords.
What is the right result_expires?
As short as the longest time between a task finishing and someone reading its result — often minutes to an hour.
How do I see whether the backend is healthy?
Watch memory and key count on the results instance, command latency, and error logs from workers failing to store results (Error while storing result). A result-store failure does not usually fail the task itself, so it can go unnoticed until a chord hangs; alert on those log lines explicitly.
What about failures and tracebacks?
Failed tasks store their exception and traceback when results are enabled, which can be large. With task_ignore_result=True, failures are still reported through signals, logs, and events — rely on those for failure visibility rather than on stored results.
Can I use the RPC backend?
The rpc:// backend sends results back over AMQP to the caller; it suits short-lived callers waiting for results, but results are not persistent and cannot be fetched by another process later.
Related
- Celery Architecture & Configuration — broker, backend, and workers.
- Setting Up Celery with a Redis Broker and RabbitMQ Backend — initial backend setup.
- Redis maxmemory Policy for Queues — eviction settings for broker vs results.
- Celery Chains, Groups, and Chords — the feature that most depends on the backend.