Celery Task Time Limits
A task with no time limit can hold a worker slot forever — stuck on a socket without a timeout, a lock, or an unexpectedly huge input. Celery's soft and hard time limits bound every task, and this guide sets them up so that stuck tasks are stopped, cleaned up, and retried or reported, as part of Celery Architecture & Configuration in Backend Frameworks & Worker Scaling.
Problem Statement
A document-conversion service runs 12 Celery workers with 8 prefork processes each. Every few days, throughput collapses: dozens of processes are busy for hours on conversions that normally take 20 seconds, stuck on a third-party converter that stopped responding without closing the connection. The queue backs up until someone restarts workers by hand. When a restart kills a task mid-write, partial output files are left behind. You want every task bounded in time, a chance to clean up before being stopped, a hard stop for tasks that ignore the first signal, limits appropriate to each task type, and stuck tasks visible in metrics.
Prerequisites
- Celery 5.3+ with the prefork pool (hard time limits require it; see Step 6 for other pools).
- Measured duration percentiles per task type.
- Task code that can clean up partial work (temporary files, database transactions).
task_acks_lateconfigured as in Celery acks_late and worker crash safety.
Step 1 — Understand Soft and Hard Limits
Celery has two limits:
- Soft time limit: when reached, Celery raises
SoftTimeLimitExceededinside the task. The task can catch it, clean up, and exit — or re-raise to fail. - Hard time limit: when reached, the pool process running the task is killed and replaced. No cleanup code runs.
# celeryconfig.py — global defaults
task_soft_time_limit = 300 # 5 min: task gets SoftTimeLimitExceeded
task_time_limit = 330 # 5.5 min: process killed if still running
Set the hard limit a little above the soft limit, leaving time for cleanup. The soft limit is the one you design around; the hard limit is the backstop for code that swallows the exception or is stuck in native code where the exception cannot be delivered.
Step 2 — Catch the Soft Limit and Clean Up
Handle SoftTimeLimitExceeded where you own resources, clean them up, and let the task fail or retry.
from celery.exceptions import SoftTimeLimitExceeded
@app.task(bind=True, acks_late=True, soft_time_limit=120, time_limit=150,
autoretry_for=(ConverterTimeout,), max_retries=3, retry_backoff=True)
def convert_document(self, doc_id: int):
tmp = tempfile.mkdtemp(prefix=f"conv-{doc_id}-")
try:
src = download(doc_id, tmp)
out = converter.convert(src, timeout=90) # client timeout below the soft limit
upload_result(doc_id, out)
except SoftTimeLimitExceeded:
log.warning("conversion exceeded soft limit", doc_id=doc_id, attempt=self.request.retries)
raise self.retry(countdown=60, exc=ConverterTimeout("soft limit"))
finally:
shutil.rmtree(tmp, ignore_errors=True) # runs on success, error, and soft limit
Two layers are visible here: a client-side timeout (90 s) that should normally fire first, and the soft limit (120 s) that catches anything the client timeout missed. Time limits are a safety net, not a substitute for timeouts on every network call — the stuck-socket problem in the scenario is best fixed at the client.
Step 3 — Set Limits per Task Type
One global limit either kills legitimately long tasks or lets short ones hang for too long. Set a conservative global default and override per task from measured durations — roughly 3–5× the p99, with an absolute cap.
# per-task via decorator
@app.task(soft_time_limit=30, time_limit=40)
def send_email(...): ...
@app.task(soft_time_limit=3600, time_limit=3660)
def build_monthly_report(...): ...
# or centrally via annotations, keeping limits in one reviewed place
task_annotations = {
"mail.send_email": {"soft_time_limit": 30, "time_limit": 40},
"docs.convert_document": {"soft_time_limit": 120, "time_limit": 150},
"reports.build_monthly_report": {"soft_time_limit": 3600, "time_limit": 3660},
}
# p99 duration per task, as input for limits
histogram_quantile(0.99, sum by (le, task) (rate(celery_task_runtime_seconds_bucket[7d])))
Recheck limits when task behaviour changes; a limit set from last year's data can start killing healthy tasks as inputs grow. Duration histograms are covered in choosing histogram buckets for job duration.
Step 4 — Understand What a Hard Kill Leaves Behind
A hard kill terminates the pool process immediately. Open transactions roll back (the database connection closes), but anything outside a transaction — files written, external API calls made, messages published — stays as it was. With late ack and task_reject_on_worker_lost, the task is requeued and runs again.
# Write output atomically so a hard kill never leaves a half-written artifact visible
def upload_result(doc_id, local_path):
tmp_key = f"converted/{doc_id}.pdf.part-{uuid4().hex}"
storage.put(tmp_key, local_path)
storage.copy(tmp_key, f"converted/{doc_id}.pdf") # visible only when complete
storage.delete(tmp_key)
# A lifecycle rule deletes stray *.part-* objects after 1 day
Design for the hard kill as if it were a crash, because it is one: atomic final writes, idempotent retries, and periodic cleanup of temporary artifacts.
Step 5 — Keep Limits Consistent with Broker Timeouts
Time limits interact with the broker's redelivery timers:
- Redis broker:
visibility_timeoutmust exceed the longesttime_limit(plus any countdown), or a still-running task is redelivered to a second worker before its own limit stops it. - RabbitMQ:
consumer_timeout(default 30 minutes) closes the channel for deliveries held longer than that; tasks with longer limits need a queue policy raising it, as covered in RabbitMQ consumer timeout for unacked messages.
LONGEST_TIME_LIMIT = 3660
broker_transport_options = {"visibility_timeout": LONGEST_TIME_LIMIT + 600} # Redis
A simple startup assertion in the worker that compares the largest configured time_limit with the broker timeout catches drift when someone raises a limit.
Step 6 — Know Which Pools Support Which Limits
Hard time limits rely on killing a child process, so they work only with the prefork pool. With solo, threads, gevent, or eventlet, the hard limit cannot be enforced, and the soft limit's exception delivery varies.
| Pool | Soft limit | Hard limit |
|---|---|---|
| prefork | Yes (signal to child) | Yes (child killed) |
| threads | No | No |
| gevent / eventlet | Timeout raised in greenlet (cooperative) | No |
| solo | No | No |
For non-prefork pools, rely on client timeouts and on gevent.Timeout or explicit deadline checks in the task, and consider moving tasks that can hang in native code to a prefork worker. Pool choice is covered in choosing Celery prefork vs gevent pools.
Step 7 — Make Timeouts Visible
Count soft-limit and hard-limit events per task; a rising rate is an early signal of a degrading dependency or growing inputs.
from celery.signals import task_failure
@task_failure.connect
def on_failure(sender=None, exception=None, **_):
if isinstance(exception, SoftTimeLimitExceeded):
TIME_LIMIT_EVENTS.labels(task=sender.name, kind="soft").inc()
elif isinstance(exception, TimeLimitExceeded):
TIME_LIMIT_EVENTS.labels(task=sender.name, kind="hard").inc()
sum by (task, kind) (increase(celery_time_limit_events_total[1h])) > 5
Hard-limit events deserve more attention than soft ones: each means cleanup code did not run and a process was replaced.
Verification
def test_soft_limit_cleans_up(celery_worker, monkeypatch):
monkeypatch.setattr(converter, "convert", lambda *a, **k: time.sleep(10))
res = convert_document.apply_async(args=[1], soft_time_limit=1, time_limit=3)
with pytest.raises(Exception):
res.get(timeout=10)
assert not glob.glob("/tmp/conv-1-*") # temp dir removed
In staging, point the converter at a black-hole endpoint and confirm that tasks stop at the soft limit, retry with backoff, and that worker throughput recovers without a manual restart.
Gotchas & Edge Cases
Catching broad exceptions. except Exception: catches SoftTimeLimitExceeded and may swallow it, leaving the task running until the hard kill. Re-raise it or handle it explicitly first.
Limits on the caller side. result.get(timeout=...) limits how long a caller waits, not how long the task runs. Both are needed.
Limits and worker_max_tasks_per_child. A hard kill replaces the process, which also resets any per-child memory growth. That is a side effect, not a strategy: use worker_max_tasks_per_child or worker_max_memory_per_child for leaks, as in fixing Celery worker memory leaks.
Retries reset the clock. Each retry gets a fresh limit. Bound total effort with max_retries.
time_limit below soft_time_limit. If the hard limit is lower, the soft limit never fires. Keep hard above soft.
FAQ
What limit should I start with? A global soft limit of a few minutes and hard limit slightly above, then per-task overrides from measured p99 durations. Anything genuinely longer than about an hour should be split into steps.
Can a task extend its own limit? No. Split long work into chained steps, each with its own limit, and checkpoint progress between them.
Why did my task hit the hard limit even though it catches the soft one? The soft limit is delivered as a signal-raised exception in the main thread of the pool process; if that thread is inside native code (a C extension, a blocking system call without a timeout), the exception is not raised until control returns to Python — which may be never. The hard limit then fires. Fix it by adding timeouts to the underlying calls so control returns regularly.
Does a hard kill count as a retry?
With late ack and reject_on_worker_lost, the message is requeued and the task runs again as a redelivery, not through self.retry, so max_retries does not bound it. Track redeliveries separately.
Related
- Celery Architecture & Configuration — the settings landscape.
- Celery Task Retry and Error Handling — retrying after a timeout.
- Celery acks_late and Worker Crash Safety — what happens to killed tasks.
- Alerting on Stuck and Stalled Jobs — catching tasks limits did not.