RabbitMQ Consumer Timeout for Unacked Messages

RabbitMQ traditionally held an unacknowledged delivery for as long as the consumer's connection lived, which made it the one major broker without a visibility timeout. Since 3.8.15 it has consumer_timeout, a delivery acknowledgement timeout that closes the channel of any consumer holding a message too long. This guide explains how it behaves and how to live with it, as part of the Visibility Timeout Deep Dive in Queue Fundamentals & Architecture.

Problem Statement

After upgrading RabbitMQ, a Celery deployment that generates large reports started failing in a strange way: every report task longer than 30 minutes was killed, the worker logged PRECONDITION_FAILED - delivery acknowledgement on channel 1 timed out, the channel closed, every other message prefetched on that channel was requeued, and the report was retried from scratch — only to hit the same limit again. Meanwhile, on another cluster, a consumer that had hung on a database lock held 200 messages for three days without anyone noticing, because the old default had no limit at all. You want long reports to run to completion, hung consumers to release their messages, and both situations to be visible in monitoring.

Prerequisites

  • RabbitMQ 3.8.15 or later (3.12+ recommended; the timeout became configurable per queue via policy in 3.12).
  • Access to rabbitmq.conf or to set policies (rabbitmqctl set_policy).
  • Knowledge of your longest legitimate time between delivery and acknowledgement per queue.
  • The rabbitmq_prometheus plugin for monitoring.

Step 1 — Understand What the Timeout Measures

consumer_timeout is the maximum time a delivery may remain unacknowledged. The default is 30 minutes (1,800,000 ms). The broker checks periodically (about once a minute), so the effective limit is the timeout plus up to a minute. When a delivery exceeds it, RabbitMQ does not requeue just that message — it closes the whole channel with PRECONDITION_FAILED, which returns every unacked message on that channel to its queue.

t = 0      worker receives report job (and, with prefetch 10, nine other jobs)
t = 30 min timeout reached for the report delivery
t ≈ 31 min broker closes channel 1: PRECONDITION_FAILED
           -> report job and all 9 prefetched jobs requeued (redelivered=true)
           -> worker must reopen a channel; the report is started again elsewhere

That blast radius is the key operational fact: one slow message costs every message sharing its channel a redelivery.

One slow delivery closes the channel A worker channel holds ten unacknowledged deliveries: one long report job and nine prefetched jobs. When the report exceeds the 30-minute consumer timeout, the broker closes the channel with PRECONDITION_FAILED. All ten deliveries return to the queue marked as redelivered, and the report starts again from the beginning on another consumer. consumer_timeout = 30 min, prefetch = 10 channel 1: 10 unacked report job: 31 min unacked 9 prefetched jobs, not started PRECONDITION_FAILED: close all 10 requeued, redelivered=true The report restarts from zero and will hit the same limit again unless something changes.

Step 2 — Set the Timeout per Queue, Not Globally

A global setting forces one value on queues with very different job durations. From RabbitMQ 3.12, set it per queue (or per group of queues) with a policy, and keep a sane global default for everything else.

# rabbitmq.conf — global default: catches hung consumers on ordinary queues
consumer_timeout = 1800000        # 30 minutes
# Long-running report queues: allow 3 hours
rabbitmqctl set_policy report-timeouts "^reports\." \
  '{"consumer-timeout": 10800000}' --apply-to queues --priority 10

# Latency-sensitive queues: catch hangs quickly
rabbitmqctl set_policy fast-timeouts "^(emails|webhooks)$" \
  '{"consumer-timeout": 300000}' --apply-to queues --priority 10

Policies apply without redeclaring queues, which matters because queue arguments are immutable. On versions before 3.12, the only per-queue option is the x-consumer-timeout queue argument at declaration (in some versions) or the global setting; check your version's documentation before relying on either.

Step 3 — Isolate Long Jobs on Their Own Channel and Prefetch

Because the timeout closes the whole channel, long jobs should never share a channel with short ones, and long-job consumers should prefetch one message so that a closure affects only the job that overran.

# celeryconfig.py — dedicated workers for the reports queue
task_routes = {"reports.*": {"queue": "reports.long"}}
worker_prefetch_multiplier = 1          # one unacked message per process
task_acks_late = True                   # ack after completion (the timeout's subject)
# run: celery -A app worker -Q reports.long --concurrency 4 --prefetch-multiplier 1

With prefetch 1 and a dedicated queue, a timeout on one report requeues only that report. The general guidance on prefetch sizing is in tuning prefetch and consumer concurrency.

Step 4 — Make Genuinely Long Jobs Checkpointed or Split

Raising the timeout indefinitely brings back the old problem: a hung consumer holds messages for hours. Prefer bounding each delivery's duration by changing the job's shape.

@app.task(bind=True, acks_late=True)
def build_report(self, report_id: str, start_chunk: int = 0):
    chunks = plan_chunks(report_id)
    deadline = time.monotonic() + 20 * 60            # stay well inside the 30-min timeout
    for i in range(start_chunk, len(chunks)):
        render_chunk(report_id, chunks[i])
        save_checkpoint(report_id, i)
        if time.monotonic() > deadline and i + 1 < len(chunks):
            build_report.apply_async(args=[report_id, i + 1])   # continue in a new delivery
            return                                               # ack this delivery now
    finalize_report(report_id)

Each delivery now lasts at most about 20 minutes; the report continues in a fresh message with its own timer. The timeout goes back to being what it should be — a detector for hung consumers — rather than a cap on job size. The same checkpoint idea appears in configuring visibility timeouts for long-running workers.

Bound each delivery, not the whole job Instead of one delivery that stays unacknowledged for three hours, the report task processes chunks for up to twenty minutes, saves a checkpoint, enqueues a continuation message starting at the next chunk, and acknowledges. Nine deliveries of twenty minutes each complete the report, and each stays well inside the thirty-minute consumer timeout. A 3-hour report as 20-minute deliveries one delivery unacked 3 h: needs a huge timeout, hides hangs chained each block: 20 min, checkpoint, continue in a new message, ack A crash loses at most one block; the 30-minute timeout still catches real hangs.

Step 5 — Distinguish Heartbeats from the Delivery Timeout

Two different timers protect different things, and they are often confused.

  • Connection heartbeats (heartbeat, default 60 s) detect a dead TCP connection. If heartbeats stop, the broker closes the connection and requeues all its unacked messages. A worker that is alive but busy keeps heartbeating (in clients that heartbeat from a separate thread or I/O loop).
  • consumer_timeout detects a delivery that has been held too long by a consumer whose connection is perfectly healthy — the hung-on-a-lock case from the problem statement.
# pika BlockingConnection: heartbeats only happen when the connection's I/O loop runs.
# Long work in the callback starves them -> use a thread and connection.add_callback_threadsafe
params = pika.ConnectionParameters(host="rabbitmq", heartbeat=60, blocked_connection_timeout=300)
Two timers, two failures Connection heartbeats catch a dead or partitioned client: if the broker receives no heartbeat for about two intervals, it closes the connection and requeues everything. The consumer timeout catches a client whose connection is healthy but which has held a delivery longer than allowed, such as a handler hung on a database lock; the broker closes that channel. What each timer detects heartbeat (60 s) dead process, network partition closes the connection consumer_timeout (30 min) live consumer, stuck delivery closes the channel Before consumer_timeout existed, the right-hand case held messages indefinitely.

A worker that blocks its client's I/O loop for longer than two heartbeat intervals is disconnected for missing heartbeats before the consumer timeout matters. Celery and aio-pika handle this; hand-written pika consumers need the thread-plus-callback pattern.

Step 6 — Monitor Timeouts and Old Unacked Messages

Track channel closures from timeouts and the age of unacknowledged messages per queue, so both overruns and hangs are visible before they cause incidents.

# Unacked messages per queue (sustained high values with low ack rate = hung consumer)
rabbitmq_queue_messages_unacked{queue=~"reports.*|emails"}

# Ack rate per queue: flat zero with unacked > 0 means consumers hold messages without progress
rate(rabbitmq_global_messages_acknowledged_total[5m])

# Channel closures (from broker logs shipped to Loki)
sum(count_over_time({app="rabbitmq"} |= "delivery acknowledgement on channel" |= "timed out" [15m]))

On the worker side, log and count PRECONDITION_FAILED channel errors per task type; a task that repeatedly hits the timeout needs Step 4, not a bigger number.

Verification

# Temporarily set a short timeout on a test queue and confirm the behaviour
rabbitmqctl set_policy test-timeout "^test\.timeout$" '{"consumer-timeout": 60000}' --apply-to queues
python publish.py --queue test.timeout --count 5
python slow_consumer.py --queue test.timeout --prefetch 5 --sleep 180   # expect channel close ~1-2 min
rabbitmqctl list_queues name messages_ready messages_unacked | grep test.timeout   # all 5 ready again

Then check that the report pipeline, after Step 4, never produces a timeout-closed channel in a week of normal operation, while the global default still catches a deliberately hung test consumer.

Gotchas & Edge Cases

Disabling the timeout. Setting a very large value (or disabling it where supported) brings back invisible hangs that hold messages for days. Prefer per-queue policies.

Quorum queue delivery limits. Each timeout-driven requeue increments the delivery count on quorum queues; with delivery-limit set, a job that always overruns is eventually dead-lettered — good, but only if someone watches the DLQ.

Timeouts during slow shutdowns. A worker draining in-flight jobs during a long graceful shutdown still holds those deliveries; if the drain outlasts the timeout, the broker closes the channel mid-drain. Keep drain windows shorter than the timeout for the affected queues.

Redelivered flag. Messages requeued by a channel closure arrive with redelivered=true. Consumers can use it as a hint to check idempotency records first.

Client reconnect storms. Many workers hitting timeouts simultaneously all reconnect at once. Use jittered reconnect backoff in clients.

FAQ

Is consumer_timeout the same as an SQS visibility timeout? Similar intent, different mechanics. SQS makes one message visible again; RabbitMQ closes the channel, affecting all its unacked messages. SQS lets consumers extend per message; RabbitMQ has no per-message extension.

Can a consumer extend its own timeout? No. Either the queue's timeout is long enough, or the job must acknowledge within it (by splitting or checkpointing).

What value should the global default be? Pick a value comfortably above the longest legitimate job on any queue that does not have its own policy — usually 15 to 60 minutes — and then give long-running queues explicit, higher policies. The default exists to catch hung consumers everywhere else, so it should be short enough that a hang is noticed within an on-call shift rather than days later.

Why did this appear after an upgrade? Older versions had no delivery timeout by default. Upgrades that introduced the 30-minute default exposed tasks that had always held messages longer.

Related