Inspecting and Revoking Tasks with the Flower API

During an incident, clicking through Flower's UI to find and stop hundreds of bad tasks is slow and error-prone. Flower exposes the same abilities through a REST API: list tasks by state and name, read a task's arguments and traceback, revoke queued tasks, terminate running ones, and change rate limits on the fly. This guide shows how to use the API safely and turn it into repeatable runbooks, as part of Flower for Celery Monitoring in Observability & Monitoring for Job Queues.

Problem Statement

A bug in a newly deployed send_campaign_email task sends each email to the wrong segment. By the time someone notices, 30,000 of these tasks are queued and 40 are running. The on-call engineer opens Flower, finds the task list, and starts revoking tasks one at a time; after ten minutes, a few hundred are revoked and several thousand more emails have gone out. There is no documented way to stop a task type in bulk, and nobody is sure whether "terminate" is safe for tasks that write to the database. The team wants a scripted, tested procedure to find tasks by name and state, stop them quickly, and understand exactly what each action does.

Prerequisites

  • Flower 2.x connected to the broker, with workers sending events (-E) so Flower knows about tasks.
  • Workers with remote control enabled (the default), since revoke and rate limits are broadcast commands.
  • API credentials: Flower's basic auth, OAuth, or a proxy that admits your tooling — see securing the Flower dashboard in production.
  • curl and jq, or a small Python script using requests.

Step 1 — Understand What Flower Knows

Flower's task list comes from events it has received, held in memory up to --max_tasks (10,000 by default). It does not read the broker's queues. That has two consequences for an incident:

  • Tasks that were queued before Flower started, or that fell out of its memory window, are not listed even though they are still in the queue.
  • Revoking works on task IDs broadcast to workers, not on the queue, so a revoked task is dropped when a worker receives it — whether or not Flower had listed it.
What Flower sees vs what the broker holds The broker queue holds all 30,000 waiting tasks. Flower only lists tasks it has received events for, capped at max_tasks in memory, so its list may show a fraction of them. A revoke request sends task IDs to every worker through a broadcast; each worker keeps a revoked set and discards a task with a revoked ID when it receives it from the queue. Flower's list is a window, not the queue broker queue: 30,000 tasks seen by Flower ≤ max_tasks not listed revoke: broadcast task IDs each worker: revoked set in memory drops matching IDs on receipt To stop tasks Flower never saw, use a kill switch (Step 4) or purge the queue.

Keep this model in mind: Flower is excellent for inspecting recent activity and for issuing commands, but for bulk removal of queued work you may need broker-level tools as well.

Step 2 — List and Filter Tasks

The /api/tasks endpoint returns tasks Flower knows about, filtered by name, state, and worker:

FLOWER=https://flower.internal
AUTH="-u $FLOWER_USER:$FLOWER_PASS"

# running and received instances of the bad task
curl -s $AUTH "$FLOWER/api/tasks?taskname=campaigns.send_campaign_email&state=STARTED" | jq 'keys | length'
curl -s $AUTH "$FLOWER/api/tasks?taskname=campaigns.send_campaign_email&state=RECEIVED&limit=5000" \
  | jq -r 'keys[]' > received_ids.txt

# what task types are failing right now
curl -s $AUTH "$FLOWER/api/tasks?state=FAILURE&limit=500" | jq -r '.[].name' | sort | uniq -c | sort -rn

RECEIVED tasks have been prefetched by a worker but not started; STARTED tasks are running, which Flower learns from the task-started event workers send when events are enabled. /api/task/types lists every task name Flower has seen, which helps when you are not sure of the exact name.

Step 3 — Inspect a Single Task

Before stopping anything, look at a few examples to confirm they are the tasks you think they are:

curl -s $AUTH "$FLOWER/api/task/info/7f3c9a1e-..." | jq '{name, state, args, kwargs, worker, received, started, retries, exception}'

The response includes the arguments, the worker that holds it, timestamps, retry count, and, for failures, the exception and traceback. Check that arguments are what you expect: in the incident, confirming that segment_id was wrong for every sampled task is what justifies stopping them all. Arguments may contain personal data; do not paste them into shared chat channels. Flower shows arguments as the string representation the worker reported in its events, which Celery truncates for large payloads, so pass IDs rather than whole documents if you want task details to stay readable.

Step 4 — Revoke Queued Tasks and Terminate Running Ones

Revoking and terminating are different operations with different risks:

Revoke vs terminate A plain revoke adds a task ID to each worker's revoked set, so the task is skipped if it has not started; a task already running continues to completion. Revoke with terminate also sends a signal, SIGTERM by default, to the pool process executing the task, stopping it mid-way, which can leave partial work behind. Terminate only works with the prefork pool. Both are held in worker memory and are lost when workers restart unless a state database is configured. Two ways to stop a task revoke • skipped if not yet started • running tasks finish normally • safe for any task • works with every pool revoke + terminate • also signals the running process • stops work mid-way • may leave partial writes • prefork pool only Revoked IDs live in worker memory; configure --statedb so they survive restarts.
# revoke queued tasks (safe)
while read id; do
  curl -s $AUTH -X POST "$FLOWER/api/task/revoke/$id" > /dev/null
done < received_ids.txt

# terminate running ones (only if stopping mid-way is safe)
curl -s $AUTH -X POST "$FLOWER/api/task/revoke/$ID?terminate=true&signal=SIGTERM"

Each revoke is a broadcast, so thousands of calls in a loop are slow and flood the broker. For bulk stops, it is usually better to stop the tasks by name in the code: deploy a guard that makes the task return immediately (if settings.CAMPAIGNS_PAUSED: return), or use a feature flag checked at the start of the task. Celery 5.3 also supports revoking by stamped headers (revoke_by_stamped_headers) if your producers stamp tasks with a campaign ID. For tasks Flower never saw, purging the queue (celery -A app purge -Q campaigns) removes everything in it — only do that if the queue contains nothing you want to keep.

Terminate only when stopping mid-way is harmless or less harmful than finishing. For the email task, terminating a task halfway through sending a batch is acceptable; for a task that moves money between accounts, it is not. Make long tasks check for cancellation at safe points instead, so you can stop them cleanly.

Step 5 — Slow Down Instead of Stopping

Sometimes the right move is to reduce pressure rather than cancel work — for example, when a downstream API is struggling. Flower's API can change a task's rate limit and resize worker pools without a deploy:

# at most 10 per minute on this worker (repeat for each worker)
curl -s $AUTH -X POST -d 'taskname=integrations.sync_crm' -d 'ratelimit=10/m' \
  "$FLOWER/api/task/rate-limit/celery@worker-3"

# shrink or grow the prefork pool
curl -s $AUTH -X POST -d 'n=2' "$FLOWER/api/worker/pool/shrink/celery@worker-3"

Celery's rate limits apply per worker, so the effective limit is the per-worker value times the number of workers. They are also lost when workers restart; make lasting changes in code. For a distributed limit, see token bucket rate limiting for Celery tasks.

Step 6 — Turn It into a Runbook

Wrap the calls in a script that is reviewed and tested before an incident, with a dry-run mode that only prints what it would do:

import os, sys, requests

FLOWER, AUTH = os.environ["FLOWER_URL"], (os.environ["FLOWER_USER"], os.environ["FLOWER_PASS"])

def stop_task_type(name: str, terminate: bool = False, dry_run: bool = True) -> None:
    for state in ("RECEIVED", "STARTED"):
        tasks = requests.get(f"{FLOWER}/api/tasks", params={"taskname": name, "state": state, "limit": 10000},
                             auth=AUTH, timeout=30).json()
        print(f"{state}: {len(tasks)} {name}")
        for task_id in tasks:
            if dry_run:
                continue
            params = {"terminate": "true"} if (terminate and state == "STARTED") else {}
            requests.post(f"{FLOWER}/api/task/revoke/{task_id}", params=params, auth=AUTH, timeout=10)

if __name__ == "__main__":
    stop_task_type(sys.argv[1], terminate="--terminate" in sys.argv, dry_run="--apply" not in sys.argv)
Choosing how to stop bad tasks Four situations and the matching action. Thousands of harmful tasks of one type: flip a kill switch in the task code so every copy returns immediately. A queue that holds nothing but bad tasks: purge that queue. A few hundred bad tasks visible in Flower: revoke them by ID through the API. Running tasks whose completion is worse than an interruption: revoke with terminate. Which tool for which situation thousands of one bad task type kill switch in task code queue holds only bad tasks purge that queue a few hundred, visible in Flower revoke by ID via the API running, and finishing is worse revoke with terminate

Link the script from the incident runbook alongside the kill-switch flag and the purge command, with a short decision guide for which to use. Writing runbooks for queue incidents covers the surrounding structure.

Verification

  • In staging, enqueue 500 slow test tasks, run the script with --apply: queued tasks are skipped by workers and appear as REVOKED in Flower.
  • With --terminate, running test tasks stop within seconds and their state becomes REVOKED.
  • After restarting a worker configured with --statedb, previously revoked IDs are still skipped.
  • A rate-limit change through the API is visible in the worker's inspect conf output and in reduced throughput.

Gotchas & Edge Cases

Revokes are forgotten on restart. Without --statedb=/var/run/celery/worker.state (on persistent storage), a worker restarted during the incident will run tasks you revoked. In Kubernetes, pods usually lose that file, so prefer the kill-switch approach for large stops.

Revoked set size. Workers keep up to 50,000 revoked IDs by default (CELERY_WORKER_REVOKES_MAX); older ones expire after three hours. Revoking more than that in one go means some are forgotten.

Terminate and acks_late. With acks_late, a terminated task's message may be redelivered depending on task_reject_on_worker_lost. Check this in staging so a terminated task does not simply come back.

API is powerful. The same credentials can shut down workers. Limit API access to the on-call group and audit calls through the proxy's logs.

FAQ

Can I revoke all tasks of a type with one API call? Flower's API revokes by ID. For a type-wide stop, use a kill switch in the task code, revoke by stamped headers, or purge the queue if it contains only that type.

Why does a task I revoked still show as STARTED? Plain revoke does not stop a running task. Use terminate=true, or wait for it to finish.

Is the API the same as celery inspect and celery control? Largely yes — Flower wraps the same remote-control commands. The CLI works without Flower and is a good fallback if Flower is down.

Related