Scraping Flower Metrics with Prometheus

Flower is best known as Celery's web dashboard, but since version 1.0 it also exposes a Prometheus endpoint. Because Flower already listens to Celery's event stream, pointing Prometheus at it is the quickest way to get task counts, runtimes, and worker status into your existing monitoring. This guide sets it up and explains where it falls short, as part of Flower for Celery Monitoring in Observability & Monitoring for Job Queues.

Problem Statement

A team runs Celery with RabbitMQ and uses Flower to look at tasks when something seems wrong. They have no alerting on Celery at all; an outage last month — workers lost their broker connection and stopped consuming for 40 minutes — was reported by a customer. Prometheus and Grafana already monitor the web tier. They want task success and failure rates, runtime percentiles, and an alert when workers go offline, and they would rather not add a new component if Flower can provide the data.

Prerequisites

  • Flower 1.2 or newer running against the same broker as the workers.
  • Celery workers started with events enabled (-E or worker_send_task_events = True).
  • Prometheus able to reach Flower over the network.
  • Flower's authentication configured, since the metrics endpoint sits behind it by default — see securing the Flower dashboard in production.

Step 1 — Turn On Task Events

Flower builds all of its knowledge from Celery events: task-received, task-started, task-succeeded, task-failed, task-retried, and worker heartbeats. Workers do not send task events by default, so without them Flower sees only heartbeats and its task metrics stay at zero.

# celeryconfig.py
worker_send_task_events = True   # same as starting workers with -E
task_send_sent_event = True      # adds task-sent from producers, for queue-time metrics

Events add a small amount of broker traffic — one message per state change per task. At thousands of tasks per second that becomes significant; Step 6 covers the trade-off.

How metrics reach Prometheus through Flower Celery workers publish task and heartbeat events to an events exchange on the broker. Flower subscribes to that exchange, keeps recent task and worker state in memory, and updates Prometheus counters and histograms as events arrive. Prometheus scrapes Flower's metrics endpoint on an interval. If events are disabled on the workers, Flower has nothing to count. Events in, metrics out Celery workers events enabled broker celeryev exchange Flower state + /metrics Prometheus scrape (pull) No events from workers means no task metrics in Flower.

Step 2 — Check the Endpoint

Flower serves metrics at /metrics on its normal port. Confirm it before configuring Prometheus:

curl -s -u "$FLOWER_USER:$FLOWER_PASS" http://flower:5555/metrics | grep '^flower_' | head
# flower_events_total{task="billing.charge",type="task-succeeded",worker="celery@w1"} 1832.0
# flower_task_runtime_seconds_bucket{le="0.5",task="billing.charge",worker="celery@w1"} 1790.0
# flower_worker_online{worker="celery@w1"} 1.0
# flower_worker_number_of_currently_executing_tasks{worker="celery@w1"} 3.0
# flower_task_prefetch_time_seconds{task="billing.charge",worker="celery@w1"} 0.004

The metrics that matter most:

Metric Type Meaning
flower_events_total counter events by task name, event type, and worker
flower_task_runtime_seconds histogram run time of succeeded tasks
flower_worker_online gauge 1 while a worker's heartbeats are arriving
flower_worker_number_of_currently_executing_tasks gauge tasks in progress per worker
flower_task_prefetch_time_seconds gauge time between a task being received and started, per task and worker

Step 3 — Configure the Scrape

Flower's authentication protects /metrics like every other page. Either give Prometheus its own credentials, or put a reverse proxy in front of Flower that allows /metrics without login only from Prometheus's network address and requires authentication for everything else. Do not turn authentication off entirely to make scraping easier: Flower's UI and API can revoke tasks and shut down workers. With basic auth:

scrape_configs:
  - job_name: flower
    scrape_interval: 30s
    metrics_path: /metrics
    basic_auth:
      username: prometheus
      password_file: /etc/prometheus/secrets/flower-password
    static_configs:
      - targets: ["flower.celery.svc:5555"]

If Flower runs with a URL prefix (--url_prefix=flower), the path becomes /flower/metrics. Run exactly one Flower instance per broker for metrics; two instances each counting the same events would double every rate in a dashboard that sums across targets.

Step 4 — Build Queries for Rates, Failures, and Runtime

Because flower_events_total counts every event type, most useful queries filter on type:

# tasks succeeded per second, by task
sum by (task) (rate(flower_events_total{type="task-succeeded"}[5m]))

# failure ratio per task
sum by (task) (rate(flower_events_total{type="task-failed"}[10m]))
  / sum by (task) (rate(flower_events_total{type=~"task-succeeded|task-failed"}[10m]))

# p95 runtime per task
histogram_quantile(0.95, sum by (task, le) (rate(flower_task_runtime_seconds_bucket[10m])))

# workers online
sum(flower_worker_online)
Which event answers which question task-succeeded events drive throughput and the runtime histogram. task-failed events give the failure ratio. task-retried events show retry pressure against a dependency. Worker heartbeats drive the worker online gauge used to detect outages. Queue depth is not in this list, because Flower does not read queue lengths from the broker. Event types and what they tell you task-succeeded throughput, runtime percentiles task-failed failure ratio per task task-retried retry pressure on a dependency worker-heartbeat worker online / offline not covered: queue depth

Step 5 — Add Alerts

Three alerts would have caught the 40-minute outage and the failure spikes that preceded it:

- alert: CeleryWorkersOffline
  expr: sum(flower_worker_online) < 2
  for: 3m
- alert: CeleryNoTasksSucceeding
  expr: sum(rate(flower_events_total{type="task-succeeded"}[10m])) == 0
  for: 10m
- alert: CeleryTaskFailureRatioHigh
  expr: |
    sum by (task) (rate(flower_events_total{type="task-failed"}[15m]))
      / sum by (task) (rate(flower_events_total{type=~"task-succeeded|task-failed"}[15m])) > 0.1
  for: 15m

Workers that lose their broker connection also stop sending heartbeats, so flower_worker_online drops within the heartbeat timeout — about two minutes by default. Alert on "no successes" as well, because a worker can be online yet stuck.

Step 6 — Know the Limits and Fill the Gaps

Flower is a convenient metrics source, not a complete one:

  • No queue depth. Flower does not read queue lengths from the broker, and backlog is the most important capacity signal. Add a broker exporter (RabbitMQ's built-in Prometheus plugin, or a Redis exporter reading list lengths) or instrument Celery with a Prometheus exporter.
  • State is in memory. When Flower restarts, counters reset (Prometheus rates cope) and events sent while it was down are never counted.
  • One process, all events. At very high task rates, a single Flower process can fall behind consuming events, delaying and distorting metrics. Beyond a few thousand tasks per second, a dedicated exporter or worker-side instrumentation scales better.
  • Label cardinality. Metrics are labelled by worker hostname; with autoscaled workers whose names change on every deploy, series accumulate. Use metric_relabel_configs to drop the worker label where you do not need it.

Also note that Flower only sees events from workers connected to the broker it watches; with several Celery apps or virtual hosts, run one Flower per broker and label targets accordingly. Treat Flower's metrics as a fast start and a good source for worker online status, and plan to add broker-level queue metrics before relying on it for capacity alerts.

Step 7 — Combine Flower with a Broker Exporter

The most complete low-effort setup pairs Flower's event metrics with the broker's own metrics. Flower tells you what workers are doing; the broker tells you what is waiting for them. Together they answer every first question during an incident.

Flower plus a broker exporter Flower supplies metrics derived from events: task rates, failures, runtime percentiles and worker online status. A broker exporter supplies what only the broker knows: queue depth, number of consumers per queue and, on RabbitMQ, message rates in and out. Grafana panels join the two sets by queue name so one dashboard shows both the work being done and the work waiting. Work being done + work waiting Flower /metrics task rates, failures, runtime workers online broker exporter queue depth, consumers message rates in and out Grafana dashboard panels joined by queue

For RabbitMQ, enable the rabbitmq_prometheus plugin and scrape port 15692; the rabbitmq_queue_messages_ready and rabbitmq_queue_consumers metrics are the ones to graph. For Redis, the redis_exporter's check-keys option can export the lengths of the Celery queue lists. Alert on a queue that has messages ready but zero consumers — that single rule would have caught the outage in the problem statement within minutes, even if Flower itself had been down. Keep task names consistent with queue names where you can, so panels from both sources line up without complex joins.

Verification

  • flower_events_total{type="task-succeeded"} increases in Prometheus as tasks run.
  • sum(flower_worker_online) matches the number of running workers, and drops within about two minutes after stopping one.
  • Stopping all workers in staging fires CeleryWorkersOffline and then CeleryNoTasksSucceeding.
  • The Grafana panel for p95 runtime per task roughly matches runtimes seen in Flower's task view.

Gotchas & Edge Cases

Events off after a config change. A deploy that drops -E or worker_send_task_events silently turns task metrics to zero while worker metrics look normal. The "no tasks succeeding" alert catches this.

Remote control disabled. If workers run with worker_enable_remote_control = False, Flower cannot inspect them but can still count events; the metrics keep working.

Worker names with process IDs. Workers started without an explicit -n name get a hostname-based default, and in containers the hostname is the pod name, which changes on every deploy. Each new name creates new series in Prometheus while the old ones go stale. Name workers by role (-n billing@%h) and drop or aggregate the worker label in recording rules so dashboards stay readable across deploys.

Clock and heartbeat interval. flower_worker_online depends on heartbeats arriving within Flower's expected interval. Workers under heavy CPU load or with long blocking tasks in the solo pool can miss heartbeats and flap between online and offline. Use for: durations on the alert, and avoid the solo pool for production workers.

Timeouts on scrape. Flower serves metrics from the same Tornado process as its UI. A heavy UI session can slow scrapes; set a scrape timeout above the default 10 seconds if you see gaps.

FAQ

Can I run Flower only for metrics, without anyone using the UI? Yes. Many teams run Flower internally just as an event-to-metrics bridge, with the UI reachable only through a port-forward.

Does Flower's runtime histogram include failed tasks? It is fed from succeeded tasks. Failure timing appears only in event counts, so if slow failures matter to you, instrument that in the worker.

Should I use Flower or celery-exporter? Both consume events. Flower gives you a UI and metrics in one process; a dedicated exporter usually adds queue length and is lighter. Pick one as the metrics source so counts are not duplicated.

How long does Flower keep task history? Flower keeps a bounded number of tasks in memory (--max_tasks, 10,000 by default) for its UI. That limit does not affect the Prometheus counters, which count every event since Flower started.

Related