Running Flower on Kubernetes
Flower is a single Python process that listens to Celery's event stream and serves a web UI, an API, and Prometheus metrics. Running it on Kubernetes is simple once you respect two facts: it keeps its state in memory, and it has powerful controls that must not be exposed to the internet. This guide builds a production-ready deployment, as part of Flower for Celery Monitoring in Observability & Monitoring for Job Queues.
Problem Statement
A team moved its Celery workers to Kubernetes and ran Flower the same way someone had shown in a blog post: a Deployment with three replicas behind a public LoadBalancer, with basic auth credentials in the container arguments. Each page load hit a different replica with different in-memory state, so the task list jumped around. Prometheus scraped all three replicas and dashboards showed three times the real task rate. The credentials were visible to anyone who could read the Deployment, and a security scan flagged the public endpoint. You want one reliable Flower instance, reachable only by the team, with credentials in Secrets and metrics scraped once.
Prerequisites
- A Kubernetes cluster where the Celery workers already run, with network access to the broker.
- Workers started with task events enabled (
-E). - An Ingress controller with authentication support (for example ingress-nginx with oauth2-proxy), or a policy of access through
kubectl port-forward. - The Prometheus Operator, if you scrape with a
ServiceMonitor.
Step 1 — Run Exactly One Replica
Flower's state lives in the process. Multiple replicas would each consume the same events independently, disagree about what they have seen, and double-count in Prometheus. Use a Deployment with one replica and the Recreate strategy, so a rollout never runs two at once:
apiVersion: apps/v1
kind: Deployment
metadata: { name: flower, namespace: celery }
spec:
replicas: 1
strategy: { type: Recreate }
selector: { matchLabels: { app: flower } }
template:
metadata: { labels: { app: flower } }
spec:
containers:
- name: flower
image: mher/flower:2.0
args: ["celery", "--broker=$(BROKER_URL)", "flower",
"--port=5555", "--persistent=True", "--db=/data/flower.db",
"--max_tasks=20000", "--url_prefix=flower"]
envFrom: [{ secretRef: { name: flower-secrets } }]
ports: [{ containerPort: 5555, name: http }]
volumeMounts: [{ name: data, mountPath: /data }]
volumes:
- name: data
persistentVolumeClaim: { claimName: flower-data }
A few seconds of downtime during a rollout is acceptable for a monitoring tool; misleading numbers are not. If you need Flower to survive node failures, the Deployment reschedules the pod automatically.
Step 2 — Keep Credentials in Secrets
The broker URL contains a password, and Flower's own authentication settings are secrets too. Put them in a Secret (ideally synced from your secret manager) and load them with envFrom, never as literal container arguments:
apiVersion: v1
kind: Secret
metadata: { name: flower-secrets, namespace: celery }
stringData:
BROKER_URL: amqp://flower:REDACTED@rabbitmq.celery.svc:5672/prod
FLOWER_BASIC_AUTH: ops:REDACTED
Flower reads any option from an environment variable prefixed with FLOWER_, so FLOWER_BASIC_AUTH sets basic auth without it appearing in kubectl describe. Give Flower its own broker user with only the permissions it needs — reading the events exchange and sending control broadcasts — rather than the workers' credentials.
Step 3 — Persist State Across Restarts
Without persistence, a pod restart forgets every task Flower had seen. --persistent=True --db=/data/flower.db writes state to a file on shutdown and reloads it on start. Back it with a small PersistentVolumeClaim (1 GiB is plenty):
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: flower-data, namespace: celery }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 1Gi } }
ReadWriteOnce is another reason for a single replica and Recreate: the volume can attach to only one pod. Persistence is best-effort — state is saved on a clean shutdown, so an OOM kill loses changes since the last start. Prometheus counters are unaffected because Prometheus stores the history.
Step 4 — Add Probes and Resource Limits
Flower's memory grows with the number of tasks it keeps (--max_tasks) and with the size of task arguments. Measure it under real load and set limits with headroom:
resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { memory: 768Mi }
readinessProbe:
httpGet: { path: /flower/healthcheck, port: http }
periodSeconds: 10
livenessProbe:
httpGet: { path: /flower/healthcheck, port: http }
periodSeconds: 30
failureThreshold: 4
The /healthcheck endpoint returns OK without authentication. If memory keeps growing toward the limit, lower --max_tasks; at high task rates, Flower's single process can also fall behind on events, which shows as CPU pinned at one core.
The figures above are typical for small task arguments; measure your own by watching container_memory_working_set_bytes for the Flower pod over a busy day. The plateau matters more than the starting point: memory rises until Flower holds --max_tasks tasks and then stays level as old tasks are evicted. Set the limit comfortably above that plateau, and lower --max_tasks rather than raising the limit if the pod is OOM-killed — history in the UI is useful, but most incident questions concern the last few minutes, and Prometheus keeps the long-term record.
Step 5 — Expose It Safely
Flower can revoke tasks, change rate limits, and shut down workers. Never expose it publicly without strong authentication. Two good options:
- Port-forward only. Keep the Service as
ClusterIPand let engineers runkubectl -n celery port-forward svc/flower 5555:5555. Kubernetes RBAC decides who can do this, and nothing is exposed at all. - Ingress with single sign-on. Put oauth2-proxy (or your platform's equivalent) in front of an internal Ingress, restricted to the on-call group. Keep Flower's basic auth as a second layer if you like.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: flower
namespace: celery
annotations:
nginx.ingress.kubernetes.io/auth-url: "https://oauth2.internal/oauth2/auth"
nginx.ingress.kubernetes.io/auth-signin: "https://oauth2.internal/oauth2/start?rd=$escaped_request_uri"
spec:
ingressClassName: internal-nginx
rules:
- host: ops.internal.example.com
http:
paths:
- path: /flower
pathType: Prefix
backend: { service: { name: flower, port: { number: 5555 } } }
--url_prefix=flower in Step 1 makes Flower's links work under the /flower path. The wider hardening checklist is in securing the Flower dashboard in production.
Step 6 — Scrape Metrics Once
With a single replica, a ServiceMonitor gives Prometheus exactly one target:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: flower, namespace: celery }
spec:
selector: { matchLabels: { app: flower } }
endpoints:
- port: http
path: /flower/metrics
interval: 30s
basicAuth:
username: { name: flower-scrape, key: username }
password: { name: flower-scrape, key: password }
Add a NetworkPolicy that allows ingress to the Flower pod only from the Ingress controller and Prometheus namespaces. What to do with the metrics is covered in scraping Flower metrics with Prometheus.
Verification
kubectl -n celery get pods -l app=flowershows one pod; a rollout replaces it without a second pod running alongside.kubectl describe deploy flowershows no credentials in plain text.- Deleting the pod brings back a new one with the previous task history loaded.
- The Flower URL returns a login redirect from outside the on-call group and is unreachable from the public internet.
- Prometheus lists exactly one
flowertarget, andflower_worker_onlinematches the running worker pods.
Gotchas & Edge Cases
Worker names change every deploy. Worker pods get new hostnames on each rollout, so Flower accumulates offline workers. They disappear after Flower restarts; set --purge_offline_workers=600 to drop workers that have been offline for ten minutes.
Events from several namespaces. Flower sees only the broker (and virtual host) it is connected to. With separate brokers per environment, run one Flower per broker and label them clearly.
Liveness probes too strict. Under heavy event load, Flower can respond slowly; an aggressive liveness probe will restart it repeatedly and lose in-memory state each time. Use generous thresholds.
Broker connection limits. Flower opens its own broker connections for events and for control commands. On managed RabbitMQ or Redis plans with tight connection caps, count Flower in the budget, and make sure its broker user is allowed to create the temporary queues that event consumers need.
Image tags. Pin the Flower image to a version, and match its Celery version to your workers where possible; mismatched Celery versions can misread event formats.
FAQ
Can I run Flower as a sidecar in a worker pod? You can, but then every worker pod has its own Flower, which brings back the multiple-replica problems. Run it as its own Deployment.
Does Flower need access to the result backend? No. It works from events. Some UI pages show results only if the backend is configured, but metrics and control do not need it.
Is a StatefulSet better than a Deployment?
Either works for a single replica with a volume. A Deployment with Recreate is simpler and sufficient.
How do I upgrade Flower safely?
Change the image tag and let the Recreate rollout replace the pod. State is saved on shutdown and reloaded, so the task history survives. Test new major versions in staging first, because option names and URL paths occasionally change between releases.
Related
- Flower for Celery Monitoring — what Flower shows and when to use it.
- Securing the Flower Dashboard in Production — the full hardening checklist.
- Scraping Flower Metrics with Prometheus — using the metrics endpoint.
- Auto-Scaling Celery Workers on Kubernetes — the workers Flower watches.