Exporting Sidekiq Metrics to Prometheus
Sidekiq keeps its statistics in Redis — processed and failed counts, queue sizes, latency, busy workers — and shows them in its Web UI. Getting them into Prometheus takes two pieces: an exporter for the shared queue state, and server middleware for per-job timings that Redis does not record. This guide sets up both, as part of Prometheus Metrics for Workers in Observability & Monitoring for Job Queues.
Problem Statement
A Rails platform runs Sidekiq across 12 pods with three queues. On-call engineers learn about problems from customers: a queue silently stopped processing for an hour after a bad deploy, and nobody noticed the retry set growing to 80,000 jobs. The Web UI shows the current state but no history and no alerting. The rest of the stack is monitored with Prometheus and Grafana. You want queue latency, backlog, retries, dead jobs, and per-job duration as Prometheus metrics, with alerts, and without every pod reporting duplicate queue numbers.
Prerequisites
- Sidekiq 7 on Ruby 3.1 or newer.
- Prometheus that can scrape pods (Kubernetes service discovery or static targets).
- The
prometheus-clientgem, or theyabeda-sidekiqgem if you prefer the Yabeda framework. - A decision on labels: queue and job class only.
Step 1 — Separate Shared State from Per-Process Work
Sidekiq metrics come in two kinds, and mixing them up is the most common mistake:
- Shared state lives in Redis and is the same no matter which process reads it: queue sizes, queue latency, the size of the retry, scheduled and dead sets, total processed and failed counts.
- Per-process work happens inside each worker: how long each job took, which jobs failed, how many threads are busy in this process.
Read shared state from one place — a small exporter deployment — and record per-process work in every pod. If every pod exported queue sizes, a dashboard that sums by queue would show twelve times the real backlog.
Step 2 — Export Queue State with a Collector
A standalone exporter reads Sidekiq's API on each scrape. With the plain prometheus-client gem, a custom collector keeps the numbers fresh without a timer:
# exporter.rb — run as its own deployment with 1 replica
require "sidekiq/api"
require "prometheus/client"
require "prometheus/middleware/exporter"
REG = Prometheus::Client.registry
Q_SIZE = REG.gauge(:sidekiq_queue_size, docstring: "Jobs enqueued", labels: [:queue])
Q_LATENCY = REG.gauge(:sidekiq_queue_latency_seconds, docstring: "Age of oldest job", labels: [:queue])
SETS = REG.gauge(:sidekiq_jobs, docstring: "Jobs in sorted sets", labels: [:set])
TOTALS = REG.gauge(:sidekiq_stats_total, docstring: "Lifetime counters", labels: [:kind])
class Refresh
def initialize(app) = @app = app
def call(env)
if env["PATH_INFO"] == "/metrics"
Sidekiq::Queue.all.each do |q|
Q_SIZE.set(q.size, labels: { queue: q.name })
Q_LATENCY.set(q.latency, labels: { queue: q.name })
end
stats = Sidekiq::Stats.new
SETS.set(stats.retry_size, labels: { set: "retry" })
SETS.set(stats.scheduled_size, labels: { set: "scheduled" })
SETS.set(stats.dead_size, labels: { set: "dead" })
TOTALS.set(stats.processed, labels: { kind: "processed" })
TOTALS.set(stats.failed, labels: { kind: "failed" })
end
@app.call(env)
end
end
Mount it with Rack (use Refresh; use Prometheus::Middleware::Exporter; run ->(_) { [404, {}, []] }) and deploy it next to Sidekiq. If you use Yabeda, yabeda-sidekiq provides the same gauges; set Yabeda::Sidekiq.config.collect_cluster_metrics = true only on the exporter and false in worker pods. Queue latency — the age of the oldest job — is the single most useful Sidekiq metric: it grows when a queue is starved even if the backlog is small.
Step 3 — Record Per-Job Metrics with Server Middleware
Sidekiq's server middleware wraps every job, which makes it the natural place to time jobs and count outcomes:
class PrometheusJobMetrics
include Sidekiq::ServerMiddleware
DURATION = Prometheus::Client.registry.histogram(
:sidekiq_job_duration_seconds, docstring: "Job run time",
labels: [:queue, :worker], buckets: [0.05, 0.1, 0.5, 1, 5, 15, 60, 300, 900])
OUTCOMES = Prometheus::Client.registry.counter(
:sidekiq_jobs_total, docstring: "Job executions", labels: [:queue, :worker, :result])
def call(job_instance, job, queue)
start = Process.clock_gettime(Process::CLOCK_MONOTONIC)
yield
OUTCOMES.increment(labels: { queue:, worker: job["class"], result: "success" })
rescue Exception
OUTCOMES.increment(labels: { queue:, worker: job["class"], result: "failure" })
raise
ensure
DURATION.observe(Process.clock_gettime(Process::CLOCK_MONOTONIC) - start,
labels: { queue:, worker: job["class"] })
end
end
Sidekiq.configure_server do |config|
config.server_middleware { |chain| chain.add PrometheusJobMetrics }
end
For ActiveJob wrappers, use job["wrapped"] || job["class"] so the label shows your job class rather than JobWrapper. Re-raising the exception is essential: middleware that swallows errors disables Sidekiq's retries.
Step 4 — Serve Metrics from Every Sidekiq Process
A Sidekiq process has no HTTP server. Start a tiny one in a thread when the server boots:
Sidekiq.configure_server do |config|
config.on(:startup) do
require "rack"
require "prometheus/middleware/exporter"
app = Rack::Builder.new do
use Prometheus::Middleware::Exporter
run ->(_) { [404, {}, ["not found"]] }
end
Thread.new { Rackup::Handler::WEBrick.run(app, Port: 9394, Host: "0.0.0.0", AccessLog: []) }
end
end
If you run Sidekiq Enterprise's multi-process mode (sidekiqswarm), each child process has its own registry. Use the prometheus-client gem's DirectFileStore so all children write to a shared directory and one endpoint serves the aggregate, or give each child its own port. Kubernetes pods with a single Sidekiq process need neither.
Step 5 — Query and Alert
The questions on-call engineers ask map to short queries:
# oldest job age per queue
max by (queue) (sidekiq_queue_latency_seconds)
# job failure ratio per class
sum by (worker) (rate(sidekiq_jobs_total{result="failure"}[10m]))
/ sum by (worker) (rate(sidekiq_jobs_total[10m]))
# p95 duration per class
histogram_quantile(0.95, sum by (worker, le) (rate(sidekiq_job_duration_seconds_bucket[10m])))
# dead set growth over the last hour
delta(sidekiq_jobs{set="dead"}[1h])
- alert: SidekiqQueueLatencyHigh
expr: max by (queue) (sidekiq_queue_latency_seconds{queue="critical"}) > 30
for: 10m
labels: { severity: page }
- alert: SidekiqStopped
expr: sum(rate(sidekiq_jobs_total[5m])) == 0 and sum(sidekiq_queue_size) > 0
for: 5m
labels: { severity: page }
Set latency thresholds per queue from its service-level objective; defining SLOs for job latency explains how.
Step 6 — Lay Out a Dashboard for On-Call
Metrics only help if the person paged at 3 a.m. can read them in seconds. Arrange the Sidekiq dashboard so it answers questions in the order they are asked: is anything wrong, where, and why.
The top row decides whether this is an incident. The middle row shows whether jobs are piling up in retries (a dependency is down), dying (a bug or bad data), or simply waiting for threads (capacity). The bottom row points at the job class responsible. Add a Grafana variable for queue so the same dashboard serves every queue, and link each alert's annotation to the dashboard with the queue pre-selected. The same structure works for any queue technology; building a Celery Grafana dashboard applies it to Celery.
Verification
- The exporter target reports one series per queue, matching the Web UI's numbers.
- Every Sidekiq pod exposes
sidekiq_job_duration_secondsandsidekiq_jobs_total, and the sum across pods matches the change insidekiq_stats_total{kind="processed"}over the same window. - A job that raises appears in
result="failure"and still lands in the retry set. - Stopping all Sidekiq pods in staging fires
SidekiqStoppedwithin about five minutes.
Gotchas & Edge Cases
Latency on empty queues is zero. A queue with no jobs reports latency 0, which looks healthy even if no workers are listening. Pair latency alerts with the "nothing processed" alert.
Lifetime counters reset. Sidekiq::Stats#processed can be reset from the Web UI. Prefer the middleware counters for rates; use the stats totals only as a cross-check.
Busy-thread metrics. Sidekiq::ProcessSet reports busy threads per process. Export it from the exporter (it is shared state) to see how saturated the fleet is; sustained 100% busy means more capacity is needed.
Exporter availability. If the single exporter pod is down, every queue gauge disappears and latency alerts go quiet rather than firing. Add an absent(sidekiq_queue_latency_seconds) alert, and run the exporter with a readiness probe so a failed Redis connection is visible instead of silently returning stale values.
Label names in middleware. Using job["queue"] from the payload rather than the queue argument can differ when jobs are moved between queues by a routing middleware. Use the argument, which is the queue the job was actually fetched from.
Scrape cost with many queues. Each queue's latency needs a Redis call. With hundreds of queues, cache results for a few seconds so a slow Redis cannot make scrapes time out.
FAQ
Is Sidekiq Enterprise's metrics feature a replacement? Sidekiq 7 includes a Metrics tab with per-job execution history, and Enterprise can send metrics to StatsD. Neither exposes Prometheus metrics directly, so the setup above is still needed for Prometheus-based alerting.
Should I use yabeda-sidekiq or write my own? Yabeda saves code and handles the multi-process case; writing your own gives full control over names and labels. Both produce the same kind of metrics — pick one and stay consistent across services.
How do I see which job class is filling the queue?
Queue size has no class label. Count by class from the enqueue side (client middleware incrementing a counter) or sample Sidekiq::Queue#each in the exporter for the top classes, capped to a fixed number.
What scrape interval should the exporter use? Fifteen to thirty seconds is enough for queues. Latency alerts use windows of several minutes, so faster scraping adds Redis load without improving detection.
Related
- Prometheus Metrics for Workers — what to measure and why.
- Instrumenting BullMQ with prom-client — the Node.js counterpart.
- Alerting on Queue Backlog with Prometheus — complete alert rules.
- Sidekiq Queue Weights and Capsules — acting on what latency shows.