Exporting Sidekiq Metrics to Prometheus

Sidekiq keeps its statistics in Redis — processed and failed counts, queue sizes, latency, busy workers — and shows them in its Web UI. Getting them into Prometheus takes two pieces: an exporter for the shared queue state, and server middleware for per-job timings that Redis does not record. This guide sets up both, as part of Prometheus Metrics for Workers in Observability & Monitoring for Job Queues.

Problem Statement

A Rails platform runs Sidekiq across 12 pods with three queues. On-call engineers learn about problems from customers: a queue silently stopped processing for an hour after a bad deploy, and nobody noticed the retry set growing to 80,000 jobs. The Web UI shows the current state but no history and no alerting. The rest of the stack is monitored with Prometheus and Grafana. You want queue latency, backlog, retries, dead jobs, and per-job duration as Prometheus metrics, with alerts, and without every pod reporting duplicate queue numbers.

Prerequisites

  • Sidekiq 7 on Ruby 3.1 or newer.
  • Prometheus that can scrape pods (Kubernetes service discovery or static targets).
  • The prometheus-client gem, or the yabeda-sidekiq gem if you prefer the Yabeda framework.
  • A decision on labels: queue and job class only.

Step 1 — Separate Shared State from Per-Process Work

Sidekiq metrics come in two kinds, and mixing them up is the most common mistake:

  • Shared state lives in Redis and is the same no matter which process reads it: queue sizes, queue latency, the size of the retry, scheduled and dead sets, total processed and failed counts.
  • Per-process work happens inside each worker: how long each job took, which jobs failed, how many threads are busy in this process.
Two sources of Sidekiq metrics Shared state such as queue sizes, latency, and the retry and dead sets is stored in Redis and read by a single exporter deployment. Per-process work such as job duration histograms and failure counters is recorded by server middleware inside each Sidekiq pod. Prometheus scrapes the exporter once and every Sidekiq pod individually, so queue numbers are not duplicated and job metrics can be summed across pods. Shared state once, per-process work everywhere Redis queues, retry, dead sets exporter (1 replica) gauges on scrape sidekiq pod 1 sidekiq pod 2 … pod 12 middleware: durations, failures, busy threads Prometheus

Read shared state from one place — a small exporter deployment — and record per-process work in every pod. If every pod exported queue sizes, a dashboard that sums by queue would show twelve times the real backlog.

Step 2 — Export Queue State with a Collector

A standalone exporter reads Sidekiq's API on each scrape. With the plain prometheus-client gem, a custom collector keeps the numbers fresh without a timer:

# exporter.rb — run as its own deployment with 1 replica
require "sidekiq/api"
require "prometheus/client"
require "prometheus/middleware/exporter"

REG = Prometheus::Client.registry
Q_SIZE    = REG.gauge(:sidekiq_queue_size, docstring: "Jobs enqueued", labels: [:queue])
Q_LATENCY = REG.gauge(:sidekiq_queue_latency_seconds, docstring: "Age of oldest job", labels: [:queue])
SETS      = REG.gauge(:sidekiq_jobs, docstring: "Jobs in sorted sets", labels: [:set])
TOTALS    = REG.gauge(:sidekiq_stats_total, docstring: "Lifetime counters", labels: [:kind])

class Refresh
  def initialize(app) = @app = app
  def call(env)
    if env["PATH_INFO"] == "/metrics"
      Sidekiq::Queue.all.each do |q|
        Q_SIZE.set(q.size, labels: { queue: q.name })
        Q_LATENCY.set(q.latency, labels: { queue: q.name })
      end
      stats = Sidekiq::Stats.new
      SETS.set(stats.retry_size,     labels: { set: "retry" })
      SETS.set(stats.scheduled_size, labels: { set: "scheduled" })
      SETS.set(stats.dead_size,      labels: { set: "dead" })
      TOTALS.set(stats.processed, labels: { kind: "processed" })
      TOTALS.set(stats.failed,    labels: { kind: "failed" })
    end
    @app.call(env)
  end
end

Mount it with Rack (use Refresh; use Prometheus::Middleware::Exporter; run ->(_) { [404, {}, []] }) and deploy it next to Sidekiq. If you use Yabeda, yabeda-sidekiq provides the same gauges; set Yabeda::Sidekiq.config.collect_cluster_metrics = true only on the exporter and false in worker pods. Queue latency — the age of the oldest job — is the single most useful Sidekiq metric: it grows when a queue is starved even if the backlog is small.

Step 3 — Record Per-Job Metrics with Server Middleware

Sidekiq's server middleware wraps every job, which makes it the natural place to time jobs and count outcomes:

class PrometheusJobMetrics
  include Sidekiq::ServerMiddleware
  DURATION = Prometheus::Client.registry.histogram(
    :sidekiq_job_duration_seconds, docstring: "Job run time",
    labels: [:queue, :worker], buckets: [0.05, 0.1, 0.5, 1, 5, 15, 60, 300, 900])
  OUTCOMES = Prometheus::Client.registry.counter(
    :sidekiq_jobs_total, docstring: "Job executions", labels: [:queue, :worker, :result])

  def call(job_instance, job, queue)
    start = Process.clock_gettime(Process::CLOCK_MONOTONIC)
    yield
    OUTCOMES.increment(labels: { queue:, worker: job["class"], result: "success" })
  rescue Exception
    OUTCOMES.increment(labels: { queue:, worker: job["class"], result: "failure" })
    raise
  ensure
    DURATION.observe(Process.clock_gettime(Process::CLOCK_MONOTONIC) - start,
                     labels: { queue:, worker: job["class"] })
  end
end

Sidekiq.configure_server do |config|
  config.server_middleware { |chain| chain.add PrometheusJobMetrics }
end

For ActiveJob wrappers, use job["wrapped"] || job["class"] so the label shows your job class rather than JobWrapper. Re-raising the exception is essential: middleware that swallows errors disables Sidekiq's retries.

Step 4 — Serve Metrics from Every Sidekiq Process

A Sidekiq process has no HTTP server. Start a tiny one in a thread when the server boots:

Sidekiq.configure_server do |config|
  config.on(:startup) do
    require "rack"
    require "prometheus/middleware/exporter"
    app = Rack::Builder.new do
      use Prometheus::Middleware::Exporter
      run ->(_) { [404, {}, ["not found"]] }
    end
    Thread.new { Rackup::Handler::WEBrick.run(app, Port: 9394, Host: "0.0.0.0", AccessLog: []) }
  end
end

If you run Sidekiq Enterprise's multi-process mode (sidekiqswarm), each child process has its own registry. Use the prometheus-client gem's DirectFileStore so all children write to a shared directory and one endpoint serves the aggregate, or give each child its own port. Kubernetes pods with a single Sidekiq process need neither.

Step 5 — Query and Alert

The questions on-call engineers ask map to short queries:

# oldest job age per queue
max by (queue) (sidekiq_queue_latency_seconds)

# job failure ratio per class
sum by (worker) (rate(sidekiq_jobs_total{result="failure"}[10m]))
  / sum by (worker) (rate(sidekiq_jobs_total[10m]))

# p95 duration per class
histogram_quantile(0.95, sum by (worker, le) (rate(sidekiq_job_duration_seconds_bucket[10m])))

# dead set growth over the last hour
delta(sidekiq_jobs{set="dead"}[1h])
Four Sidekiq alerts and their severity Queue latency above its target for ten minutes pages on-call, because users are waiting. Zero jobs processed for five minutes while a queue has work also pages, because processing has stopped. A retry set that grows steadily for thirty minutes raises a warning, because a dependency is failing. Any increase in the dead set creates a ticket, because jobs have exhausted their retries and need a human decision. What to alert on, and how loudly page: latency over target 10 min users are waiting page: nothing processed 5 min while the queue has work warn: retry set growing 30 min a dependency is failing ticket: dead set increased jobs need a human decision
- alert: SidekiqQueueLatencyHigh
  expr: max by (queue) (sidekiq_queue_latency_seconds{queue="critical"}) > 30
  for: 10m
  labels: { severity: page }
- alert: SidekiqStopped
  expr: sum(rate(sidekiq_jobs_total[5m])) == 0 and sum(sidekiq_queue_size) > 0
  for: 5m
  labels: { severity: page }

Set latency thresholds per queue from its service-level objective; defining SLOs for job latency explains how.

Step 6 — Lay Out a Dashboard for On-Call

Metrics only help if the person paged at 3 a.m. can read them in seconds. Arrange the Sidekiq dashboard so it answers questions in the order they are asked: is anything wrong, where, and why.

A Sidekiq dashboard in three rows The top row answers whether users are affected, with queue latency per queue, throughput, and failure ratio. The middle row answers where the trouble is, with retry set size, dead set size, and busy threads as a share of total. The bottom row answers why, with p95 duration and failure count broken down by job class. Read top to bottom: impact, location, cause latency per queue throughput failure ratio retry set size dead set size busy threads % p95 duration by job class failures by job class

The top row decides whether this is an incident. The middle row shows whether jobs are piling up in retries (a dependency is down), dying (a bug or bad data), or simply waiting for threads (capacity). The bottom row points at the job class responsible. Add a Grafana variable for queue so the same dashboard serves every queue, and link each alert's annotation to the dashboard with the queue pre-selected. The same structure works for any queue technology; building a Celery Grafana dashboard applies it to Celery.

Verification

  • The exporter target reports one series per queue, matching the Web UI's numbers.
  • Every Sidekiq pod exposes sidekiq_job_duration_seconds and sidekiq_jobs_total, and the sum across pods matches the change in sidekiq_stats_total{kind="processed"} over the same window.
  • A job that raises appears in result="failure" and still lands in the retry set.
  • Stopping all Sidekiq pods in staging fires SidekiqStopped within about five minutes.

Gotchas & Edge Cases

Latency on empty queues is zero. A queue with no jobs reports latency 0, which looks healthy even if no workers are listening. Pair latency alerts with the "nothing processed" alert.

Lifetime counters reset. Sidekiq::Stats#processed can be reset from the Web UI. Prefer the middleware counters for rates; use the stats totals only as a cross-check.

Busy-thread metrics. Sidekiq::ProcessSet reports busy threads per process. Export it from the exporter (it is shared state) to see how saturated the fleet is; sustained 100% busy means more capacity is needed.

Exporter availability. If the single exporter pod is down, every queue gauge disappears and latency alerts go quiet rather than firing. Add an absent(sidekiq_queue_latency_seconds) alert, and run the exporter with a readiness probe so a failed Redis connection is visible instead of silently returning stale values.

Label names in middleware. Using job["queue"] from the payload rather than the queue argument can differ when jobs are moved between queues by a routing middleware. Use the argument, which is the queue the job was actually fetched from.

Scrape cost with many queues. Each queue's latency needs a Redis call. With hundreds of queues, cache results for a few seconds so a slow Redis cannot make scrapes time out.

FAQ

Is Sidekiq Enterprise's metrics feature a replacement? Sidekiq 7 includes a Metrics tab with per-job execution history, and Enterprise can send metrics to StatsD. Neither exposes Prometheus metrics directly, so the setup above is still needed for Prometheus-based alerting.

Should I use yabeda-sidekiq or write my own? Yabeda saves code and handles the multi-process case; writing your own gives full control over names and labels. Both produce the same kind of metrics — pick one and stay consistent across services.

How do I see which job class is filling the queue? Queue size has no class label. Count by class from the enqueue side (client middleware incrementing a counter) or sample Sidekiq::Queue#each in the exporter for the top classes, capped to a fixed number.

What scrape interval should the exporter use? Fifteen to thirty seconds is enough for queues. Latency alerts use windows of several minutes, so faster scraping adds Redis load without improving detection.

Related