Tracing Sidekiq Jobs with OpenTelemetry

In a Rails application, a slow user action often turns out to be a Sidekiq job that ran seconds or minutes after the request finished. Without tracing across the queue, the request trace ends at perform_async, and the job's own trace starts from nothing. OpenTelemetry's Sidekiq instrumentation carries the trace context in the job payload so the two join up. This guide configures it and explains the choices that matter, as part of Distributed Tracing for Async Jobs in Observability & Monitoring for Job Queues.

Problem Statement

A Rails marketplace confirms orders in a controller, then enqueues ChargeCardJob, SendReceiptJob, and NotifySellerJob. Customers sometimes receive receipts 20 minutes late. The team has OpenTelemetry on the web tier; traces show POST /orders finishing in 150 ms. Sidekiq has its own logs and metrics but no connection to the request that caused each job. When a receipt is late, nobody can see whether SendReceiptJob waited in the queue, retried after an SMTP error, or waited on ChargeCardJob. You want the three jobs to appear with the order request in tracing, including retries and time spent waiting.

Prerequisites

  • Rails 7 with Sidekiq 7 (ActiveJob or native Sidekiq::Job).
  • The OpenTelemetry Ruby SDK and an OTLP exporter, sending to a Collector or tracing backend.
  • The same OTEL_SERVICE_NAME convention for web and worker processes (for example shop-web, shop-sidekiq).

Step 1 — Install and Configure the Instrumentation

Add the SDK, the exporter, and the instrumentations — the Sidekiq one provides both client (enqueue) and server (execute) middleware:

# Gemfile
gem "opentelemetry-sdk"
gem "opentelemetry-exporter-otlp"
gem "opentelemetry-instrumentation-rails"
gem "opentelemetry-instrumentation-sidekiq"
gem "opentelemetry-instrumentation-active_job"
gem "opentelemetry-instrumentation-pg"
gem "opentelemetry-instrumentation-net_http"
# config/initializers/opentelemetry.rb
require "opentelemetry/sdk"
require "opentelemetry/exporter/otlp"

OpenTelemetry::SDK.configure do |c|
  c.service_name = ENV.fetch("OTEL_SERVICE_NAME", "shop")
  c.use "OpenTelemetry::Instrumentation::Rails"
  c.use "OpenTelemetry::Instrumentation::PG"
  c.use "OpenTelemetry::Instrumentation::Net::HTTP"
  c.use "OpenTelemetry::Instrumentation::Sidekiq", {
    span_naming: :job_class,              # "SendReceiptJob publish" / "SendReceiptJob process"
    propagation_style: :link,             # see Step 2
    trace_launcher_heartbeat: false,
    trace_poller_enqueue: false,
    trace_poller_wait: false,
  }
end

The same initializer runs in web and Sidekiq processes; set OTEL_SERVICE_NAME differently in each deployment so traces show which side a span came from. Turning off the launcher and poller spans avoids a stream of background spans that have nothing to do with jobs.

Step 2 — Choose Between Child Spans and Links

The instrumentation supports three propagation_style values, and the choice shapes how traces look:

Sidekiq propagation styles With the child style, the job's process span becomes a child of the enqueue span, so the request and the job share one trace that can last as long as the job waited. With the link style, the job starts its own trace whose root span carries a link to the enqueue span, so traces stay short but are navigable in both directions. With none, the job's trace has no connection to the request. child, link, or none child POST /orders SendReceipt process one long trace link POST /orders SendReceipt process span link two short traces none POST /orders SendReceipt process no connection time →
  • :child makes the job part of the request's trace. It is the most intuitive view for short waits: one trace, everything nested. But a job that waits 20 minutes or retries for hours produces a trace that spans hours, which some backends truncate or cannot search well.
  • :link (the gem's default) starts a new trace for each job and adds a span link back to the enqueueing span. Traces stay short and bounded; most backends show the link as a clickable reference in both directions.
  • :none disables propagation; use it only for jobs where the enqueue context is irrelevant, such as scheduled maintenance.

For request-driven jobs where you want the full story in one view and waits are short, :child works well. For most production systems with retries and variable waits, :link keeps traces manageable. The same trade-off, in more depth, is covered in span links for batch and fan-out jobs.

Step 3 — Understand What Goes Into the Payload

The client middleware injects W3C trace context into the job hash before it is pushed to Redis:

{
  "class" => "SendReceiptJob",
  "args" => [48213],
  "jid" => "b4f1…",
  "enqueued_at" => 1758189012.4,
  "traceparent" => "00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01"
}

The server middleware extracts it when the job starts. Because the context is in the job itself, it survives retries (a retried job keeps its payload) and scheduled jobs (perform_in stores the same hash in the schedule set). Jobs enqueued from places without an active span — Rails console, rake tasks, cron — carry no context and start fresh traces.

Step 4 — Handle ActiveJob

With ActiveJob on the Sidekiq adapter, two layers are involved: ActiveJob wraps your job in Sidekiq::ActiveJob::Wrapper. Enable both instrumentations and the Sidekiq one names spans after the wrapped class when span_naming: :job_class is set, so traces show SendReceiptJob rather than the wrapper. The ActiveJob instrumentation adds its own spans for perform and callbacks; if the nested spans look redundant, disable one of the two for execution and keep the Sidekiq one for propagation.

Step 5 — Add Attributes for Queue Questions

The instrumentation records messaging.system, the queue name, the job ID and the class. Add what you need to answer "why was this late?":

class TraceQueueAttributes
  include Sidekiq::ServerMiddleware
  def call(_job_instance, job, queue)
    span = OpenTelemetry::Trace.current_span
    span.set_attribute("sidekiq.retry_count", job["retry_count"].to_i)
    span.set_attribute("sidekiq.latency_ms", ((Time.now.to_f - job["enqueued_at"].to_f) * 1000).round)
    span.set_attribute("order.id", job["args"].first.to_s) if job["class"].to_s.end_with?("ReceiptJob")
    yield
  end
end

Sidekiq.configure_server do |config|
  config.server_middleware { |chain| chain.add TraceQueueAttributes }
end

Add this middleware after the OpenTelemetry one so the current span is the job's span. With sidekiq.latency_ms on every job span, a tracing query for "SendReceiptJob with latency above 60 s" finds late receipts directly, and each one shows whether the time was queue wait or retries.

Step 6 — Sample Consistently

Use parent-based sampling in both web and Sidekiq processes, so a sampled request produces sampled job spans. With :link propagation, the job's trace is a new root, and parent-based samplers treat it as a new decision — so a request sampled at 10% has its linked jobs sampled independently at 10%. If you need linked jobs to follow the request's decision, either use :child, or configure tail-based sampling in the Collector to keep traces with errors or high latency regardless of ratio.

How sampling interacts with propagation style With child propagation and a parent-based sampler, a sampled request always has its jobs sampled, and an unsampled one never does. With link propagation, the job starts a new root, so its sampling decision is independent and a sampled request may have unsampled jobs. Adding tail-based sampling in the Collector keeps every trace containing an error or exceeding a latency threshold, regardless of the head decision. Who decides whether a job is sampled child job follows the request's decision link job decides independently + tail sampling keep all errors and slow traces anyway Late receipts are exactly the traces tail sampling should keep.

Step 7 — Diagnose a Late Receipt

With tracing in place, the late-receipt question from the problem statement becomes a lookup. Search for SendReceiptJob process spans with sidekiq.latency_ms above 60,000, open one, and follow the link back to the order request:

Following one late receipt through its spans The order request completes in 150 milliseconds and enqueues the receipt job. The job starts within a second, but its first attempt spends 30 seconds waiting on the SMTP server and fails with a timeout. Sidekiq schedules a retry with exponential backoff, and the second attempt runs about 18 minutes later, succeeding in 400 milliseconds. The delay came from retry backoff, not from queue wait or capacity. Where 20 minutes went POST /orders 150 ms attempt 1: 30 s SMTP timeout retry backoff ≈ 18 min attempt 2: 0.4 s Not queue wait, not capacity: a timeout plus the retry schedule. Fix: shorter SMTP timeout, a faster first retry, a circuit breaker.

In the marketplace example, most late receipts looked like this: the first attempt hit a slow SMTP relay, waited 30 seconds for a timeout, and Sidekiq's default backoff scheduled the next attempt many minutes later. Adding workers would not have helped at all. The fix was a shorter SMTP timeout, a custom sidekiq_retry_in that retries the first failure after 30 seconds, and a fallback provider — changes that only made sense once traces showed where the time went. Metrics alone would have shown rising latency without saying whether it came from waiting, retries, or slow execution.

Verification

  • An order request in staging produces spans named SendReceiptJob publish in the web service and SendReceiptJob process in the Sidekiq service, connected by parent-child or a link depending on the style.
  • A job that fails once and succeeds on retry shows two process spans, the first with an exception event.
  • A job scheduled with perform_in(5.minutes) keeps its connection to the request.
  • Jobs enqueued from a rake task start their own traces without errors.
  • The latency attribute matches Sidekiq's reported queue latency for the same jobs.

Gotchas & Edge Cases

Initializer order in Sidekiq. Sidekiq loads Rails, so the initializer runs, but custom boot files that load job classes before Rails can bypass it. Check that OpenTelemetry.tracer_provider is configured in a Sidekiq console.

Exporter on shutdown. Sidekiq processes that exit without flushing lose the last batch of spans. Call OpenTelemetry.tracer_provider.shutdown in a config.on(:shutdown) hook.

Payload size. The traceparent string adds about 60 bytes per job — insignificant, but worth knowing for very high-volume queues with strict Redis memory budgets.

Unique-job digests. Uniqueness gems compute digests from the job's arguments, not the whole payload, so the added trace field does not break deduplication — but custom digest code that hashes the full job hash would. See Sidekiq unique jobs and deduplication.

FAQ

Do I need the ActiveJob instrumentation if I use Sidekiq directly? No. Native Sidekiq::Job classes need only the Sidekiq instrumentation.

Can I see Sidekiq queue latency without tracing? Yes, from metrics — see exporting Sidekiq metrics to Prometheus. Tracing adds the per-job explanation that aggregate metrics cannot.

Does this work with Sidekiq batches? Each job in a batch is traced individually. The batch itself has no span; link the callback job to the jobs that triggered it, as described in the span-links guide.

Can one job's trace show the jobs it enqueues in turn? Yes. When a job enqueues another job, the client middleware runs inside the first job's span, so the second job is connected to it in the same way the first was connected to the request. Chains of jobs become chains of linked or nested traces you can follow step by step.

Related