Sidekiq Quiet and Shutdown Timeouts
Sidekiq's shutdown is a two-phase sequence — stop fetching, then give running jobs a bounded time to finish — and deploys go wrong when that sequence and the orchestrator's timing disagree. This guide aligns them, as part of Graceful Shutdown & Worker Deployments in Backend Frameworks & Worker Scaling.
Problem Statement
A Rails app runs Sidekiq on Kubernetes with the default 30-second termination grace period and Sidekiq's default 25-second shutdown timeout. After each deploy, support sees a handful of customers with duplicated exports and a few "stuck" reports that never finished. Logs show Sidekiq pushing unfinished jobs back to Redis at shutdown, and some pods being SIGKILLed before that push completed. Long export jobs (up to 10 minutes) never finish during a deploy and restart from scratch every time. You want deploys that finish short jobs, hand back long ones promptly and cleanly, never lose a job to SIGKILL, and avoid restarting long work from zero.
Prerequisites
- Sidekiq 7.x (OSS, Pro, or Enterprise — behaviour differs in Step 3).
- Access to the Kubernetes Deployment spec (
terminationGracePeriodSeconds,preStop). - Job duration percentiles per queue.
- Idempotent jobs, since interrupted jobs run again.
Step 1 — Understand the Signals
Sidekiq responds to two signals during shutdown:
- TSTP (quiet): stop fetching new jobs; keep running current ones. The process stays up.
- TERM (shutdown): stop fetching, wait up to the shutdown timeout (
-t, default 25 seconds) for running jobs, then push any still-running jobs back to Redis and exit.
# config/sidekiq.yml
:concurrency: 10
:timeout: 25 # the -t shutdown timeout: seconds to wait for running jobs on TERM
:queues:
- critical
- default
- exports
The shutdown timeout must be shorter than the orchestrator's grace period, with room for Sidekiq to push unfinished jobs back and exit; otherwise SIGKILL arrives during the push.
Step 2 — Align the Grace Period with the Timeout
Set the Kubernetes grace period well above Sidekiq's timeout. Add a short preStop so load balancers and autoscalers settle, and remember that preStop time counts against the grace period.
spec:
terminationGracePeriodSeconds: 45 # > timeout (25) + preStop (5) + push-back margin
containers:
- name: sidekiq
command: ["bundle", "exec", "sidekiq", "-C", "config/sidekiq.yml"]
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"]
Make sure Sidekiq is PID 1 (or started with exec from an entrypoint script), or the TERM signal may never reach it and every pod is SIGKILLed at the end of the grace period. The general drain pattern is described in Graceful Shutdown & Worker Deployments.
Step 3 — Know What Happens to Unfinished Jobs
What "push back" means depends on the edition:
| Edition | Fetch mechanism | Job still running at timeout | Process SIGKILLed |
|---|---|---|---|
| Sidekiq OSS | BRPOP (basic fetch) |
Pushed back to queue | Job lost |
| Sidekiq Pro | super_fetch |
Pushed back to queue | Recovered from private working queue on restart |
| Enterprise | super_fetch + more |
Same as Pro | Same as Pro |
With OSS basic fetch, a job popped by a process that is SIGKILLed exists nowhere — the same at-most-once behaviour described in Redis Streams vs Redis Lists for job queues. Pro's super_fetch keeps in-progress jobs in a per-process working list and recovers them after a crash. If you run OSS, getting the timeout and grace period right is the only protection against losing jobs on deploy.
# Sidekiq Pro: enable reliable fetch
Sidekiq.configure_server do |config|
config.super_fetch!
end
Step 4 — Quiet Early for Long-Running Queues
A 25-second timeout cannot accommodate 10-minute exports. Instead of stretching the timeout (and every deploy), quiet the export process long before stopping it, so it stops taking new exports and finishes the ones it has.
# Separate deployment for the exports queue, with a long grace period and early quiet
spec:
terminationGracePeriodSeconds: 660 # exports can take 10 min
containers:
- name: sidekiq-exports
command: ["bundle", "exec", "sidekiq", "-q", "exports", "-c", "3", "-t", "600"]
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "kill -TSTP 1 && sleep 20"] # quiet, then allow TERM
Isolating long jobs in their own deployment means only that deployment rolls slowly; the main workers still deploy in seconds. For very long work, checkpointing (Step 5) is better than any timeout.
Step 5 — Checkpoint Long Jobs So Interruption Is Cheap
Even with early quiet, node failures and emergency deploys interrupt long jobs. Make them resumable: record progress and skip completed work on retry.
class ExportJob
include Sidekiq::Job
sidekiq_options queue: "exports", retry: 5
def perform(export_id)
export = Export.find(export_id)
batches = export.row_batches(size: 5_000)
batches.each_with_index do |batch, i|
next if i < export.completed_batches # resume point
export.append_csv(batch) # idempotent per batch index
export.update_column(:completed_batches, i + 1)
return if Sidekiq::CLI.instance&.stopping? # stop between batches on shutdown
end
export.finalize!
end
end
Checking stopping? between batches lets the job return promptly when shutdown begins; because it returns without error, re-enqueue it explicitly (ExportJob.perform_async(export_id)) before returning, or let the push-back handle it if the timeout is reached. Either way, the next run resumes at the recorded batch instead of starting over.
Step 6 — Watch Deploys for Lost or Repeated Jobs
Measure what shutdown does in practice: jobs pushed back at shutdown, jobs retried after a deploy, and pods that exit with SIGKILL.
Sidekiq.configure_server do |config|
config.on(:quiet) { Rails.logger.info(event: "sidekiq.quiet") }
config.on(:shutdown) { Rails.logger.info(event: "sidekiq.shutdown", busy: Sidekiq::WorkSet.new.size) }
end
# Pods killed rather than exiting cleanly (exit code 137)
sum(increase(kube_pod_container_status_last_terminated_exitcode{container=~"sidekiq.*"}[1d]) == 137)
Any SIGKILL of a Sidekiq OSS pod is a potential lost job; treat it as a bug in the timing configuration.
Step 7 — Keep Rollouts from Draining Capacity
During a rolling deploy, quieting pods still count as running but take no new work. If the rollout replaces many pods at once, effective capacity drops sharply while old pods drain. Keep maxUnavailable low and add surge capacity so new pods start before old ones quiet.
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 25% # start replacements first
maxUnavailable: 0 # never remove capacity before its replacement is ready
With readiness gated on Sidekiq having started (a simple probe that checks the process heartbeat in Redis), the rollout proceeds only as fast as new pods come up, and queue latency stays flat throughout the deploy. The broader rollout mechanics are in zero-downtime worker deploys on Kubernetes.
Verification
Run a deploy under load in staging with a mix of 1-second and 5-minute jobs:
kubectl rollout restart deployment/sidekiq
kubectl rollout restart deployment/sidekiq-exports
kubectl get pods -l app=sidekiq -w # all old pods should terminate with exit code 0
Then reconcile: every job enqueued during the test completed exactly once (idempotency checks record duplicates), and every export finished without restarting from batch zero.
Gotchas & Edge Cases
Shell entrypoints swallowing TERM. sh -c "bundle exec sidekiq" makes the shell PID 1; it does not forward signals. Use exec or run Sidekiq directly.
Timeout larger than grace. A -t 60 with a 30-second grace period guarantees SIGKILL mid-shutdown.
Autoscaler scale-down. Scale-down evicts pods the same way as deploys; long jobs need the same protection, or scale-down stabilization windows long enough to avoid churn.
Scheduled and retry sets are safe. Jobs waiting in the scheduled or retry sets live in Redis, not in the process, and are unaffected by shutdown.
FAQ
Is TSTP necessary if TERM already stops fetching? For short jobs, TERM alone is fine. TSTP is useful when you want a long quiet period before the stop — for long-running queues — or during manual maintenance.
Should I raise the timeout for all workers? No. It slows every deploy and hides long jobs. Split long jobs into their own deployment and checkpoint them.
How do I pick the timeout for the main deployment? Take the p99 duration of jobs on the queues that deployment serves and add a margin; jobs that exceed it are pushed back and rerun, which is acceptable if they are rare and idempotent. If the p99 is above a minute, the queue probably mixes long and short jobs and should be split, as in Step 4.
What happens to jobs pushed back at shutdown? They go to the front of their queue and are picked up by another process almost immediately, with no retry count consumed — the interrupted run simply did not happen from Sidekiq's point of view. Any side effects it performed before interruption did happen, which is why idempotency matters.
Can Kubernetes tell Sidekiq to quiet on its own?
Only through the preStop hook, as shown in Step 4. There is no built-in signal for "stop taking work but keep running"; TSTP from preStop is the standard way to get that phase.
Does super_fetch make timing irrelevant? It prevents loss after SIGKILL, but interrupted jobs still restart. Timing still decides how often that happens.
Related
- Graceful Shutdown & Worker Deployments — cross-framework principles.
- Zero-Downtime Worker Deploys on Kubernetes — rollout strategy.
- Sidekiq Performance Tuning — concurrency and queues.
- Draining RQ Workers Safely — the equivalent for RQ.