Running Rails Solid Queue in Production
Solid Queue is the default Active Job backend in Rails 8, and this guide takes it from the generator's defaults to a production configuration, as part of Database-Backed Job Queues in Backend Frameworks & Worker Scaling. The defaults are sensible for a small app; at real traffic you need to decide where the queue tables live, size workers against database connections, and keep finished-job tables from growing forever.
Problem Statement
A Rails 8 application migrated from Sidekiq to Solid Queue to drop its Redis dependency. In development everything worked. In production three problems appeared within a week: an ActiveRecord::ConnectionTimeoutError storm when the worker pod scaled up, jobs occasionally referencing records that did not exist yet, and a solid_queue_jobs table that grew to 40 million rows. You want a configuration where workers never exhaust the connection pool, enqueues are safe relative to the business transaction, recurring tasks run once, and table growth is bounded.
Prerequisites
- Rails 7.1+ with the
solid_queuegem (Rails 8 includes it by default), Postgres or MySQL 8. bin/rails solid_queue:installrun once, producingconfig/queue.yml,config/recurring.yml, anddb/queue_schema.rb.- A process manager or container platform to run
bin/jobs(the Solid Queue supervisor) alongside the web process. - Familiarity with Active Job APIs (
perform_later,retry_on,discard_on).
Step 1 — Decide Where the Queue Tables Live
The Rails 8 generator configures a separate queue database (queue: in database.yml) with its own connection. That isolates queue churn from application tables, but it also means perform_later inside an application transaction writes to a different database: the job row commits immediately even if the surrounding transaction later rolls back. That is the source of "jobs for records that don't exist".
# config/database.yml — option A: separate queue database (the generator default)
production:
primary:
<<: *default
database: app_production
queue:
<<: *default
database: app_production_queue
migrations_paths: db/queue_migrate
# config/environments/production.rb — option A
config.active_job.queue_adapter = :solid_queue
config.solid_queue.connects_to = { database: { writing: :queue } }
# Rails 7.2+: enqueue only after the surrounding transaction commits
config.active_job.enqueue_after_transaction_commit = :default
With a separate database, turn on enqueue_after_transaction_commit so Active Job defers the insert until the application transaction commits. That closes the rollback case but leaves a small window: a crash between commit and enqueue loses the job. Option B — queue tables in the primary database, no connects_to — makes the job insert part of the same transaction and closes both gaps, at the cost of putting queue write load on the primary. Choose B when job volume is modest (hundreds per second or less) and correctness of enqueue matters; choose A when volume is high or the primary is already busy.
Step 2 — Configure Dispatchers and Workers in queue.yml
bin/jobs starts a supervisor that forks three kinds of processes: dispatchers move scheduled jobs into the ready set when they become due; workers poll ready jobs and run them in a thread pool; a scheduler enqueues recurring tasks. Configure them per environment:
# config/queue.yml
production:
dispatchers:
- polling_interval: 1 # how often scheduled jobs are promoted to ready
batch_size: 500
concurrency_maintenance_interval: 300
workers:
- queues: [critical, default] # order = priority: critical drained first
threads: 5
processes: 2
polling_interval: 0.1 # seconds between polls when idle
- queues: [mailers]
threads: 3
processes: 1
- queues: [reports]
threads: 2 # heavy, memory-hungry jobs: few threads
processes: 1
polling_interval: 2
Listing several queues in one worker makes it drain them in order, so critical always empties before default is touched. Separate worker entries give queues isolated capacity: a surge of reports cannot occupy the threads that send password-reset emails. The general approach is covered in routing high-priority jobs, which applies the same idea in Celery.
Step 3 — Size Threads Against the Connection Pool
Each worker thread needs a database connection while it runs a job (for the job's own queries), and each Solid Queue process also uses connections for polling, heartbeats, and claiming. If the pool is smaller than threads plus overhead, threads block on checkout and time out.
# config/database.yml — pool large enough for the busiest process
production:
primary:
<<: *default
pool: <%= ENV.fetch("DB_POOL") { 12 } %> # >= max threads per process + 2
queue:
<<: *default
pool: <%= ENV.fetch("QUEUE_DB_POOL") { 12 } %>
Then check the total against the server's limit. The configuration above runs four worker processes with 5, 5, 3, and 2 threads, one dispatcher, and one scheduler per bin/jobs instance — roughly 25 connections per instance to each database. Three pods of it is 75 connections, before counting the web tier. That arithmetic is what caused the timeout storm in the problem statement: the autoscaler added pods faster than max_connections allowed. Cap worker pod count against the connection budget, or put a pooler in front — the same sizing discipline as in tuning the Sidekiq Redis connection pool, applied to Postgres.
Step 4 — Use Concurrency Controls Instead of External Locks
Solid Queue can limit how many jobs with the same key run at once, which replaces ad hoc Redis locks for "only one sync per account at a time".
class SyncAccountJob < ApplicationJob
queue_as :default
# At most one running sync per account; others wait (blocked) up to 15 minutes
limits_concurrency to: 1, key: ->(account) { account.id }, duration: 15.minutes
retry_on Net::ReadTimeout, wait: :polynomially_longer, attempts: 8
discard_on ActiveJob::DeserializationError
def perform(account)
AccountSync.new(account).run
end
end
Blocked jobs wait in solid_queue_blocked_executions and are released when the running job finishes. duration is a safety valve: if a worker dies without releasing its semaphore, it expires after that period. Set it longer than the job's worst-case runtime, or two syncs can overlap after an expiry.
Concurrency controls are also the cleanest way to protect a fragile downstream. A partner API that allows five concurrent requests becomes limits_concurrency to: 5, key: "partner-api" on every job that calls it — a fleet-wide cap enforced by the database, independent of how many worker pods or threads exist. Compared with a Redis-based semaphore, there is no separate TTL bookkeeping and no risk of the lock store and the job store disagreeing, because both are rows in the same database.
Two behaviours are worth knowing before relying on it. Blocked jobs do not consume worker threads — they sit in the blocked table and are promoted to ready when a slot frees — so a long queue of blocked jobs costs rows, not capacity. And the key is evaluated at enqueue time from the job's arguments, so it must be derivable from them; a key that depends on database state read inside perform will not work. For controlling rate rather than concurrency (requests per second instead of requests in flight), use the approaches in rate limiting and throttling jobs.
Step 5 — Define Recurring Tasks Once
Recurring tasks live in config/recurring.yml and are enqueued by the scheduler process. Solid Queue records each run with a unique key per task and time, so running bin/jobs on several pods does not enqueue duplicates.
# config/recurring.yml
production:
purge_finished_jobs:
command: "SolidQueue::Job.clear_finished_in_batches(sleep_between_batches: 0.3)"
schedule: every hour at minute 12
nightly_reconcile:
class: ReconcilePaymentsJob
queue: reports
schedule: at 2:30am every day # evaluated in the app's time zone
refresh_exchange_rates:
class: RefreshRatesJob
args: [ "ECB" ]
schedule: every 15 minutes
Schedules use Fugit syntax, and times follow config.time_zone. Daylight-saving transitions still apply to wall-clock schedules like 2:30am; the pitfalls are covered in timezone-safe job scheduling across DST.
Step 6 — Bound Table Growth
By default Solid Queue keeps finished jobs (preserve_finished_jobs = true) for clear_finished_jobs_after (one day). At a million jobs a day that is a million extra rows at steady state — fine — but if the cleanup task is missing, the table grows without limit, which is how it reached 40 million rows.
# config/application.rb
config.solid_queue.preserve_finished_jobs = true # keep for debugging...
config.solid_queue.clear_finished_jobs_after = 12.hours # ...but not for long
config.solid_queue.silence_polling = true # keep polling SQL out of the log
Keep the purge_finished_jobs recurring task from Step 5 — Solid Queue does not delete finished jobs on its own. On Postgres, tune autovacuum on the busiest tables (solid_queue_ready_executions, solid_queue_claimed_executions, solid_queue_jobs) as described in Database-Backed Job Queues.
Step 7 — Add Mission Control for Visibility
mission_control-jobs provides a web UI for queues, failed jobs, retries, and discards.
# Gemfile
gem "mission_control-jobs"
# config/routes.rb — mount behind your admin authentication
authenticate :user, ->(u) { u.admin? } do
mount MissionControl::Jobs::Engine, at: "/jobs"
end
Failed jobs stay in solid_queue_failed_executions with their error until retried or discarded from the UI; they are the dead-letter set. Export failure counts to your metrics system as well — a dashboard nobody opens does not page anyone.
Verification
# Processes registered and heartbeating
bin/rails runner 'puts SolidQueue::Process.pluck(:kind, :name, :last_heartbeat_at).inspect'
# Ready backlog and oldest ready job per queue
bin/rails runner '
SolidQueue::ReadyExecution.group(:queue_name).count.each { |q, n| puts "#{q}: #{n}" }
puts SolidQueue::ReadyExecution.minimum(:created_at)'
Then prove transactional behaviour: in a console, open a transaction, perform_later a job, raise to roll back, and confirm no job row exists (option B) or that the job was never inserted (option A with deferred enqueue).
Gotchas & Edge Cases
Running bin/jobs in the web process. The Puma plugin (plugin :solid_queue) is convenient for small apps but ties job capacity to web scaling and makes restarts disrupt both. Run a separate job deployment in production.
Forgetting the queue schema. Option A needs db/queue_schema.rb loaded into the queue database (bin/rails db:prepare handles it). Option B needs the queue tables migrated into the primary; convert the schema into a normal migration.
Long jobs and heartbeats. Worker processes heartbeat every 60 seconds by default; a process considered dead (process_alive_threshold, default 5 minutes) has its claimed jobs released. A thread blocked in native code can starve the heartbeat thread; keep jobs that hold the GVL for minutes out of the main worker pool.
MySQL and SKIP LOCKED. MySQL 8.0 is required; MySQL 5.7 and older MariaDB lack SKIP LOCKED and workers will serialize.
FAQ
Is Solid Queue fast enough to replace Sidekiq? For most applications, yes. It handles thousands of jobs per second on a well-provisioned database. Sidekiq on Redis remains faster at very high rates and has a deeper ecosystem of middleware; see Sidekiq Performance Tuning.
Can I keep using Sidekiq for some queues?
Yes. Active Job lets you set the adapter per job class with self.queue_adapter = :sidekiq, which is a practical way to migrate gradually.
Where do failed jobs go?
Into solid_queue_failed_executions after retry_on attempts are exhausted. Retry or discard them from Mission Control or with SolidQueue::FailedExecution#retry.
Related
- Database-Backed Job Queues — the design behind Solid Queue and its limits.
- Migrating from Redis to a Postgres Job Queue — moving an existing Sidekiq workload over.
- Sidekiq Quiet and Shutdown Timeouts — the Sidekiq shutdown model for comparison.
- Building a Postgres Job Queue with SKIP LOCKED — the mechanics Solid Queue implements.