SQS vs RabbitMQ for AWS Workloads
On AWS the default answer for a job queue is SQS, and it is usually the right one — but not always. This guide works through the decision for a concrete workload as part of Message Broker Comparison in Queue Fundamentals & Architecture, comparing SQS with RabbitMQ run either on EC2/EKS or as Amazon MQ, on semantics, features, latency, limits, cost, and operations.
Problem Statement
A company migrating a Celery application from its own data centre to AWS runs RabbitMQ today, with 40 queues, per-message priorities on two of them, topic-exchange routing for event fan-out, and about 1,500 messages per second at peak. The platform team wants to use SQS to eliminate broker operations; the application team worries about losing priorities, routing, and the low latency they rely on for user-facing jobs. You need a decision per workload — not a blanket one — backed by the specific features each queue uses, measured latency requirements, and a cost estimate at the real volume.
Prerequisites
- An inventory of queues with their features: priorities, routing keys and exchange types, TTLs, dead-lettering, message sizes, and consumer patterns.
- Latency requirements per queue (time from publish to start of processing).
- Peak and average message rates per queue.
- The framework constraints: Celery supports both brokers, with some features (like
task_routesto priority queues) behaving differently on SQS.
Step 1 — Map the Semantics Side by Side
Both deliver at least once with acknowledgement-based redelivery; the mechanics differ in ways that shape application code.
| Aspect | SQS (standard) | SQS FIFO | RabbitMQ |
|---|---|---|---|
| Delivery | At least once, occasional duplicates | Exactly-once processing window (5 min dedup) | At least once |
| Ordering | Best effort | Per message group | Per queue (single consumer) |
| Redelivery trigger | Visibility timeout expiry | Visibility timeout expiry | Channel close or nack |
| Consumer model | Poll (long polling up to 20 s) | Poll | Push with prefetch |
| Per-message priority | No (use separate queues) | No | Yes (x-max-priority) |
| Routing / fan-out | Via SNS or EventBridge | Via SNS FIFO | Exchanges: direct, topic, headers |
| Max message size | 256 KB | 256 KB | Configurable (default 128 MB) |
| Delayed delivery | Up to 15 min per message | Queue-level delay only | Plugin or TTL + DLX |
The redelivery trigger is the difference teams feel first. RabbitMQ redelivers the moment a consumer's connection drops; SQS waits until the visibility timeout expires, so a crashed worker delays its in-flight jobs by up to that timeout. Size it carefully — see the visibility timeout deep dive.
Step 2 — Classify Each Queue by the Features It Needs
Go through the inventory and mark each queue. Most turn out to use nothing SQS lacks.
# queue-inventory.yaml (excerpt)
- name: emails
rate_peak: 400/s
features: [] # plain work queue -> SQS standard
latency_need: 30s
- name: report_generation
rate_peak: 5/s
features: [long_jobs] # 20-minute jobs -> SQS with heartbeat, or RabbitMQ
latency_need: 5m
- name: user_actions
rate_peak: 300/s
features: [priority] # x-max-priority 10 -> split into 2-3 SQS queues
latency_need: 1s
- name: domain_events
rate_peak: 600/s
features: [topic_routing, fanout] # 12 consumers bind with patterns -> SNS + SQS per consumer
latency_need: 2s
Priorities on SQS become separate queues (user_actions_high, user_actions_low) with workers polling high first or with more workers on high. Topic routing becomes SNS topics with filter policies and one SQS queue per subscriber. Both translations work well; they change infrastructure, not application logic. The Celery routing side is covered in routing high-priority jobs in Celery.
Step 3 — Measure Latency Where It Matters
End-to-end latency differs mostly in the consumer model. RabbitMQ pushes messages into a prefetch buffer, so an idle consumer starts within a millisecond or two. SQS consumers long-poll; a waiting poll returns as soon as a message arrives, typically within tens of milliseconds, but each receive is an HTTPS call.
# latency_probe.py — publish with a timestamp, record consumer-side delay
import time, json, statistics
samples = []
def on_message(body):
sent = json.loads(body)["sent_at"]
samples.append((time.time() - sent) * 1000)
# ... run 10,000 messages at 50/s through each broker, then:
q = statistics.quantiles(samples, n=100)
print(f"p50={q[49]:.1f}ms p95={q[94]:.1f}ms p99={q[98]:.1f}ms")
Typical results in the same region: RabbitMQ p50 around 1–3 ms and p99 under 20 ms; SQS standard p50 around 15–30 ms and p99 around 100–200 ms. For queues whose latency need is seconds or more, the difference is irrelevant. For interactive jobs where a user waits on a spinner, 200 ms at p99 may matter.
Step 4 — Estimate Cost at Your Volume
SQS charges per request; RabbitMQ costs instances (and engineer time). At 1,500 messages per second peak and roughly 50 million messages per month:
# Rough monthly estimate, us-east-1 list prices (verify current pricing)
msgs = 50_000_000
# SQS standard: send + receive + delete per message, batched by 10 on each call
requests = msgs * 3 / 10
sqs_cost = max(0, requests - 1_000_000) / 1_000_000 * 0.40 # ~ $5.60
# plus SNS for fan-out queues: 12 subscribers on domain_events (~15M msgs) -> deliveries to SQS are free,
# publishes ~15M -> ~$7.50
# Amazon MQ for RabbitMQ: 3-node cluster, mq.m5.large
amazon_mq_cost = 3 * 0.288 * 730 + 3 * 200 * 0.10 # instances + storage ~ $690
print(f"SQS+SNS ~ ${sqs_cost + 7.5:.0f}/month; Amazon MQ ~ ${amazon_mq_cost:.0f}/month")
At this volume SQS is dramatically cheaper and needs no capacity planning. The comparison narrows only at very high sustained volume (billions of messages per month) or when batching is impossible. Include operations in the comparison: a self-managed RabbitMQ cluster on EKS costs less in instances than Amazon MQ but more in on-call time, upgrades, and incident risk.
Step 5 — Check the Limits That Bite
Each option has limits that turn into incidents if discovered late:
- SQS message size (256 KB). Larger payloads need the claim-check pattern — store the body in S3 and send a reference, as in claim-check pattern for large payloads. Celery task arguments occasionally exceed this.
- SQS in-flight limits. 120,000 in-flight messages per standard queue, 20,000 per FIFO queue. Aggressive prefetching by many consumers can hit it.
- SQS delay (15 minutes). Celery
countdown/etabeyond 15 minutes on SQS is handled by Celery re-publishing, which works but adds churn; long schedules belong in a scheduler. - RabbitMQ memory and disk alarms. When a node crosses its memory high watermark or disk free limit, it blocks all publishers. Queue depth must be bounded or monitored closely.
- Amazon MQ instance limits. Connection and channel limits per broker size; large worker fleets with many channels can exhaust smaller sizes.
Step 6 — Decide per Queue, Then Simplify
For the scenario, the classification leads to a mostly-SQS design with one exception:
emails -> SQS standard (+ DLQ)
report_generation -> SQS standard with visibility heartbeat
user_actions -> SQS standard x2 (high/low), workers poll high first
domain_events -> SNS topic + 12 SQS queues with filter policies
realtime_ui_jobs -> RabbitMQ (Amazon MQ, small) — p99 latency need 50 ms, user waiting
Running two brokers has a cost of its own. If only one small queue needs RabbitMQ, consider whether that workload could avoid a queue altogether (a synchronous call with a timeout) before keeping a broker for it. The single-broker design is simpler to operate; the split design is justified only when a real requirement forces it.
Verification
Before migrating, run the production job mix through SQS in staging and compare: redelivery rate, p99 latency per queue, DLQ arrivals, and the behaviour when a worker is killed mid-job. After migrating each queue, watch ApproximateAgeOfOldestMessage and NumberOfMessagesReceived / NumberOfMessagesDeleted (a ratio well above 1 means redeliveries, often from a too-short visibility timeout).
aws cloudwatch get-metric-statistics --namespace AWS/SQS \
--metric-name ApproximateAgeOfOldestMessage --dimensions Name=QueueName,Value=emails \
--statistics Maximum --period 300 --start-time "$(date -u -d '-1 day' +%FT%TZ)" --end-time "$(date -u +%FT%TZ)"
Gotchas & Edge Cases
Celery on SQS lacks some features. No remote control (celery inspect broadcast), no events for Flower by default, and polling rather than push. Plan monitoring accordingly — metrics come from CloudWatch and your own instrumentation.
FIFO throughput. FIFO queues have lower per-queue limits unless high-throughput mode is enabled; ordering requirements do not come free.
Cross-region. SQS queues are regional; RabbitMQ federation or shovels can bridge regions. Multi-region job processing needs a design either way.
Hidden costs of empty polls. Consumers with short polling (WaitTimeSeconds=0) on idle queues generate enormous request counts. Always long-poll.
FAQ
Is Amazon MQ a good middle ground? It removes patching and failover work while keeping RabbitMQ semantics, at a significant instance cost and with broker-size limits. It fits when RabbitMQ features are genuinely needed and the team does not want to run the cluster.
Can Celery use SQS in production? Yes, widely. Configure long polling, a visibility timeout above your longest task, and predefined queues (to avoid Celery creating queues at runtime), and monitor through CloudWatch.
What about Amazon MQ for ActiveMQ? It is a different protocol family (JMS/OpenWire, AMQP 1.0). For Python and Ruby job frameworks, RabbitMQ or SQS are the practical choices.
Related
- Message Broker Comparison — the full broker landscape.
- How to Choose Between RabbitMQ and Redis for Async Tasks — the self-hosted comparison.
- Managed Cloud Queues for Background Jobs — SQS alongside Google and Azure services.
- Scaling Queue Partitions in AWS SQS — throughput beyond one queue.