Managed Cloud Queues for Background Jobs

Managed queues trade control for the absence of servers to patch, and this guide covers using them as the backbone of background processing, as part of Backend Frameworks & Worker Scaling. AWS SQS, Google Cloud Tasks and Pub/Sub, and Azure Service Bus all deliver messages at least once, scale without capacity planning, and bill per request — but they differ sharply in how work reaches your code (you pull it, or they push it to an HTTP endpoint or function), in what concurrency controls they give you, and in how failures surface.

The appeal is operational: no broker cluster, no failover drills, no disk alarms at 3 a.m. The cost is that the queue's semantics are fixed by the provider. You cannot tune a Redis eviction policy or patch RabbitMQ's requeue behaviour; you design around what the service offers. This guide is about knowing those semantics well enough to design around them on purpose.

The Scenario: A Worker Fleet Nobody Wants to Run

A media company runs RabbitMQ on three VMs for background jobs: transcoding requests, notification fan-out, and partner webhooks. The cluster works, but it has needed two emergency interventions in six months (a disk-full alarm that blocked publishers, and a network partition that split the cluster), and the only engineer who understood its configuration has moved teams. Leadership asks whether moving to the cloud provider's managed queues would remove the operational burden — and what it would cost in money, latency, and control.

The honest answer depends on the workload. Transcoding requests (long jobs, bursty, tolerant of seconds of latency) fit a managed queue with container workers. Partner webhooks (rate-limited per partner, need controlled retry) fit a push-style task service with per-queue rate limits. Notification fan-out (one event, several consumers) fits pub/sub. One RabbitMQ cluster was serving three different shapes of work, and the managed world has a service for each.

One broker, three managed shapes A single RabbitMQ cluster carried three workloads. In the managed design, transcoding requests go to a pull queue such as SQS consumed by autoscaled container workers; partner webhooks go to a push task queue such as Cloud Tasks that calls an HTTP handler at a controlled rate; notification events go to a pub/sub topic with a subscription per consuming service. Match each workload to a service shape transcode requests partner webhooks notification events pull queue (SQS) push tasks (Cloud Tasks) pub/sub topic container workers HTTP handler, rate cap subscription per team

Architectural Overview: Pull, Push, and Event-Source Delivery

Managed services deliver work in three ways, and the delivery model shapes everything else — concurrency control, retries, and how you scale.

Pull. Your workers call the service to receive messages, process them, and delete them. SQS, Service Bus receivers, and Pub/Sub pull subscriptions work this way. You control concurrency completely (it is your worker count), and scaling means scaling your fleet — usually on queue depth, as in scaling workers with KEDA on queue length.

Push. The service calls your HTTP endpoint with each task and treats a 2xx response as success. Cloud Tasks and Pub/Sub push subscriptions work this way. The service controls the dispatch rate and concurrency (you configure it), retries on non-2xx, and your handler scales like any web service.

Event-source mapping. A serverless platform polls the queue on your behalf and invokes a function with batches. Lambda with SQS, Azure Functions with Service Bus, and Cloud Run with Eventarc triggers work this way. You write only the handler; concurrency is bounded by function scaling limits and a per-source maximum.

# The same job, three delivery models (conceptual)
pull:
  who_polls: your worker process
  concurrency_control: worker count × threads
  success_signal: DeleteMessage / Complete
push:
  who_polls: the provider
  concurrency_control: queue config (max dispatches/s, max concurrent)
  success_signal: HTTP 2xx within the deadline
event_source:
  who_polls: the serverless runtime
  concurrency_control: function concurrency + source max concurrency
  success_signal: function returns without error (or partial batch response)

The success signal row is where most production bugs live. A push handler that returns 200 before finishing its work loses the task on a crash. A Lambda that throws on one bad record fails the whole batch unless partial-batch responses are enabled — the subject of handling partial batch failures in Lambda with SQS.

Who initiates delivery? In pull delivery, workers call receive on the queue and delete messages after processing. In push delivery, the queue service sends an HTTP request to the handler and retries on non-2xx responses. In event-source mapping, a managed poller reads batches from the queue and invokes a function, deleting the batch when the function succeeds. Three delivery models pull worker: receive, delete you own concurrency SQS, Service Bus push service: POST /task provider owns rate Cloud Tasks, Pub/Sub push event source runtime invokes fn(batch) runtime owns scaling Lambda + SQS, Functions

Implementation 1: SQS with Autoscaled Container Workers

For long or resource-heavy jobs — transcoding, PDF rendering, ML inference — a pull queue consumed by containers gives the most control. The worker loop is short; the important parts are long polling, a visibility timeout that covers the job, and deleting only after success.

# worker.py — SQS consumer for long jobs
import json, signal, boto3

sqs = boto3.client("sqs")
QUEUE_URL = "https://sqs.eu-west-1.amazonaws.com/123456789012/transcode"
running = True
signal.signal(signal.SIGTERM, lambda *_: globals().__setitem__("running", False))

while running:
    resp = sqs.receive_message(
        QueueUrl=QUEUE_URL,
        MaxNumberOfMessages=1,          # long jobs: one at a time per worker
        WaitTimeSeconds=20,             # long polling: fewer empty receives, lower cost
        VisibilityTimeout=900,          # 15 min; extend with a heartbeat for longer jobs
        MessageAttributeNames=["All"],
    )
    for msg in resp.get("Messages", []):
        job = json.loads(msg["Body"])
        transcode(job)                  # idempotent: writes deterministic output keys
        sqs.delete_message(QueueUrl=QUEUE_URL, ReceiptHandle=msg["ReceiptHandle"])

Scale the worker deployment on ApproximateNumberOfMessagesVisible divided by a target backlog per worker, with a minimum of one or zero replicas. Dead-lettering is configured on the queue with a redrive policy, as in configuring an SQS redrive policy. For short jobs, the same queue can instead be consumed by Lambda, which removes the worker fleet entirely — processing SQS with AWS Lambda covers batch sizes, concurrency caps, and timeouts.

Implementation 2: Push-Based Tasks with Rate Limits

For outbound calls to rate-limited partners, a push task queue is often the simplest correct design, because the rate limit is enforced by the service before your code runs. Google Cloud Tasks lets you set dispatch rate and concurrency per queue and schedules retries with backoff on failure.

# One queue per partner, each with its own rate and retry policy
gcloud tasks queues create webhooks-acme \
  --location=europe-west1 \
  --max-dispatches-per-second=5 \       # partner allows 5 req/s
  --max-concurrent-dispatches=10 \
  --max-attempts=12 \
  --min-backoff=5s --max-backoff=1h \
  --max-doublings=6
# enqueue.py — create an HTTP task targeting our own handler
from google.cloud import tasks_v2
import json

client = tasks_v2.CloudTasksClient()
parent = client.queue_path("my-project", "europe-west1", "webhooks-acme")

def enqueue_webhook(event_id: str, payload: dict) -> None:
    client.create_task(parent=parent, task={
        "name": f"{parent}/tasks/{event_id}",       # named task: duplicate creates are rejected
        "http_request": {
            "http_method": tasks_v2.HttpMethod.POST,
            "url": "https://hooks.internal.example.com/deliver/acme",
            "headers": {"Content-Type": "application/json"},
            "body": json.dumps(payload).encode(),
            "oidc_token": {"service_account_email": "tasks-invoker@my-project.iam.gserviceaccount.com"},
        },
        "dispatch_deadline": {"seconds": 60},       # handler must answer within this
    })

Named tasks give deduplication: creating a task with a name that exists (or existed within about an hour) fails, so a retried enqueue does not double-deliver. Google Cloud Tasks for HTTP workers covers the handler side, authentication, and the deadline semantics. Azure Service Bus offers a different set of guarantees — sessions for ordered processing and built-in dead-letter subqueues — detailed in Azure Service Bus sessions and dead-lettering.

Identity and Access: Who May Enqueue, Who May Consume

Self-hosted brokers usually rely on network isolation and a shared password. Managed queues put every call behind the provider's identity system, which is a real security improvement if you use it deliberately — and a source of confusing failures if you do not.

Apply least privilege in both directions. Producers get permission to send to specific queues and nothing else; consumers get receive, delete, and visibility-change on their queue and nothing else. Separate roles per service mean a compromised API pod cannot drain or purge the queue, and a compromised worker cannot inject messages into another team's queue.

# Terraform: producer may only send; consumer may only receive/delete/extend
data "aws_iam_policy_document" "producer" {
  statement {
    actions   = ["sqs:SendMessage", "sqs:GetQueueUrl"]
    resources = [aws_sqs_queue.transcode.arn]
  }
}

data "aws_iam_policy_document" "consumer" {
  statement {
    actions = ["sqs:ReceiveMessage", "sqs:DeleteMessage",
               "sqs:ChangeMessageVisibility", "sqs:GetQueueAttributes"]
    resources = [aws_sqs_queue.transcode.arn]
  }
  statement {                                    # if the queue is KMS-encrypted
    actions   = ["kms:Decrypt"]
    resources = [aws_kms_key.queues.arn]
  }
}

Push delivery reverses the direction of trust: the provider calls you, so your handler must verify that the request really came from the queue. Cloud Tasks and Pub/Sub push attach an OIDC token signed by Google for a service account you choose; the handler (or the platform in front of it, such as Cloud Run with authentication required) verifies the token's issuer and audience before doing any work. An unauthenticated task endpoint is an open door for anyone who can guess the URL to trigger your background work.

Encryption at rest is on by default for all three providers; customer-managed keys add a second permission (like the kms:Decrypt above) that consumers need, which is the most common cause of "the worker suddenly gets access denied" after a security hardening change.

Separate roles in both directions The API service's role may only send messages to the queue. The worker's role may only receive, delete, and change visibility on that queue, plus decrypt with the queue's key. For push delivery, the queue service calls the handler with a signed OIDC token, and the handler verifies the issuer and audience before processing. Least privilege around one queue API role SendMessage only queue (KMS key) worker role receive, delete, decrypt push: handler verifies OIDC issuer + audience first

Observability Across Providers

Each provider exposes queue metrics in its own monitoring system: CloudWatch for SQS, Cloud Monitoring for Cloud Tasks and Pub/Sub, Azure Monitor for Service Bus. The names differ but the questions do not, and mapping each provider's metric onto the same four signals keeps dashboards and alerts consistent across a multi-cloud estate.

Signal SQS (CloudWatch) Pub/Sub (Cloud Monitoring) Service Bus (Azure Monitor)
Backlog ApproximateNumberOfMessagesVisible subscription/num_undelivered_messages ActiveMessages
Queue time ApproximateAgeOfOldestMessage subscription/oldest_unacked_message_age (derive from enqueue timestamps)
In flight ApproximateNumberOfMessagesNotVisible subscription/num_outstanding_messages LockedMessages (via SDK)
Dead letters DLQ ApproximateNumberOfMessagesVisible dead-letter topic backlog DeadletteredMessages

Queue time is the signal to alert on first; it is what users experience and what an SLO should be defined on, as described in defining SLOs for job latency. Pull all four into one place — the Prometheus CloudWatch exporter, Stackdriver exporter, or Grafana's native data sources for each cloud — so a single dashboard answers "is any queue unhealthy?" regardless of provider. Visualizing SQS metrics in Grafana walks through the AWS side.

Trade-off Analysis

Capability AWS SQS Google Cloud Tasks Google Pub/Sub Azure Service Bus
Delivery model Pull (or Lambda) Push to HTTP Push or pull Pull (or Functions)
Ordering FIFO queues per group None Ordering keys Sessions
Per-queue rate limiting No (consumer-side) Yes, native Flow control per subscriber No (consumer-side)
Scheduling a single message Up to 15 min delay Arbitrary schedule time (up to 30 days) No Scheduled enqueue time
Dead-lettering Redrive policy Max attempts, then dropped (log it) Dead-letter topic Built-in DLQ subqueue
Max message size 256 KB (1 MB extended by some SDKs via S3) 1 MB task 10 MB 256 KB standard, 100 MB premium
Fan-out to many consumers Via SNS No Native Topics and subscriptions

Two rows deserve emphasis. Cloud Tasks drops a task after its final attempt with no dead-letter destination, so the handler must record permanent failures itself (in a database or a log-based alert). And SQS's 15-minute delay limit means longer delays need a scheduler elsewhere — a database row or EventBridge Scheduler — rather than the queue itself. The broader broker comparison is in SQS vs RabbitMQ for AWS workloads.

Failure Modes & Recovery

The handler that acknowledges too early. A push handler returns 202 immediately and does the work in a background thread. If the instance is scaled in or crashes, the work is lost and the service will not retry — it saw success. Remediation: do the work synchronously within the dispatch deadline, or record it durably before returning 2xx.

Retry storms against a recovering dependency. When a downstream outage ends, every queued message is delivered at the service's maximum rate. Push queues with rate caps absorb this; pull consumers need their own limits. Remediation: cap concurrency at the consumer or the queue, and see preventing retry storms after an outage.

Visibility or ack deadlines shorter than the job. SQS visibility timeouts and Pub/Sub ack deadlines both redeliver work that takes longer than configured. Remediation: set them from p99 job duration and extend them with heartbeats for long jobs.

Poison messages in event-source batches. A single malformed record makes a Lambda batch fail repeatedly, and every other record in the batch is retried with it until the poison one reaches the DLQ. Remediation: partial-batch responses, so only the failing record is retried.

Draining a backlog after an outage After a thirty-minute downstream outage, 40,000 messages are waiting. An uncapped consumer fleet that autoscaled on backlog sends thousands of requests per second the moment the dependency returns, knocking it over again. A push queue capped at 50 dispatches per second drains the backlog in about fourteen minutes without exceeding the dependency's capacity. 40,000 queued messages when the dependency recovers uncapped fleet 3,000 req/s spike, dependency falls over again capped at 50/s steady drain over ~14 minutes Autoscaling on backlog makes the storm worse; the cap belongs next to the dependency.

Performance Tuning and Cost

Managed queues bill per request, so throughput tuning and cost tuning are the same exercise. The main knobs:

  • Long polling. An SQS consumer with WaitTimeSeconds=20 makes at most three receive calls a minute when idle instead of hundreds. It is the single largest cost lever for idle-heavy queues.
  • Batching. SQS SendMessageBatch, ReceiveMessage with MaxNumberOfMessages=10, and DeleteMessageBatch cut request count up to tenfold. Lambda event-source batch size and a batching window do the same for functions.
  • Payload size. SQS bills in 64 KB chunks; a 200 KB message costs four requests. Keep messages small and put payloads in object storage — the claim-check pattern.
  • Concurrency caps. Lambda's ScalingConfig.MaximumConcurrency on the event source caps how many concurrent invocations one queue can drive, protecting downstream databases from a burst.
  • Worker right-sizing. For container workers, match CPU and memory to the job and scale on backlog per worker, not CPU.
Rough monthly cost, 50M jobs/month on SQS standard (us-east-1 list prices):
  requests per job: send 1 + receive 1 + delete 1 = 3  (unbatched)
  150M requests  -> ~$60/month
  batched by 10 on send/delete, receive 10 at a time -> ~15M requests -> ~$6/month
  (Compute for the workers dominates either way.)

The request bill is rarely what matters; the compute that processes the messages is. Put the tuning effort where the money is — worker efficiency and idle capacity.

Moving a Workload off a Self-Hosted Broker

For the media company in the scenario, the move from RabbitMQ is a sequence of small, reversible steps rather than one cutover:

  1. Classify each queue by shape — pull, push, or fan-out — using the delivery-model comparison above. Queues that rely on RabbitMQ-specific features (per-message priority, header exchanges, x-single-active-consumer) need a design decision, not a translation.
  2. Create the managed queues and their dead-letter destinations in infrastructure code, with least-privilege roles, before any application change.
  3. Wrap enqueue behind an interface (publish(queue, payload)) if it is not already, so one flag decides which transport a producer uses.
  4. Deploy consumers for the new queue alongside the old ones. They idle until traffic arrives.
  5. Flip producers per queue, lowest-risk first, and let the RabbitMQ backlog drain on the existing consumers — the same dual-run approach as in migrating from Redis to a Postgres job queue.
  6. Move alerts to the managed queue's metrics (backlog, age, DLQ depth) before the old queue goes quiet, so there is no window without monitoring.
  7. Decommission the RabbitMQ queue once its enqueue rate has been zero for a full business cycle.

Each step is independently reversible until the last, which is what makes the migration safe to run during normal working hours.

FAQ

Is a managed queue always cheaper than running RabbitMQ or Redis? In direct cost, usually yes at low to moderate volume, and it removes on-call load. At very high sustained volume (billions of messages a month), self-hosted brokers can be cheaper in raw spend, but the comparison must include engineer time for operations.

Can I get exactly-once processing from a managed queue? No provider guarantees end-to-end exactly-once for your side effects. SQS FIFO and Pub/Sub exactly-once delivery reduce duplicates at the transport, but a crash after your side effect and before acknowledgement still causes redelivery. Keep handlers idempotent.

Push or pull for background jobs? Push when you want the provider to enforce rate limits and your handlers already run as HTTP services; pull when jobs are long, resource-heavy, or need fine-grained control over concurrency and batching. Cloud Pub/Sub push vs pull for workers compares them on one service.

How do I avoid vendor lock-in? Keep a thin interface in your code (enqueue, handle, acknowledge) and keep job payloads provider-neutral. The provider-specific parts — redrive policies, IAM, rate configuration — live in infrastructure code, which is where they belong anyway.

Related