Cloud Pub/Sub Push vs Pull for Workers

Google Cloud Pub/Sub can deliver the same messages to your workers in two ways — pushing them to an HTTP endpoint or letting workers pull them — and the choice changes how you control concurrency, handle slow jobs, and scale. This guide compares the two on a concrete workload as part of Managed Cloud Queues for Background Jobs in Backend Frameworks & Worker Scaling, then configures each correctly: ack deadlines, flow control, dead-letter topics, retry policy, and exactly-once delivery.

Problem Statement

A logistics company publishes ShipmentScanned events to a Pub/Sub topic. Three teams consume them: a tracking service updates customer-visible status (fast, ~50 ms per event), a billing service computes surcharges (moderate, ~300 ms), and an analytics job enriches events with route data (slow, up to 20 seconds, calls an external mapping API). All three started with push subscriptions to Cloud Run because it needed no worker code. Tracking works perfectly. Billing occasionally double-charges. Analytics times out constantly, and its retries hammer the mapping API. Each team needs to know whether push or pull fits its workload, and how to configure it.

Prerequisites

  • A Pub/Sub topic and permission to create subscriptions on it.
  • For push: an HTTPS endpoint (Cloud Run, GKE, or App Engine) and a service account for Pub/Sub to authenticate as.
  • For pull: a long-running worker (GKE, Compute Engine, or Cloud Run jobs/worker pools) using the google-cloud-pubsub client with streaming pull.
  • A dead-letter topic per subscription, and the Pub/Sub service agent granted publisher on it and subscriber on the source subscription.

Step 1 — Understand How Each Model Delivers

A subscription is the unit of delivery: each subscription gets its own copy of every message published to the topic, so the three teams each create their own. Within a subscription, messages are distributed to consumers, and each must be acknowledged within the ack deadline or it is redelivered.

Push: Pub/Sub sends an HTTPS POST per message. A 2xx response is an ack; anything else, or no response within the ack deadline, is a nack and triggers redelivery after the retry policy's backoff. Pub/Sub adjusts the push rate itself, slowing down when your endpoint returns errors ("push backoff") and speeding up when it succeeds.

Pull (streaming): your worker holds a long-lived gRPC stream; Pub/Sub sends messages down it up to your client's flow-control limits, and your code acks each one explicitly. The client library extends ack deadlines automatically while messages are being processed.

Three subscriptions, two delivery models The ShipmentScanned topic fans out to three subscriptions, each receiving every message. The tracking subscription pushes each message to a Cloud Run endpoint. The billing and analytics subscriptions are consumed by worker processes over streaming pull, with client-side flow control and automatic ack-deadline extension. One topic, one subscription per consumer topic ShipmentScanned sub: tracking (push) sub: billing (pull) sub: analytics (pull) POST to Cloud Run workers: stream + ack workers: max 8 in flight

Step 2 — Match the Workload to the Model

The three teams' results follow directly from the delivery mechanics:

Workload Push fits? Why
Tracking: 50 ms, stateless, high rate Yes Short handlers finish well inside the ack deadline; Cloud Run autoscaling tracks the push rate
Billing: 300 ms, must not double-charge Either, with care Duplicates came from non-idempotent code, not the model — fix with exactly-once delivery or idempotency (Step 5)
Analytics: up to 20 s, rate-limited external API Pull Needs a hard cap on in-flight work to respect the API's limit; push concurrency is governed by Pub/Sub and Cloud Run scaling, not by you

The general rule: push when handlers are short and you are happy to let the platform decide concurrency; pull when you need to cap concurrency precisely, process long jobs, or batch acknowledgements for throughput.

Step 3 — Configure a Push Subscription Correctly

For tracking, push is right; the configuration just needs an ack deadline above the handler's worst case, authentication, a dead-letter topic, and a retry policy with backoff.

gcloud pubsub subscriptions create tracking-push \
  --topic=shipment-scanned \
  --push-endpoint=https://tracking-abc123-ew.a.run.app/pubsub \
  --push-auth-service-account=pubsub-pusher@my-project.iam.gserviceaccount.com \
  --ack-deadline=30 \                        # seconds Pub/Sub waits for the HTTP response
  --min-retry-delay=5s --max-retry-delay=300s \
  --dead-letter-topic=shipment-scanned-dlq \
  --max-delivery-attempts=10
# Cloud Run handler: the body wraps the message; 2xx acks, anything else nacks
import base64, json
from fastapi import FastAPI, Request, Response

app = FastAPI()

@app.post("/pubsub")
async def receive(request: Request) -> Response:
    envelope = await request.json()
    msg = envelope["message"]
    event = json.loads(base64.b64decode(msg["data"]))
    attempt = int(envelope.get("deliveryAttempt", 1))       # present when a DLQ is configured
    try:
        await update_tracking(event)                       # idempotent upsert keyed by scan id
    except TransientError:
        return Response(status_code=503)                   # nack: retried with backoff
    return Response(status_code=204)                       # ack

With Cloud Run, also cap --max-instances and --concurrency so a burst cannot scale the service past what its database can take. That is the push-model equivalent of the concurrency cap in processing SQS with AWS Lambda.

Who controls how much runs at once With push delivery, the number of messages in flight is the product of Pub/Sub's adaptive push rate and Cloud Run's instance count and per-instance concurrency; you bound it only indirectly through max instances and concurrency. With pull delivery, the client's flow control settings set an exact maximum of outstanding messages per worker. Concurrency: indirect vs exact push adaptive push rate x max instances x concurrency bounded, but not exact pull FlowControl(max_messages=8) x worker count exact ceiling for a rate-limited API Analytics needs the exact ceiling: 3 workers x 8 = 24 concurrent mapping calls, never more.

Step 4 — Configure Streaming Pull with Flow Control

For analytics, the worker pulls with flow control so that at most eight messages are outstanding per worker process. The client library holds leases on those messages and extends their ack deadlines up to a maximum while they are processed.

# analytics_worker.py
from concurrent.futures import ThreadPoolExecutor
from google.cloud import pubsub_v1

subscriber = pubsub_v1.SubscriberClient()
path = subscriber.subscription_path("my-project", "analytics-pull")

def callback(message: pubsub_v1.subscriber.message.Message) -> None:
    try:
        enrich_with_route(json.loads(message.data))        # up to ~20 s, calls mapping API
        message.ack()
    except RateLimited:
        message.nack()                                     # redelivered after retry-policy backoff
    except Exception:
        message.nack()

flow = pubsub_v1.types.FlowControl(
    max_messages=8,                                        # hard cap on in-flight per process
    max_bytes=16 * 1024 * 1024,
    max_lease_duration=600,                                # stop extending after 10 minutes
)
future = subscriber.subscribe(path, callback=callback, flow_control=flow,
                              scheduler=pubsub_v1.subscriber.scheduler.ThreadScheduler(
                                  ThreadPoolExecutor(max_workers=8)))
future.result()                                            # block; cancel() on SIGTERM

Three worker replicas give an exact ceiling of 24 concurrent mapping calls, which the team sizes to the API's limit. Scale replicas on subscription/num_undelivered_messages with a maximum replica count derived from the same limit. The general rate-limiting approach is in rate limiting third-party API calls from workers.

Step 5 — Stop Double Charges with Exactly-Once Delivery

Pub/Sub's default is at-least-once: a message can be redelivered even after a successful ack if the ack is lost or arrives after the deadline. For billing, enable exactly-once delivery on a pull subscription. With it, Pub/Sub will not redeliver a message once an ack succeeds, and the client reports whether each ack actually succeeded.

gcloud pubsub subscriptions create billing-pull \
  --topic=shipment-scanned \
  --enable-exactly-once-delivery \
  --ack-deadline=60 \
  --dead-letter-topic=shipment-scanned-dlq --max-delivery-attempts=10
from google.cloud.pubsub_v1.subscriber import exceptions as sub_exceptions

def billing_callback(message) -> None:
    with db.transaction():
        if already_billed(message.message_id):            # still keep a cheap guard
            message.ack()
            return
        compute_and_record_surcharge(json.loads(message.data), message.message_id)
    try:
        message.ack_with_response().result()               # confirms the ack was accepted
    except sub_exceptions.AcknowledgeError as e:
        log.warning("ack failed; message may be redelivered", extra={"code": e.error_code})

Exactly-once delivery covers the transport — it prevents Pub/Sub redelivering an acknowledged message within the subscription. It does not cover a crash after compute_and_record_surcharge commits but before the ack, so keep the idempotency check keyed by message id, as described in idempotent consumers with Postgres unique constraints. Exactly-once delivery is available for pull subscriptions only and adds some latency.

Two duplicate paths, two fixes Path one: the handler commits and acks, but the ack is lost or late, and Pub/Sub redelivers. Exactly-once delivery closes this path by confirming acks. Path two: the handler commits and the process crashes before acking, so Pub/Sub correctly redelivers. Only an idempotency check keyed by message id makes the second delivery harmless. Why billing charged twice commit, ack lost or late fix: exactly-once delivery commit, crash before ack fix: idempotency by message_id Billing needs both; tracking's idempotent upsert needs neither.

Step 6 — Wire Dead-Letter Topics and Watch Them

A dead-letter topic receives messages that exceeded max-delivery-attempts (between 5 and 100). Unlike SQS, it is a topic, so you need a subscription on it to keep and inspect the messages.

PROJECT_NUMBER=$(gcloud projects describe my-project --format='value(projectNumber)')
SA="service-${PROJECT_NUMBER}@gcp-sa-pubsub.iam.gserviceaccount.com"
gcloud pubsub topics add-iam-policy-binding shipment-scanned-dlq --member="serviceAccount:$SA" --role=roles/pubsub.publisher
gcloud pubsub subscriptions add-iam-policy-binding analytics-pull --member="serviceAccount:$SA" --role=roles/pubsub.subscriber
gcloud pubsub subscriptions create shipment-scanned-dlq-hold --topic=shipment-scanned-dlq \
  --message-retention-duration=7d

The IAM bindings are the step most often missed: without them, messages exceed their attempts and are not forwarded, and the subscription keeps retrying. Alert on subscription/num_undelivered_messages for the hold subscription, in line with alerting on dead-letter queue growth.

Verification

# Oldest unacked message age per subscription should stay near the handler's duration
gcloud monitoring time-series list \
  --filter='metric.type="pubsub.googleapis.com/subscription/oldest_unacked_message_age"' \
  --interval-start-time="$(date -u -d '-15 min' +%FT%TZ)"

# Publish a poison message and confirm it lands on the DLQ hold subscription
gcloud pubsub topics publish shipment-scanned --message='{"broken": true}'
gcloud pubsub subscriptions pull shipment-scanned-dlq-hold --limit=1 --auto-ack

For analytics, check the mapping API's request metrics: concurrent requests should never exceed replicas × max_messages.

Gotchas & Edge Cases

Ack deadline vs handler time on push. A push handler slower than the ack deadline is nacked even if it succeeds, then retried — duplicate work plus wasted API calls. This was analytics' original failure.

Push endpoints and cold starts. A push burst to a scaled-to-zero Cloud Run service returns errors during cold starts, triggering push backoff and retries. Keep a minimum instance count for latency-sensitive push subscriptions.

Ordering keys reduce parallelism. Enabling message ordering delivers messages with the same ordering key in order, one at a time; a failing message blocks its key until it succeeds or is dead-lettered.

Subscription expiration. Subscriptions with no activity expire after 31 days by default. Set --expiration-period=never for subscriptions that may be idle but must keep existing.

FAQ

Is pull more expensive than push? Pricing is by data volume, not delivery model, so direct cost is similar. Pull needs always-on workers, while push can scale to zero; that compute cost is usually the real difference.

Can one subscription use both push and pull? A subscription has one delivery type at a time, though you can change it. For different consumers, create separate subscriptions.

How does Pub/Sub compare with Cloud Tasks for background jobs? Pub/Sub fans one event out to many subscribers; Cloud Tasks targets one handler per task with explicit rate limits and scheduling. See Google Cloud Tasks for HTTP workers.

Related