Azure Service Bus Sessions and Dead-Lettering

Azure Service Bus is the most feature-rich of the managed queues — sessions, a built-in dead-letter subqueue, scheduled messages, and transactions — and this guide covers the features that matter for background jobs, as part of Managed Cloud Queues for Background Jobs in Backend Frameworks & Worker Scaling. The focus is on the two that most teams need and most often misconfigure: sessions for ordered per-entity processing, and the dead-letter subqueue for poison messages.

Problem Statement

An insurance platform processes policy change events — endorsements, cancellations, reinstatements — through a Service Bus queue consumed by an Azure Container Apps worker. Two problems keep recurring. Events for the same policy are sometimes applied out of order when several worker replicas receive them concurrently, producing a cancelled policy that is also marked active. And malformed events retry indefinitely until someone notices the queue is not draining. You want events for each policy processed strictly in order while different policies proceed in parallel, poison events quarantined automatically after a few attempts, and a safe procedure to inspect and resubmit them.

Prerequisites

  • A Service Bus namespace (Standard tier supports sessions and dead-lettering; Premium adds larger messages and isolation).
  • The azure-servicebus Python SDK (v7+) or the equivalent .NET/Java/JS SDK.
  • Managed identity for the worker with the Azure Service Bus Data Receiver role on the queue, and Data Sender for any process that resubmits.
  • An entity id (policy number) on every message, used as the session id.

Step 1 — Create a Session-Enabled Queue with a Delivery Limit

Sessions must be enabled when the queue is created; the setting cannot be changed later. MaxDeliveryCount sets how many delivery attempts a message gets before Service Bus moves it to the queue's dead-letter subqueue.

az servicebus queue create \
  --resource-group insurance-rg --namespace-name insurance-sb \
  --name policy-events \
  --enable-session true \               # FIFO per session id
  --max-delivery-count 6 \              # 6 attempts, then dead-letter
  --lock-duration PT1M \                # peek-lock duration; renew for longer work
  --default-message-time-to-live P14D \
  --enable-dead-lettering-on-message-expiration true   # expired messages go to DLQ, not void

The dead-letter subqueue exists automatically at policy-events/$DeadLetterQueue; there is nothing to create or wire up. That is a real difference from SQS, where the DLQ is a separate queue joined by a redrive policy, as in configuring an SQS redrive policy.

Sessions and the dead-letter subqueue Messages with session id P-100 are locked by receiver 1, messages for P-200 by receiver 2, and P-300 waits for a free receiver. Within each session messages are delivered in order. Messages that exceed the maximum delivery count move to the queue's built-in dead-letter subqueue. policy-events (sessions enabled) session P-100: 3 msgs session P-200: 1 msg session P-300: waiting $DeadLetterQueue after 6 deliveries receiver 1 locks P-100 receiver 2 locks P-200 one receiver per session at a time

Step 2 — Send Messages with a Session Id

Every message sent to a session-enabled queue must carry a session_id; sending without one fails. Use the entity whose events must stay ordered.

# producer.py
from azure.identity import DefaultAzureCredential
from azure.servicebus import ServiceBusClient, ServiceBusMessage
import json

client = ServiceBusClient("insurance-sb.servicebus.windows.net", DefaultAzureCredential())

def publish(event: dict) -> None:
    with client.get_queue_sender("policy-events") as sender:
        sender.send_messages(ServiceBusMessage(
            json.dumps(event),
            session_id=event["policy_number"],       # ordering scope
            message_id=event["event_id"],            # enables duplicate detection if configured
            content_type="application/json",
            application_properties={"event_type": event["type"], "schema": "v2"},
        ))

If the queue has duplicate detection enabled (a creation-time setting with a time window), message_id doubles as the dedup key, similar to SQS FIFO's deduplication id.

Step 3 — Receive by Session with Peek-Lock

A session receiver locks an entire session: while it holds the lock, no other receiver gets messages for that policy. Messages are received in peek-lock mode — locked and invisible to others but not removed — and removed only when you complete them.

# worker.py — process any available session, one at a time per task
from azure.servicebus import NEXT_AVAILABLE_SESSION, ServiceBusReceiveMode
from azure.servicebus.exceptions import OperationTimeoutError

def process_next_session() -> None:
    try:
        with client.get_queue_receiver(
            "policy-events",
            session_id=NEXT_AVAILABLE_SESSION,       # take whichever session is free
            receive_mode=ServiceBusReceiveMode.PEEK_LOCK,
            max_wait_time=30,                        # close the session after 30 s idle
        ) as receiver:
            for msg in receiver:                     # messages of ONE session, in order
                try:
                    apply_policy_event(json.loads(str(msg)))
                    receiver.complete_message(msg)
                except PermanentEventError as exc:
                    receiver.dead_letter_message(msg, reason="permanent", error_description=str(exc))
                except Exception:
                    receiver.abandon_message(msg)    # delivery count +1, retried in order
                    break                            # stop this session; don't skip ahead
    except OperationTimeoutError:
        pass                                         # no session available right now

Run several of these loops concurrently per worker replica (one per thread or asyncio task) to process several policies in parallel. The break after abandon is what preserves order: the abandoned message is redelivered as the next message of the session, and nothing later is processed first. The ordering model is the same lane concept described in Message Ordering Guarantees.

Step 4 — Renew Locks for Long Work

The message lock (and the session lock) expires after LockDuration — one minute here. If processing takes longer, the lock lapses, the message is delivered again, and for sessions another receiver can take the session. The SDK's AutoLockRenewer renews locks in the background.

from azure.servicebus import AutoLockRenewer

renewer = AutoLockRenewer(max_lock_renewal_duration=600)    # renew for up to 10 minutes

with client.get_queue_receiver("policy-events", session_id=NEXT_AVAILABLE_SESSION,
                               auto_lock_renewer=renewer) as receiver:
    for msg in receiver:
        run_long_underwriting_check(msg)             # can take several minutes
        receiver.complete_message(msg)

The renewal ceiling should exceed your longest expected processing time, and processing that hangs forever should still be bounded — a stuck handler holding a session lock blocks that policy until the renewal limit. This is the Service Bus form of the lease-renewal pattern in configuring visibility timeouts for long-running workers.

Renew the lock or lose it With a one-minute lock and no renewal, a four-minute job loses its lock after one minute and the message is delivered to another receiver while the first is still working. With AutoLockRenewer, the lock is renewed before each expiry and the job completes the message at four minutes with no redelivery. A 4-minute job with LockDuration = 1 minute no renewal locked lock lost: redelivered while still running AutoLockRenewer renewed renewed renewed complete Cap total renewal time so a hung handler cannot hold a session forever.

Step 5 — Read and Resubmit Dead-Lettered Messages

Dead-lettered messages keep their body, properties, and two system properties explaining why: DeadLetterReason and DeadLetterErrorDescription. Inspect them with a receiver on the dead-letter subqueue; resubmit by sending a copy to the main queue and completing the dead-lettered original.

from azure.servicebus import ServiceBusSubQueue

def triage(limit: int = 50) -> None:
    with client.get_queue_receiver("policy-events", sub_queue=ServiceBusSubQueue.DEAD_LETTER,
                                   receive_mode=ServiceBusReceiveMode.PEEK_LOCK) as dlq:
        for msg in dlq.receive_messages(max_message_count=limit, max_wait_time=5):
            print(msg.message_id, msg.session_id, msg.dead_letter_reason,
                  msg.dead_letter_error_description, msg.delivery_count)
            dlq.abandon_message(msg)                 # look, don't touch

def resubmit(message_ids: set[str]) -> None:
    with client.get_queue_receiver("policy-events", sub_queue=ServiceBusSubQueue.DEAD_LETTER) as dlq, \
         client.get_queue_sender("policy-events") as sender:
        for msg in dlq.receive_messages(max_message_count=100, max_wait_time=5):
            if msg.message_id not in message_ids:
                dlq.abandon_message(msg)
                continue
            sender.send_messages(ServiceBusMessage(
                b"".join(msg.body), session_id=msg.session_id, message_id=msg.message_id,
                application_properties={**(msg.application_properties or {}), "resubmitted": True}))
            dlq.complete_message(msg)                # remove from DLQ only after the send

Resubmitted messages join the end of their session. If later events for the same policy were already processed, the resubmitted one arrives out of order — so the consumer should check event sequence numbers, or the operator should resubmit only after confirming no later events exist for that policy. Service Bus Explorer in the Azure portal offers the same operations interactively.

Order the resubmit steps deliberately: send the copy first, then complete the dead-lettered original. If the process dies between the two, the worst case is a duplicate in the main queue (which an idempotent consumer absorbs), never a message lost from both places. Doing it the other way round — complete, then send — can lose the event entirely.

Send first, complete second The operator's resubmit tool reads a message from the dead-letter subqueue under a lock, sends a copy with the same session id and message id to the main queue, and only then completes the original. A crash after the send leaves a duplicate that duplicate detection or idempotent handling absorbs; a crash before the send leaves the original safely in the dead-letter subqueue. Crash-safe resubmission 1. peek-lock in DLQ inspect reason 2. send copy same session + message id 3. complete original removed from DLQ Crash between 2 and 3: a duplicate, absorbed by dedup or idempotency. Never a loss.

Replay safety in general is covered in replaying dead-letter messages in RabbitMQ.

Step 6 — Schedule Messages Instead of Sleeping

For "try again in an hour" or "send the renewal reminder in 30 days", schedule a message rather than holding a worker or abandoning repeatedly. Scheduled messages are stored but invisible until their enqueue time.

import datetime as dt
with client.get_queue_sender("policy-events") as sender:
    seq = sender.schedule_messages(
        ServiceBusMessage(json.dumps({"type": "RenewalReminder", "policy_number": "P-100"}),
                          session_id="P-100"),
        dt.datetime.now(dt.timezone.utc) + dt.timedelta(days=30))
    # sender.cancel_scheduled_messages(seq)  # cancellable by sequence number

Keep the returned sequence number if you might need to cancel — for example, if the policy is cancelled before the reminder is due.

Verification

# Queue and DLQ counts, including active vs scheduled
az servicebus queue show -g insurance-rg --namespace-name insurance-sb -n policy-events \
  --query "countDetails.{active:activeMessageCount, deadletter:deadLetterMessageCount, scheduled:scheduledMessageCount}"

Then test the two guarantees: publish ten sequenced events for one policy with a handler that sleeps randomly, run three replicas, and assert they were applied in order; publish one malformed event and confirm it appears in the dead-letter subqueue after six deliveries with a meaningful DeadLetterReason.

Gotchas & Edge Cases

Idle sessions hold receivers. A receiver waiting on a session with no messages blocks that slot until max_wait_time. Keep it short (seconds) so receivers move on to sessions with work.

Hot sessions. A policy with a flood of events is processed by one receiver at a time; its throughput is one handler's throughput. Keep session ids at the finest granularity that still covers the ordering requirement.

Abandon in a loop. Abandoning a message immediately redelivers it with no backoff. For transient errors, prefer scheduling a retry message with a delay and completing the original, or accept fast retries up to MaxDeliveryCount.

Dead letters expire too. Messages in the dead-letter subqueue do not expire by TTL, but they do count toward the namespace's storage quota. Alert on DLQ count and triage promptly.

FAQ

Do I need sessions if I don't need ordering? No. Sessions reduce parallelism to the number of active sessions. For independent jobs, a non-session queue with competing receivers is simpler and faster.

How is MaxDeliveryCount different from SQS maxReceiveCount? They play the same role. Service Bus increments the delivery count on each lock expiry or abandon, and dead-letters when it exceeds the limit — into a built-in subqueue instead of a separate queue.

Can Azure Functions consume a session-enabled queue? Yes; set isSessionsEnabled on the Service Bus trigger. The Functions host manages session locks and concurrency (maxConcurrentSessions) for you.

Related