Azure Service Bus Sessions and Dead-Lettering
Azure Service Bus is the most feature-rich of the managed queues — sessions, a built-in dead-letter subqueue, scheduled messages, and transactions — and this guide covers the features that matter for background jobs, as part of Managed Cloud Queues for Background Jobs in Backend Frameworks & Worker Scaling. The focus is on the two that most teams need and most often misconfigure: sessions for ordered per-entity processing, and the dead-letter subqueue for poison messages.
Problem Statement
An insurance platform processes policy change events — endorsements, cancellations, reinstatements — through a Service Bus queue consumed by an Azure Container Apps worker. Two problems keep recurring. Events for the same policy are sometimes applied out of order when several worker replicas receive them concurrently, producing a cancelled policy that is also marked active. And malformed events retry indefinitely until someone notices the queue is not draining. You want events for each policy processed strictly in order while different policies proceed in parallel, poison events quarantined automatically after a few attempts, and a safe procedure to inspect and resubmit them.
Prerequisites
- A Service Bus namespace (Standard tier supports sessions and dead-lettering; Premium adds larger messages and isolation).
- The
azure-servicebusPython SDK (v7+) or the equivalent .NET/Java/JS SDK. - Managed identity for the worker with the
Azure Service Bus Data Receiverrole on the queue, andData Senderfor any process that resubmits. - An entity id (policy number) on every message, used as the session id.
Step 1 — Create a Session-Enabled Queue with a Delivery Limit
Sessions must be enabled when the queue is created; the setting cannot be changed later. MaxDeliveryCount sets how many delivery attempts a message gets before Service Bus moves it to the queue's dead-letter subqueue.
az servicebus queue create \
--resource-group insurance-rg --namespace-name insurance-sb \
--name policy-events \
--enable-session true \ # FIFO per session id
--max-delivery-count 6 \ # 6 attempts, then dead-letter
--lock-duration PT1M \ # peek-lock duration; renew for longer work
--default-message-time-to-live P14D \
--enable-dead-lettering-on-message-expiration true # expired messages go to DLQ, not void
The dead-letter subqueue exists automatically at policy-events/$DeadLetterQueue; there is nothing to create or wire up. That is a real difference from SQS, where the DLQ is a separate queue joined by a redrive policy, as in configuring an SQS redrive policy.
Step 2 — Send Messages with a Session Id
Every message sent to a session-enabled queue must carry a session_id; sending without one fails. Use the entity whose events must stay ordered.
# producer.py
from azure.identity import DefaultAzureCredential
from azure.servicebus import ServiceBusClient, ServiceBusMessage
import json
client = ServiceBusClient("insurance-sb.servicebus.windows.net", DefaultAzureCredential())
def publish(event: dict) -> None:
with client.get_queue_sender("policy-events") as sender:
sender.send_messages(ServiceBusMessage(
json.dumps(event),
session_id=event["policy_number"], # ordering scope
message_id=event["event_id"], # enables duplicate detection if configured
content_type="application/json",
application_properties={"event_type": event["type"], "schema": "v2"},
))
If the queue has duplicate detection enabled (a creation-time setting with a time window), message_id doubles as the dedup key, similar to SQS FIFO's deduplication id.
Step 3 — Receive by Session with Peek-Lock
A session receiver locks an entire session: while it holds the lock, no other receiver gets messages for that policy. Messages are received in peek-lock mode — locked and invisible to others but not removed — and removed only when you complete them.
# worker.py — process any available session, one at a time per task
from azure.servicebus import NEXT_AVAILABLE_SESSION, ServiceBusReceiveMode
from azure.servicebus.exceptions import OperationTimeoutError
def process_next_session() -> None:
try:
with client.get_queue_receiver(
"policy-events",
session_id=NEXT_AVAILABLE_SESSION, # take whichever session is free
receive_mode=ServiceBusReceiveMode.PEEK_LOCK,
max_wait_time=30, # close the session after 30 s idle
) as receiver:
for msg in receiver: # messages of ONE session, in order
try:
apply_policy_event(json.loads(str(msg)))
receiver.complete_message(msg)
except PermanentEventError as exc:
receiver.dead_letter_message(msg, reason="permanent", error_description=str(exc))
except Exception:
receiver.abandon_message(msg) # delivery count +1, retried in order
break # stop this session; don't skip ahead
except OperationTimeoutError:
pass # no session available right now
Run several of these loops concurrently per worker replica (one per thread or asyncio task) to process several policies in parallel. The break after abandon is what preserves order: the abandoned message is redelivered as the next message of the session, and nothing later is processed first. The ordering model is the same lane concept described in Message Ordering Guarantees.
Step 4 — Renew Locks for Long Work
The message lock (and the session lock) expires after LockDuration — one minute here. If processing takes longer, the lock lapses, the message is delivered again, and for sessions another receiver can take the session. The SDK's AutoLockRenewer renews locks in the background.
from azure.servicebus import AutoLockRenewer
renewer = AutoLockRenewer(max_lock_renewal_duration=600) # renew for up to 10 minutes
with client.get_queue_receiver("policy-events", session_id=NEXT_AVAILABLE_SESSION,
auto_lock_renewer=renewer) as receiver:
for msg in receiver:
run_long_underwriting_check(msg) # can take several minutes
receiver.complete_message(msg)
The renewal ceiling should exceed your longest expected processing time, and processing that hangs forever should still be bounded — a stuck handler holding a session lock blocks that policy until the renewal limit. This is the Service Bus form of the lease-renewal pattern in configuring visibility timeouts for long-running workers.
Step 5 — Read and Resubmit Dead-Lettered Messages
Dead-lettered messages keep their body, properties, and two system properties explaining why: DeadLetterReason and DeadLetterErrorDescription. Inspect them with a receiver on the dead-letter subqueue; resubmit by sending a copy to the main queue and completing the dead-lettered original.
from azure.servicebus import ServiceBusSubQueue
def triage(limit: int = 50) -> None:
with client.get_queue_receiver("policy-events", sub_queue=ServiceBusSubQueue.DEAD_LETTER,
receive_mode=ServiceBusReceiveMode.PEEK_LOCK) as dlq:
for msg in dlq.receive_messages(max_message_count=limit, max_wait_time=5):
print(msg.message_id, msg.session_id, msg.dead_letter_reason,
msg.dead_letter_error_description, msg.delivery_count)
dlq.abandon_message(msg) # look, don't touch
def resubmit(message_ids: set[str]) -> None:
with client.get_queue_receiver("policy-events", sub_queue=ServiceBusSubQueue.DEAD_LETTER) as dlq, \
client.get_queue_sender("policy-events") as sender:
for msg in dlq.receive_messages(max_message_count=100, max_wait_time=5):
if msg.message_id not in message_ids:
dlq.abandon_message(msg)
continue
sender.send_messages(ServiceBusMessage(
b"".join(msg.body), session_id=msg.session_id, message_id=msg.message_id,
application_properties={**(msg.application_properties or {}), "resubmitted": True}))
dlq.complete_message(msg) # remove from DLQ only after the send
Resubmitted messages join the end of their session. If later events for the same policy were already processed, the resubmitted one arrives out of order — so the consumer should check event sequence numbers, or the operator should resubmit only after confirming no later events exist for that policy. Service Bus Explorer in the Azure portal offers the same operations interactively.
Order the resubmit steps deliberately: send the copy first, then complete the dead-lettered original. If the process dies between the two, the worst case is a duplicate in the main queue (which an idempotent consumer absorbs), never a message lost from both places. Doing it the other way round — complete, then send — can lose the event entirely.
Replay safety in general is covered in replaying dead-letter messages in RabbitMQ.
Step 6 — Schedule Messages Instead of Sleeping
For "try again in an hour" or "send the renewal reminder in 30 days", schedule a message rather than holding a worker or abandoning repeatedly. Scheduled messages are stored but invisible until their enqueue time.
import datetime as dt
with client.get_queue_sender("policy-events") as sender:
seq = sender.schedule_messages(
ServiceBusMessage(json.dumps({"type": "RenewalReminder", "policy_number": "P-100"}),
session_id="P-100"),
dt.datetime.now(dt.timezone.utc) + dt.timedelta(days=30))
# sender.cancel_scheduled_messages(seq) # cancellable by sequence number
Keep the returned sequence number if you might need to cancel — for example, if the policy is cancelled before the reminder is due.
Verification
# Queue and DLQ counts, including active vs scheduled
az servicebus queue show -g insurance-rg --namespace-name insurance-sb -n policy-events \
--query "countDetails.{active:activeMessageCount, deadletter:deadLetterMessageCount, scheduled:scheduledMessageCount}"
Then test the two guarantees: publish ten sequenced events for one policy with a handler that sleeps randomly, run three replicas, and assert they were applied in order; publish one malformed event and confirm it appears in the dead-letter subqueue after six deliveries with a meaningful DeadLetterReason.
Gotchas & Edge Cases
Idle sessions hold receivers. A receiver waiting on a session with no messages blocks that slot until max_wait_time. Keep it short (seconds) so receivers move on to sessions with work.
Hot sessions. A policy with a flood of events is processed by one receiver at a time; its throughput is one handler's throughput. Keep session ids at the finest granularity that still covers the ordering requirement.
Abandon in a loop. Abandoning a message immediately redelivers it with no backoff. For transient errors, prefer scheduling a retry message with a delay and completing the original, or accept fast retries up to MaxDeliveryCount.
Dead letters expire too. Messages in the dead-letter subqueue do not expire by TTL, but they do count toward the namespace's storage quota. Alert on DLQ count and triage promptly.
FAQ
Do I need sessions if I don't need ordering? No. Sessions reduce parallelism to the number of active sessions. For independent jobs, a non-session queue with competing receivers is simpler and faster.
How is MaxDeliveryCount different from SQS maxReceiveCount? They play the same role. Service Bus increments the delivery count on each lock expiry or abandon, and dead-letters when it exceeds the limit — into a built-in subqueue instead of a separate queue.
Can Azure Functions consume a session-enabled queue?
Yes; set isSessionsEnabled on the Service Bus trigger. The Functions host manages session locks and concurrency (maxConcurrentSessions) for you.
Related
- Managed Cloud Queues for Background Jobs — how Service Bus compares with SQS and Google's services.
- Single Active Consumer in RabbitMQ — a self-hosted equivalent of sessions.
- Dead-Letter Queues & Poison Messages — DLQ design across brokers.
- Handling Out-of-Order Events with Sequence Numbers — protecting resubmitted messages from reordering.