Handling Partial Batch Failures in Lambda with SQS

By default a Lambda function processing an SQS batch succeeds or fails as a unit, and this guide replaces that with per-message outcomes, as part of Managed Cloud Queues for Background Jobs in Backend Frameworks & Worker Scaling. Partial batch responses are a small API — one setting and one return shape — with a handful of traps that silently turn them back into all-or-nothing, or worse, delete messages that failed.

Problem Statement

A notification function consumes batches of 25 messages. About one message in 5,000 references a user who has since been deleted, and the handler raises on it. Each time, the whole batch of 25 returns to the queue, all 25 are reprocessed — sending 24 duplicate push notifications — and after five rounds all 25 land in the DLQ. Support sees users complaining about duplicate notifications, and on-call sees a DLQ full of perfectly valid messages. You want the one bad message retried and eventually dead-lettered on its own, and the other 24 processed exactly once.

Prerequisites

  • An SQS event source mapping on a Lambda function (see processing SQS with AWS Lambda).
  • Permission to update the mapping's FunctionResponseTypes.
  • A handler that processes records individually inside a loop.
  • For FIFO queues, knowledge of which records share a MessageGroupId.

Step 1 — Enable ReportBatchItemFailures on the Mapping

The feature is off by default. Without the setting, the mapping ignores whatever your function returns and treats any successful invocation as "delete everything".

resource "aws_lambda_event_source_mapping" "notify" {
  event_source_arn        = aws_sqs_queue.notify.arn
  function_name           = aws_lambda_function.notify.arn
  batch_size              = 25
  function_response_types = ["ReportBatchItemFailures"]    # the switch
}
# Or on an existing mapping
aws lambda update-event-source-mapping --uuid "$MAPPING_UUID" \
  --function-response-types ReportBatchItemFailures

This ordering matters when rolling out: deploy the mapping setting first, then the code that returns failures. The reverse order briefly runs code that catches exceptions and returns failure lists that the mapping ignores — so failed messages are deleted as if they succeeded.

One bad record, 24 good ones Without partial batch responses, the handler raises on one record and all 25 messages are returned to the queue, reprocessed, and eventually dead-lettered together. With partial batch responses, the handler reports only the failing message id; the mapping deletes the other 24 and returns only the one failure to the queue. Batch of 25, record 17 fails all-or-nothing all 25 return, 24 duplicate sends, all 25 reach the DLQ partial response 24 deleted after success 1 Only the reported message returns to the queue and counts a receive toward the DLQ.

Step 2 — Return the Right Shape

The function must return an object with a batchItemFailures list, each entry carrying the failed record's messageId as itemIdentifier. An empty list means "everything succeeded".

# handler.py
import json

def handler(event, context):
    failures = []
    for record in event["Records"]:
        try:
            send_notification(json.loads(record["body"]))
        except UserDeleted:
            # permanent: log and treat as handled so it is not retried pointlessly
            log.info("skip notification for deleted user", extra={"msg": record["messageId"]})
        except Exception:
            log.exception("notification failed", extra={"msg": record["messageId"]})
            failures.append({"itemIdentifier": record["messageId"]})
    return {"batchItemFailures": failures}
// handler.ts — the same contract with the AWS types
import type { SQSEvent, SQSBatchResponse } from "aws-lambda";

export const handler = async (event: SQSEvent): Promise<SQSBatchResponse> => {
  const batchItemFailures: SQSBatchResponse["batchItemFailures"] = [];
  for (const record of event.Records) {
    try {
      await sendNotification(JSON.parse(record.body));
    } catch (err) {
      batchItemFailures.push({ itemIdentifier: record.messageId });
    }
  }
  return { batchItemFailures };
};

Note the first except clause: a message that can never succeed (the user is gone) is better acknowledged and logged than retried five times into the DLQ. Distinguishing permanent from transient failures is covered in retrying only transient errors by exception type.

Step 3 — Avoid the Traps That Fail the Whole Batch

Several responses make Lambda treat the entire batch as failed, even with the feature enabled:

# 1. Raising an exception — the whole batch fails
raise RuntimeError("database down")        # all records return to the queue

# 2. An itemIdentifier that is empty, null, or not in the batch — whole batch fails
{"batchItemFailures": [{"itemIdentifier": ""}]}

# 3. A malformed response (wrong key name, not a dict) — whole batch fails
{"failures": [...]}

The first is sometimes what you want: if the database is unreachable, every record will fail, and raising avoids processing 25 records that are doomed. Make that a deliberate choice by checking dependency health before the loop, rather than letting an unexpected exception escape. The second and third are always bugs; test the response shape (Step 6).

Conversely, returning None, an empty dict, or {"batchItemFailures": []} means all succeeded. A handler that swallows every exception and returns nothing therefore deletes failed messages — the most dangerous trap because nothing looks wrong.

What each return value means Raising an exception or returning a malformed response fails the entire batch. Returning a list of valid message ids retries only those messages. Returning an empty list, an empty object, or nothing deletes every message in the batch, including any that silently failed. Handler result and mapping action raise / malformed response whole batch returns [{"itemIdentifier": id}, ...] only listed ids return [] / {} / None everything deleted A handler that swallows every error and returns nothing silently deletes failures.

Step 4 — Handle FIFO Queues Without Breaking Order

On a FIFO queue, a batch contains messages from one or more message groups, in order. If message 3 of group A fails and you report only message 3, the mapping deletes messages 4 and 5 of group A — which then ran before message 3 will be retried. To preserve order, stop processing a group at its first failure and report that message and every later message of the same group.

def handler(event, context):
    failures, failed_groups = [], set()
    for record in event["Records"]:
        group = record["attributes"]["MessageGroupId"]
        if group in failed_groups:
            failures.append({"itemIdentifier": record["messageId"]})   # skip: keep order
            continue
        try:
            apply_event(json.loads(record["body"]))
        except Exception:
            failed_groups.add(group)
            failures.append({"itemIdentifier": record["messageId"]})
    return {"batchItemFailures": failures}

The reported messages return to the queue together and are redelivered in their original order.

Groups that did not fail are unaffected: in the example below, group B's messages are processed and deleted even though group A stopped at its second message. That is the property that makes FIFO with partial responses efficient — one broken aggregate does not stall the others sharing the batch — while still never applying a later event for an aggregate before an earlier one. If you use Powertools (Step 5), its SqsFifoPartialProcessor implements exactly this rule; with a hand-written handler, keep the failed_groups set and write a test for it, because the mistake produces no error, only a silently reordered aggregate.

Stop the group, not the batch A FIFO batch holds A1, A2, A3 from group A and B1, B2 from group B. A1 succeeds. A2 fails, so A2 and the later A3 are both reported as failures and return to the queue in order. B1 and B2 are processed and deleted normally. FIFO batch: report the failure and the rest of its group group A A1 ok, deleted A2 fails: reported A3 skipped: reported group B B1 ok, deleted B2 ok, deleted Reporting only A2 would delete A3, which then took effect before A2 on retry.

The ordering rules behind this are in FIFO ordering with SQS message group IDs.

Step 5 — Use Powertools to Remove the Boilerplate

AWS Lambda Powertools (Python, TypeScript, Java, .NET) implements the loop, the response shape, FIFO short-circuiting, and the "all failed means raise" rule.

from aws_lambda_powertools.utilities.batch import (
    BatchProcessor, EventType, process_partial_response)
from aws_lambda_powertools.utilities.batch import SqsFifoPartialProcessor

processor = BatchProcessor(event_type=EventType.SQS)       # standard queues
# processor = SqsFifoPartialProcessor()                    # FIFO: stops at first failure per group

def record_handler(record):
    send_notification(record.json_body)                   # raise to mark this record failed

def handler(event, context):
    return process_partial_response(
        event=event, record_handler=record_handler, processor=processor, context=context)

When every record in a batch fails, Powertools raises instead of returning a full failure list, which makes the invocation count as an error in metrics and alarms — a batch in which nothing succeeded is usually a dependency outage, and you want it to look like one.

Step 6 — Test the Contract

Unit-test the response shape with constructed events; integration-test the end-to-end behaviour against a real queue in a sandbox account or LocalStack.

def make_event(bodies):
    return {"Records": [
        {"messageId": f"m{i}", "body": json.dumps(b),
         "attributes": {"MessageGroupId": b.get("group", "g")}} for i, b in enumerate(bodies)]}

def test_reports_only_failing_record(monkeypatch):
    monkeypatch.setattr(mod, "send_notification",
                        lambda b: (_ for _ in ()).throw(RuntimeError()) if b.get("bad") else None)
    resp = mod.handler(make_event([{}, {"bad": True}, {}]), None)
    assert resp == {"batchItemFailures": [{"itemIdentifier": "m1"}]}

def test_fifo_reports_rest_of_group():
    resp = fifo_handler(make_event([{"group": "A"}, {"group": "A", "bad": True},
                                    {"group": "A"}, {"group": "B"}]), None)
    assert [f["itemIdentifier"] for f in resp["batchItemFailures"]] == ["m1", "m2"]

Verification

In staging, send a batch containing one poison message and watch the metrics:

aws sqs send-message-batch --queue-url "$Q" --entries file://batch-with-one-poison.json
# After ~5 visibility cycles:
aws sqs get-queue-attributes --queue-url "$DLQ" --attribute-names ApproximateNumberOfMessages
# expect 1, not 25

In production, compare the DLQ arrival rate before and after the change — it should drop to the true poison rate — and check that duplicate-send metrics from the notification service fall to near zero.

Gotchas & Edge Cases

Timeouts fail the whole batch. If the function times out mid-loop, no response is returned and every record returns to the queue. Size the batch so that a full batch finishes well within the timeout.

Idempotency is still required. Partial responses stop batch-induced duplicates, not SQS's at-least-once duplicates or retries after a crash. Keep side effects idempotent.

Concurrency inside the handler. Processing records in parallel threads is fine for standard queues, but collect results carefully — a record whose thread raised must appear in the failure list, and FIFO queues should stay sequential per group.

Metrics hide partial failures. An invocation that returns a failure list counts as a success in Lambda's Errors metric. Emit your own counter of failed records per invocation, or failures become invisible.

FAQ

Does partial batch response work for Kinesis and DynamoDB Streams? Yes, with a different meaning: for stream sources, the reported item is a checkpoint, and processing resumes from the lowest failed sequence number. Everything after it is retried.

Is a smaller batch size an alternative? Batch size 1 avoids the problem but multiplies invocation count and cost. Partial responses give the isolation of size 1 at the efficiency of larger batches.

Should I raise when the whole batch fails? Yes. If every record fails — typically because a shared dependency is down — raising makes the invocation an error in Lambda metrics, which feeds alarms and lets the mapping back off. Returning a failure list with every id has the same retry effect but reports the invocation as a success, hiding the outage from dashboards built on the Errors metric.

What happens to a message reported as failed? It stays in the queue, becomes visible again after the visibility timeout, and its receive count increases. After maxReceiveCount it moves to the DLQ like any other failing message.

Related