Visualizing SQS Metrics in Grafana
Amazon SQS publishes its metrics to CloudWatch automatically, and Grafana can read CloudWatch directly. The difficulty is not getting the data but reading it correctly: SQS metrics are approximate, arrive at one-minute resolution with a delay, and mean different things depending on the statistic you choose. This guide builds a dashboard that avoids those traps, as part of Grafana Dashboards for Queues in Observability & Monitoring for Job Queues.
Problem Statement
A company runs 25 SQS queues feeding Lambda functions and ECS workers, plus a dead-letter queue for each. Engineers look at queue depth in the AWS console one queue at a time. Last month an ECS service scaled to zero by mistake and messages sat in a queue for five hours; the queue depth barely moved because the producer's volume was small, so nobody noticed. Messages also piled up in a dead-letter queue for two weeks without anyone looking. The team already uses Grafana for application metrics and wants SQS on the same screens, with panels that show when messages are getting old and when anything lands in a dead-letter queue.
Prerequisites
- Grafana 10 or newer with the CloudWatch data source plugin (built in).
- An IAM role or user Grafana can assume, with
cloudwatch:GetMetricData,cloudwatch:ListMetrics, andsqs:ListQueuespermissions. - A naming convention linking each queue to its DLQ (for example,
ordersandorders-dlq). - Optionally, Prometheus with a CloudWatch exporter, if you prefer PromQL.
Step 1 β Connect the CloudWatch Data Source
In Grafana, add a CloudWatch data source using an IAM role (preferred on EC2, ECS, or EKS with IRSA) rather than long-lived keys:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["cloudwatch:GetMetricData", "cloudwatch:ListMetrics", "sqs:ListQueues", "tag:GetResources"],
"Resource": "*"
}]
}
Set the default region to where your queues live, and add a second data source per region or account if needed. Grafana queries CloudWatch on every dashboard refresh, and each metric returned counts toward CloudWatch API charges β Step 6 covers how to keep that under control.
Step 2 β Choose the Metrics and Statistics That Matter
SQS publishes about ten metrics. Five of them answer the operational questions:
| Metric | Statistic | Question it answers |
|---|---|---|
ApproximateAgeOfOldestMessage |
Maximum | How long has the oldest message been waiting? |
ApproximateNumberOfMessagesVisible |
Maximum | How many messages are waiting? |
ApproximateNumberOfMessagesNotVisible |
Maximum | How many are in flight (received, not deleted)? |
NumberOfMessagesSent |
Sum | How fast are producers adding work? |
NumberOfMessagesDeleted |
Sum | How fast are consumers finishing work? |
The statistic matters. The Approximateβ¦ metrics are gauges β use Maximum (or Average) per period, never Sum, which adds up samples and produces meaningless numbers. NumberOfMessagesSent and NumberOfMessagesDeleted are counts per period β use Sum. NumberOfMessagesReceived counts every receive, including redeliveries after a visibility timeout, so it can be much higher than sent; compare it to deleted to spot messages being received repeatedly without completing.
ApproximateAgeOfOldestMessage is the most valuable of them. It is the SQS equivalent of queue latency: it rises steadily when consumers stop, regardless of how many messages there are, which is exactly the case that went unnoticed for five hours.
Step 3 β Build the Panels
Use a dashboard variable to select queues, then repeat panels for each selected queue or show all queues on shared panels. A layout that works for most teams:
In the CloudWatch query editor, use the Metric Search mode for the per-queue panels so new queues appear automatically:
Namespace: AWS/SQS
Metric: ApproximateAgeOfOldestMessage
Statistic: Maximum
Dimensions: QueueName = *
Period: 60
For the DLQ panels, use a Metrics Insights query, which filters by name pattern and aggregates in one request:
SELECT MAX(ApproximateNumberOfMessagesVisible)
FROM SCHEMA("AWS/SQS", QueueName)
WHERE QueueName LIKE '%-dlq'
GROUP BY QueueName
Show age in human units (Grafana's duration (s) unit), and add threshold lines at each queue's target so the panel reads as "within target" or "not" without mental arithmetic.
Step 4 β Make Dead-Letter Queues Impossible to Ignore
A message in a DLQ is a job that failed every attempt. It needs a human decision β fix and redrive, or discard β and it will not go away by itself until the retention period deletes it (four days by default, up to 14). Give DLQs a prominent stat panel that is red whenever any DLQ has messages, and a table that shows which ones.
Note that ApproximateNumberOfMessagesVisible on a DLQ is the right metric; NumberOfMessagesSent does not count messages moved there by a redrive policy, so a panel of "sent to DLQ" stays at zero while the DLQ fills. Also note that CloudWatch stops publishing metrics for a queue that has been inactive for about six hours, so a DLQ that received messages long ago and has seen no activity since may show "no data" rather than a number. Configure panels to treat "no data" as a warning colour rather than green, and alert with "treat missing data as breaching" for DLQs. Redriving is covered in dead-letter queues & poison messages.
Step 5 β Account for Delay and Resolution
CloudWatch SQS metrics are published at one-minute resolution and can arrive a few minutes late. On a dashboard, that means the most recent point may be missing or partial. Set the panel's relative time shift or "hide last N minutes" option, or simply avoid drawing conclusions from the last two minutes. For alerts, use evaluation windows of at least 5 minutes with "missing data" handled explicitly.
If you need faster or more precise numbers for a critical queue, have consumers record their own metrics β processing time, wait time from the SentTimestamp attribute β and export them to Prometheus. The pattern is in measuring queue wait time with enqueue timestamps.
Step 6 β Control CloudWatch Costs
Each dashboard refresh calls GetMetricData, which is billed per metric requested. A dashboard with 25 queues Γ 5 metrics refreshed every 10 seconds by several viewers adds up. Keep costs in check:
- Set the dashboard's default refresh to 1 minute or longer; SQS data does not change faster.
- Use Metrics Insights queries that return many queues in one request instead of one query per queue.
- For wall displays and heavy use, export SQS metrics once into Prometheus with YACE (Yet Another CloudWatch Exporter) or CloudWatch metric streams, and point Grafana at Prometheus instead.
The exporter route also brings SQS into PromQL, so SQS queues can share alert rules and recording rules with the rest of your queues. The trade-off is one more component to run and another minute or two of delay on top of CloudWatch's own. For a team with a handful of queues and occasional viewers, direct CloudWatch queries are simpler; for dozens of queues on a shared wall display, an exporter pays for itself quickly.
Verification
- Every production queue appears on the age panel, including queues created after the dashboard was built.
- Stopping a consumer in staging makes that queue's age line climb within about five minutes.
- Sending a message to a test DLQ turns the DLQ stat red and lists the queue in the table.
- A DLQ with no recent activity shows as "no data" in a warning colour, not as green zero.
- CloudWatch API usage for the Grafana role stays within the expected budget.
Gotchas & Edge Cases
FIFO queue names. FIFO queues end in .fifo, so a DLQ pattern of %-dlq misses orders-dlq.fifo. Use %-dlq% or a tag-based filter.
Approximate means approximate. SQS is distributed; the visible and in-flight counts are estimates and can briefly disagree with reality, especially for small numbers. Do not alert on "exactly zero" for depth.
Lambda consumers and in-flight counts. With Lambda event source mappings, messages are in flight while Lambda polls in batches; a high in-flight count is normal and not a sign of stuck work.
Cross-account queues. Queues in other accounts need a data source with a role in that account, or CloudWatch cross-account observability.
FAQ
Should I use CloudWatch alarms or Grafana alerts? Either works. CloudWatch alarms integrate with SNS and autoscaling; Grafana alerts sit next to your other alerts. Pick one per queue so alerts are not duplicated.
What target should I set for age of oldest message? Derive it from how long users can wait for the work, then subtract typical processing time. For interactive work, often under a minute; for batch pipelines, hours. Defining SLOs for job latency walks through it.
Can I see per-message detail in Grafana? No. CloudWatch provides aggregate metrics only. For individual messages, use logs or traces from the consumers.
Related
- Grafana Dashboards for Queues β shared panel patterns.
- Managed Cloud Queues β SQS and its alternatives.
- Dead-Letter Queues & Poison Messages β handling what lands in DLQs.
- Building a Celery Grafana Dashboard β the same layout for Celery.