Shipping Worker Logs to Loki
Grafana Loki stores logs cheaply by indexing only a few labels and scanning the rest at query time, which suits worker logs well — provided the labels are chosen carefully. This guide ships structured worker logs from Kubernetes to Loki and builds the queries used during job incidents, as part of Structured Logging for Workers in Observability & Monitoring for Job Queues.
Problem Statement
A platform team moved worker logs from Elasticsearch to Loki to cut cost. The first attempt promoted every JSON field to a label, including job_id and order_id. Within a day Loki was rejecting writes with "maximum active stream limit exceeded", ingesters were running out of memory, and queries that used to take a second timed out. The second attempt used no JSON-derived labels at all, so every query scanned all worker logs for the time range. You want a labelling scheme that keeps the stream count bounded, queries that find one job's history in seconds, sensible retention, and dashboards that link from metric alerts to the relevant log lines.
Prerequisites
- Kubernetes with worker pods writing one JSON object per line to stdout (see Structured Logging for Workers).
- Loki 3.x (for structured metadata and bloom-filter acceleration) deployed via Helm or Grafana Cloud.
- Grafana Alloy (the successor to Grafana Agent and Promtail) as a DaemonSet, or Promtail if already deployed.
- Grafana for querying, with a Loki data source configured.
Step 1 — Understand Streams and Cardinality
Loki groups log lines into streams, one per unique combination of label values. Each active stream costs memory in the ingesters and a chunk in storage. The index covers labels only; everything else is scanned. So labels must be few, bounded, and useful for narrowing queries — and high-cardinality identifiers must not be labels.
Good labels (bounded, narrow queries):
namespace, app/service, container, env, level, queue, job_name
-> e.g. 40 services x 3 envs x 4 levels x ~20 queues = a few thousand streams at most
Never labels (unbounded):
job_id, order_id, customer_id, trace_id, pod name if pods churn fast
-> one stream per job: millions of streams, ingesters fall over
job_name and queue are bounded by your codebase, so they are safe and extremely useful: most job investigations start with "this job type" or "this queue". Pod name is borderline — label it only if pods are long-lived, and prefer app for filtering.
Step 2 — Collect with Alloy and Promote Only Bounded Fields
Alloy discovers pods, tails their container logs, parses JSON, and promotes a chosen set of fields to labels. Everything else stays in the line.
// alloy/config.alloy
discovery.kubernetes "pods" { role = "pod" }
discovery.relabel "workers" {
targets = discovery.kubernetes.pods.targets
rule { source_labels = ["__meta_kubernetes_pod_label_tier"] regex = "worker" action = "keep" }
rule { source_labels = ["__meta_kubernetes_namespace"] target_label = "namespace" }
rule { source_labels = ["__meta_kubernetes_pod_label_app"] target_label = "service" }
rule { source_labels = ["__meta_kubernetes_pod_container_name"] target_label = "container" }
}
loki.source.kubernetes "workers" {
targets = discovery.relabel.workers.output
forward_to = [loki.process.json.receiver]
}
loki.process "json" {
stage.json {
expressions = { level = "level", queue = "queue", job_name = "job_name",
job_id = "job_id", trace_id = "trace_id" }
}
stage.labels { values = { level = "", queue = "", job_name = "" } } // bounded only
stage.structured_metadata { values = { job_id = "", trace_id = "" } } // indexed-ish, not streams
forward_to = [loki.write.default.receiver]
}
loki.write "default" {
endpoint { url = "http://loki-gateway.monitoring.svc/loki/api/v1/push" }
}
Structured metadata (Loki 3.x) attaches key-value pairs to each line without creating streams. Filtering on it is much faster than parsing JSON in every line, which makes job_id and trace_id lookups cheap while keeping cardinality out of the index.
Step 3 — Set Limits and Retention per Tenant
Protect Loki from the next labelling mistake with explicit limits, and set retention to match how long job incidents are actually investigated.
# loki values.yaml (excerpt)
loki:
limits_config:
max_global_streams_per_user: 20000 # fails fast if someone adds job_id as a label
max_label_names_per_series: 12
ingestion_rate_mb: 20
ingestion_burst_size_mb: 40
allow_structured_metadata: true
retention_period: 336h # 14 days of worker logs
compactor:
retention_enabled: true
delete_request_store: s3
schemaConfig:
configs:
- from: "2026-01-01"
store: tsdb
object_store: s3
schema: v13
index: { prefix: index_, period: 24h }
Different retention per stream is possible with retention_stream rules — for example 30 days for job_name=~"charge.*|payout.*" and 7 days for high-volume analytics jobs. The volume-reduction techniques in log sampling for high-volume queues reduce what reaches Loki in the first place.
Step 4 — Write the Queries Incidents Need
With labels narrowing the streams and structured metadata carrying ids, the standard investigation queries are fast.
# Everything for one job, across attempts (structured metadata filter: no JSON parsing)
{namespace="prod", job_name="send_confirmation"} | job_id="4f1c9e2a-7b3d"
# Everything for one business entity (JSON field in the line)
{namespace="prod", service="notifications"} | json | order_id="8812"
# Final failures per job type over time (for a Grafana panel)
sum by (job_name) (count_over_time({namespace="prod", level="error"} | json | final="true" [5m]))
# Slowest completions in the last hour
topk(20, max_over_time({namespace="prod", job_name="build_export"} | json
| msg="job succeeded" | unwrap duration_ms [1h]) by (job_id))
# Everything linked to one trace
{namespace="prod"} | trace_id="7a1f3c9e4b2d8f60a1c3e5f7b9d1e3a5"
The order of operations inside a query determines its cost. Loki first uses the stream selector to pick which chunks to read at all; then line filters and structured-metadata filters discard lines cheaply; only then do parsers such as | json run, on whatever is left. A query that starts with a broad selector and parses JSON before filtering reads and parses everything; the same question with a narrow selector and an early metadata filter touches a tiny fraction of the data.
The first query is the one to make instantaneous: put job_name in the stream selector whenever you know it, because it cuts the scanned data to one job type before the metadata filter runs.
Step 5 — Link Alerts and Traces to Logs in Grafana
Configure the Loki data source so trace_id values become links to Tempo (or another tracing backend), and add data links from metric panels to pre-filtered log queries.
# grafana provisioning: datasources/loki.yaml
apiVersion: 1
datasources:
- name: Loki
type: loki
url: http://loki-gateway.monitoring.svc
jsonData:
derivedFields:
- name: TraceID
matcherType: label # structured metadata key
matcherRegex: trace_id
datasourceUid: tempo
url: "$${__value.raw}"
On the job-failures panel of the queue dashboard, add a data link to /explore with the query {namespace="prod", job_name="${__field.labels.job_name}", level="error"} | json | final="true". An engineer who sees the failure rate spike clicks once and sees the failing jobs, with the trace link on each line. The dashboard side is covered in Grafana Dashboards for Queues.
Step 6 — Alert on Log Patterns Loki Sees First
Some conditions show up in logs before metrics: a new error class, or a specific message such as "payload schema version unknown". Loki's ruler evaluates LogQL alert rules and sends them to Alertmanager.
# loki ruler rule group
groups:
- name: worker-log-alerts
rules:
- alert: JobsFailingPermanently
expr: |
sum by (job_name) (count_over_time({namespace="prod", level="error"} | json | final="true" [10m])) > 5
for: 5m
labels: { severity: page }
annotations:
summary: "{{ $labels.job_name }}: jobs giving up permanently"
- alert: UnknownPayloadSchema
expr: sum(count_over_time({namespace="prod"} |= "payload schema version unknown" [5m])) > 0
labels: { severity: ticket }
Keep log-based alerts for patterns metrics cannot express; for rates and latencies, prefer the Prometheus alerts in alerting on queue backlog with Prometheus, which are cheaper to evaluate.
Verification
# Stream count stays bounded after rollout
logcli series '{namespace="prod"}' --analyze-labels | head -20
# A job lookup by id returns in about a second
time logcli query '{namespace="prod", job_name="send_confirmation"} | job_id="4f1c9e2a-7b3d"' --since=24h
--analyze-labels lists each label and its number of values; any label with thousands of values (other than none) is a mistake to fix before it reaches the stream limit. Watch loki_ingester_memory_streams in Loki's own metrics after each change to the Alloy pipeline.
Gotchas & Edge Cases
Multi-line output breaks parsing. A worker that prints raw stack traces produces lines that fail stage.json. Fix it at the source with a JSON exception field; do not rely on multiline stages as a workaround for worker logs.
Level label mismatches. Libraries emit warn, warning, and WARNING. Normalise in the pipeline (stage.template or a stage.replace) so the level label has four values, not ten.
Out-of-order timestamps. Loki accepts somewhat out-of-order writes per stream, but workers with skewed clocks can still be rejected. Use the container runtime timestamp or ensure NTP on nodes.
Query cost of | json over wide ranges. Parsing every line for a week of all workers is slow. Narrow with labels and structured metadata first; parse last.
FAQ
Promtail or Alloy? Alloy for new deployments — it replaces Promtail and the Grafana Agent and adds OpenTelemetry pipelines. The labelling principles are identical in both.
Should job_id be structured metadata or stay in the line? Structured metadata, if your Loki version supports it. It makes id lookups much faster without adding streams; the value can stay in the JSON line too.
How does Loki compare with Elasticsearch for worker logs? Loki is far cheaper to store and operate because it indexes little; Elasticsearch answers arbitrary field queries faster. For job investigations that start from a job type or an id, Loki with good labels and structured metadata is usually fast enough.
Related
- Structured Logging for Workers — the log schema this pipeline carries.
- Log Sampling for High-Volume Queues — reduce volume before it reaches Loki.
- Grafana Dashboards for Queues — dashboards that link into these queries.
- Distributed Tracing for Async Jobs — the traces linked from each log line.