Cutting Worker Costs with Spot Instances

Queue workers are among the best workloads for spot and preemptible instances — stateless, horizontally scaled, and built around redelivery — and this guide shows how to move them there without turning interruptions into incidents, as part of Capacity Planning for Job Queues in Observability & Monitoring for Job Queues. Spot capacity is typically 60–90% cheaper than on-demand; the price is that the cloud provider can reclaim a node with about two minutes' notice (AWS, Azure) or 30 seconds (Google Cloud).

Problem Statement

A media company runs 120 worker nodes on on-demand instances for thumbnail generation, transcoding, and metadata extraction; compute is its largest infrastructure cost. Finance asks for a 40% reduction. The engineering concern is interruptions: transcodes run up to 20 minutes, and an early experiment that moved a transcoding pool to spot caused jobs to restart repeatedly, some never finishing during a period of heavy reclamation. You want most worker capacity on spot, interruptions handled so that short jobs finish and long jobs checkpoint, a guaranteed on-demand floor for critical queues, and a measured cost saving net of redelivered work.

Prerequisites

  • Kubernetes with node groups or node pools that can use spot capacity (EKS managed node groups or Karpenter, GKE spot node pools, AKS spot pools).
  • Workers that already shut down gracefully on SIGTERM and whose jobs are idempotent — see graceful shutdown for Go workers or handling SIGTERM in Celery workers.
  • An interruption handler: AWS Node Termination Handler (or Karpenter's native handling), GKE's built-in graceful node shutdown, or Azure's scheduled events handling.
  • Job duration percentiles per job type.

Step 1 — Classify Jobs by Interruption Tolerance

Not every job belongs on spot. The deciding factor is how much work an interruption throws away compared with the notice period.

Job type p99 duration Idempotent Checkpointable Placement
Thumbnails 4 s Yes Not needed Spot
Metadata extraction 30 s Yes Not needed Spot
Transcode 20 min Yes Yes (per segment) Spot with checkpoints
Payment reconciliation 2 min Yes, with keys No On-demand
Customer-facing exports 5 min Yes Partially On-demand floor + spot burst

Jobs that finish well within the notice period lose nothing: the worker stops fetching on the notice and completes what it has. Long jobs lose everything since their last checkpoint, so they need checkpoints to qualify. Jobs whose redelivery is expensive or visible (payments, anything with a strict latency objective) keep an on-demand floor.

Duration versus the notice period A timeline marks the two-minute interruption notice. Thumbnail and metadata jobs finish well inside it and run on spot. Transcodes run twenty minutes but checkpoint per segment, so an interruption loses at most one segment, and they also run on spot. Payment reconciliation stays on on-demand because redelivery is costly and latency matters. Where each job type runs 2-min notice thumbnails, metadata: spot transcode 20 min, checkpoint per segment: spot payments: on-demand The question is not duration alone but work lost per interruption.

Step 2 — Create Separate Spot and On-Demand Pools

Separate node pools let you place each queue's workers deliberately. With Karpenter on EKS, two NodePools express it cleanly; the spot pool allows many instance types so the provider has more capacity pools to draw from, which reduces interruption rates.

apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: workers-spot }
spec:
  template:
    metadata: { labels: { capacity: spot } }
    spec:
      requirements:
        - { key: karpenter.sh/capacity-type, operator: In, values: ["spot"] }
        - { key: karpenter.k8s.aws/instance-category, operator: In, values: ["c", "m", "r"] }
        - { key: karpenter.k8s.aws/instance-generation, operator: Gt, values: ["5"] }
        - { key: kubernetes.io/arch, operator: In, values: ["amd64", "arm64"] }   # widen the pool
      taints: [{ key: capacity, value: spot, effect: NoSchedule }]
  disruption: { consolidationPolicy: WhenEmptyOrUnderutilized }
---
apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: workers-ondemand }
spec:
  template:
    metadata: { labels: { capacity: on-demand } }
    spec:
      requirements:
        - { key: karpenter.sh/capacity-type, operator: In, values: ["on-demand"] }
  limits: { cpu: "200" }                   # the floor is deliberate and bounded

Diversifying instance types and architectures matters more than any other spot setting: a pool restricted to one instance type in one zone is reclaimed together when that capacity pool tightens.

Step 3 — Place Workers with Tolerations and a Floor

Spot-tolerant workers get a toleration and a preference for spot; critical workers are pinned to on-demand. For queues that need both, run two deployments — a small on-demand floor and a spot deployment that scales with backlog.

# Thumbnail workers: spot only
spec:
  template:
    spec:
      tolerations: [{ key: capacity, value: spot, effect: NoSchedule }]
      nodeSelector: { capacity: spot }
      terminationGracePeriodSeconds: 110          # < 120 s notice minus handler overhead
---
# Export workers: on-demand floor (2 replicas) ...
spec:
  replicas: 2
  template:
    spec:
      nodeSelector: { capacity: on-demand }
---
# ... plus a spot burst deployment scaled by KEDA on the same queue
spec:
  template:
    spec:
      tolerations: [{ key: capacity, value: spot, effect: NoSchedule }]
      nodeSelector: { capacity: spot }

The floor guarantees progress even during a wave of reclamation; the burst deployment provides cheap capacity most of the time. Scaling the burst deployment on queue depth is covered in scaling workers with KEDA on queue length.

Step 4 — Turn the Interruption Notice into a Graceful Drain

The notice reaches the node, not your process. An interruption handler (Karpenter or AWS Node Termination Handler) watches for it, cordons the node, and evicts pods — which sends SIGTERM to your worker with the pod's grace period. The worker must stop fetching immediately and finish or hand back in-flight jobs within the remaining time.

# celery worker settings for spot nodes
worker_prefetch_multiplier = 1        # don't hold messages you may not get to finish
task_acks_late = True                 # unacked work returns to the queue if we're killed
task_reject_on_worker_lost = True
worker_cancel_long_running_tasks_on_connection_loss = True

# In long tasks: check a shutdown flag between checkpoints
import signal, threading
shutting_down = threading.Event()
signal.signal(signal.SIGTERM, lambda *_: shutting_down.set())

@app.task(bind=True, acks_late=True)
def transcode(self, video_id: str):
    for seg in pending_segments(video_id):             # resumes from the last checkpoint
        if shutting_down.is_set():
            raise self.retry(countdown=5)              # hand back promptly, keep progress
        encode_segment(video_id, seg)
        mark_segment_done(video_id, seg)               # durable checkpoint
    finalize(video_id)

Set terminationGracePeriodSeconds slightly below the notice period so the pod exits before the node disappears. On GKE, where the notice is about 30 seconds, only short jobs and aggressively checkpointed ones fit; plan the placement table accordingly.

Two minutes, used well At time zero the provider issues the interruption notice. Within seconds the handler cordons the node and evicts worker pods, sending SIGTERM. Workers stop fetching new jobs immediately. Short jobs finish within seconds; long jobs stop at their next checkpoint and are handed back to the queue. Pods exit by 110 seconds, before the node is reclaimed at 120 seconds. From notice to reclaim notice t=0 cordon + SIGTERM stop fetch, finish short jobs long jobs checkpoint, retry exit ≤110 s node reclaimed at 120 s; nothing still running loses more than one segment Prefetched but unstarted messages must be released too, or they wait out the visibility timeout.

Step 5 — Watch Interruption Rates and Redelivered Work

Savings are real only if interruptions do not waste much work. Track interruptions, redelivered jobs, and work thrown away, per queue.

# Interruptions per hour (from Karpenter or the termination handler)
sum(increase(karpenter_nodes_terminated_total{reason="interruption"}[1h]))

# Share of jobs that are redeliveries (retries caused by worker loss)
sum(rate(jobs_redelivered_total{cause="worker_lost"}[1h])) / sum(rate(jobs_completed_total[1h]))

# Wasted compute: seconds of work discarded by interruptions
sum(increase(job_discarded_work_seconds_total[1h]))

If redelivery rates for a queue exceed a few percent, the job type is too long or poorly checkpointed for the current interruption rate — move it to on-demand or add checkpoints. A sudden rise across all pools means the chosen instance types are under pressure; widen the diversification in Step 2.

Step 6 — Compute the Net Saving

Compare cost per thousand completed jobs before and after, including the redelivered work.

def cost_per_1k_jobs(node_hours: float, price_per_hour: float, jobs_completed: int) -> float:
    return node_hours * price_per_hour / jobs_completed * 1000

before = cost_per_1k_jobs(node_hours=120 * 720, price_per_hour=0.34, jobs_completed=410_000_000)
after  = cost_per_1k_jobs(node_hours=24 * 720, price_per_hour=0.34, jobs_completed=410_000_000) \
       + cost_per_1k_jobs(node_hours=104 * 720, price_per_hour=0.11, jobs_completed=410_000_000)
print(f"{before:.4f} -> {after:.4f} per 1k jobs, saving {(1 - after / before):.0%}")
# 0.0716 -> 0.0344 per 1k jobs, saving 52%
Cost per thousand jobs, before and after Before, 120 on-demand nodes cost about 0.072 dollars per thousand completed jobs. After, a 24-node on-demand floor plus 104 spot nodes cost about 0.034 dollars per thousand jobs, including the extra capacity needed for redelivered work, a saving of about 52 percent. $ per 1,000 completed jobs all on-demand 0.072: 120 on-demand nodes floor + spot floor 104 spot 0.034: -52% Includes the extra spot nodes that absorb redelivered work and churn.

The spot pool needed slightly more nodes (104 spot plus a 24-node on-demand floor, instead of 120) to cover redelivered work and churn, and the saving still exceeded the 40% target. Revisit the calculation quarterly: spot prices and interruption rates change.

Verification

Before moving production queues, run an interruption drill in staging: use the provider's fault-injection tooling (AWS FIS aws:ec2:send-spot-instance-interruptions) or simply drain spot nodes with the same grace period, under load.

aws fis start-experiment --experiment-template-id "$SPOT_INTERRUPT_TEMPLATE"
# Then check: no job failed permanently, redeliveries only for in-flight jobs, transcodes resumed

Pass criteria: zero jobs lost, redeliveries limited to jobs in flight at the notice, long jobs resumed from checkpoints rather than restarting, and queue wait staying within its objective thanks to the on-demand floor.

Gotchas & Edge Cases

Prefetched messages. A worker holding prefetched messages releases them only when its connection closes or the visibility timeout expires. With prefetch_multiplier=1 and late acks, loss of a node costs at most one message per slot.

Zone concentration. Spot capacity can vanish across a whole zone. Spread spot pools across zones and keep the on-demand floor multi-zone too.

Daemon overhead. Every node runs log agents and exporters; many small spot nodes multiply that overhead. Prefer medium-sized instances.

Long grace periods. A pod with a 600-second grace period on a node reclaimed at 120 seconds simply dies. Keep grace periods within the notice window on spot pools.

FAQ

Are spot interruptions frequent? It varies by instance type, region, and time; diversified pools commonly see a few percent of nodes interrupted per day. Measure your own rate — it drives the placement table.

Should the broker or database run on spot? No. Stateful components with failover costs belong on on-demand or reserved capacity; only stateless workers should run on spot.

What about Savings Plans or reserved instances? Use them for the on-demand floor, which runs all the time. Spot covers the variable part above the floor.

Related