Home / Insights / Kubernetes
Kubernetes

Scale from the queue, not the clock: Karpenter and KEDA on Amazon EKS.

For bursty work such as security scans and data jobs, scale workers from queue depth and let nodes follow the pods. What we configure, with examples.

Amazon EKSKarpenterKEDAAmazon SQS

Some workloads are idle most of the day and then all arrive at once: a customer connects fifty repositories to a security scanner, or a nightly export drops a few hundred thousand records into a queue. Sizing for the peak wastes money all day. Sizing for the average makes the burst wait.

On two platforms with exactly this shape, an application-security SaaS with 94 scanner tools and an ad-tech identity platform with batch data jobs, we used the same approach: scale workers from the queue, and let nodes follow the workers.

Two autoscalers, two jobs

  • KEDA decides how many workers should exist, based on the work waiting: messages in an Amazon SQS queue, items in a Redis list.
  • Karpenter decides which nodes should exist to run the pods that KEDA created, and removes them when they are no longer needed.

Neither looks at the clock. Nothing runs at 3 a.m. because a schedule said so; it runs because there is work.

Workers from queue depth

For work that comes in discrete units, a ScaledJob starts one job per batch of messages and lets each job exit when it is done:

apiVersion: keda.sh/v1alpha1
kind: ScaledJob
metadata:
  name: data-export
spec:
  jobTargetRef:
    template:
      spec:
        restartPolicy: Never
        containers:
          - name: worker
            image: <account>.dkr.ecr.eu-west-1.amazonaws.com/data-export:1.4.2
  pollingInterval: 15
  maxReplicaCount: 50
  triggers:
    - type: aws-sqs-queue
      authenticationRef:
        name: keda-aws
      metadata:
        queueURL: https://sqs.eu-west-1.amazonaws.com/<account>/data-export
        queueLength: "5"
        awsRegion: eu-west-1

KEDA reads the queue through an IAM role for its service account, so there are no keys in the cluster. On the security platform the same pattern runs on Redis queues, with one scaled job per tool family.

Nodes per workload role

The mistake we see most often is one big node group for everything, where a burst of scan jobs evicts the ingress controller. We give each kind of workload its own Karpenter NodePool with a taint, so it can only land on capacity meant for it:

apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: scanners
spec:
  template:
    spec:
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: default
      taints:
        - key: dedicated
          value: scanner
          effect: NoSchedule
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]
        - key: karpenter.k8s.aws/instance-category
          operator: In
          values: ["c", "m"]
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m
  limits:
    cpu: "1000"

The security platform ended up with eleven node pools: scanners, workers, control services, add-ons and several variants. Add-ons run on on-demand capacity only; interruptible spot capacity is reserved for work that can be retried.

What we tune, and why

  • Queue length per job. Small batches start faster; large batches waste less time on container start-up. Measure the job's start-up cost first.
  • Maximum replicas. A ceiling protects downstream systems such as databases and third-party APIs from your own burst.
  • CPU limits per NodePool. The cost ceiling you can explain to finance.
  • Consolidation. Removing empty or underused nodes is where the savings come from; give it a short delay so nodes are not recycled between two bursts.
  • A fallback. On one cluster we kept Cluster Autoscaler installed but disabled during the Karpenter rollout, so going back was a values change.

When not to do this

If your workload is steady, a well-sized managed node group is simpler and just as cheap. Queue-driven scaling earns its keep when the gap between the quiet hours and the busy minutes is large, which is exactly when a fixed fleet hurts the most.

What’s next for
your business?

Let’s talk