Autoscale a GPU inference deployment with KEDA based on external queue length

domain: keda.sh · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗

Steps

  1. Define a ScaledObject with apiVersion keda.sh/v1alpha1, spec.scaleTargetRef pointing at your GPU inference Deployment
  2. Set spec.minReplicaCount and spec.maxReplicaCount to bound the scale range
  3. Add a trigger such as type: rabbitmq with metadata queueName, mode: QueueLength, and a value threshold (or a prometheus trigger with serverAddress and query)
  4. Tune spec.pollingInterval (default 30s) and spec.cooldownPeriod (default 300s) to match how quickly the queue and GPU workload change
  5. Apply the ScaledObject and confirm KEDA creates and manages the underlying HPA for the target Deployment

Known gotchas

Related routes

Configure KEDA to autoscale GPU inference pods on Kubernetes using NVIDIA DCGM Exporter metrics
keda.sh · 6 steps · unrated
Create a KEDA ScaledObject to autoscale a GPU inference Deployment based on a custom metrics trigger
keda.sh · 5 steps · unrated
Pause and resume KEDA autoscaling on a GPU inference workload using annotations, without deleting the ScaledObject
keda.sh · 5 steps · unrated

Give your agent this knowledge — and 15,500+ more routes

One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus: claude mcp add --transport http waymark https://mcp.waymark.network/mcp

Need this verified for your stack — or a route we don't have yet?

We author + individually verify a route for your exact task within 24h. Custom route — $25 · Teams: Pilot — $750/mo · all plans