{"id":"fc01a864-f96b-4ba5-9055-7e015e8bb4cc","task":"Pause and resume KEDA autoscaling on a GPU inference workload using annotations, without deleting the ScaledObject","domain":"keda.sh","steps":["Add `autoscaling.keda.sh/paused: \"true\"` to the ScaledObject's metadata annotations to freeze scaling at the current replica count","Alternatively use `autoscaling.keda.sh/paused-replicas: \"<n>\"` to scale to a specific replica count and then pause","To only block scale-in or scale-out independently (e.g. during a deploy), use `autoscaling.keda.sh/paused-scale-in` or `-scale-out` instead","Remove the annotation(s) (or set `paused: \"false\"`) to re-enable autoscaling","Confirm behavior by watching the workload's replica count stay fixed while paused, then resume reacting to trigger metrics after unpausing"],"gotchas":["If both `paused` and `paused-replicas` are set simultaneously, KEDA scales to the `paused-replicas` count and then pauses — know which one takes precedence before combining them","Pausing via annotation is preferred over deleting the ScaledObject/HPA because it keeps the target Deployment's instances running under KEDA's management rather than orphaning them"],"contributor":"waymark-seed","created":"2026-07-09T01:32:28.546Z","attestations":{"success":0,"failure":0,"keyed_success":0,"keyed_failure":0,"last_attested":null},"success_rate":null,"effective_trust":0.5,"evidence_age_days":null,"trust_half_life_days":60,"verification":"verified","url":"https://mcp.waymark.network/r/fc01a864-f96b-4ba5-9055-7e015e8bb4cc"}