When N = 0, the current WD versions can still get scaled down to 0, even when the task queue is genuinely busy
KEDA's temporal scaler measures task queue backlog depth and not the throughput
In the reporter's environment, a worker that's actually busy and is keeping up, reads ~0 tasks at almost every poll
As a result, KEDA scales it down to 0, once the trigger has read inactive for a continuous cooldownPeriod (5m by default). Any single poll in between that sees a non-zero backlog resets the cool down clock
So a queue with steady, bursty traffic still sizes the k8s deployment to 0 if every burst drains between the polls. i.e. KEDA never sees the tasks getting processed
The next task then faces a cold start of up to one pollingInterval to notice the backlog, plus pod scheduling and worker startup latency
When N = 0, the current WD versions can still get scaled down to 0, even when the task queue is genuinely busy
KEDA's temporal scaler measures task queue backlog depth and not the throughput
In the reporter's environment, a worker that's actually busy and is keeping up, reads ~0 tasks at almost every poll
As a result, KEDA scales it down to 0, once the trigger has read inactive for a continuous cooldownPeriod (5m by default). Any single poll in between that sees a non-zero backlog resets the cool down clock
So a queue with steady, bursty traffic still sizes the k8s deployment to 0 if every burst drains between the polls. i.e. KEDA never sees the tasks getting processed
The next task then faces a cold start of up to one pollingInterval to notice the backlog, plus pod scheduling and worker startup latency