Skip to content

KEDA can scale a worker lane to zero while its model is actively loading #293

Description

@igor-raits

Problem

With generation workers configured for scale-to-zero, KEDA can terminate a worker while SGLang is still downloading or initializing a model. Subsequent requests start the same cold-load cycle again, so clients alternate between provisioning and MODEL_LOADING without the model ever becoming ready.

This is more than the expected first-request MODEL_LOADING response: the worker itself loses the opportunity to finish loading.

Observed behavior

Observed with SIE v0.7.2 on Kubernetes/KEDA using an NVIDIA L4 generation lane:

autoscaling:
  cooldownPeriod: 600
workers:
  pools:
    l4:
      bundles:
        sglang:
          minReplicas: 0
          maxReplicas: 1

A cold request correctly activated the lane. GPU-node provisioning, the 16.1 GB worker image pull, and model initialization consumed the cooldown. SGLang then began loading, but KEDA deleted the pod seconds later because every scaling signal was inactive. There was no OOM, application crash, authentication error, or invalid model configuration.

Typical client responses were:

{"error":{"code":"provisioning","message":"No worker available for bundle 'sglang'. Provisioning in progress.","type":"server_error"}}

followed by:

{"error":{"code":"MODEL_LOADING","message":"Model 'Qwen/Qwen3-4B-Instruct-2507' is still loading; retry later.","type":"server_error"}}

Why the lane becomes inactive

The relevant behavior is unchanged in v0.7.3 and current main (c5d9468):

  1. Gateway pending demand expires after 120 seconds: demand_tracker.rs.
  2. Durable publication clears the earlier demand marker: finish_dispatch_handoff.
  3. For generation, both LoadingStarted and LoadingInProgress produce a terminal MODEL_LOADING response and ACK the JetStream delivery: dispatcher.rs.
  4. Queue depth and active leases therefore become zero. The five worker KEDA triggers cover pending demand, queue depth, active-lease GPUs, warm floor, and rejected-request rate, but no current model-loading state: keda-scaledobject.yaml.

After the cooldown, KEDA sees an idle lane even though the worker is actively loading a model.

Reproduction

  1. Configure an SGLang lane with minReplicas: 0 and a cooldown shorter than the full cold-node/image/model startup path.
  2. Ensure the lane and its model cache are cold.
  3. Send one generation request.
  4. Observe the lane scale to one and return provisioning, then MODEL_LOADING on retry.
  5. Stop client retries or let the pending-demand window expire while the model is loading.
  6. Observe the queued delivery become acknowledged, every KEDA trigger become inactive, and the worker scale to zero before readiness.

Expected behavior

An active model load should keep its lane above zero until the load reaches Ready, Failed, or Cancelled, independently of client retry behavior and the original work item's ACK state.

Proposed durable design

The Python model registry already owns the authoritative loading lifecycle: it inserts into _loading before starting the task and clears it in finally (registry implementation). Propagate that state through the existing health path:

  1. Extend engine PingResponse with loading_models.
  2. Include it in worker-sidecar health heartbeats.
  3. Store it in the gateway worker registry.
  4. During the existing capacity-snapshot reconciliation, emit a complete lane-scoped gauge such as sie.gateway.model_loads_in_progress{pool,machine_profile,bundle}. Emit explicit zero for configured lanes with no live workers and use the existing snapshot freshness guard so dead/stale workers cannot pin a lane up.
  5. Add a sixth worker KEDA trigger with threshold 1.
  6. Publish health promptly when loading state changes, in addition to periodic heartbeats.

This avoids relying on a response counter, a refreshed short TTL, or continued client retries. Keeping the generation delivery unacknowledged would also change the current terminal-response contract and could execute a stale request after the caller retries.

Acceptance criteria

  • An ACKed generation request can have queue depth zero while the lane remains active because one or more models are loading.
  • Loading remains active without client retries.
  • Ready, Failed, and Cancelled all clear the loading state.
  • A stale or dead worker is excluded after the normal heartbeat timeout.
  • Zero workers produces an explicit lane value of zero rather than a stale value or scaler error.
  • The Helm chart renders and validates the sixth freshness-gated KEDA trigger.
  • An integration test covers scale-to-zero demand expiry while model loading remains in progress.

Current workaround

Increasing autoscaling.cooldownPeriod to 3600 and configuring lane-specific preloadModels allowed the observed cold start to complete and was verified with a successful generation request. This is a useful operational mitigation, but it is still a time allowance rather than lifecycle-aware scaling. minReplicas: 1 also avoids the race at the cost of a permanently warm GPU.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions