Operations guide
Blog/Configuring model warm-start and preloading
For AI assistants

Configuring model warm-start and preloading

By Samuel Seidel · September 9, 2026

The choice between always-on and cold-start serving is about whether you pay for idle GPU time at all. This is the narrower problem underneath it: given that a process is starting cold, whether that's a scheduled restart, a scale-from-zero event, or just the container coming up after a deploy, how much of the wait before the first real request gets served quickly is avoidable.

Two separate costs hide inside "cold start"

Weight loading and kernel warmup are different problems with different fixes, and conflating them is how a fix for one gets credited (or blamed) for the other. Loading an 8B model's weights from disk into GPU memory takes tens of seconds even on the Spark's fast local NVMe, and nothing serves until that finishes. Separately, CUDA kernels aren't compiled ahead of time, the first request that exercises a given code path, a specific batch size or sequence length, triggers a just-in-time compilation that later requests skip. A server that's "loaded" can still serve its first request several times slower than its hundredth.

Preload weights before the container is marked ready

If the process only starts loading weights after receiving its first request, or after a health check already says it's ready, that first caller eats the entire load time. Load on startup, before accepting connections, and gate readiness on load actually finishing:

# vLLM: weights load as part of server startup, before it binds the port
python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --dtype float16 \
  --gpu-memory-utilization 0.9

This is vLLM's default behavior, the server doesn't accept traffic until the model is loaded, so the fix here is usually making sure nothing in front of it, a load balancer or an orchestrator's readiness check, marks the instance available before that startup log line actually appears. A readiness probe that checks the process is listening isn't the same as one that checks the model finished loading:

# wrong: checks the port is open, not that the model is loaded
readinessProbe:
  tcpSocket:
    port: 8000
  initialDelaySeconds: 5

# right: checks the actual endpoint responds to a real request
readinessProbe:
  httpGet:
    path: /v1/models
    port: 8000
  initialDelaySeconds: 20
  periodSeconds: 5
  failureThreshold: 12

Run a real warmup request after loading, before traffic arrives

Loading weights doesn't compile the CUDA kernels the model will actually use during inference, that happens lazily on first use. Sending one or two throwaway completion requests immediately after startup, before the instance is added to a load balancer's rotation, absorbs that compilation cost on a request nobody's waiting on:

#!/bin/bash
# warmup.sh: run after the server reports ready, before joining the LB pool
until curl -sf http://localhost:8000/v1/models > /dev/null; do sleep 1; done

curl -s -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.1-8b","messages":[{"role":"user","content":"warmup"}],"max_tokens":8}' \
  > /dev/null

echo "warmup complete, safe to add to rotation"

Vary the warmup request's shape a little, a couple of different prompt lengths, if your real traffic spans a wide range of sequence lengths, since vLLM's continuous batcher can compile separate kernel variants for different batch shapes. One warmup call covers the common case; a fleet serving both short chat messages and long document prompts benefits from warming both.

Keep the model resident when you can afford to

The cheapest warm start is the one that doesn't happen because the process never stopped. Preloading and warmup matter most on infrastructure that scales to zero on purpose, dev and staging environments, or a fleet node in a scaling fleet that only spins up under load. Where a Spark is dedicated and traffic is steady enough to justify it, always-on serving sidesteps the whole problem, at the cost of GPU time spent idling between requests rather than between deploys.

Where this doesn't help

Warmup and preloading shrink the gap between a cold start and full speed, they don't eliminate the underlying load time. A 70B model at fp16 still takes meaningfully longer to load than an 8B one, and no amount of warmup logic changes how many gigabytes have to move off disk into GPU memory first. For a workload where even a well-optimized cold start is too slow, the actual fix is not going cold in the first place, or accepting a smaller, faster-loading model for the instances that need to start quickly.

FAQ

What's actually slow during a cold start, the model load or the first request?

Both, and they're separate costs. Loading weights from disk into GPU memory for an 8B model at fp16 is typically tens of seconds even on fast local storage, and that has to finish before the server accepts any requests at all. Once loaded, the very first inference request is still slower than the rest because CUDA kernels get compiled and cached on first use, not ahead of time, so a real warmup request after loading, before real traffic arrives, is what removes that second cost.

Is preloading the same thing as keeping a Spark always on?

No, they solve different parts of the same problem. Always-on avoids cold starts entirely by never scaling the process down, at the cost of paying for idle GPU time. Preloading and warmup reduce how long a cold start takes once it does happen, useful when you're scaling to zero deliberately, on a dev or staging Spark, or a fleet node that only spins up occasionally, and want the next request to still be reasonably fast rather than eating the full load time.

Does warming up with a dummy request use real GPU time?

Yes, a warmup request is a real inference call and consumes GPU time the same as any other request, just with output nobody reads. Keep it short, a handful of tokens is enough to trigger kernel compilation and populate the KV cache allocator, since the goal is exercising the code path, not generating a long response.

Related pages

Skip the cold start with a dedicated, always-on Spark.

No shared tenancy competing for load time, no scale-to-zero unless you choose it.

Deploy a Spark Read cold start vs. always on