Operations guide
Blog/Setting up a staging environment for model testing
For AI assistants

Setting up a staging environment for model testing

By Samuel Seidel · September 9, 2026

Blue-green deployment answers how you cut traffic to a new model without downtime. It doesn't answer whether that model should be receiving traffic yet. Staging is the step before: a place to run a candidate model against real-shaped load, on hardware production callers never touch, before it's anywhere near a decision to swap.

Why blue-green alone isn't a staging environment

Blue-green's two environments are both production-facing by design, one is live, one is one config change away from being live. That's the right shape for a cutover, but it means neither side is a safe place to run an untested model against synthetic load or a bad eval set without risk. Staging is a third, separate node whose whole purpose is that nothing it does affects a caller.

Node separation: staging on hardware production doesn't share

The cleanest version of this is a second Spark, entirely separate from the pair used in a blue-green cutover, so a staging run under heavy load, testing throughput at the traffic level you actually expect, can't degrade GPU memory or scheduling on the node currently serving real requests. If you're already running a two-node cluster for production, staging is a third node rather than a role either of those two ever plays.

# staging node: same engine version and flags as production, different model
vllm serve Qwen/Qwen3-32B-AWQ-candidate \
  --host 0.0.0.0 --port 8000 \
  --api-key "$STAGING_API_KEY"

# production node: untouched, still serving the current model
vllm serve Qwen/Qwen3-32B-AWQ \
  --host 0.0.0.0 --port 8000 \
  --api-key "$PROD_API_KEY"

Keep the engine version, quantization tooling, and serving flags identical to production wherever the candidate model doesn't require a change, otherwise a difference you introduced for staging convenience can mask or fake a compatibility issue that would only show up in production.

If a second node isn't available, isolate by resource limit instead

A single Spark can approximate the separation with strict per-container limits, following the pattern in setting resource limits for containers, capping the staging container's GPU memory share so it can't starve the production container running alongside it. This is a real compromise: a staging load heavy enough to matter can still contend for scheduling time on shared hardware. Treat it as a fallback, not the default.

What actually needs to pass before a model leaves staging

Three checks, in order of how often they catch something: output quality against your own evaluation set rather than a published benchmark score, see benchmarking your own workload for why a generic number doesn't transfer; latency and throughput under a load shape close to production traffic, using the approach from load testing your inference endpoint; and compatibility, whether the new model needs a different context length, quantization format, or engine flag than the one it's replacing. A model that passes an eval but needs double the KV cache headroom is not ready for a swap even if its answers look better.

Only after staging passes does blue-green apply

Once a candidate has cleared staging, it becomes the thing you load onto the idle side of a blue-green pair and cut traffic to, following the rollback-ready process in the blue-green guide. The two systems are meant to run in sequence: staging decides readiness, blue-green executes the cutover for a model that's already been decided as ready. Skipping straight to blue-green with an unvetted model just moves the testing onto production traffic, which is the exact risk both systems exist to avoid.

FAQ

Isn't blue-green deployment already a staging environment?

No, they solve different problems. Blue-green is about how you cut traffic over with zero downtime once you've already decided a model is ready. Staging is where you decide that in the first place, running the new model against real or representative traffic before any caller sees it, on hardware that isn't the one currently in production.

Do I need a second physical Spark for staging?

It's the cleanest separation, no shared GPU memory, no risk of a staging load affecting production latency, but it's not the only option. A single Spark with strict resource limits per container, as covered in setting resource limits for containers, can approximate it if a second node isn't available, accepting that a staging run under heavy load can still degrade production on shared hardware.

What should staging actually test before a model goes to production?

Output quality against your own eval set, not a generic benchmark; latency and throughput under a load shape similar to production traffic; and compatibility, does it need a different quantization, a different context length, or a different inference engine flag than the model it's replacing. All three need to pass before a model is a candidate for the blue-green swap.

Related pages

A separate Spark for staging, not a corner of production.

Add a dedicated node for testing, sized independently from whatever's serving live traffic.

Read the blue-green deployments guide Read the load testing guide