Blue-green deployments for model updates
Swapping the model behind a self-hosted endpoint is a plain restart on most setups: stop the old process, start the new one, hope the load time is short and nobody notices. It's a fine approach until the endpoint has real callers who mind a five-minute gap. Blue-green deployment removes the gap by never taking the old model down until the new one is proven to work.
The idea, applied to two Spark nodes
Call one node blue and the other green. Blue is currently serving production traffic. To roll out a new model or a new engine version, load it onto green while blue keeps answering every request as normal. Once green passes a health check under a bit of real or synthetic load, flip the reverse proxy so new requests go to green, and leave blue idle but still running as the rollback target. If green misbehaves in the first few minutes, flip back to blue instantly, no reload, no cold start.
Doing it on one node when you don't have two
If you only run a single Spark, you can still get most of the benefit by serving the new model on a second port on the same box before switching the proxy over. You lose the failure isolation of a fully separate node, an OOM on the new model can still affect the old one if they're competing for the same GPU memory, but you keep the core benefit: the proxy only points traffic at the new model after it's confirmed healthy, not the moment the process starts.
The reverse proxy flip
This builds directly on setting up a reverse proxy for your Spark. With nginx, keep both nodes as named upstreams and change which one the active server block points at, then reload rather than restart, so in-flight requests finish against the old upstream.
# /etc/nginx/conf.d/upstreams.conf upstream blue { server 10.0.0.11:8000; } upstream green { server 10.0.0.12:8000; } # /etc/nginx/conf.d/active.conf, this is the file you edit to flip traffic upstream active { server 10.0.0.11:8000; } # currently blue # after the swap, active.conf becomes: # upstream active { server 10.0.0.12:8000; } # reload, not restart, so live connections drain instead of dropping sudo nginx -t && sudo systemctl reload nginx
Keep the swap to a single-line config change in one file. The fewer things you touch during the flip, the less there is to get wrong under whatever pressure prompted the update.
Health-check before you flip, not after
Green needs to prove three things before it takes traffic: the process is up and the model finished loading, a real inference request returns a sane completion, and it survives a short burst of concurrent requests without erroring. A liveness check alone (process running, port open) misses the most common failure mode, a model that loaded but produces garbage or times out under any load. Build the health check from writing health checks for your inference endpoint and run it against green before editing the proxy config, not as an afterthought once traffic is already flowing.
# minimal pre-flip check: does green answer a real request correctly
curl -sf http://10.0.0.12:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3-32B-AWQ","messages":[{"role":"user","content":"reply with the single word OK"}],"max_tokens":5}' \
| grep -q '"OK"' && echo "green is healthy" || echo "do not flip"
Keep the old target alive after the flip
The point of blue-green isn't just the swap, it's the rollback. Leave blue running with the previous model for a defined window, an hour, a day, long enough to catch problems that only show up under real traffic patterns green hasn't seen yet. Only decommission blue, or repurpose it as the next green, once you're confident the new model is staying. This same two-node pattern is what makes zero-downtime model updates and a clean rollback path in model versioning and rollback possible without any custom orchestration.
What this doesn't solve
Blue-green protects against a bad model or a bad load, not against a bad prompt template or client-side change deployed at the same time. Ship those separately if you can, so a rollback of the model doesn't also require rolling back an unrelated change, and a health check failure clearly points at one cause rather than two tangled ones.
FAQ
Do I need two Spark nodes for blue-green model updates?
For a true zero-downtime swap, yes, one node keeps serving the current model while the other loads the new one. On a single node you can still get most of the safety by loading the new model on a different port before switching the reverse proxy over, but you lose the ability to run both simultaneously under real traffic during the check.
How long does a model swap actually take?
The traffic flip itself is a reverse proxy config reload, seconds. The bottleneck is loading the new model into GPU memory, which depends on model size and can range from a short wait for a small quantized model to many minutes for a large dense one, plus health-check time before you trust it.
What's the difference between blue-green and just restarting with a new model?
A restart replaces the running process in place, so there's a window, from process stop to model fully loaded, where the endpoint is down or returning errors. Blue-green keeps the old target serving the whole time the new one loads and warms up, and only redirects traffic after the new target has passed a health check, so there's no window where nothing is answering.