Updating or swapping models without downtime
A Spark is one machine. That's the whole constraint behind this page: if the same machine is serving requests and also needs to load a new model, something in that sequence has a gap. Whether that gap matters depends on what you're running, and there are two honest ways to handle it.
Why a single machine can't fully avoid it
Updating a model in place, a new fine-tune, a version bump, a different base model, means stopping the process holding the old weights in GPU memory before the new one can load. Even if the new weights are already sitting on the 1 TB NVMe, loading them into the unified 128 GB memory and warming up the serving process takes some amount of time, and requests arriving during that window fail or queue. There's no way to make that window exactly zero on one machine; the honest options are minimizing it or avoiding it by using two machines.
Option one: accept a restart window
For most internal or low-traffic deployments, a short restart is a reasonable tradeoff against the cost and complexity of a second machine. A few things shrink the window without eliminating it:
- Pre-download the new model's weights to
/workspacebefore the swap, so the restart only pays for loading into memory, not a network transfer. - Script the swap (stop the old process, point the config at the new weights, start the new process) so the window is however long that sequence takes to run, not however long it takes a person to do it by hand.
- Schedule the swap for low-traffic hours if your usage pattern has any, since a two-minute gap at 3am costs less than the same gap at peak use.
This is the right call if a brief interruption, seconds to a couple of minutes depending on model size, is genuinely tolerable for your users.
Option two: a second Spark to cut over to
If you need requests served continuously through the swap, that requires a second machine: bring the new model up and warmed on a second Spark, redirect traffic to it once it's confirmed healthy, then update the first machine at leisure (or leave it as the new standby for next time). This is a genuinely different use of a second instance than GPUwerk's two-node cluster, which pools two Sparks' memory together over the NVLink-class interconnect to serve one larger model that wouldn't fit in a single node's 128 GB. Here, the two nodes stay independent, each capable of serving the full workload on its own, and the second one exists purely so you can cut over to it. That's simply two separate on-demand Sparks at $0.79/hour each, run in parallel during the cutover, rather than the $1.79/hour two-node cluster mode.
Whatever sits in front of the two Sparks to redirect traffic, a load balancer, a DNS change, or a reverse proxy you point at whichever node is live, is your own setup; GPUwerk doesn't provide a built-in traffic-switching layer between separate instances.
Which one to pick
The honest framing: a second machine buys you continuity, not speed. The new model still takes exactly as long to load and warm up either way. What a second Spark changes is whether users notice while that happens. If they don't, and cost or complexity matters more than a short gap, staying on one machine and scripting a clean restart is the simpler and cheaper choice.
Practical checklist
- Pre-download new weights before any swap, so the downtime window is only the memory load, not a network transfer.
- Script the swap end to end if you're doing it on one machine, to keep the window as short and repeatable as possible.
- Use a second Spark only if continuous availability during updates genuinely matters, since it doubles the running cost for the overlap period.
- Don't confuse a cutover pair with the two-node cluster mode, since a cluster pools memory for one bigger model, it doesn't give you two independent standbys.
- Build your own health check and traffic switch between two Sparks, since GPUwerk doesn't provide one.
See the two-node cluster guide for how that memory-pooling mode differs, or a first engagement if you want help designing an update process for a specific workload.