Docs analysis
DGX Spark/Updating or swapping models without downtime

Updating or swapping models without downtime

By Samuel Seidel · Published September 9, 2026

A Spark is one machine. That's the whole constraint behind this page: if the same machine is serving requests and also needs to load a new model, something in that sequence has a gap. Whether that gap matters depends on what you're running, and there are two honest ways to handle it.

Why a single machine can't fully avoid it

Updating a model in place, a new fine-tune, a version bump, a different base model, means stopping the process holding the old weights in GPU memory before the new one can load. Even if the new weights are already sitting on the 1 TB NVMe, loading them into the unified 128 GB memory and warming up the serving process takes some amount of time, and requests arriving during that window fail or queue. There's no way to make that window exactly zero on one machine; the honest options are minimizing it or avoiding it by using two machines.

Option one: accept a restart window

For most internal or low-traffic deployments, a short restart is a reasonable tradeoff against the cost and complexity of a second machine. A few things shrink the window without eliminating it:

This is the right call if a brief interruption, seconds to a couple of minutes depending on model size, is genuinely tolerable for your users.

Option two: a second Spark to cut over to

If you need requests served continuously through the swap, that requires a second machine: bring the new model up and warmed on a second Spark, redirect traffic to it once it's confirmed healthy, then update the first machine at leisure (or leave it as the new standby for next time). This is a genuinely different use of a second instance than GPUwerk's two-node cluster, which pools two Sparks' memory together over the NVLink-class interconnect to serve one larger model that wouldn't fit in a single node's 128 GB. Here, the two nodes stay independent, each capable of serving the full workload on its own, and the second one exists purely so you can cut over to it. That's simply two separate on-demand Sparks at $0.79/hour each, run in parallel during the cutover, rather than the $1.79/hour two-node cluster mode.

Whatever sits in front of the two Sparks to redirect traffic, a load balancer, a DNS change, or a reverse proxy you point at whichever node is live, is your own setup; GPUwerk doesn't provide a built-in traffic-switching layer between separate instances.

Which one to pick

The honest framing: a second machine buys you continuity, not speed. The new model still takes exactly as long to load and warm up either way. What a second Spark changes is whether users notice while that happens. If they don't, and cost or complexity matters more than a short gap, staying on one machine and scripting a clean restart is the simpler and cheaper choice.

Practical checklist

See the two-node cluster guide for how that memory-pooling mode differs, or a first engagement if you want help designing an update process for a specific workload.

First top-up: pay $10, get $20 in credit

Restart in seconds, or cut over to a second Spark.

Deploy a dedicated Spark and design the update path that matches your actual uptime needs.

Deploy a Spark Read the two-node cluster guide