Migrating an existing self-hosted AI deployment to a DGX Spark
If you're already self-hosting on a workstation with a consumer GPU, or on a rented cloud GPU instance elsewhere, moving to a Spark is not a from-scratch deployment. It's a migration with its own specific gotchas: the memory architecture is different, the quantization that worked on your old box may not be the right one here, and your data needs a plan to move that doesn't depend on the old machine staying up while you test the new one.
Why the move isn't a straight lift-and-shift
A Spark's 128 GB is unified memory, shared between CPU and GPU, not a discrete VRAM pool sitting next to system RAM the way a workstation with an RTX card or a cloud instance with a discrete NVIDIA GPU is built. GPUwerk's DGX Spark vs Mac Studio comparison and the memory-bandwidth arithmetic on the benchmarks page cover why that architecture changes the numbers that actually govern serving speed: decode speed is bandwidth-bound, and a Spark's bandwidth figure is not the same as whatever card you're moving off of. Don't assume a model and quantization that ran well on your existing hardware will hit the same tokens per second here; the two boxes are shaped differently even when the total memory number looks similar on paper.
Step 1: re-check quantization for the new memory budget
If you were memory-constrained on your old hardware, for example running a 4-bit build on a 24 GB consumer card because full precision didn't fit, the Spark's 128 GB may open room you didn't have before. Conversely, if your old setup was a multi-GPU cloud instance with more total VRAM than 128 GB, a model that ran at full precision there may not fit on a single Spark at all.
- Recalculate what fits. Parameters times bytes per parameter, plus KV cache headroom for your expected context and concurrency, against the 128 GB ceiling (or 256 GB if you're moving to a two-node cluster). GPUwerk's quantization reference covers the format options and which serving engine each belongs to.
- Don't assume the same quantized checkpoint file just works. A GGUF build made for llama.cpp on your old box still runs under llama.cpp on a Spark; that part transfers directly. But if your old deployment used a format tied to different hardware, for example an INT8 path tuned for a specific card's tensor cores, check whether an equivalent build exists for the Spark's Blackwell architecture (MXFP4 or NVFP4) rather than assuming it ports unchanged.
- Re-run your own eval set after switching formats. If moving hardware also means moving quantization format, treat it as a new decision, not a rubber stamp of the old one. GPUwerk's production quantization guide covers building that eval set.
Step 2: move the data before you move the traffic
Model weights, fine-tuned adapters, vector indexes, prompt templates, and any application state need to land on the new instance before you cut traffic over, and this is usually the slowest part of a migration, not the model swap itself.
- Weights. If they're already hosted on Hugging Face or a similar registry, pulling them fresh onto the Spark is often simpler and faster than copying from your old box, especially if the old box's uplink is the bottleneck. If they're custom fine-tunes that only exist on the old machine,
rsyncorscpover SSH direct to the new instance's/workspace, per GPUwerk's SSH guide, works for a one-time bulk transfer. - Fine-tuned adapters and LoRA weights. Usually small enough that transfer time isn't the concern, correctness is: confirm the adapter's base model and quantization format match what you're serving on the Spark before assuming it loads correctly.
- Vector indexes and other application state. If these live in a database or service external to the GPU box itself, this step may be nothing more than pointing the new instance's config at the same external service. If they were colocated on the old GPU machine, they move the same way the weights do.
- Do the transfer while the old deployment is still serving traffic. Don't take the old system offline to free up bandwidth or attention for the copy; run the migration in parallel so you have a fallback if something on the new box doesn't work as expected.
Step 3: validate before cutover
Stand up the new deployment on the Spark, run it against the same eval set or a slice of real traffic, and compare output quality and latency against the old system side by side. This is the point to catch a quantization regression from step 1 or a subtly wrong config from step 2, while the old system is still the one actually serving users.
Check latency characteristics specifically, not just correctness: a Spark's single-stream decode speed and its behavior under concurrent load are governed by the bandwidth and engine arithmetic on GPUwerk's benchmarks page, and those numbers won't match a different GPU architecture even for the same model and similar quantization. If your workload has a latency SLA, measure it on the new box before committing to it, don't assume parity from the spec sheet.
Step 4: cut over deliberately
- Point traffic at the new endpoint gradually if you can, a canary slice of requests or a percentage-based router switch, rather than an all-at-once DNS cutover, so a problem shows up on a fraction of traffic instead of all of it.
- Keep the old deployment running until the new one has handled real traffic without issues, not just passed the eval set. An eval set catches correctness regressions; it doesn't always catch load-related issues that only show up under production concurrency.
- Decommission the old system on your own schedule, once you're confident, rather than the Spark's billing cycle or a self-imposed deadline forcing a premature cutover.
What doesn't need to change
If your existing deployment already runs behind vLLM or llama.cpp, the serving layer itself is largely the same; GPUwerk's vLLM doc and llama.cpp doc cover Spark-specific setup, but the framework, its API surface, and anything your application built against that API surface should carry over with minimal change. The migration work is concentrated in the memory and quantization layer underneath, not the serving interface above it.
A first engagement can help plan the cutover and validate the eval set before you commit production traffic to a new box.