Migration
DGX Spark/Migrating an existing self-hosted AI deployment to a DGX Spark

Migrating an existing self-hosted AI deployment to a DGX Spark

By Samuel Seidel · Published September 9, 2026

If you're already self-hosting on a workstation with a consumer GPU, or on a rented cloud GPU instance elsewhere, moving to a Spark is not a from-scratch deployment. It's a migration with its own specific gotchas: the memory architecture is different, the quantization that worked on your old box may not be the right one here, and your data needs a plan to move that doesn't depend on the old machine staying up while you test the new one.

Why the move isn't a straight lift-and-shift

A Spark's 128 GB is unified memory, shared between CPU and GPU, not a discrete VRAM pool sitting next to system RAM the way a workstation with an RTX card or a cloud instance with a discrete NVIDIA GPU is built. GPUwerk's DGX Spark vs Mac Studio comparison and the memory-bandwidth arithmetic on the benchmarks page cover why that architecture changes the numbers that actually govern serving speed: decode speed is bandwidth-bound, and a Spark's bandwidth figure is not the same as whatever card you're moving off of. Don't assume a model and quantization that ran well on your existing hardware will hit the same tokens per second here; the two boxes are shaped differently even when the total memory number looks similar on paper.

Step 1: re-check quantization for the new memory budget

If you were memory-constrained on your old hardware, for example running a 4-bit build on a 24 GB consumer card because full precision didn't fit, the Spark's 128 GB may open room you didn't have before. Conversely, if your old setup was a multi-GPU cloud instance with more total VRAM than 128 GB, a model that ran at full precision there may not fit on a single Spark at all.

Step 2: move the data before you move the traffic

Model weights, fine-tuned adapters, vector indexes, prompt templates, and any application state need to land on the new instance before you cut traffic over, and this is usually the slowest part of a migration, not the model swap itself.

Step 3: validate before cutover

Stand up the new deployment on the Spark, run it against the same eval set or a slice of real traffic, and compare output quality and latency against the old system side by side. This is the point to catch a quantization regression from step 1 or a subtly wrong config from step 2, while the old system is still the one actually serving users.

Check latency characteristics specifically, not just correctness: a Spark's single-stream decode speed and its behavior under concurrent load are governed by the bandwidth and engine arithmetic on GPUwerk's benchmarks page, and those numbers won't match a different GPU architecture even for the same model and similar quantization. If your workload has a latency SLA, measure it on the new box before committing to it, don't assume parity from the spec sheet.

Step 4: cut over deliberately

What doesn't need to change

If your existing deployment already runs behind vLLM or llama.cpp, the serving layer itself is largely the same; GPUwerk's vLLM doc and llama.cpp doc cover Spark-specific setup, but the framework, its API surface, and anything your application built against that API surface should carry over with minimal change. The migration work is concentrated in the memory and quantization layer underneath, not the serving interface above it.

A first engagement can help plan the cutover and validate the eval set before you commit production traffic to a new box.

Validate your migration before you cut traffic over.

Deploy a Spark alongside your existing setup, run it side by side, and cut over on your own schedule. First top-up: pay $10, get $20 in credit.

Deploy a Spark Read the quantization reference