What is unified memory?
Unified memory means the CPU and GPU share one physical pool of memory, addressed the same way by both, instead of each having its own separate memory with data copied between them. The alternative, the traditional discrete-GPU model, gives the GPU a dedicated VRAM pool that's fast but fixed in size, usually much smaller than the system's main RAM, and moving data in or out means an explicit transfer over PCIe. Unified memory removes that split: whatever memory exists on the machine is one pool, and both processors read and write it directly.
Why discrete VRAM is normally the limit
On a workstation with a discrete GPU, the GPU's own VRAM is what has to hold a model's weights, and that pool is fixed at whatever the card shipped with, commonly 16 to 80 GB on cards built for this. A model larger than that pool doesn't fit, full stop, regardless of how much system RAM the machine has, because system RAM isn't directly usable by the GPU without copying data across the PCIe bus, which is far slower than VRAM itself and impractical to lean on for active inference. That ceiling is what pushes larger models toward splitting across multiple GPUs, each contributing its own separate VRAM pool.
How the DGX Spark's GB10 does it differently
The DGX Spark's GB10 Grace Blackwell superchip pairs a Grace CPU and a Blackwell GPU with 128 GB of LPDDR5x memory shared between them at 273 GB/s of bandwidth, all specs published on the DGX Spark hardware page. There's no separate, smaller VRAM pool the GPU is limited to; the full 128 GB is available to the GPU for model weights and KV cache, the same way it's available to the CPU. That's the specific mechanism behind why a Spark can hold a 120B-parameter model, at a reasonable quantization, that would need to be split across two or more discrete data-center GPUs to fit any other way.
The tradeoff: capacity for bandwidth
Unified memory on the Spark isn't free of cost, it trades one constraint for another. LPDDR5x at 273 GB/s is well below the bandwidth of a discrete data-center GPU's VRAM, which is why the benchmarks page shows dense models decoding relatively slowly on a single stream: tok/s is bounded by bandwidth divided by active parameter bytes, and 273 GB/s is a modest number to divide by. The Spark's advantage is capacity, holding weights that simply don't fit elsewhere at this price point, not raw single-stream speed. The DGX Spark vs H100 comparison lays out exactly where that trade lands in either direction.
Why this matters for a 120B-parameter model specifically
A 120B-parameter model at a practical quantization, plus room for KV cache to serve more than one request, needs on the order of 100+ GB just for weights before cache is even considered. A discrete GPU anywhere near the Spark's price sits at a fraction of that in VRAM, so running such a model means either model-parallel splitting across several cards or not running it at all. Because the Spark's 128 GB is one pool rather than a small VRAM tier backed by a larger, slower system-RAM tier, the model loads whole, on one machine, with no cross-GPU communication overhead. That's the direct link between unified memory as an architecture and the model sizes GPUwerk can actually serve; see which models fit in 128 GB for what that looks like model by model.