Memory
DGX Spark/GPU memory management for multiple models on one Spark

GPU memory management for multiple models on one Spark

By Samuel Seidel · Published September 9, 2026

A single DGX Spark has 128 GB of memory shared between CPU and GPU, roughly 121.6 GiB of it addressable once the OS and drivers take their share. That's one pool, not a GPU pool plus a system pool the way a discrete card works, and it changes how you think about running more than one model on the box. This page is about the budgeting: what fits, what doesn't, and the point where the honest answer is a second node rather than a tighter squeeze.

Start from what each thing actually costs

Three categories eat the 128 GB, and they don't shrink just because you're running several models instead of one.

Add up weights plus a KV cache budget per model you intend to keep resident simultaneously, and compare the total against roughly 121.6 GiB. If it doesn't fit with headroom to spare, something has to give before you find out the hard way in production.

Two ways to run more than one model

The engines behave differently once you're past a single model, and the choice matters more here than it does when serving one thing.

Separate serving processes, each pinned to its own memory slice. Run two vLLM containers, or a vLLM container next to a llama.cpp instance, each started with an explicit memory ceiling: vLLM's --gpu-memory-utilization is a fraction of total memory, not free memory, so if you're running two vLLM processes, the fractions have to be planned together, not each set as if it owned the box. A 0.5 and a 0.35 on two separate containers is a plan; two containers each independently set to 0.85 will have the second one fail to start once the first has already claimed its share.

One gateway routing to models loaded on demand. LiteLLM in front of multiple backends, or an engine with its own model-swapping support, keeps only the requested model's weights resident and evicts or lazy-loads the rest. This trades a load-time delay on a cold model for not paying every model's memory cost at once, which is usually the right trade when the models are used unevenly, one hit constantly and the others occasionally.

Static partitioning is simpler to reason about and easier to debug when something goes wrong. On-demand loading fits more models than the box could hold simultaneously, at the cost of a cold-start delay whenever traffic shifts between them.

A concrete budgeting example

Say you want a 27B coding model available at low latency and a 70B model for occasional heavier requests, both on one Spark. The 27B at 4-bit is roughly 15 GB of weights; the 70B at 4-bit is roughly 37 GB. That's 52 GB of weights before either serves a single token, leaving on the order of 65-70 GB for KV cache and headroom across both. Give the 27B most of the remaining cache, since it's the one taking constant interactive traffic, and set the 70B's --max-num-seqs low, since it's the occasional-use model and doesn't need to serve several concurrent requests well. Set both engines' memory fractions explicitly rather than trusting a default meant for a single-model box.

Run that same math for your actual model choices before deploying; the numbers above are illustrative, not a promise about any specific checkpoint's footprint on your build.

When the budget doesn't close

If weights plus a workable KV cache for your intended model lineup don't fit in roughly 121.6 GiB with headroom left over, that's the signal to stop tuning fractions and add capacity instead of chasing a smaller number that will bite you under load.

The distinction that matters: a cluster solves "one model doesn't fit on one node." A fleet solves "several independent workloads are fighting each other for the same node's memory and bandwidth." Multi-tenant memory pressure on a single Spark is usually the second problem, not the first, so check the fleet guide before reaching for clustering.

Practical checklist

See the vLLM install and flags guide for the memory math behind a single deployment, or a first engagement if you want help sizing a specific multi-model lineup.

First top-up: pay $10, get $20 in credit

128GB, budgeted your way.

Deploy a dedicated Spark and set your own memory fractions across as many models as fit.

Deploy a Spark Read the vLLM guide