GPU memory management for multiple models on one Spark
A single DGX Spark has 128 GB of memory shared between CPU and GPU, roughly 121.6 GiB of it addressable once the OS and drivers take their share. That's one pool, not a GPU pool plus a system pool the way a discrete card works, and it changes how you think about running more than one model on the box. This page is about the budgeting: what fits, what doesn't, and the point where the honest answer is a second node rather than a tighter squeeze.
Start from what each thing actually costs
Three categories eat the 128 GB, and they don't shrink just because you're running several models instead of one.
- Weights, roughly parameters × bytes per parameter. A 70B model at FP8 is about 70 GB on its own; a 27B at 4-bit is closer to 15 GB. Run the arithmetic per model before you commit to a lineup, since this is the part that doesn't compress at request time.
- KV cache, which scales with context length times concurrent sequences, and which is the thing that actually gets squeezed when memory is tight. This is where a multi-model setup pays its real cost: every model you keep loaded has its own KV cache reservation, even the one nobody is calling this minute.
- Headroom for the CUDA context, activations, and framework overhead. Leave more of it than feels necessary; a box that runs fine in testing at 95% utilization can OOM the first time two requests land on different models at once.
Add up weights plus a KV cache budget per model you intend to keep resident simultaneously, and compare the total against roughly 121.6 GiB. If it doesn't fit with headroom to spare, something has to give before you find out the hard way in production.
Two ways to run more than one model
The engines behave differently once you're past a single model, and the choice matters more here than it does when serving one thing.
Separate serving processes, each pinned to its own memory slice. Run two vLLM containers, or a vLLM container next to a llama.cpp instance, each started with an explicit memory ceiling: vLLM's --gpu-memory-utilization is a fraction of total memory, not free memory, so if you're running two vLLM processes, the fractions have to be planned together, not each set as if it owned the box. A 0.5 and a 0.35 on two separate containers is a plan; two containers each independently set to 0.85 will have the second one fail to start once the first has already claimed its share.
One gateway routing to models loaded on demand. LiteLLM in front of multiple backends, or an engine with its own model-swapping support, keeps only the requested model's weights resident and evicts or lazy-loads the rest. This trades a load-time delay on a cold model for not paying every model's memory cost at once, which is usually the right trade when the models are used unevenly, one hit constantly and the others occasionally.
Static partitioning is simpler to reason about and easier to debug when something goes wrong. On-demand loading fits more models than the box could hold simultaneously, at the cost of a cold-start delay whenever traffic shifts between them.
A concrete budgeting example
Say you want a 27B coding model available at low latency and a 70B model for occasional heavier requests, both on one Spark. The 27B at 4-bit is roughly 15 GB of weights; the 70B at 4-bit is roughly 37 GB. That's 52 GB of weights before either serves a single token, leaving on the order of 65-70 GB for KV cache and headroom across both. Give the 27B most of the remaining cache, since it's the one taking constant interactive traffic, and set the 70B's --max-num-seqs low, since it's the occasional-use model and doesn't need to serve several concurrent requests well. Set both engines' memory fractions explicitly rather than trusting a default meant for a single-model box.
Run that same math for your actual model choices before deploying; the numbers above are illustrative, not a promise about any specific checkpoint's footprint on your build.
When the budget doesn't close
If weights plus a workable KV cache for your intended model lineup don't fit in roughly 121.6 GiB with headroom left over, that's the signal to stop tuning fractions and add capacity instead of chasing a smaller number that will bite you under load.
- Drop a model, or serve it via API instead of hosting it locally, if it's the one contributing the least to your workload but taking a disproportionate memory share.
- Move to a two-node cluster if the actual problem is one model too large for a single node, per the two-node cluster guide. Two Sparks link over 200 GbE into a 256 GB pool, which changes what fits, though single-user generation speed doesn't move much from the same bandwidth arithmetic that governs one node; a cluster buys capacity, not raw speed.
- Run a second independent node rather than a cluster if the problem is contention between separate workloads rather than one model too big to fit. The fleet scaling guide covers when a second, independently rented Spark with its own memory budget makes more sense than trying to fit everything on one box.
The distinction that matters: a cluster solves "one model doesn't fit on one node." A fleet solves "several independent workloads are fighting each other for the same node's memory and bandwidth." Multi-tenant memory pressure on a single Spark is usually the second problem, not the first, so check the fleet guide before reaching for clustering.
Practical checklist
- Add up weights plus a working KV cache for every model you want resident at once, and compare against roughly 121.6 GiB before deploying, not after.
- Set every engine's memory fraction explicitly when running more than one serving process; don't let two defaults each assume they own the whole box.
- Use a gateway with on-demand loading when models are used unevenly, and static partitioning when you need predictable latency on all of them at once.
- If the budget doesn't close with reasonable headroom, add a node before you shrink the KV cache past the point where concurrency suffers.
See the vLLM install and flags guide for the memory math behind a single deployment, or a first engagement if you want help sizing a specific multi-model lineup.