Capacity
DGX Spark/GPU utilization monitoring and right-sizing your deployment

GPU utilization monitoring and right-sizing your deployment

By Samuel Seidel · Published September 9, 2026

Deciding whether to stay on one Spark, move to a two-node cluster, or add a second independent node is a question the actual utilization numbers should answer, not a guess made from how busy the workload feels. This page covers what to watch and what each pattern of numbers implies; for the mechanics of collecting the numbers in the first place, see the monitoring and observability guide.

The three numbers that matter

GPU compute utilization, from nvidia-smi or its Prometheus exporter, tells you what fraction of time the GPU is actively computing versus idle waiting on the next request. Sustained high utilization (consistently above roughly 80-90%) during your busy periods means the box is compute-bound and requests are likely queuing; low utilization with high memory use means the box is memory-constrained rather than compute-constrained, a different problem with a different fix.

Memory utilization, both how much is allocated to resident model weights and KV cache, and how close that sits to the ceiling. A box sitting near its memory limit is one traffic spike away from an OOM regardless of how much spare compute it has; see the memory management guide for the budgeting math behind this number.

Request queue depth and latency, from your serving engine's own metrics (vLLM and llama.cpp both expose queue length and time-to-first-token). This is the number closest to what users actually experience, and it can look bad even when GPU utilization looks fine, if queue depth is climbing while compute utilization sits well under 100%, the bottleneck is more likely scheduling or a memory ceiling limiting concurrent sequences than raw compute.

Reading the pattern

Four utilization patterns come up in practice, and they point to different actions:

Scale up vs. scale out

The two-node cluster ($1.79/hour per pricing, linking two Sparks into a 256 GB pool over 200 GbE) is the right move when one model's memory footprint, not overall traffic, is the constraint: a model too large to fit in one node's roughly 121.6 GiB of usable memory, or a multi-model lineup whose combined footprint genuinely needs the larger pool. It's the wrong move when the actual problem is several independent workloads competing for one node's resources; that's what the fleet scaling guide covers, adding a second, independently rented Spark rather than clustering the two together. Clustering two nodes doesn't fix contention between unrelated workloads, it just gives them a bigger shared pool to contend over.

Before committing to either, check whether the actual bottleneck is something cheaper to fix: enabling prompt caching to reduce compute per request, tightening the memory budget per the multi-tenant guide, or moving a batch-shaped workload off the interactive box entirely per the batch processing guide. Utilization data that looks like a capacity problem is sometimes a configuration problem wearing a capacity problem's clothes.

How long to watch before deciding

A single busy hour or a single quiet weekend isn't enough signal to resize on. Watch utilization across at least one full weekly cycle, including your actual peak period, whatever drives it (business hours, a batch job that runs overnight, a launch spike), before concluding the pattern is real rather than a one-off. A dedicated rental bills by the hour regardless of load, so the cost of watching a week before deciding is small next to the cost of resizing on a false signal and having to walk it back.

Practical checklist

See monitoring and observability for how to collect these metrics, and scaling to a fleet for when independent capacity beats a bigger cluster.

First top-up: pay $10, get $20 in credit

Size the deployment to the numbers, not the guess.

Deploy a dedicated Spark and watch real utilization before deciding whether to scale up or out.

Deploy a Spark See pricing