GPU utilization monitoring and right-sizing your deployment
Deciding whether to stay on one Spark, move to a two-node cluster, or add a second independent node is a question the actual utilization numbers should answer, not a guess made from how busy the workload feels. This page covers what to watch and what each pattern of numbers implies; for the mechanics of collecting the numbers in the first place, see the monitoring and observability guide.
The three numbers that matter
GPU compute utilization, from nvidia-smi or its Prometheus exporter, tells you what fraction of time the GPU is actively computing versus idle waiting on the next request. Sustained high utilization (consistently above roughly 80-90%) during your busy periods means the box is compute-bound and requests are likely queuing; low utilization with high memory use means the box is memory-constrained rather than compute-constrained, a different problem with a different fix.
Memory utilization, both how much is allocated to resident model weights and KV cache, and how close that sits to the ceiling. A box sitting near its memory limit is one traffic spike away from an OOM regardless of how much spare compute it has; see the memory management guide for the budgeting math behind this number.
Request queue depth and latency, from your serving engine's own metrics (vLLM and llama.cpp both expose queue length and time-to-first-token). This is the number closest to what users actually experience, and it can look bad even when GPU utilization looks fine, if queue depth is climbing while compute utilization sits well under 100%, the bottleneck is more likely scheduling or a memory ceiling limiting concurrent sequences than raw compute.
Reading the pattern
Four utilization patterns come up in practice, and they point to different actions:
- High compute utilization, memory headroom left. The box is doing real, sustained work and there's room to add a second model or increase concurrency before memory becomes the constraint. Usually the healthiest pattern; no action needed beyond continuing to watch it.
- High compute utilization, memory near the ceiling. Both resources are tight at once. This is the pattern that justifies scaling up, moving to the two-node cluster if the constraint is one model needing more room, per the memory management guide's distinction between a cluster and a fleet.
- Low compute utilization, memory near the ceiling. Memory is the binding constraint, not compute, adding more compute (a cluster) doesn't fix this; the fix is reducing what's resident (fewer simultaneous models, smaller KV cache budget, or a smaller model) or moving to a cluster specifically for its larger memory pool rather than for its compute.
- Low compute utilization, memory headroom left. The box is underused. If this holds over days, not just a quiet hour, it's a signal you're paying for more than the workload needs; a single node or even a scheduled-stop pattern may be the cheaper fit than what you're currently running.
Scale up vs. scale out
The two-node cluster ($1.79/hour per pricing, linking two Sparks into a 256 GB pool over 200 GbE) is the right move when one model's memory footprint, not overall traffic, is the constraint: a model too large to fit in one node's roughly 121.6 GiB of usable memory, or a multi-model lineup whose combined footprint genuinely needs the larger pool. It's the wrong move when the actual problem is several independent workloads competing for one node's resources; that's what the fleet scaling guide covers, adding a second, independently rented Spark rather than clustering the two together. Clustering two nodes doesn't fix contention between unrelated workloads, it just gives them a bigger shared pool to contend over.
Before committing to either, check whether the actual bottleneck is something cheaper to fix: enabling prompt caching to reduce compute per request, tightening the memory budget per the multi-tenant guide, or moving a batch-shaped workload off the interactive box entirely per the batch processing guide. Utilization data that looks like a capacity problem is sometimes a configuration problem wearing a capacity problem's clothes.
How long to watch before deciding
A single busy hour or a single quiet weekend isn't enough signal to resize on. Watch utilization across at least one full weekly cycle, including your actual peak period, whatever drives it (business hours, a batch job that runs overnight, a launch spike), before concluding the pattern is real rather than a one-off. A dedicated rental bills by the hour regardless of load, so the cost of watching a week before deciding is small next to the cost of resizing on a false signal and having to walk it back.
Practical checklist
- Track GPU compute utilization, memory utilization, and request queue depth together; any one number alone is misleading.
- Match the pattern to the fix: compute-bound justifies more compute, memory-bound doesn't, no matter how it feels from the outside.
- Distinguish "one model needs more room" (cluster) from "several workloads are fighting for one node" (fleet) before choosing which to scale.
- Rule out cheaper fixes, caching, tighter memory budgeting, moving batch work off the interactive box, before committing to added hardware cost.
- Watch at least a full weekly cycle including peak load before resizing, not a single busy or quiet stretch.
See monitoring and observability for how to collect these metrics, and scaling to a fleet for when independent capacity beats a bigger cluster.