Run Nemotron-3-Super-120B on a DGX Spark: measured tok/s
Nemotron-3-Super-120B-A12B at NVFP4 runs on one 128 GB Spark under vLLM at 22.7 to 23.7 tok/s decode for a single user. Prefill is close to 1,900 tok/s, which is the more useful number if your workload is long prompts with short answers.
Measured on one Spark
| Engine | Quantization | Decode, single user | Prefill |
|---|---|---|---|
| vLLM | NVFP4 | 22.7 to 23.7 tok/s | ~1,900 tok/s |
From the vLLM project's own DGX Spark writeup, quoted on GPUwerk's benchmarks page, updated August 26, 2026. The source does not state an exact run date, and no concurrency sweep for this model is published, so aggregate throughput at scale is not measured.
Memory footprint
Weights: about 67 GB, MoE with 120B total parameters and 12B active, per GPUwerk's models page, which labels this an approximate figure checked against Hugging Face repository sizes rather than a live measurement. Against roughly 121.6 GiB of addressable memory on a Spark, that leaves about 54.6 GiB for KV cache, activations and the CUDA context, a GPUwerk calculation subtracting the stated weight figure from the Spark's addressable total, not a separate measurement. The models page calls this combination the one vLLM's own Spark testing used, with room for real context.
How to run it
Download the NVFP4 weights and serve them with vLLM's CUDA 13 container, following the pattern in the vLLM doc and the download steps in the models doc:
hf download nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 \
--local-dir /workspace/models/nemotron-3-super-120b
docker run -d --name vllm \
--gpus all --ipc=host --restart unless-stopped \
-p 8000:8000 \
-v /workspace/models:/workspace/models \
vllm/vllm-openai:cu130-nightly \
vllm serve /workspace/models/nemotron-3-super-120b \
--served-model-name nemotron-3-super-120b \
--gpu-memory-utilization 0.85 \
--max-model-len 32768
What it is good for on a Spark
22.7 to 23.7 tok/s decode is usable for one person in a chat interface, comfortably above the roughly 10 tok/s line where a model stops feeling responsive, but it is a step down from the smaller MoE models GPUwerk stages by default.
The prefill number is the stronger case for this model: at close to 1,900 tok/s, it processes a long system prompt or a document to summarize quickly, before decode ever becomes the bottleneck. That makes it a reasonable pick for contract review or ticket classification workloads with large context and modest output length.
Without a published concurrency sweep for this model, GPUwerk cannot say how it behaves once several people share the endpoint; do not assume it plateaus the way the MoE models in GPUwerk's own concurrency tests do, since that behavior has not been measured here.
Try it
A Spark rents for $0.79/hour, billed per minute against a prepaid balance, on the pricing page. Download the weights and run the command above to check whether 22.7 to 23.7 tok/s actually suits your workload before committing to anything larger. GPUwerk's first engagement is the option if you would rather hand off the setup and tuning.
FAQ
How fast is Nemotron-3-Super-120B on a DGX Spark?
22.7 to 23.7 tok/s decode under vLLM at NVFP4, single user, with prefill around 1,900 tok/s. Figures are from the vLLM project's own DGX Spark writeup, quoted on GPUwerk's benchmarks page.
Does Nemotron-3-Super-120B fit in 128 GB with room to spare?
Yes. The NVFP4 build is about 67 GB of weights against roughly 121.6 GiB of addressable memory on a Spark, leaving headroom for real context and some concurrency.
Was Nemotron-3-Super-120B tested at multiple concurrency levels?
Not on GPUwerk's benchmarks page. Only the single-user decode and prefill numbers are reported there; no concurrency sweep for this model is published, so treat aggregate throughput at scale as not measured.