Run Nemotron Super 49B on a DGX Spark: measured tok/s
Nemotron Super 49B v1.5 at NVFP4 runs under vLLM but not comfortably for one person: 5.79 tok/s single-stream. It is a dense model, so every token reads all 49B parameters, and it only earns its place on this hardware batched at high concurrency.
Measured on one Spark
| Engine | Quantization | Decode, single user | Prefill |
|---|---|---|---|
| vLLM | NVFP4 | 5.79 tok/s | not reported |
| Concurrency | Aggregate throughput | Per-request speed |
|---|---|---|
| 1 | 5.79 tok/s | 5.79 tok/s |
| 256 | 695.1 tok/s | 2.85 tok/s |
Single-stream figure from the vLLM project's own DGX Spark writeup. Concurrency figures from Dendro Logic's concurrency benchmark, running nvcr.io/nvidia/vllm:26.03-py3 on a single GB10 with 128 GB, on a character-extraction workload of roughly 1,500 prompt tokens and up to 400 output tokens. Both are quoted on GPUwerk's benchmarks page, updated August 26, 2026; neither source states an exact run date.
Memory footprint
Not measured. Neither GPUwerk's benchmarks page nor its models page states a weight size in gigabytes for Nemotron Super 49B v1.5 at NVFP4, so no footprint figure is given here. What both pages do establish is that it is a dense 49B model, meaning decode reads the full parameter count on every token, which is the direct cause of the 5.79 tok/s single-user number.
How to run it
Serve the NVFP4 build with vLLM's CUDA 13 container, following the pattern in the vLLM doc. Because this model only makes sense batched, raise --max-num-seqs above the interactive default of 4 rather than leaving it low:
docker run -d --name vllm \
--gpus all --ipc=host --restart unless-stopped \
-p 8000:8000 \
-e HF_TOKEN="$HF_TOKEN" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:cu130-nightly \
vllm serve nvidia/Nemotron-Super-49B-v1.5-NVFP4 \
--gpu-memory-utilization 0.85 \
--max-model-len 32768 \
--max-num-seqs 64
Sweep vllm bench serve --num-prompts --request-rate against your own prompt shape before committing to a concurrency level, per the vLLM doc's verify-and-measure guidance.
What it is good for on a Spark
Not a chat model on this hardware. 5.79 tok/s is well under the roughly 10 tok/s line where a chat window feels usable, and no flag or quantization change fixes a dense model's bandwidth ceiling.
It is a batch model. At 256 concurrent requests, aggregate throughput reaches 695.1 tok/s, a 120-fold rise over the single-user figure, which suits overnight document processing, batch classification or an agent fleet where no single request needs to return quickly.
Read the per-request column before planning around the aggregate number: at that same concurrency, each request crawls at 2.85 tok/s. Find the concurrency level for your own workload that keeps per-request speed tolerable while capturing most of the aggregate gain, rather than assuming 256 is the right number for you.
Try it
A Spark is $0.79/hour on the pricing page, billed per minute against a prepaid balance, cheap enough to run your own concurrency sweep before deciding whether 5.79 tok/s single-user and 695.1 tok/s aggregate fit your batch workload. GPUwerk's first engagement covers picking and tuning a concurrency level for you instead.
FAQ
Is Nemotron Super 49B fast enough for a chat interface on a Spark?
No. It measures 5.79 tok/s single user, well under the roughly 10 tok/s line where a chat window starts to feel usable. It is a dense model, so every token reads all 49B parameters, and no configuration fixes that on this hardware.
What throughput does Nemotron Super 49B reach at high concurrency?
695.1 tok/s aggregate at 256 concurrent requests under vLLM, a 120-fold rise from the single-user number. Per-request speed at that concurrency is 2.85 tok/s, so this is a throughput result for batch work, not a latency one.
How much memory does Nemotron Super 49B need on a Spark?
Not measured. Neither GPUwerk's benchmarks page nor its models page states a weight size in gigabytes for this NVFP4 build, so no footprint number is given here rather than an estimate.