Docs/Models/Nemotron Super 49B
Models

Run Nemotron Super 49B on a DGX Spark: measured tok/s

Samuel Seidel · Published September 7, 2026

Nemotron Super 49B v1.5 at NVFP4 runs under vLLM but not comfortably for one person: 5.79 tok/s single-stream. It is a dense model, so every token reads all 49B parameters, and it only earns its place on this hardware batched at high concurrency.

Measured on one Spark

Engine Quantization Decode, single user Prefill
vLLM NVFP4 5.79 tok/s not reported
Concurrency Aggregate throughput Per-request speed
15.79 tok/s5.79 tok/s
256695.1 tok/s2.85 tok/s

Single-stream figure from the vLLM project's own DGX Spark writeup. Concurrency figures from Dendro Logic's concurrency benchmark, running nvcr.io/nvidia/vllm:26.03-py3 on a single GB10 with 128 GB, on a character-extraction workload of roughly 1,500 prompt tokens and up to 400 output tokens. Both are quoted on GPUwerk's benchmarks page, updated August 26, 2026; neither source states an exact run date.

Memory footprint

Not measured. Neither GPUwerk's benchmarks page nor its models page states a weight size in gigabytes for Nemotron Super 49B v1.5 at NVFP4, so no footprint figure is given here. What both pages do establish is that it is a dense 49B model, meaning decode reads the full parameter count on every token, which is the direct cause of the 5.79 tok/s single-user number.

How to run it

Serve the NVFP4 build with vLLM's CUDA 13 container, following the pattern in the vLLM doc. Because this model only makes sense batched, raise --max-num-seqs above the interactive default of 4 rather than leaving it low:

docker run -d --name vllm \
  --gpus all --ipc=host --restart unless-stopped \
  -p 8000:8000 \
  -e HF_TOKEN="$HF_TOKEN" \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:cu130-nightly \
  vllm serve nvidia/Nemotron-Super-49B-v1.5-NVFP4 \
    --gpu-memory-utilization 0.85 \
    --max-model-len 32768 \
    --max-num-seqs 64

Sweep vllm bench serve --num-prompts --request-rate against your own prompt shape before committing to a concurrency level, per the vLLM doc's verify-and-measure guidance.

What it is good for on a Spark

Not a chat model on this hardware. 5.79 tok/s is well under the roughly 10 tok/s line where a chat window feels usable, and no flag or quantization change fixes a dense model's bandwidth ceiling.

It is a batch model. At 256 concurrent requests, aggregate throughput reaches 695.1 tok/s, a 120-fold rise over the single-user figure, which suits overnight document processing, batch classification or an agent fleet where no single request needs to return quickly.

Read the per-request column before planning around the aggregate number: at that same concurrency, each request crawls at 2.85 tok/s. Find the concurrency level for your own workload that keeps per-request speed tolerable while capturing most of the aggregate gain, rather than assuming 256 is the right number for you.

Try it

A Spark is $0.79/hour on the pricing page, billed per minute against a prepaid balance, cheap enough to run your own concurrency sweep before deciding whether 5.79 tok/s single-user and 695.1 tok/s aggregate fit your batch workload. GPUwerk's first engagement covers picking and tuning a concurrency level for you instead.

FAQ

Is Nemotron Super 49B fast enough for a chat interface on a Spark?

No. It measures 5.79 tok/s single user, well under the roughly 10 tok/s line where a chat window starts to feel usable. It is a dense model, so every token reads all 49B parameters, and no configuration fixes that on this hardware.

What throughput does Nemotron Super 49B reach at high concurrency?

695.1 tok/s aggregate at 256 concurrent requests under vLLM, a 120-fold rise from the single-user number. Per-request speed at that concurrency is 2.85 tok/s, so this is a throughput result for batch work, not a latency one.

How much memory does Nemotron Super 49B need on a Spark?

Not measured. Neither GPUwerk's benchmarks page nor its models page states a weight size in gigabytes for this NVFP4 build, so no footprint number is given here rather than an estimate.

See alsoFull benchmark numbers across engines See alsoWhich models fit in 128 GB

Run it yourself, not our numbers.

A dedicated Spark, deployed in minutes, with $20 in credit for your first $10 top-up.

Deploy a Spark