Run gpt-oss-20b on a DGX Spark: reported tok/s and memory
gpt-oss-20b is the smaller sibling in OpenAI's open-weight pair, and on a Spark it reports almost the same single-stream decode speed as the 120B version, because both share the same MXFP4 active-parameter footprint. Every figure below is a third-party report, cited from its source. GPUwerk has not run this model on its own fleet yet.
Reported figures
Third-party numbers, not GPUwerk measurements. Source and date are given because that distinction matters more than the figures themselves.
| Model and format | Engine | Decode | Prefill |
|---|---|---|---|
| gpt-oss-20B, MXFP4 | llama.cpp | ~60.9 tok/s | ~2,009 tok/s |
| gpt-oss-120b, MXFP4 | llama.cpp | ~60.5 tok/s | ~1,956 tok/s |
Reported by user eugr in the llama.cpp DGX Spark performance discussion, quoted on GPUwerk's benchmarks page. No exact run date is stated in the source; the benchmarks page was last updated 14 September 2026. GPUwerk has not reproduced these numbers on its own fleet and has no vLLM figure for this model to cite.
The near-tie in decode speed against the 120B model is not a coincidence. Both models activate roughly the same order of magnitude of parameters per token, so the bandwidth arithmetic below lands in the same range for both, even though gpt-oss-20b's total parameter count, and therefore its weight footprint, is a fraction of the 120B's.
Memory footprint arithmetic
Neither source above states an exact MXFP4 weight size in GB for gpt-oss-20b, so treat that figure as not measured, not reported. What the reported decode speed lets you back out is the active-parameter side of the decode-speed formula on the benchmarks page:
tokens/sec ≈ memory bandwidth ÷ active parameter bytes # Spark: 273 GB/s # ~60.9 tok/s reported, real engines land 50-80% of ceiling # implies a ceiling of roughly 76 to 122 tok/s # implies roughly 2.2 to 3.6 GB of active weight reads per token
That range brackets the same active-parameter read as gpt-oss-120b, whose benchmarks-page arithmetic puts it at roughly 3 GB per token for a ~5B active-parameter count. Nothing here tells you the total weight size that has to fit in the Spark's 121.6 GiB of addressable memory before generation starts, only the per-token read once it is loaded. Verify the actual download size against the repository's own file listing before planning KV cache headroom around it.
gpt-oss-20b versus gpt-oss-120b
Pick 20b when memory is the constraint, not speed. The two models report the same single-stream decode speed on this hardware, so the 120B's usual advantage, more total capability, has to be worth the extra weight it occupies. A smaller total footprint leaves more room for KV cache, a second model resident alongside it, or a faster cold start when a Spark boots a fresh container.
Pick 120b when the workload needs the larger model's capability and you are not fighting for memory. Since decode speed is not the differentiator here, the decision comes down to what each model gets right on your actual prompts, which this page cannot answer for you. The models overview covers the general arithmetic for sizing either one against a specific memory budget.
Neither model's total weight footprint has a GPUwerk-published number, so budget headroom from the file sizes in each model's own repository rather than a figure quoted here.
How to run it
For one user, build llama.cpp for GB10 and serve the GGUF directly, following the llama.cpp doc's build and serve pattern:
cd ~/llama.cpp/build ./bin/llama-server \ -hf <gpt-oss-20b-gguf-repo> \ --host 0.0.0.0 \ --port 30000
Swap in the GGUF repository for gpt-oss-20b MXFP4 in place of the placeholder; the flags are otherwise the same pattern the llama.cpp doc uses for the 120B model. For a shared endpoint, use vLLM's CUDA 13 container from the vLLM doc, adjusting --max-model-len and --max-num-seqs to the memory this model actually leaves you once you have checked its repository file sizes:
docker run -d --name vllm \
--gpus all --ipc=host --restart unless-stopped \
-p 8000:8000 \
-e HF_TOKEN="$HF_TOKEN" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:cu130-nightly \
vllm serve openai/gpt-oss-20b \
--gpu-memory-utilization 0.85 \
--max-model-len 32768 \
--max-num-seqs 4
No vLLM figure for gpt-oss-20b appears in either source above, so treat the flags here as a starting point from the 120B pattern rather than a tuned configuration for this model.
Try it
Rent a Spark at $0.79/hour, billed per minute against a prepaid balance, on the pricing page, and run the command above against your own prompts. GPUwerk has not measured this model on its own fleet, so your first run is the only reliable number until that changes. If you would rather have someone else stand up the endpoint, GPUwerk's first engagement covers exactly that.
FAQ
What decode speed is reported for gpt-oss-20b on a DGX Spark?
Roughly 60.9 tok/s decode under llama.cpp at MXFP4, single-stream, reported by user eugr in the llama.cpp DGX Spark performance discussion and quoted on GPUwerk's benchmarks page. GPUwerk has not measured this figure on its own fleet.
Should I run gpt-oss-20b or gpt-oss-120b on a Spark?
gpt-oss-20b if you need headroom for a second model, more KV cache, or faster load times, since its reported weight footprint is roughly a sixth of gpt-oss-120b's. gpt-oss-120b if raw capability matters more, since the two report near-identical single-stream decode speed on this hardware.
Has GPUwerk measured gpt-oss-20b on its own fleet?
No. Every figure on this page is a third-party report cited from its original source. GPUwerk's own measured figures, for Qwen3-Coder 30B-A3B AWQ and Llama 3.3 70B AWQ, are on the benchmarks page and the respective model pages.