Docs/Models/gpt-oss-120b
Models

Run gpt-oss-120b on a DGX Spark: measured tok/s

Samuel Seidel · Published September 7, 2026

gpt-oss-120b at MXFP4 runs on one 128 GB Spark and is fast enough to sit in a chat window: 60.5 tok/s decode under llama.cpp, single user. Under vLLM it drops to 33.5 tok/s single-stream but scales to 862.8 tok/s aggregate once 256 requests share the box.

Measured on one Spark

Engine Quantization Decode, single user Prefill
llama.cpp MXFP4 ~60.5 tok/s ~1,956 tok/s
vLLM MXFP4 33.5 tok/s not reported

llama.cpp figures reported by user eugr in the llama.cpp DGX Spark performance discussion. vLLM single-stream and concurrency figures from Dendro Logic's concurrency benchmark, running nvcr.io/nvidia/vllm:26.03-py3 on a single GB10 with 128 GB. Neither source states an exact run date; both are quoted on GPUwerk's benchmarks page, updated August 26, 2026.

Concurrency (vLLM) Aggregate throughput Per-request speed
133.5 tok/s33.6 tok/s
256862.8 tok/s3.62 tok/s

From Dendro Logic's character-extraction workload, roughly 1,500 prompt tokens in and up to 400 out with structured JSON. At 256 concurrent, aggregate throughput is 25.8 times the single-user number, but each individual request crawls at 3.62 tok/s. That is a batch result, not a chat-latency one.

Memory footprint

Neither the benchmarks page nor the models page states a total MXFP4 weight size for gpt-oss-120b, so that figure is not measured. What both pages do give is the active-parameter side of the calculation: gpt-oss-120b activates roughly 5B parameters per token, about 3 GB of weight reads, which is GPUwerk's own arithmetic on the benchmarks page rather than a measurement. That is the reason decode speed lands near 60 tok/s rather than the low single digits a dense 120B model would produce, and it is also why headroom for KV cache cannot be stated here without a real weight number to subtract from the Spark's 121.6 GiB addressable memory.

How to run it

For one user, build llama.cpp for GB10 and serve the GGUF directly, following the llama.cpp doc's build and serve pattern:

cd ~/llama.cpp/build
./bin/llama-server \
  -hf <gpt-oss-120b-gguf-repo> \
  --host 0.0.0.0 \
  --port 30000

Swap in the GGUF repository for gpt-oss-120b MXFP4 in place of the placeholder; the flags are otherwise unchanged from the llama.cpp doc's serving example. For a shared endpoint, use vLLM's CUDA 13 container from the vLLM doc:

docker run -d --name vllm \
  --gpus all --ipc=host --restart unless-stopped \
  -p 8000:8000 \
  -e HF_TOKEN="$HF_TOKEN" \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:cu130-nightly \
  vllm serve openai/gpt-oss-120b \
    --gpu-memory-utilization 0.85 \
    --max-model-len 32768 \
    --max-num-seqs 4

What it is good for on a Spark

Single-user, llama.cpp: 60.5 tok/s is comfortably above reading speed, so it works as a local coding assistant or chat interface for one person, with the caveat that llama.cpp's batching is weaker than vLLM's once a second person joins.

Shared endpoint, vLLM: the single-request speed drops to 33.5 tok/s, still usable for one person, but the value case is the 862.8 tok/s aggregate at 256 concurrent, which suits overnight batch classification, document extraction or an agent fleet where no individual request needs to feel instant.

Prefill at roughly 1,900 to 1,950 tok/s across both engines makes it a good fit for long-prompt, short-answer work: summarizing a contract or classifying a ticket against a large system prompt finishes quickly regardless of which engine serves the decode.

Try it

Rent a Spark at $0.79/hour, billed per minute against a prepaid balance, on the pricing page, and run either command above against your own prompts before trusting any of the numbers on this page. If you would rather have someone else stand up the endpoint and tune the flags, GPUwerk's first engagement covers exactly that.

FAQ

What single-user speed does gpt-oss-120b get on a DGX Spark?

About 60.5 tok/s decode under llama.cpp, or 33.5 tok/s under vLLM. Both are single-stream, MXFP4, measured on one GB10 and reported on GPUwerk's benchmarks page.

Which engine should I use for gpt-oss-120b on a Spark?

llama.cpp for one user: it is roughly 80% faster single-stream. vLLM once more than one person shares the endpoint, since its aggregate throughput reaches 862.8 tok/s at 256 concurrent against 33.5 tok/s for one request.

Does gpt-oss-120b fit in 128 GB alongside other work?

GPUwerk has not measured or published its total MXFP4 weight footprint, so treat that as not measured rather than assume a number. Its active parameter count is roughly 5B, which is why decode speed is closer to a small dense model than a 120B one.

See alsoFull benchmark numbers across engines See alsoWhich models fit in 128 GB

Run it yourself, not our numbers.

A dedicated Spark, deployed in minutes, with $20 in credit for your first $10 top-up.

Deploy a Spark