Best open-weight models to run on a DGX Spark
GPUwerk has published measured or reported tok/s figures for six models on a DGX Spark, each with its own page. This is the navigation layer: which one fits which job, using only the numbers already on those six pages, nothing new.
The one-line verdict per model
- Qwen3-Coder-30B-A3B (AWQ, GPUwerk-measured): 80.9 tok/s single-user, 3,235 tok/s aggregate at 256 concurrent. The default pick for chat and coding, and it saturates rather than collapses under load, so it also holds up as a shared endpoint.
- gpt-oss-120b (MXFP4): 60.5 tok/s single-user on llama.cpp, or 33.5 tok/s on vLLM scaling to 862.8 tok/s aggregate at 256 concurrent. Best single-user pick if you want llama.cpp specifically; best batch and agent-fleet pick by aggregate throughput among these six.
- Nemotron-3-Super-120B-A12B (NVFP4): 22.7 to 23.7 tok/s single-user, prefill near 1,900 tok/s. Usable for chat, a step down from the two above, and its prefill number makes it a reasonable pick for long-prompt, short-answer work like contract review.
- DeepSeek-V4-Flash (reported, not GPUwerk-measured, patched fork): 14 to 21 tok/s at 2-bit quantization. Not a default for anything on one Spark; it exists on this list as the dense-model reality check at the extreme end, a 304B-total model forced onto one node by dropping below the 4-bit floor GPUwerk's quantization guidance treats as the point where quality starts to visibly degrade.
- Nemotron Super 49B v1.5 (NVFP4): 5.79 tok/s single-user, 695.1 tok/s aggregate at 256 concurrent. Not a chat model, a batch model: it only earns its place at high concurrency.
- Llama 3.3 70B (AWQ, GPUwerk-measured): 6.0 tok/s single-user, peaking at 485 tok/s aggregate at 128 concurrent before collapsing to 300 tok/s at 256. The clearest dense-model reality check in GPUwerk's own numbers: a 70B dense model reads all 70B parameters every token, and no amount of concurrency tuning turns it into a chat model on this hardware.
Chat, single person
Qwen3-Coder-30B-A3B at 80.9 tok/s is the fastest single-stream figure across all six pages, well above the roughly 10 tok/s line GPUwerk's benchmarks page draws for a chat interface to feel usable. gpt-oss-120b on llama.cpp at 60.5 tok/s is the other real option if you specifically want llama.cpp's engine. Neither Nemotron Super 49B at 5.79 tok/s nor Llama 3.3 70B at 6.0 tok/s clears that bar; both are dense models where every token reads the full parameter count.
Coding
Qwen3-Coder-30B-A3B again, both by name, it is a coding-tuned model, and by number: 80.9 tok/s single-user, with prefill at 5,705 tok/s on an 8,192-token prompt for fast handling of a large file or diff as context. It is also the model GPUwerk stages by default alongside Llama 3.3 70B, so it deploys with no download step.
Batch and agent-fleet work
This is where the dense-versus-MoE split matters most. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent under vLLM, the highest aggregate figure among these six. Nemotron Super 49B, despite a poor single-user number, reaches 695.1 tok/s aggregate at the same concurrency, a 120-fold rise from its single-stream figure, exactly the shape of a model that only makes sense batched. Qwen3-Coder-30B-A3B also holds up here at 3,235 tok/s aggregate. Llama 3.3 70B is the outlier to watch: its aggregate throughput peaks at 485 tok/s at 128 concurrent, then falls to 300 tok/s at 256 as the dense model collapses past its knee, so batch jobs on it need to stay near that 128-concurrency point rather than assume more concurrency is always better.
The dense-model reality check
Two models on this list make the same point from different directions. Llama 3.3 70B, a 70B dense model, decodes at 6.0 tok/s single-user and its aggregate throughput actively falls once concurrency passes its knee, worse at everything at once rather than trading latency for throughput. DeepSeek-V4-Flash, at 304B total parameters, does not fit on one Spark at any quality-preserving quantization at all; the 14 to 21 tok/s figure reported for it required a patched, non-stock vLLM fork and 2-bit expert planes, a research configuration rather than something to build a service on, and GPUwerk has not reproduced those numbers on its own fleet. Read together, they are the honest counterweight to the mixture-of-experts models above: total parameter count does not predict speed on this hardware, active parameter count does, and a dense model much past 30 to 50B starts fighting the 273 GB/s bandwidth ceiling that GPUwerk's benchmarks page and glossary both explain.
How to pick
Start from the job, not the model name. One person typing prompts: Qwen3-Coder-30B-A3B. Many people or an agent fleet, tolerant of per-request latency: gpt-oss-120b or Nemotron Super 49B, sized to your concurrency. Long documents with short answers: Nemotron-3-Super-120B for its prefill number. Anything that reaches for a 70B-plus dense model, or DeepSeek-V4-Flash specifically, deserves a second look at whether an MoE model of similar total size does the same job faster on this hardware, per the numbers above.
Rent a Spark at $0.79/hour and run your own prompts against whichever model this points at before trusting the numbers on any of these seven pages, including this one. A first engagement covers picking and tuning a model for a specific workload if you would rather hand that off.