Buying analysis
DGX Spark/Best open-weight models to run on a DGX Spark

Best open-weight models to run on a DGX Spark

By Samuel Seidel · Published September 9, 2026

GPUwerk has published measured or reported tok/s figures for six models on a DGX Spark, each with its own page. This is the navigation layer: which one fits which job, using only the numbers already on those six pages, nothing new.

The one-line verdict per model

Chat, single person

Qwen3-Coder-30B-A3B at 80.9 tok/s is the fastest single-stream figure across all six pages, well above the roughly 10 tok/s line GPUwerk's benchmarks page draws for a chat interface to feel usable. gpt-oss-120b on llama.cpp at 60.5 tok/s is the other real option if you specifically want llama.cpp's engine. Neither Nemotron Super 49B at 5.79 tok/s nor Llama 3.3 70B at 6.0 tok/s clears that bar; both are dense models where every token reads the full parameter count.

Coding

Qwen3-Coder-30B-A3B again, both by name, it is a coding-tuned model, and by number: 80.9 tok/s single-user, with prefill at 5,705 tok/s on an 8,192-token prompt for fast handling of a large file or diff as context. It is also the model GPUwerk stages by default alongside Llama 3.3 70B, so it deploys with no download step.

Batch and agent-fleet work

This is where the dense-versus-MoE split matters most. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent under vLLM, the highest aggregate figure among these six. Nemotron Super 49B, despite a poor single-user number, reaches 695.1 tok/s aggregate at the same concurrency, a 120-fold rise from its single-stream figure, exactly the shape of a model that only makes sense batched. Qwen3-Coder-30B-A3B also holds up here at 3,235 tok/s aggregate. Llama 3.3 70B is the outlier to watch: its aggregate throughput peaks at 485 tok/s at 128 concurrent, then falls to 300 tok/s at 256 as the dense model collapses past its knee, so batch jobs on it need to stay near that 128-concurrency point rather than assume more concurrency is always better.

The dense-model reality check

Two models on this list make the same point from different directions. Llama 3.3 70B, a 70B dense model, decodes at 6.0 tok/s single-user and its aggregate throughput actively falls once concurrency passes its knee, worse at everything at once rather than trading latency for throughput. DeepSeek-V4-Flash, at 304B total parameters, does not fit on one Spark at any quality-preserving quantization at all; the 14 to 21 tok/s figure reported for it required a patched, non-stock vLLM fork and 2-bit expert planes, a research configuration rather than something to build a service on, and GPUwerk has not reproduced those numbers on its own fleet. Read together, they are the honest counterweight to the mixture-of-experts models above: total parameter count does not predict speed on this hardware, active parameter count does, and a dense model much past 30 to 50B starts fighting the 273 GB/s bandwidth ceiling that GPUwerk's benchmarks page and glossary both explain.

How to pick

Start from the job, not the model name. One person typing prompts: Qwen3-Coder-30B-A3B. Many people or an agent fleet, tolerant of per-request latency: gpt-oss-120b or Nemotron Super 49B, sized to your concurrency. Long documents with short answers: Nemotron-3-Super-120B for its prefill number. Anything that reaches for a 70B-plus dense model, or DeepSeek-V4-Flash specifically, deserves a second look at whether an MoE model of similar total size does the same job faster on this hardware, per the numbers above.

Rent a Spark at $0.79/hour and run your own prompts against whichever model this points at before trusting the numbers on any of these seven pages, including this one. A first engagement covers picking and tuning a model for a specific workload if you would rather hand that off.

First top-up: pay $10, get $20 in credit

Stop reading model pages. Run one.

A dedicated Spark, deployed in minutes, with two models staged by default.

Deploy a Spark Browse all models