How GPUwerk benchmarks the DGX Spark
Every model page and the benchmarks page cites a tok/s figure. This page is the answer to "how was that measured", written once so it doesn't need repeating on every page that cites a number.
What GPUwerk measures itself, versus what it cites
Two models are staged on every node in the fleet: Qwen3-Coder-30B-A3B (AWQ) and Llama 3.3 70B (AWQ). Those are the only two GPUwerk measures on its own hardware, using vllm bench serve against vLLM 0.27.1, because they're the two it can re-run at will rather than take on faith. Every other model figure on the site, gpt-oss-120b, Nemotron-3-Super-120B, Nemotron Super 49B, DeepSeek-V4-Flash, is a published third-party number with a named source and a link, not something GPUwerk has reproduced on its own fleet. The benchmarks page keeps those two categories visually separate for exactly this reason.
Single-stream first, then concurrency
A single-user figure, one request in flight, answers whether a model feels usable typing into a chat window. It's the number most reviews stop at, and it's also the worst case for this hardware: memory bandwidth is fixed at 273 GB/s, so a lone request pays the full latency cost of every token with nothing to share it with.
That's why GPUwerk also sweeps concurrency, from 1 request up through 256, and reports aggregate throughput and per-request speed at each step rather than a single point. Batch work, agent fleets, and shared endpoints behave nothing like single-stream chat: on gpt-oss-120b aggregate throughput rises from 33.5 tok/s at concurrency 1 to 862.8 tok/s at 256, a real number, but the same sweep shows per-request speed falling to under 4 tok/s at that point, which is why both numbers are published together and neither is presented alone as "the" DGX Spark speed.
Fixed inputs, so the number means something
GPUwerk's own two-model benchmark run uses 512 input tokens and 256 output tokens per request, three requests per concurrency level, on a single GB10 node with nothing else scheduled on its GPU. Prefill is measured separately, median time to first token on an 8,192-token prompt, because prefill and decode are bound by different resources (compute versus memory bandwidth) and blending them into one number would hide which workload the figure actually predicts. A model that prefills fast and decodes slowly suits long-document, short-answer work; the reverse suits short-prompt, long-answer chat. Reporting them separately is the only way a reader can tell which one they're getting.
Raw logs, not just the summary table
The unedited vllm bench serve logs for both fleet models are linked from the benchmarks page: Qwen3-Coder 30B-A3B and Llama 3.3 70B. Each log opens with the GPU's process list and free memory at the instant the run started, so a reader can check that the node was actually idle rather than trust a summary number. That header exists because of a mistake GPUwerk made once and is not trying to bury: a table published on 2 September 2026 turned out to have been measured on a node that was serving another model of GPUwerk's own the entire time, on the same GPU. A single nvidia-smi reading before each run had looked clean, but that's a point sample that can land between decode steps; the other service's own request log showed 20 to 42 generations per minute running throughout. The table was withdrawn the same day, and the whole set was re-run on an emptied node rather than corrected in place, because the concurrency curve had changed shape, not just height, and annotating old numbers wouldn't have surfaced that.
What "GPUwerk-measured" and "reported" mean on model pages
Model pages and the best-models guide mark each figure as either GPUwerk-measured or reported. GPUwerk-measured means the number above: this fleet, this engine version, this run, with a linked log. Reported means a figure pulled from a published source elsewhere, credited by name, and not independently reproduced. DeepSeek-V4-Flash's 14 to 21 tok/s figure is the clearest example: it comes from a patched, non-stock vLLM fork run by someone else on the model's Hugging Face discussion thread, and GPUwerk has not reproduced it. Treating that distinction as visible on the page, rather than folding every number into one undifferentiated table, is the point of the methodology.
What this doesn't cover
Quantization format, kernel version, unified-memory settings and --max-num-seqs all move the result, and GPUwerk's own numbers will not exactly match a different kernel release or a different quantization of the same model, a point the benchmarks page flags directly rather than implying its numbers are universal constants. The only benchmark that answers your specific question is one run against your own prompts and your own output lengths, and both engines used here, vLLM and llama.cpp, ship a harness for exactly that.
Rent a Spark at $0.79/hour and run the same vllm bench serve or llama-bench commands against your own workload before trusting any number on this site, including this page's description of how the numbers were made.