Run Qwen3-Coder 30B-A3B on a DGX Spark: measured tok/s
Qwen3-Coder 30B-A3B AWQ (4-bit) runs on one 128 GB Spark under vLLM at 80.9 tok/s decode for a single user, measured by GPUwerk on 2 September 2026. It is a mixture-of-experts model, so it keeps scaling under load instead of falling over: 3,235 tok/s aggregate at 256 concurrent requests.
Measured on one Spark
| Concurrency | Aggregate throughput | Per-request speed |
|---|---|---|
| 1 | 80.9 tok/s | 80.9 tok/s |
| 8 | 845 tok/s | 35.2 tok/s |
| 32 | 1,679 tok/s | 18.4 tok/s |
| 64 | 2,421 tok/s | 13.1 tok/s |
| 128 | 3,010 tok/s | 8.2 tok/s |
| 256 | 3,235 tok/s | 4.5 tok/s |
Engine: vLLM 0.27.1, quantization: AWQ, 512 tokens in and 256 out, three requests per concurrency slot, zero failed requests at every level. Measured on GPUwerk's own fleet, one GB10 node with nothing else on its GPU, 2 September 2026. Raw log: qwen3-coder-30b-awq-spark.txt. Full context on GPUwerk's benchmarks page.
Prefill on an 8,192-token prompt runs at 5,705 tok/s, dropping to 586 tok/s filling the model's full 262,144-token context, a single request that took 447 seconds, per the benchmarks page. Prefill rate is not constant in prompt length because attention cost grows quadratically, so budget time-to-first-token against the context length you actually send.
Memory footprint
GPUwerk has not published an exact weight size in GB for this AWQ build. What is published: it activates roughly 3B parameters per token (the A3B in the name), about 1.7 GB of weight reads, per the arithmetic on the benchmarks page. That gives a 273 ÷ 1.7 ≈ 160 tok/s theoretical ceiling, and the measured 80.9 tok/s sits at roughly half of it, inside the 50 to 80% band that real engines land in on this hardware. On memory alone it is the roomiest model GPUwerk has measured: vLLM's own doc reports a 924,912-token KV cache at --gpu-memory-utilization 0.85 on an otherwise empty Spark, rising to 994,064 at 0.90.
Q8_0 on llama.cpp, a third-party figure
Everything above is GPUwerk's own measurement, taken under vLLM with an AWQ (4-bit) build. A separate, higher-precision Q8_0 GGUF build of this model has also been benchmarked on a Spark under llama.cpp, cited from its source rather than measured here:
| Format | Engine | Decode | Prefill |
|---|---|---|---|
| AWQ, 4-bit | vLLM | 80.9 tok/s (GPUwerk, measured) | 5,705 tok/s at 8k prompt (GPUwerk, measured) |
| Q8_0 | llama.cpp | 44.3 tok/s (third party, reported) | ~1,654 tok/s (third party, reported) |
Q8_0 figures reported by user eugr in the llama.cpp DGX Spark performance discussion, quoted on GPUwerk's benchmarks page, updated 14 September 2026. AWQ figures are GPUwerk's own measurement from the table above, restated for comparison. GPUwerk has not run the Q8_0 build itself.
The gap is not a contradiction. Q8_0 holds roughly double the bits per weight of the AWQ 4-bit build, which raises memory reads per token and lowers decode speed on this bandwidth-bound hardware; it is a quality-for-speed tradeoff, not a measurement discrepancy between the two engines. Pick AWQ under vLLM for the speed and concurrency behavior on this page. Pick the Q8_0 GGUF under llama.cpp, per the llama.cpp doc, if the extra precision matters more than throughput for a single-user workload.
How to run it
Serve the AWQ build with vLLM's CUDA 13 container, following the pattern in the vLLM doc. GPUwerk's own run used the model's full context length and a high --max-num-seqs, since this model saturates rather than collapses under load:
docker run -d --name vllm \
--gpus all --ipc=host --restart unless-stopped \
-p 8000:8000 \
-e HF_TOKEN="$HF_TOKEN" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:cu130-nightly \
vllm serve <qwen3-coder-30b-a3b-awq-repo> \
--served-model-name qwen3-coder \
--gpu-memory-utilization 0.85 \
--max-model-len 262144 \
--max-num-seqs 256
Sweep vllm bench serve --num-prompts --request-rate against your own prompt shape, per the vLLM doc's verify-and-measure guidance, though on this model there is less risk in overshooting the concurrency setting than on a dense model.
What it is good for on a Spark
A genuine default for both cases. Single-stream, 80.9 tok/s is fast enough for a coding assistant or chat interface for one person. Shared, it keeps climbing to 3,235 tok/s aggregate at 256 concurrent, up only 7.5% from the 128-concurrent figure while per-request latency roughly doubles, so there is little left to win past 128 but also nothing punishing you for pushing further.
That saturate-not-collapse behavior is the practical difference from a dense model of similar total size, such as the Llama 3.3 70B on the same hardware, which falls apart past its own knee instead of leveling off.
Long-context work is a strong secondary case given the prefill numbers, provided the context is not pushed all the way to the model's 262,144-token ceiling, where prefill rate drops by an order of magnitude.
Try it
A Spark rents for $0.79/hour, billed per minute against a prepaid balance, on the pricing page. This is one of the two models staged on every GPUwerk node, so you can deploy and test it with no download step. GPUwerk's first engagement is the option for a managed setup and tuning pass instead.
FAQ
Is Qwen3-Coder 30B-A3B fast enough for a chat interface on a Spark?
Yes. GPUwerk measured 80.9 tok/s single-user under vLLM, comfortably above reading speed and the roughly 10 tok/s line where a chat window starts to feel usable.
What happens to Qwen3-Coder 30B-A3B at high concurrency on a Spark?
It saturates rather than collapses. Aggregate throughput rises to 3,235 tok/s at 256 concurrent requests, up only 7.5% from the 128-concurrent figure, while per-request latency roughly doubles. Nothing is left to win past that point, but nothing breaks either.
How much memory does Qwen3-Coder 30B-A3B need on a Spark?
GPUwerk's pages do not state an exact weight size in GB for this AWQ build. It is small enough that vLLM's own doc reports a 924,912-token KV cache at --gpu-memory-utilization 0.85 on an otherwise empty Spark, the largest cache figure GPUwerk has published for any model.