Run Llama 3.3 70B on a DGX Spark: measured tok/s
Llama 3.3 70B AWQ (4-bit) runs on one 128 GB Spark under vLLM at 6.0 tok/s decode for a single user, measured by GPUwerk on 2 September 2026. It is a dense model, so it is a batch tool here: aggregate throughput peaks at 485 tok/s at 128 concurrent requests, then falls apart at 256.
Measured on one Spark
| Concurrency | Aggregate throughput | Per-request speed |
|---|---|---|
| 1 | 6.0 tok/s | 6.0 tok/s |
| 8 | 125 tok/s | 5.2 tok/s |
| 32 | 298 tok/s | 3.9 tok/s |
| 64 | 466 tok/s | 2.8 tok/s |
| 128 | 485 tok/s | 1.4 tok/s |
| 256 | 300 tok/s | 0.4 tok/s |
Engine: vLLM 0.27.1, quantization: AWQ, 512 tokens in and 256 out, three requests per concurrency slot, zero failed requests at every level. Measured on GPUwerk's own fleet, one GB10 node with nothing else on its GPU, 2 September 2026. Raw log: llama33-70b-spark.txt. Full context on GPUwerk's benchmarks page.
At an 8,192-token prompt, prefill runs at 391 tok/s, against the same model's 5,705 tok/s figure for the smaller Qwen3-Coder 30B-A3B on the same hardware, per the benchmarks page.
Memory footprint
About 37 GiB of weights at AWQ, per GPUwerk's vLLM doc. At --gpu-memory-utilization 0.85 on an otherwise empty Spark that leaves a 207,120-token KV cache; at 0.90 it leaves 227,712 tokens, against roughly 121.6 GiB of addressable memory on the box. It is a dense 70B model, meaning decode reads the full parameter count on every token, which is the direct cause of the 6.0 tok/s single-user figure: 273 GB/s of bandwidth divided by roughly 37 GB of weight reads per token gives a 7.4 tok/s ceiling, and the measured 6.0 sits at about four fifths of it, per the arithmetic on GPUwerk's benchmarks page.
How to run it
Serve the AWQ build with vLLM's CUDA 13 container, following the pattern in the vLLM doc. Raise --max-num-seqs toward the concurrency you plan to run at rather than leaving it at the interactive default of 4:
docker run -d --name vllm \
--gpus all --ipc=host --restart unless-stopped \
-p 8000:8000 \
-e HF_TOKEN="$HF_TOKEN" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:cu130-nightly \
vllm serve <llama-3.3-70b-awq-repo> \
--served-model-name llama-3.3-70b \
--gpu-memory-utilization 0.85 \
--max-model-len 32768 \
--max-num-seqs 128
Sweep vllm bench serve --num-prompts --request-rate against your own prompt shape before picking a concurrency setting, per the vLLM doc's verify-and-measure guidance; the knee on a dense model is a cliff, not a plateau.
What it is good for on a Spark
Not a chat model here. 6.0 tok/s single-stream is under the roughly 10 tok/s line where a chat window feels usable, and no configuration change fixes a dense model's bandwidth ceiling on this hardware.
It is a batch model, and specifically a 128-concurrency one on this box: aggregate throughput climbs to 485 tok/s at 128 requests, which is where GPUwerk's own run peaked. Past that point it collapses rather than plateaus. Going to 256 concurrent drops aggregate throughput to 300 tok/s and pushes median time to first token from around 8 seconds to 40, so the model gets worse at everything at once rather than trading latency for throughput.
Use it for overnight document processing or batch classification sized to stay near the 128-concurrency knee, not past it. If a person is waiting on the output, the smaller Qwen3-Coder 30B-A3B on the same hardware is the better default: it saturates instead of collapsing.
Try it
A Spark is $0.79/hour, billed per minute against a prepaid balance, on the pricing page. Run your own concurrency sweep near the 128-request knee before committing a production workload to this model. GPUwerk's first engagement covers finding and tuning that concurrency level for you.
FAQ
Is Llama 3.3 70B usable as a chat model on a DGX Spark?
Not comfortably. GPUwerk measured 6.0 tok/s single-user under vLLM, under the roughly 10 tok/s line where a chat window feels responsive. It is a dense model, so every token reads all 70B parameters, and no flag changes that.
What is the best concurrency level for Llama 3.3 70B on a Spark?
128 concurrent requests, where GPUwerk measured the peak aggregate of 485 tok/s at 1.4 tok/s per request. Pushing to 256 collapses aggregate throughput to 300 tok/s and pushes median time to first token from 8 seconds to 40, so 256 is worse on every axis, not just slower.
How much memory does Llama 3.3 70B AWQ need on a Spark?
About 37 GiB of weights, per GPUwerk's vLLM doc. At --gpu-memory-utilization 0.85 that leaves a 207,120-token KV cache on an otherwise empty Spark; at 0.90 it leaves 227,712 tokens.