Run Qwen3-235B-A22B on a DGX Spark: one node vs a two-node cluster
Qwen3-235B-A22B is a mixture-of-experts model, but its total size still overwhelms a single Spark. One published report ran it two ways: crawling at roughly 1.8 tok/s on one node, and roughly 7 times faster clustering two nodes over llama.cpp's RPC backend. Both figures are third-party, cited below. GPUwerk has not run this model on its own fleet.
Single node versus two-node RPC
| Configuration | Engine | Decode |
|---|---|---|
| One Spark, Q4_K_XL | llama.cpp | ~1.8 tok/s |
| Two Sparks, RPC over TCP/IP, layers split manually | llama.cpp | ~12.5 tok/s |
Both figures from an NVIDIA developer forum report on single- and multi-node Qwen3-235B inference, quoted on GPUwerk's benchmarks page, updated 14 September 2026. GPUwerk has not reproduced either figure on its own fleet.
The roughly 7× speedup is a real, published result. It is also, by the report's own description, a workaround rather than a production configuration: llama.cpp's RPC backend splits model layers across nodes manually over a TCP/IP connection, not a native tensor-parallel pipeline. The report's author estimated that a proper multi-node setup using vLLM with NCCL could do better still, but that configuration was not the one tested. Read the 12.5 tok/s figure as evidence that clustering helps on a model this size, not as a ceiling on what clustering can do.
Why one Spark struggles here
Qwen3-235B-A22B activates 22B parameters per token by name, which by the decode-speed arithmetic on the benchmarks page should land somewhere in the tens of tokens per second at 4-bit, similar to a dense 22B model. The measured 1.8 tok/s is far below that, which points to a second constraint: at 235B total parameters, even Q4_K_XL leaves the model's weights consuming most of a single Spark's 121.6 GiB of addressable unified memory, with little headroom left for KV cache, activations or the CUDA context. When a model this close to the memory ceiling runs on unified memory shared with the host, paging and reduced batching room can suppress throughput well below what the active-parameter count alone predicts. The report does not break out how much of the shortfall is bandwidth versus memory pressure, so treat that as the likely explanation rather than a confirmed one.
What this says about giant MoE models
The general MoE advantage on this hardware, explained on the models overview, is that decode speed tracks active parameters, not total ones. That holds at moderate scale: a 30B-A3B model runs close to the speed of a dense 3B model. It stops holding once total size pushes the model's weights to the edge of what a single Spark can hold, because then the constraint shifts from bandwidth to raw capacity, and a model that barely fits behaves worse than the active-parameter arithmetic predicts.
Qwen3-235B-A22B is the demonstration case for that limit on this hardware. It is not proof that MoE stops mattering at large scale in general, only that a single 128 GB Spark is not enough headroom to realize the MoE speed advantage once the model itself is this large.
The 256 GB cluster option
GPUwerk pairs two Sparks into a combined 256 GB pool at $1.79/hour, described in the two-node cluster guide. The extra memory is the point for a model like this: more room for weights and KV cache before the box is fighting itself for headroom, which is the more likely explanation for the 1.8 tok/s single-node figure above than pure bandwidth. GPUwerk has not published a measured two-node inference benchmark of its own for any model, this one included, so the 12.5 tok/s figure above is the best published reference point for what pairing can do, not a GPUwerk result.
Try it
A single Spark rents at $0.79/hour, and the paired 256 GB configuration at $1.79/hour combined, both billed per minute against a prepaid balance on the pricing page. Given the single-node figure above, start with a smaller model on one Spark before committing to a cluster rental for this one. If you want someone else to size the configuration first, GPUwerk's first engagement covers exactly that.
FAQ
How fast is Qwen3-235B-A22B on a single DGX Spark?
Roughly 1.8 tok/s at Q4_K_XL, reported in an NVIDIA developer forum report on single- and multi-node Qwen3-235B inference. That is well under the roughly 10 tok/s line where a chat interface feels usable.
Does clustering two Sparks make Qwen3-235B-A22B usable?
The same report clustered two Sparks with llama.cpp's RPC backend over TCP/IP and reached roughly 12.5 tok/s, about 7 times faster, by splitting layers manually across both nodes. That is a real result but a workaround rather than a native tensor-parallel pipeline, and the author estimated a proper vLLM-plus-NCCL setup could do better.
Has GPUwerk measured Qwen3-235B-A22B on its own fleet?
No. Both figures on this page are third-party numbers from a single published report. GPUwerk has not run this model, single-node or clustered, and has not published a measured two-node inference benchmark for any model.