DGX Spark two-node cluster: setup and what it actually buys you
Two DGX Sparks link natively over 200 GbE into a 256 GB pool. Before you cable anything, know what that pairing does and does not change: memory capacity goes up a lot, and single-user generation speed should not move much, by the same bandwidth arithmetic that governs a single node. This page covers both halves, the honest expectations and the practical cabling.
What two Sparks buy you
GPUwerk has not published a measured two-node inference benchmark, so this section states what is established elsewhere on the site and is explicit about where it is prediction rather than measurement.
- Combined memory. Two 128 GB units pool to roughly 256 GB of addressable unified memory over the ConnectX-7 link, per GPUwerk's cloud providers comparison, which lists the two-node option at $1.79/hour, and its fine-tuning guide, which puts a 405B model's QLoRA memory need within reach of a two-Spark pool where it is not on one node.
- Generation speed should not meaningfully change. Decode speed is bandwidth-bound, set by how fast the node holding a model's active weights can read them out of memory, per the arithmetic on GPUwerk's benchmarks page. Networking a second node in does not raise that per-node bandwidth figure, so single-stream tokens per second for one user should look close to the single-node number for the same model and quantization. GPUwerk has not run this pairing itself to confirm it, and states it as a prediction from the published arithmetic rather than a measurement.
- Prompt processing: not measured for a pair. Prefill is compute-bound rather than bandwidth-bound, and it is plausible that splitting a large prompt's work across two nodes' compute could help. GPUwerk has no measured figure for this and is not going to publish one it has not run.
The practical read: pair two Sparks to fit a model, or a fine-tune, that does not fit on one node. Do not pair them expecting your chat interface to feel faster, because the number that governs single-user decode speed, per-node bandwidth, is unlikely to change when a second machine joins over the network, though GPUwerk has not measured that directly.
Cabling
Each DGX Spark ships with two QSFP ports on a ConnectX-7 adapter, rated at 200 GbE, per NVIDIA's own DGX Spark specifications, also listed on GPUwerk's hardware page. Linking two units at the physical layer means:
- A QSFP cable rated for the port speed, connected directly between one ConnectX-7 port on each Spark. A direct-attach copper cable is the usual choice for a short run between two boxes on the same desk or rack shelf; use fiber if the units are farther apart than copper reliably handles at 200 GbE.
- Both ports need to be free. If either QSFP port is already in use for a separate network connection, either free it or use the second port pair instead.
- Power and boot both units normally before connecting the cable is not a strict requirement, but bringing the link up after both machines have finished booting avoids diagnosing a network problem that is actually a boot-order issue.
What NVIDIA's own documentation covers
NVIDIA ships DGX OS with the networking stack and clustering tooling needed to bring the ConnectX-7 link up and present the pair to a serving framework as one target. In general terms, per NVIDIA's own DGX Spark and DGX OS documentation, that process covers: confirming the physical link is up and negotiating at the expected speed, assigning the interface an address on a point-to-point or small subnet dedicated to the inter-node link, and configuring the distributed inference or training framework, such as vLLM's multi-node serving mode, to use that interface for tensor or pipeline parallelism across the two nodes. GPUwerk has not independently reproduced NVIDIA's own clustering walkthrough end to end and is not publishing a copy-paste command sequence here; follow NVIDIA's current DGX OS documentation for the exact commands, since flag names and defaults can change between DGX OS releases.
Once the interface is up, which framework you point at the pair matters more than the network settings. GPUwerk's vLLM doc covers single-node serving in detail; a two-node deployment is the same framework configured for multi-node tensor or pipeline parallelism rather than a different tool.
Should you pair two Sparks?
Pair them if a model, or a fine-tune, you want to run does not fit in 128 GB. Skip it if the goal is faster single-user chat: that number is set by per-node bandwidth, and pairing is unlikely to move it, per the arithmetic on GPUwerk's benchmarks page. The DGX Spark vs Mac Studio comparison goes further into single-node decode speed if that is the axis you actually care about.
If you are unsure which side of that line your model and workload land on, rent a single Spark first at $0.79/hour, confirm it does not fit, then rent a second at $1.79/hour combined to test the paired configuration before buying anything.
A first engagement can help scope whether your workload actually needs a second node before you cable one up.