One DGX Spark or a two-node cluster: which do you need
People reach for a second Spark expecting it to make things faster. That's the wrong reason to buy one, and it's worth saying plainly before the pricing page does the persuading for you. The question that actually decides this is whether your model fits in 128 GB, not how impatient you are watching it generate.
The number that matters: memory, not speed
A single DGX Spark carries 128 GB of unified memory at $0.79 an hour. Two Sparks linked over their 200 GbE ConnectX-7 ports pool into 256 GB at $1.79 an hour combined. That's the entire physical difference between the two configurations: double the memory ceiling, double the hourly rate. Nothing about the link changes how fast either node computes on its own.
Single-user decode speed on this hardware is bandwidth-bound: how many tokens per second one person sees is set by how fast a node can move weights through memory for each generated token, which is a function of that node's own memory bandwidth, not of how many nodes are wired together. Pairing a second Spark doesn't add bandwidth to the node doing the generating for one user; it adds a second pool of memory and a second node's worth of compute available to whatever framework you point at the pair. If the goal is a snappier single-user chat, a second node is the wrong lever, and GPUwerk's DGX Spark vs Mac Studio comparison goes into that decode-speed arithmetic directly if that's the axis you actually care about.
When one Spark is enough
If your model, at whatever quantization you're planning to run, fits comfortably in 128 GB with headroom for context and KV cache, one Spark is the right unit. That covers most of the models GPUwerk has measured directly: gpt-oss-120b at MXFP4, Qwen3-Coder 30B-A3B at AWQ, Llama 3.3 70B at AWQ, and Nemotron-3-Super-120B at NVFP4 all run on a single node, per the figures on GPUwerk's model pages. Start here even if you're not certain, because at $0.79 an hour it's the cheaper way to find out, and it tells you directly whether you need to go further rather than guessing from a spec sheet.
When you need the second node
Pair two Sparks when a model, or a fine-tune of one, genuinely does not fit in 128 GB. That's a memory-capacity problem, and 256 GB pooled is the only way this hardware line solves it, short of a different class of machine entirely. A large dense model, a large mixture-of-experts model with many active parameters, or a fine-tuned checkpoint that's grown past the original model's footprint are the situations where this actually applies.
It's worth being honest that GPUwerk has not published a measured two-node inference benchmark. What's on this page is arithmetic and vendor documentation, not a number we ran ourselves, and the practical setup detail, cabling, the ConnectX-7 link, and pointing a framework like vLLM's multi-node serving mode at the pair, is covered in full in GPUwerk's two-node cluster guide rather than repeated here.
A short framework
Ask: does the model fit in 128 GB at the quantization you're planning to use? If yes, one Spark, and the $1.79 rate buys you nothing you need. If no, a two-node cluster is very likely the only configuration on this hardware line that fits it at all, and the extra cost is the price of the memory ceiling, not of extra speed. If you're not sure which side of that line you're on, rent a single Spark first, try the model, and only add a second node once you've confirmed it doesn't fit rather than budgeting for both up front.
The mistake to avoid is buying the second node to fix a complaint about slowness. If a model already fits on one Spark and generation still feels slow for one user, the fix is a smaller model, a more aggressive quantization, or accepting the bandwidth ceiling that single-node hardware has, not a second Spark sitting next to the first one.
What actually changes once the second node is in
Once you do need the pair, it's worth knowing what you're signing up for beyond the extra memory. Two Sparks linked over their 200 GbE ConnectX-7 ports need a serving framework that understands multi-node tensor or pipeline parallelism, which in practice means vLLM configured for that mode rather than a different tool entirely. That's a real setup step, not a checkbox: the interface has to come up, negotiate its expected speed, and get assigned an address on a link dedicated to the pair, before the framework can treat the two nodes as one target. GPUwerk's two-node cluster guide covers the cabling and the general shape of that process, and points to NVIDIA's own DGX OS documentation for the exact commands, since flag names and defaults shift between DGX OS releases.
It's also worth setting expectations about what you get once it's running: 256 GB of usable memory for weights, context, and KV cache combined, not double the tokens-per-second of a single node. If your plan for the cluster involves both fitting a larger model and expecting a proportional speed jump, only the first half of that plan is grounded in how this hardware behaves.
A cheaper way to find out
Because the second node only pays for itself when the model genuinely doesn't fit, the lowest-risk way to make this decision is to rent, not buy. A single Spark at $0.79 an hour is cheap enough to load the model and watch it fail to fit, or fit comfortably, within minutes rather than guessing from a parameter count and a quantization format on a spec sheet. If it doesn't fit, adding a second Spark at $1.79 an hour combined for a short test run costs far less than committing to hardware or a longer contract before you've confirmed the pairing actually solves your problem.