Scaling
DGX Spark/Scaling from one DGX Spark to a fleet

Scaling from one DGX Spark to a fleet

By Samuel Seidel · Published September 9, 2026

A single rented Spark runs one model, or a small set of them, for as many concurrent requests as its memory and bandwidth allow. Past a point, more concurrency means more nodes, not a bigger one. This guide covers when that point arrives, how it differs from linking two Sparks into a 256 GB cluster, and the practical work of routing requests across a small fleet of independent rentals.

Fleet scaling is a different problem than clustering

GPUwerk's two-node cluster guide covers linking two Sparks over ConnectX-7 into one 256 GB memory pool for a single model that does not fit on one node. That is a capacity problem: the model itself is too big. Fleet scaling is a different problem: the model fits fine on one Spark, but one Spark cannot serve enough concurrent requests fast enough. The fix there is not a bigger pool, it is more independent copies of the same model, each fully loaded on its own node, with requests spread across them.

Do not confuse the two and cable Sparks together when what you actually need is parallel capacity. A linked pair still serves one workload from two nodes acting as one memory space; it does not double your request throughput for a model that already fits on a single 128 GB node. For that, rent Sparks independently and put a router in front of them.

When to move from one node to several

The trigger is queueing, not a fixed request count, since throughput depends heavily on model size and quantization. Watch your serving framework's own queue depth and time-to-first-token. If requests are waiting behind an occupied GPU rather than being rejected outright, you are past the point where a single node's aggregate throughput covers your traffic. GPUwerk's benchmarks page publishes aggregate throughput at various concurrency levels for models it has measured, for example Qwen3-Coder-30B-A3B at roughly 3,235 tok/s aggregate at 256 concurrent requests on one Spark under vLLM; if your traffic pattern is already saturating a number like that, a second node is the next step, not a bigger prompt budget or a different quantization.

A second, unrelated trigger is availability: a single Spark is a single point of failure. If a stopped or terminated instance taking your only endpoint down is not acceptable, running two nodes behind a router gives you redundancy even before you need the extra throughput.

Renting the second and third node

Each additional Spark is a separate deploy in the console, billed independently at the same $0.79/hour on-demand rate as the first. There is no fleet discount published, and no shared billing entity across instances; each is its own instance with its own SSH port, its own workspace, and its own bill. Confirm capacity in the console before assuming a second node is available at the region you want, the way you would for the first.

Deploy the same model image to each node so they are interchangeable from the router's point of view. If you built a custom image on the first Spark, either export and redeploy it or script the setup so the second and third nodes match exactly; a router that assumes identical backends will misbehave against nodes running different model versions or quantizations.

Load balancing across independent nodes

Each Spark keeps its own public HTTPS endpoint and its own SSH port, per GPUwerk's SSH docs, so nothing about the platform gives you a single fleet-wide address out of the box. You put that layer in front yourself.

Whichever router you pick, health-check the backend endpoints directly rather than trusting that an instance is up because you deployed it. A stopped instance still resolves its hostname; it will not answer a health check, and a router that skips this step will keep routing traffic at a dead node until requests start failing.

Two-node cluster vs a two-Spark fleet: the actual decision

Put plainly: cluster if one copy of your model does not fit in 128 GB. Fleet if it fits but you need more concurrent capacity or redundancy than one node gives you. The two are not mutually exclusive at larger scale, a fleet of clustered pairs is a legitimate shape for a workload that needs both a large model and high concurrency, but most GPUwerk customers land on one side or the other, not both.

If you are not sure which side you are on, rent one Spark first, run your model, and check whether it fits and whether one node's throughput covers your traffic before committing to either path. The two-node cluster guide has the memory math for the clustering side of that decision.

A first engagement can help size a fleet before you rent nodes you do not end up needing.

First top-up: pay $10, get $20 in credit

Rent your next node in minutes.

Each Spark is $0.79/hour on demand, deployed independently, so you scale a fleet one node at a time.

Deploy a Spark See pricing