Buying analysis

What to know before buying a DGX Spark

By the GPUwerk fleet team · Updated August 26, 2026 · 11 min read

We sell these machines, we rent these machines, and we run a fleet of them. That gives us an obvious commercial interest, so this guide leads with the numbers that argue against buying one. If the Spark is wrong for your workload, we would rather you find out here than after the invoice clears.

The one number that decides everything: 273 GB/s

The DGX Spark pairs a GB10 Grace Blackwell superchip with 128 GB of coherent unified LPDDR5x memory. NVIDIA markets it at "1 petaFLOP of AI performance", that figure is sparse FP4, a best-case marketing number, not something you will see in a serving log. The number that actually governs your day-to-day experience is memory bandwidth: 273 GB/s, shared between the CPU and the GPU.

Autoregressive token generation is memory-bound. Every token requires streaming the active weights out of memory, so single-stream decode speed is roughly bandwidth divided by active model size. For context: an RTX 5090 has about 1.79 TB/s, the M5 Ultra Mac Studio announced August 25, 2026 has 1.2 TB/s (Apple; the outgoing M3 Ultra was about 800 GB/s), and an H100 about 3.35 TB/s of HBM3. The Spark is an order of magnitude behind the datacenter part and less than a quarter of a top Mac Studio.

Do the arithmetic before you buy. A 70B model at FP8 is ~70 GB of weights; 273 GB/s divided by 70 GB puts the theoretical ceiling under 4 tokens/sec, and LMSYS measured 2.7 tokens/sec decode on Llama 3.1 70B FP8 under SGLang at batch 1. That is not a benchmarking artifact or a driver bug. It is physics, and no firmware update will change it.

Published benchmark numbers

These are figures from third-party reviews, not our own marketing. Where sources disagree we say so below the table.

WorkloadPrefill (tok/s)Decode (tok/s)Source
GPT-OSS 20B MXFP4, Ollama, batch 12,05349.7LMSYS
Llama 3.1 8B FP8, SGLang, batch 17,99120.5LMSYS
Llama 3.1 8B FP8, SGLang, batch 327,949368 (aggregate)LMSYS
DeepSeek-R1 14B FP8, SGLang, batch 82,07483.5 (aggregate)LMSYS
Llama 3.1 70B FP8, SGLang, batch 18032.7LMSYS
gpt-oss 120B MXFP4, single stream33.5Dendro Logic
gpt-oss 120B MXFP4, 256 concurrent862.8 (aggregate)Dendro Logic
Nemotron Super 49B, 256 concurrent695 (aggregate), vs 5.8 at batch 1Dendro Logic

Where the sources disagree. Single-stream reviews and concurrency reviews reach opposite verdicts on the same box, and both are correct. LMSYS's batch-1 numbers make the Spark look slow; Dendro Logic's 120× throughput scaling from batch 1 to 256 concurrent streams makes it look like a bargain server. The disagreement is entirely about what you are measuring. If you are one person typing into a chat window, believe the batch-1 numbers. If you are running a batch extraction pipeline or serving a team, believe the concurrency numbers.

Third-party 70B-class Q4 figures also vary widely, some roundups quote 35–45 tok/s for the Spark on 70B Q4, which is hard to reconcile with LMSYS's 2.7 tok/s on 70B FP8. The gap is quantization (Q4 halves the bytes moved versus FP8) plus, in some cases, aggregate rather than per-stream reporting. Treat any 70B number above ~10 tok/s per stream with suspicion unless the article states the quantization and the batch size. Ours: expect single-digit to low-teens tok/s per stream on dense 70B, and 30s on MoE models like gpt-oss 120B where only a fraction of parameters are active per token.

The honest weaknesses

What it is genuinely good at

Where to actually buy one in Europe

"DGX Spark" is the NVIDIA Founders Edition. The same GB10 board ships from OEM partners at different prices, and in the EU the OEM units have generally been cheaper and easier to get than the Founders Edition.

ModelStreet price (Aug 2026)Notes
NVIDIA DGX Spark Founders Edition$4,699 US / ~€4,800 EUUp from the $3,999 announcement price. Reference design, 4 TB SSD.
ASUS Ascent GX10~€3,780 typical, seen from ~€2,900Widely stocked in the EU (LDLC, idealo listings). Usually the cheapest route in; 1 TB SSD on the base SKU.
Dell Pro Max with GB10 (FCM1253)~$5,650 (4 TB) US, up from ~$3,999 at launch; approaching £6,000 UKBest build quality and enterprise support; EU/UK pricing is markedly worse than US.
MSI EdgeExpert / other GB10 partnersVaries, roughly in the ASUS bandSame silicon, differing SSD sizes and warranty terms.

Two things buyers routinely miss. First, SSD capacity differs by SKU (1 TB to 4 TB) and model weights are large, a 1 TB unit fills faster than you expect. Second, EU list prices carry VAT that US comparisons do not; compare ex-VAT to ex-VAT or you will misjudge the gap by ~20%.

Total cost of ownership: buy vs rent vs build a 5090 rig

Three-year view, one machine's worth of capacity, EU electricity at roughly €0.25/kWh. Power assumes a realistic duty cycle rather than 24/7 full load.

Buy a SparkRent in the cloudBuild a 5090 rig
Up-front€3,800–€4,800€0€4,200–€5,600 (5090 ~€3,800 street + platform)
Power, 3 yr~€260 (≈40 W idle / 150 W avg mixed use)Included~€700–€1,000 (600 W board power, higher idle)
Ongoing / opexYour time: OS, drivers, updates, backups$2.90/hr on-demand, or $1,490/mo reservedYour time, plus more of it, you own the driver stack
3-year total (heavy use)~€4,100–€5,100~$53,600 reserved 24/7~€4,900–€6,600
3-year total (8 hrs/day, weekdays)Same, you own it~$18,100 at $2.90/hrSame, you own it
Max model that fits120B-class, residentWhatever you rent~32B Q4; beyond that, PCIe offload
Single-stream speedSlow (2.7–50 tok/s by model)Whatever you rentFast, 205 tok/s on 20B MXFP4, if it fits in 32 GB
Data residencyYour buildingDepends entirely on providerYour building

How to read that table. Owning is dramatically cheaper than 24/7 rental, a Spark pays for itself against $2.90/hr in roughly nine weeks of continuous use. That is the honest case for buying, and it is a strong one. But almost nobody runs a dev box continuously. If your real usage is a few hours a day of bursty experimentation, on-demand rental at per-minute billing costs less than the capital and never leaves you holding depreciating silicon when the GB10 successor lands.

The 5090 rig is the right answer more often than Spark vendors like to admit. If everything you run fits in 32 GB, 7B to 32B models, image generation, fine-tuning small models, a 5090 is roughly 4× faster on both prefill and decode (LMSYS: 8,519 tok/s prefill and 205 tok/s decode on GPT-OSS 20B, versus 2,053 / 49.7 on the Spark) at a comparable or higher total system price: Tom's Hardware's GPU price tracker had 5090s at a $4,381–$4,829 street range in August 2026, well above the $1,999 MSRP, which puts a complete rig level with or above a Spark. Buy the Spark when the 128 GB is the point. Buy the 5090 when speed is the point. Anyone telling you the Spark beats a 5090 at things that fit in 32 GB is selling.

Two costs the table cannot show. Cloud GPU pricing is volatile, H100 on-demand rates move by double-digit percentages within a single quarter, and RTX 5090 rentals span $0.25 to $1.49/hr across providers, so a rental budget built on today's spot rate is a guess. And owned hardware carries an operational tax: firmware, driver upgrades on a rare ARM+CUDA-13 combination, backups, and the afternoon you lose when a vLLM release breaks sm_121. That tax is real whether or not it appears on an invoice; it is the reason our on-premise offering is managed rather than a box on a pallet.

So: is a DGX Spark worth it?

If you do buy, our setup guide covers first boot through serving a model, including the parts the quick-start skips.

Test it before you spend €4,000.

Rent a dedicated DGX Spark in EU-Central by the minute, run your actual workload, and decide with data. Your first $100 top-up gets $50 free. If it's the right machine, we'll deliver and install one.

Deploy a Spark Buy it installed