DGX Spark specs
A DGX Spark pairs an NVIDIA GB10 Grace Blackwell superchip with 128 GB of unified LPDDR5x memory at 273 GB/s of shared bandwidth, a 20-core Arm CPU, 1 TB of NVMe storage, and 200 Gb/s networking, rated at 1 PFLOP of sparse FP4 AI performance. Reading the spec sheet is one thing; knowing which numbers govern your actual workload is another, and that is what this page adds.
The full spec sheet
This is the table from our DGX Spark hardware page, unchanged, for reference:
| Spec | Value |
|---|---|
| Superchip | GB10 Grace Blackwell |
| AI performance | 1 PFLOP (sparse FP4) |
| Memory | 128 GB unified LPDDR5x |
| Memory bandwidth | 273 GB/s |
| CPU | 20-core Arm (10x X925 + 10x A725) |
| Storage | 1 TB NVMe |
| Network | ConnectX-7, 200 Gb/s |
| Power | 240 W max |
| OS | DGX OS 7, Ubuntu 24.04 |
| CUDA | Full stack preinstalled |
What each spec means in practice
273 GB/s memory bandwidth: the ceiling on single-stream speed
This is the number that bounds how fast one conversation generates tokens. Autoregressive generation is memory-bound: producing each token means streaming the model's active weights out of memory once, so decode speed is roughly bandwidth divided by active model size. Our published benchmarks spell out the arithmetic: 273 GB/s divided by roughly 3 GB of active weights for a mixture-of-experts model like gpt-oss-120B predicts about 90 tokens/sec as a theoretical ceiling, and real engines land at 50 to 80% of that, matching the measured 60 tok/s we recorded on llama.cpp. A dense 49B model, where every token reads the full parameter count, measured only 5.79 tok/s on the same hardware, and the bandwidth math predicts that result before you run it.
128 GB unified memory: a capacity number, not a speed number
This is what decides whether a model fits at all, independent of how fast it then runs. It is enough to hold 70B-class models without splitting them across multiple GPUs, which is the headline reason to consider a Spark over a consumer card in the first place. It is a capacity ceiling, not a speed advantage: a model that fits in 128 GB still generates tokens at whatever rate the 273 GB/s bandwidth allows.
1 PFLOP sparse FP4: for batch and fine-tuning, not chat latency
NVIDIA's headline "1 petaFLOP of AI performance" figure is sparse FP4 compute, a best-case number you will not see reflected in a single chat session's response time. It matters when many requests are processed together, batch inference serving several users at once, or fine-tuning, where the GPU is doing dense parallel work rather than waiting on memory for one token at a time. Our benchmarks page shows this directly: concurrency numbers climb into the hundreds and low thousands of tokens/sec in aggregate even where single-stream decode sits in the tens, because batching lets compute, not just bandwidth, do the work.
20-core Arm CPU: orchestration, not the bottleneck
The Grace CPU side handles the OS, data loading and orchestration around the GPU workload rather than the inference math itself. It is capable enough that it is rarely the limiting factor in an LLM serving workload; the GPU and memory bandwidth are what to watch.
1 TB NVMe storage: enough for models, not a data lake
This is what every GPUwerk rental instance includes as well, per our pricing page. It comfortably holds several large model checkpoints plus working data for a fine-tuning run, but it is local, single-node storage: GPUwerk keeps only a periodic recovery copy of it for hardware failure, not a backup service, so anything you need to keep has to leave the node before you stop trusting that copy.
200 Gb/s ConnectX-7 networking: node-to-node, not your internet uplink
This is the interconnect used to link two Sparks into a 256 GB cluster for models too large for one node, not a measure of how fast your own internet connection to the machine will be. On GPUwerk's rental service it is included in the hourly rate with no separate egress charge.
240 W max power: a system peak, not a running average
NVIDIA states this as the whole-system peak, not a sustained draw figure. Owner reports referenced in our before-you-buy guide put idle draw around 40 W, though we have not seen an instrumented measurement we would publish as fact ourselves. If you are buying rather than renting, that gap between idle and peak power is worth budgeting for; if you rent, power is included in the hourly rate.
How the Spark compares to other hardware
Specs on their own do not answer whether a Spark is the right machine for a given workload; that depends on what you are comparing it to. We wrote head-to-head pages for the machines people actually cross-shop against: DGX Spark vs RTX 5090 for raw single-GPU speed against capacity, DGX Spark vs H100 and DGX Spark vs H200 for the datacenter comparison, plus DGX Spark vs Mac Studio for the other unified-memory option people consider.
FAQ
What does 273 GB/s of memory bandwidth actually limit?
Single-stream decode speed, the rate at which one conversation generates tokens. Autoregressive generation is memory-bound: every token means streaming the active weights out of memory, so decode speed is roughly bandwidth divided by active model size. It does not limit how large a model fits or how well batched, multi-user serving scales.
Is 128 GB of unified memory the same as 128 GB of VRAM?
Functionally yes for capacity: the CPU and GPU share the same 128 GB pool over the GB10's coherent memory, so it behaves like one large VRAM pool for a model that fits. It is not as fast as a discrete GPU's dedicated VRAM, that is what the 273 GB/s bandwidth figure describes.
What is the 1 PFLOP figure for and when doesn't it matter?
1 PFLOP is NVIDIA's sparse FP4 compute figure, a best-case marketing number you will not see in a serving log. It matters for batch throughput and fine-tuning, where many tokens are processed in parallel. It has almost no bearing on single-stream chat latency, which is bounded by memory bandwidth instead.
Related reading: what a Spark costs to buy or rent, current stock and availability in Europe, where to buy one in Europe, and how to rent one by the minute.