The DGX Spark, measured.

Everything NVIDIA's spec sheet says, plus what a fleet operator would tell you about the parts it leaves out. Full specs, what actually fits in 128 GB, and where the machine loses.

Full specifications

Per node. Every GPUwerk instance is one whole machine, these numbers are all yours.

NVIDIA DGX Spark
Superchip
GB10 Grace Blackwell
AI performance
1 PFLOP (sparse FP4)
Memory
128 GB unified LPDDR5x
Memory bandwidth
273 GB/s
CPU
20-core Arm (10× X925 + 10× A725)
Storage
1 TB NVMe
Network
ConnectX-7 · 200 Gb/s
Power
240 W max
OS
DGX OS 7 · Ubuntu 24.04
CUDA
Full stack preinstalled

273 GB/s memory bandwidth is the constraint that matters for single-stream LLM generation speed. The Spark's strength is capacity: models that simply don't fit on consumer GPUs, plus FP4 compute for batch serving and fine-tuning, where throughput scales far better than the single-stream numbers suggest.

Fleet benchmarks

We would rather show you an empty table than someone else's numbers. Our first full benchmark run is still going on the fleet, and results land here with methodology and raw logs when it finishes. Until then, the published third-party figures we trust are collected and argued with in the buying guide.

ModelQuantizationFrameworkPrefill tok/sGenerate tok/sMax context
Qwen3 32BQ4vLLMmeasuring…measuring…measuring…
Llama 3.3 70BQ4vLLMmeasuring…measuring…measuring…
DeepSeek-R1 70B distillQ4Ollamameasuring…measuring…measuring…

Want to be notified when the numbers land, or run your own? Your first $100 top-up comes with $50 free, plenty for an evaluation day on a dedicated node.

What should I run on it?

Opinionated guidance from operating a fleet of these.

Team coding assistant

Qwen3-Coder 32B at Q4 leaves room for long contexts and several concurrent users. Serve with vLLM behind an OpenAI-compatible endpoint and point your editors at it.

Private chat + RAG

A 70B-class instruct model at Q4 fits in one node with headroom for an embedding model and a vector store on the same box, a whole private assistant on one machine. See the private ChatGPT setup →

Fine-tuning

LoRA fine-tunes of 70B-class models fit in unified memory without model-parallel gymnastics. Slower per step than a datacenter GPU, but it runs where consumer cards simply can't.

How it compares

Verdict-first comparisons against the machines buyers actually cross-shop. We say where the rival wins.

Guides & analysis

For anyone evaluating, buying, or setting up a DGX Spark, written from the fleet room, not a press kit.

Stop reading benchmarks. Run your own.

Deploy a dedicated Spark in under a minute, your first $100 top-up comes with $50 free, enough for a full evaluation day.

Deploy a Spark