Comparison

DGX Spark vs Strix Halo

By Samuel Seidel · Updated September 4, 2026 · 5 min read

Both machines pack 128 GB of unified memory into a mini-PC. On raw single-user token generation they are close enough that the difference doesn't matter. What actually separates them is prompt processing speed, the software stack underneath, and roughly $1,000.

The verdict, up front

Buy Strix Halo if… you're running one model at a time in llama.cpp, don't need CUDA, and the price gap matters. The Register's hands-on testing found the two boxes "churn out tokens at a similar pace" in single-batch llama.cpp, with Strix Halo taking a narrow lead on its Vulkan backend, for two-thirds to half the Spark's price.

Choose the Spark if… you feed it long prompts or documents, run concurrent or batched jobs, fine-tune anything, or need the software to just work. The same testing measured the Spark's GPU at 2-3x faster time-to-first-token on a short prompt, with the gap widening on longer ones, and roughly 2x faster full fine-tuning. That's before counting the software: anything built on CUDA runs on the Spark without porting; AMD's ROCm/HIP stack is closing the gap but isn't there yet.

Side-by-side specs

Comparing Nvidia's Founders Edition DGX Spark against HP's Z2 Mini G1a, the Strix Halo box The Register tested. Other Strix Halo OEMs (Framework, Minisforum, ASUS) sell 128 GB configurations from a little over $2,000, cheaper than the G1a's tested configuration.

 NVIDIA DGX SparkAMD Strix Halo (HP Z2 Mini G1a)
PlatformGB10 Grace Blackwell superchipRyzen AI Max+ Pro 395 APU
CPU20-core Arm (10× X925 + 10× A725)16-core Zen 5, up to 5.1 GHz
GPUBlackwell, FP4/FP8 tensor coresRadeon 8060S, RDNA 3.5, no low-precision tensor path
Unified memory128 GB LPDDR5x, 273 GB/s128 GB LPDDR5x, 256 GB/s
Peak AI compute~500 dense TFLOPS realistic (1 PFLOP sparse FP4 claimed)~56 TFLOPS BF16 (AMD doesn't publish a peak figure)
NPUNoneXDNA 2, 50 TOPS, limited software support today
Networking10 GbE + ConnectX-7, 200 Gb/s for clustering2.5 GbE, no clustering path
OSDGX OS (Ubuntu-based, Linux only)Windows 11 Pro or Ubuntu 24.04, your choice
Software stackCUDA, TensorRT-LLM, NeMo, vLLM: the same as datacenter GPUsROCm, HIP, Vulkan, llama.cpp, Ollama
Storage1-4 TB NVMe depending on config2× user-serviceable M.2 NVMe slots
Power240 W adapter, external brick300 W PSU, integrated
Price$3,999 MSRP (Founders Edition)$2,949 as tested; 128 GB OEM boxes from ~$2,000

What the numbers actually showed

The Register ran both machines through single-user inference, multi-batch inference, fine-tuning and image generation. The pattern that emerges is consistent: the closer a workload gets to compute-bound, the further the Spark pulls ahead.

Source for every figure above: The Register's hands-on test, published December 25, 2025, testing a Founders Edition Spark against HP's Z2 Mini G1a.

The software question

This is the part a spec sheet can't show. The Spark runs DGX OS, Nvidia's Ubuntu-based distribution, with the same CUDA, TensorRT-LLM and NCCL stack as a datacenter DGX. A vLLM config or fine-tuning script validated on a Spark deploys to rented datacenter GPUs unchanged. Strix Halo runs ROCm and HIP, which have closed real ground over the past year but still lag CUDA's near two-decade head start: expect the occasional missing wheel, fork, or workaround that a CUDA workflow doesn't have. If your stack is already CUDA-first, that's the whole decision.

Which to choose

First top-up: pay $10, get $20 in credit

Don't buy on a spec sheet. Rent by the hour first.

Load your actual model on a dedicated DGX Spark in EU-Central and measure the prefill and decode speed yourself before you commit to hardware.

Deploy a Spark Read the setup guide