DGX Spark vs Strix Halo
Both machines pack 128 GB of unified memory into a mini-PC. On raw single-user token generation they are close enough that the difference doesn't matter. What actually separates them is prompt processing speed, the software stack underneath, and roughly $1,000.
The verdict, up front
Buy Strix Halo if… you're running one model at a time in llama.cpp, don't need CUDA, and the price gap matters. The Register's hands-on testing found the two boxes "churn out tokens at a similar pace" in single-batch llama.cpp, with Strix Halo taking a narrow lead on its Vulkan backend, for two-thirds to half the Spark's price.
Choose the Spark if… you feed it long prompts or documents, run concurrent or batched jobs, fine-tune anything, or need the software to just work. The same testing measured the Spark's GPU at 2-3x faster time-to-first-token on a short prompt, with the gap widening on longer ones, and roughly 2x faster full fine-tuning. That's before counting the software: anything built on CUDA runs on the Spark without porting; AMD's ROCm/HIP stack is closing the gap but isn't there yet.
Side-by-side specs
Comparing Nvidia's Founders Edition DGX Spark against HP's Z2 Mini G1a, the Strix Halo box The Register tested. Other Strix Halo OEMs (Framework, Minisforum, ASUS) sell 128 GB configurations from a little over $2,000, cheaper than the G1a's tested configuration.
| NVIDIA DGX Spark | AMD Strix Halo (HP Z2 Mini G1a) | |
|---|---|---|
| Platform | GB10 Grace Blackwell superchip | Ryzen AI Max+ Pro 395 APU |
| CPU | 20-core Arm (10× X925 + 10× A725) | 16-core Zen 5, up to 5.1 GHz |
| GPU | Blackwell, FP4/FP8 tensor cores | Radeon 8060S, RDNA 3.5, no low-precision tensor path |
| Unified memory | 128 GB LPDDR5x, 273 GB/s | 128 GB LPDDR5x, 256 GB/s |
| Peak AI compute | ~500 dense TFLOPS realistic (1 PFLOP sparse FP4 claimed) | ~56 TFLOPS BF16 (AMD doesn't publish a peak figure) |
| NPU | None | XDNA 2, 50 TOPS, limited software support today |
| Networking | 10 GbE + ConnectX-7, 200 Gb/s for clustering | 2.5 GbE, no clustering path |
| OS | DGX OS (Ubuntu-based, Linux only) | Windows 11 Pro or Ubuntu 24.04, your choice |
| Software stack | CUDA, TensorRT-LLM, NeMo, vLLM: the same as datacenter GPUs | ROCm, HIP, Vulkan, llama.cpp, Ollama |
| Storage | 1-4 TB NVMe depending on config | 2× user-serviceable M.2 NVMe slots |
| Power | 240 W adapter, external brick | 300 W PSU, integrated |
| Price | $3,999 MSRP (Founders Edition) | $2,949 as tested; 128 GB OEM boxes from ~$2,000 |
What the numbers actually showed
The Register ran both machines through single-user inference, multi-batch inference, fine-tuning and image generation. The pattern that emerges is consistent: the closer a workload gets to compute-bound, the further the Spark pulls ahead.
- Single-user token generation, short prompt: a near tie. Both are memory-bandwidth-bound here (273 GB/s vs 256 GB/s), and bandwidth is close enough that Strix Halo edged ahead on its Vulkan backend in llama.cpp.
- Time to first token: the Spark's GPU processed a 256-token prompt 2-3x faster, and the gap grows with prompt length because prefill is compute-bound, not bandwidth-bound. For long documents or big system prompts, this is the difference customers actually feel.
- Multi-batch inference (vLLM, Qwen3-30B-A3B, batch 1-64): the Spark's faster GPU wins, though The Register notes the advantage mostly matters if you're running batch jobs routinely rather than overnight.
- Fine-tuning: a full fine-tune of Llama 3.2 3B took the Spark about two-thirds the time of the G1a. QLoRA on Llama 3.1 70B (both machines can hold it, thanks to 128 GB) took the Spark roughly 20 minutes against the G1a's 50+.
- Image generation (FLUX.1 Dev, ComfyUI): the Spark's clearest win, roughly 2.5x the G1a's measured throughput, because diffusion models are compute-bound rather than bandwidth-bound.
Source for every figure above: The Register's hands-on test, published December 25, 2025, testing a Founders Edition Spark against HP's Z2 Mini G1a.
The software question
This is the part a spec sheet can't show. The Spark runs DGX OS, Nvidia's Ubuntu-based distribution, with the same CUDA, TensorRT-LLM and NCCL stack as a datacenter DGX. A vLLM config or fine-tuning script validated on a Spark deploys to rented datacenter GPUs unchanged. Strix Halo runs ROCm and HIP, which have closed real ground over the past year but still lag CUDA's near two-decade head start: expect the occasional missing wheel, fork, or workaround that a CUDA workflow doesn't have. If your stack is already CUDA-first, that's the whole decision.
Which to choose
- You run one model at a time in llama.cpp, no CUDA dependency: Strix Halo, for two-thirds to half the price, with token generation close enough not to matter.
- You feed it long prompts, documents, or run things concurrently: the Spark's prefill and batch advantage is real and grows with load.
- You fine-tune regularly, or generate images/video: the Spark wins clearly, roughly 2x on fine-tuning and 2.5x on FLUX.1.
- Your existing tooling is CUDA-based: the Spark, full stop. Porting a working CUDA pipeline to ROCm is real engineering time, not a checkbox.
- Not sure which workload you actually have: rent a Spark by the hour and measure your own model rather than guessing from someone else's benchmark.