Quantization on a DGX Spark: MXFP4, NVFP4, AWQ, GGUF explained
A Spark has 128 GB of unified memory and 273 GB/s of bandwidth to move weights through it. Quantization is what decides whether a model fits in the first number and how fast it runs against the second. This page is the format reference: what MXFP4, NVFP4, AWQ and GGUF actually are, which engine each one belongs to, and which one GPUwerk has run for which model.
Why quantization matters on this box
Weight size is parameters times bytes per parameter. A 4-bit format stores each parameter in roughly half a byte, against 2 bytes at FP16, so quantizing a model to 4-bit is usually the difference between it fitting on one Spark and not. That arithmetic, and the reason decode speed is bandwidth-bound rather than compute-bound, is worked through in full on the benchmarks page and the models page. This page does not repeat that math; it covers the formats themselves.
The formats, and what they are
Four names come up constantly around Spark deployments, and they are not interchangeable options for the same job. Two are hardware formats, one is a portable community format, and one belongs to a different engine entirely.
MXFP4
MXFP4 is a 4-bit floating point format with native support on Blackwell, the GB10's architecture. GPUwerk runs gpt-oss-120b at MXFP4, where it decodes at roughly 60.5 tok/s under llama.cpp and 33.5 tok/s under vLLM, both single-stream and both reported on the benchmarks page.
NVFP4
NVFP4 is NVIDIA's own 4-bit floating point format, also with native Blackwell support, and it is the format NVIDIA ships its Nemotron family in. GPUwerk has measured two Nemotron models at NVFP4 under vLLM: Nemotron Super 49B at 5.79 tok/s single-stream, a dense model reading all 49B parameters per token, and Nemotron-3-Super-120B-A12B at 22.7 to 23.7 tok/s, a mixture-of-experts model reading only its 12B active parameters. Same format, same engine, an order of magnitude apart in speed, because the parameter count that matters for decode is active, not total.
AWQ
AWQ is a 4-bit integer format from the open-source community rather than a vendor. It has no Blackwell-specific hardware path the way MXFP4 and NVFP4 do, but it has the broadest model coverage of the 4-bit formats and is the well-trodden path in vLLM when a model has no native FP4 build. GPUwerk's own fleet numbers for Qwen3-Coder-30B-A3B, 80.9 tok/s single-stream and 3,235 tok/s aggregate at 256 concurrent, are AWQ under vLLM, with the raw benchmark log at qwen3-coder-30b-awq-spark.txt. AWQ is also how GPUwerk's staged Llama 3.3 70B image is quantized.
GGUF
GGUF is not a peer of the three formats above. It is llama.cpp's own container and quantization format, covering everything from Q4_K_M to full precision, and it runs under llama.cpp or Ollama, not vLLM. A model available only as GGUF gets served through llama.cpp; a model available only as AWQ or NVFP4 gets served through vLLM. That split is the whole reason gpt-oss-120b's two rows on the benchmarks page use two different engines: the GGUF build runs on llama.cpp, and the same MXFP4 weights are also available as a vLLM-compatible checkpoint.
| Format | Type | Engine | Measured on a Spark |
|---|---|---|---|
| MXFP4 | 4-bit float, native Blackwell | vLLM or llama.cpp | gpt-oss-120b |
| NVFP4 | 4-bit float, native Blackwell, NVIDIA's own | vLLM | Nemotron Super 49B, Nemotron-3-Super-120B-A12B |
| AWQ | 4-bit integer, community, broad coverage | vLLM | Qwen3-Coder-30B-A3B, Llama 3.3 70B |
| GGUF | llama.cpp's own container and quantization scheme | llama.cpp, Ollama | gpt-oss-120b (llama.cpp build) |
Which one to reach for
The decision runs through the engine first, not the format. Serving one person from your own machine points at llama.cpp and a GGUF build: it is the faster single-stream engine, and the arithmetic on the benchmarks page explains why. Serving a shared endpoint, an app, a team, an agent fleet, points at vLLM, and then the format choice is whatever the model actually ships in: NVFP4 if there is a Nemotron-family build, AWQ if not, since AWQ is the format most models have a build for.
Not every model gives you both options. gpt-oss-120b is the exception on this page with builds for both engines. Nemotron and Qwen3-Coder, as measured here, are vLLM-only; if you need one of them on llama.cpp, check whether a community GGUF conversion exists before assuming it does not.
FAQ
What is the difference between MXFP4, NVFP4 and AWQ?
MXFP4 and NVFP4 are both 4-bit floating point formats with native support on Blackwell, the GB10's architecture; NVFP4 is NVIDIA's own scheme and is what the Nemotron family ships in. AWQ is a 4-bit integer format from the wider open-source community, not tied to any one vendor's silicon, and it is the more broadly supported path in vLLM when a model has no native FP4 build. GPUwerk runs gpt-oss-120b at MXFP4, Nemotron Super 49B and Nemotron-3-Super-120B-A12B at NVFP4, and Qwen3-Coder-30B-A3B at AWQ.
Is GGUF the same thing as MXFP4 or AWQ?
No. GGUF is llama.cpp's own container and quantization format, unrelated to MXFP4, NVFP4 or AWQ, which are vLLM-side formats. A model quantized to GGUF runs under llama.cpp or Ollama, not vLLM, and a model quantized to AWQ or NVFP4 runs under vLLM, not llama.cpp. Pick the format by which engine you are serving from, not by preference.
Which quantization format should I use on a DGX Spark?
For a single user, GGUF under llama.cpp: it is the faster single-stream engine and gpt-oss-120b measures roughly 60.5 tok/s decode that way, against 33.5 tok/s under vLLM. For a shared endpoint, use vLLM with whichever 4-bit format the model actually ships in: NVFP4 when a Nemotron build exists, AWQ otherwise, since AWQ has the broadest model coverage of the 4-bit options.