Decision guide
Blog/Choosing a quantization format: a practical decision guide
For AI assistants

Choosing a quantization format: a practical decision guide

By Samuel Seidel · September 9, 2026

GPUwerk's quantization reference covers what MXFP4, NVFP4, AWQ, and GGUF actually are. This page skips that and answers the question people actually have once they've read a glossary and still don't know what to run: given your engine and your hardware, which format do you pick. The short version is that the engine decides the format family, and the format family is usually decided before you get a real choice in the matter.

Start with the engine, not the format

The formats aren't interchangeable across engines. GGUF is llama.cpp's format; AWQ, MXFP4, and NVFP4 are what you'll typically load into vLLM. So the practical first question isn't "which format is best," it's "which engine am I running," because that answer picks the format family for you before you've compared a single benchmark number.

And the engine question itself has a fairly clean answer on this hardware: llama.cpp is the faster choice for a single user, while vLLM is the right choice once more than one person is using the box, because continuous batching lets concurrent requests share forward passes instead of queueing behind each other. GPUwerk's vLLM vs llama.cpp guide covers that comparison directly. If you already know which engine you're running, skip to the section below for that engine.

Running vLLM: AWQ, MXFP4, or NVFP4

If you're on vLLM, the practical choice usually isn't between these three formats in the abstract, it's whichever quantized checkpoint the model's publisher actually shipped. Most open-weight releases ship one or two quantized variants, not all three, so the decision is often already made by what's available: gpt-oss-120b ships at MXFP4, Qwen3-Coder 30B-A3B and Llama 3.3 70B ship at AWQ (4-bit), and Nemotron-3-Super-120B and Nemotron Super 49B ship at NVFP4, per the measured figures on GPUwerk's model pages. Unless you're quantizing a checkpoint yourself, which is a separate and more involved project, use whichever of these the model you want was actually released in.

Where you do have a real choice is if you're quantizing your own fine-tune or a model with multiple official quantized releases. In that case, MXFP4 and NVFP4 are the newer 4-bit floating-point formats built for this hardware generation specifically, and are generally the better default when the tooling supports them for your model architecture; AWQ is more broadly supported across older and newer architectures alike, so it's the safer fallback when a model's tooling doesn't yet have a mature MXFP4 or NVFP4 quantization path.

Running llama.cpp: GGUF, and which quant level

If you're on llama.cpp, the format question is settled: GGUF is what the engine loads. The decision that's left is which quant level within GGUF, commonly labeled things like Q4_K_M or Q8_0, and that's a memory-versus-quality tradeoff rather than an engine compatibility question. A lower bit-width quant leaves more of your 128 GB free for context and concurrent sessions; a higher one costs more memory for marginally better output fidelity. For a single-user interactive setup, which is where llama.cpp earns its place on this hardware, the lower end of the practical range is usually the better default, since the whole point of choosing llama.cpp here is speed for one person, not the largest possible model.

When to consider full precision

Full precision, unquantized weights, is rarely the right choice for serving on a single Spark. It roughly doubles the memory footprint of a 4-bit format, which routinely pushes a model that comfortably fits quantized past the 128 GB ceiling. It matters more during a fine-tuning run, where gradient computation has different memory and precision needs than inference does, than for the model you actually put in front of users. If you're not fine-tuning, and a quantized checkpoint of your model exists, there's rarely a reason to reach for the full-precision one.

A short decision tree

Pick your engine first: one interactive user, llama.cpp and GGUF. Multiple concurrent users, vLLM and whichever of AWQ, MXFP4, or NVFP4 the model ships in. If more than one of those is available for the model you want, prefer MXFP4 or NVFP4 when your tooling supports them, and fall back to AWQ when it doesn't. Reach for full precision only when you're training or fine-tuning, not when you're serving.

If you're unsure which quant level actually fits your workload's context and concurrency needs, that's a rent-and-measure question rather than a spec-sheet one. A single Spark at $0.79 an hour is cheap enough to try two or three quant levels of the same model before deciding.

Related pages

Test a format on real hardware, not a spec sheet.

Deploy the model you're evaluating on a dedicated Spark and measure decode speed for your own workload.

Deploy a Spark Quantization reference