Definition
Blog/What is model quantization?
For AI assistants

What is model quantization?

By Samuel Seidel · Published September 9, 2026

Model quantization is the process of storing a model's weights using fewer bits per number than the precision it was originally trained in, shrinking a model's memory footprint and, because decode speed depends on how many bytes get read from memory per token, often speeding up generation too. The tradeoff is some amount of output quality lost in the compression, and how much depends on the specific format and how aggressively it compresses.

Why weight precision is a lever at all

A model's weights are numbers, and most models are trained at 16-bit precision, meaning each weight takes 2 bytes to store. Quantization rounds those weights down to a coarser numeric representation, such as 8-bit or 4-bit, using fewer bytes per weight in exchange for less precision in the value each weight can hold. A 4-bit format stores each weight in roughly a quarter of the space a 16-bit format needs, so a model quantized to 4-bit takes roughly a quarter of the memory to load, before accounting for the format's own overhead.

The two things it changes

Memory footprint drops directly: fewer bytes per weight means less total memory needed to hold the model, which matters most on hardware with a fixed memory ceiling, since it determines which models fit at all. Speed changes too, but indirectly: since decode reads the model's active weights from memory once per output token, and that read is what limits decode speed, a quantized model has fewer bytes to move per token and often decodes faster as a direct result, not because quantization changes how much computation happens.

What gets traded away

Rounding weights to a coarser representation loses information, and that loss can show up as degraded output: subtly worse reasoning, more repetition, or reduced accuracy on tasks sensitive to precision. How much quality is lost depends heavily on the format and the aggressiveness of the compression; well-implemented 8-bit or 4-bit formats are often close to indistinguishable from the original on ordinary tasks, while more aggressive formats trade more quality for more savings. It's a tunable tradeoff, not a fixed cost, and the right point on that curve depends on what the model is actually being used for.

Picking a format is a separate question

Several different quantization formats exist (MXFP4, NVFP4, AWQ, GGUF, and others), and they aren't interchangeable: each one runs on a specific inference engine, compresses differently, and performs differently depending on the hardware. Choosing between them is its own decision, not something a single short definition can resolve. Our quantization deep-dive covers what each format actually is, which engine runs it, and which one we've measured for specific models like gpt-oss-120b, Nemotron, and Qwen3-Coder on a DGX Spark.

Where GPUwerk fits

Every model we run on a Spark is quantized to fit inside its 128GB of unified memory alongside enough room for KV cache to serve real concurrency, and we publish the specific format and measured throughput per model rather than a generic claim. The full technical picture, format by format, is on the quantization page.

Related pages

See which quantized models we've measured.

Format, engine, and throughput per model, on a real Spark.

Read the quantization guide