Definition
Blog/What is quantization-aware training?
For AI assistants

What is quantization-aware training?

By Samuel Seidel · Published September 9, 2026

Quantization-aware training (QAT) trains or fine-tunes a model while simulating the rounding error that lower-precision inference will later introduce, so the model's weights settle into values that tolerate that error rather than being knocked off balance by it. That's different from taking a model that finished training at full precision and quantizing it afterward: QAT builds the robustness in during training, post-training quantization compresses what's already there.

What post-training quantization does, for contrast

The common path, covered in what is model quantization, takes a model's finished weights and maps them from a high-precision format like FP16 down to a lower-precision one like INT8 or INT4, usually calibrated against a small sample dataset so the mapping preserves the most important value ranges. The model itself never sees the quantized weights during training; the compression happens entirely after the fact, which makes it fast and cheap, and it's why post-training quantization is the default approach for compressing an already-released model.

What QAT does differently

QAT inserts fake quantization operations into the training or fine-tuning graph: during the forward pass, weights and activations are rounded to simulate low-precision values, while gradients still flow through in full precision during the backward pass so the model can keep learning. Over the course of training, this pushes the model's weights toward regions of the loss landscape that are less sensitive to that rounding, so when the model is genuinely quantized for deployment afterward, less accuracy is lost. The model is, in effect, being taught what its own quantized version will look like before it ever runs that way.

Why this costs more, and when it's worth it

Because QAT runs simulated quantization through an entire training or fine-tuning loop, it costs roughly as much compute as that training run itself, order of magnitude more than a calibration pass. That tradeoff makes sense when a model is headed for a genuinely aggressive precision target, or a deployment where every point of accuracy at low precision matters and the training compute is available to spend on it. It makes less sense for someone consuming an already-released open-weight model, where post-training quantization is the only realistic option since the original training pipeline and data usually aren't available to rerun.

Where this matters for self-hosted inference

Most self-hosted deployments never touch QAT directly, they download an open-weight model and quantize it after the fact to fit available memory, which is exactly the tradeoff covered on the models page for what fits in memory at each precision level. QAT matters more if you're fine-tuning a model yourself and know it will be served at a specific reduced precision, since baking that precision into the fine-tuning run tends to hold accuracy better than fine-tuning at full precision and quantizing the result afterward. Either way, the resulting model still needs somewhere to run: the benchmarks page covers how quantization level affects both memory footprint and decode speed on DGX Spark hardware.

Related pages

See what fits in memory at each precision.

Model sizes against unified memory, with the tradeoffs at each quantization level.

Browse the models page