Quantized vs full-precision models in production
GPUwerk's quantization reference covers what MXFP4, NVFP4, AWQ and GGUF actually are and which engine each runs on. This page is about the decision that reference does not make for you: whether to run a quantized build or full precision once you are serving real traffic, not just testing a model on your own laptop.
The memory constraint comes first
On a single Spark's 128 GB of unified memory, full precision (FP16, 2 bytes per parameter) caps you at roughly 60 to 65B parameters once you leave room for KV cache and framework overhead, while a 4-bit build fits a model with roughly four times as many parameters in the same footprint. That arithmetic is worked through on GPUwerk's benchmarks page and is why gpt-oss-120b, at 120B parameters, only fits on one Spark at MXFP4 and not at full precision. If the model you want does not fit at full precision on the hardware you are renting, the decision is already made; the question is which quantized format, not whether to quantize at all.
If the model does fit at full precision with headroom to spare, you have an actual choice to make, and that is what the rest of this page is about.
Where quantization quality risk actually shows up
A 4-bit format is not free; it is a bet that the precision you are discarding does not matter for your task. That bet pays off differently depending on what you are asking the model to do, and GPUwerk has not run a controlled accuracy study across its own catalog to put numbers on this, so the following is task-shape guidance to test against your own data, not a measured result.
- Open-ended generation, chat, summarization, drafting. The task with the most slack. Small shifts in token probabilities rarely change whether an answer is useful, and this is where a 4-bit format is least likely to cost you anything you would notice.
- Multi-step reasoning and code generation. More sensitive. A single wrong token partway through a reasoning chain or a generated function can cascade into a wrong answer or code that does not run, in a way a summarization error does not. If you are running an agent loop or a coding assistant in production, treat quantization level as something to validate against your own eval set, not assume is safe because it looked fine in a demo.
- Structured extraction and classification with a fixed answer set. Often robust to quantization because the output space is narrow, but worth a quick accuracy check against a held-out set before trusting it, especially if the categories are similar to each other.
- Anything with a regulatory, financial, or safety-relevant output. Treat full precision as the safer default here and require your own evaluation before dropping to 4-bit, since the cost of an error is asymmetric with the cost of the extra memory.
A practical decision process
- Check whether full precision fits at all, on the node size you are renting, with room for your expected context length and concurrency. If it does not fit, quantize and move to format selection on GPUwerk's quantization reference.
- If it fits, build a small eval set from your actual task, not a generic benchmark. A few dozen real prompts with known-good answers or a rubric is enough to catch a quantization regression that matters for your use case.
- Run both. Deploy the full-precision model and a 4-bit build of the same model against your eval set. Because the constraint is memory, not compute, you can usually do this sequentially on one rented Spark rather than needing two nodes at once: run one, record results, swap the model, run the other.
- Decide on the gap, not the format name. If the 4-bit build's answers are indistinguishable from full precision on your eval set, ship the smaller one; it leaves you more memory headroom for concurrency or a second model. If there is a real gap on the tasks you actually run, full precision earns its memory cost.
Production considerations beyond the model weights themselves
Two things people forget to budget for when comparing footprints:
- KV cache scales with context length and concurrency, not with the quantization format of the weights. A 4-bit model still needs full-precision-scale KV cache room per active request unless you are also using a quantized KV cache scheme, so the memory saved by quantizing weights buys you concurrency headroom or context length, not a proportional cut in total memory use.
- Engine choice is coupled to format, not incidental. As covered on the quantization reference, AWQ and NVFP4 run under vLLM; GGUF runs under llama.cpp or Ollama. Picking a quantization format also picks your serving engine, which has its own throughput and single-stream latency characteristics independent of the weights' precision.
Revisit the decision when the model changes, not just once
A quantization decision made for one model version doesn't automatically carry over to the next one. A vendor's updated checkpoint can shift how much headroom you have before a 4-bit build starts costing you accuracy on your specific tasks, and a new model family entirely may not have a mature quantized build yet, forcing full precision as a stopgap until one lands. Treat the eval-set comparison in this guide as something to re-run on model upgrades, not a one-time gate you clear and forget.
It's also worth separating a temporary constraint from a permanent one. If you're running full precision only because a good 4-bit build doesn't exist yet for a model you need, that's different from a considered decision that your task is quantization-sensitive. Track which reason applies, because the first one is worth revisiting every time the model's ecosystem catches up, and the second one isn't going to change on its own.
A first engagement can help build that eval set and run the comparison on real hardware before you commit a production deployment to one format.