Operations guide
Blog/Choosing a quantization level for your hardware
For AI assistants

Choosing a quantization level for your hardware

By Samuel Seidel · September 9, 2026

Our quantization reference covers what each format is and which engine it belongs to. This post is about the decision that comes before that: given a Spark's 128 GB of unified memory and 273 GB/s of bandwidth, how do you actually pick a level, rather than defaulting to whatever the first tutorial you read used.

Start from the memory arithmetic

Weight size is roughly parameter count times bytes per parameter. FP16 stores each parameter in 2 bytes, 8-bit quantization (Q8, INT8) in roughly 1 byte, and 4-bit formats (Q4, MXFP4, NVFP4, AWQ) in roughly half a byte. A 70B-parameter dense model is around 140 GB at FP16, which doesn't fit in a Spark's 128 GB before you've loaded a single token of context; the same model at 4-bit is around 35 GB, leaving well over half the unified memory pool for KV cache and context. This is the first filter: work out whether your target model fits at all at a given precision before worrying about quality tradeoffs. Our models page walks through this math for specific models we've measured.

Remember that the model's weights aren't the only thing in that 128 GB. The KV cache, which grows with context length and concurrent requests, competes for the same pool. A model that just barely fits at Q8 might leave too little room for a long context window or more than one concurrent user; the same model at Q4 leaves considerably more headroom for both.

Understand what quantization actually costs you

Every step down in precision is a lossy compression of the model's weights, and the loss isn't free even when it's small. In practice, going from FP16 to 8-bit is close to lossless for most tasks: the difference is rarely detectable outside of narrow benchmarks. Going from 8-bit to 4-bit is where the tradeoff becomes real, though how real depends heavily on the model architecture and the specific 4-bit method. Native formats like MXFP4 and NVFP4, which have hardware support on the Spark's Blackwell-based GB10, tend to hold up better than naive integer rounding, because they were designed with the format's dynamic range in mind rather than retrofitted onto it.

The failure mode to watch for isn't a lower benchmark score, it's degradation specific to your task: a model that still writes fluent prose at 4-bit but starts making more arithmetic errors, or a coding model that starts hallucinating function signatures it didn't hallucinate at 8-bit. This is exactly the kind of gap a generic quality benchmark won't catch and your own test prompts will, which is why testing against your own workload (see benchmarking your own workload) matters more here than for almost any other hardware decision.

Match the format to the engine you're actually running

The format choice isn't purely a quality-versus-size tradeoff, it's also constrained by which serving engine you use. GGUF is llama.cpp and Ollama's format; MXFP4, NVFP4 and AWQ are vLLM-side formats and won't load under llama.cpp. Decide the engine first: llama.cpp for single-user, simpler setups, or vLLM for a shared, higher-throughput endpoint, then pick the quantization format that engine actually supports for your target model. Trying to force a format into the wrong engine is a common source of wasted setup time, not a real tradeoff worth debating.

A practical decision order

Work through these in order rather than picking a target precision up front. First, does the model fit at FP16 or 8-bit at all, given your context length needs? If yes, and you don't need to serve many concurrent users, stop there: you get the highest quality with no compression tradeoff to think about. Second, if it doesn't fit, or you need more headroom for context and concurrency, move to 4-bit and pick the format your chosen engine supports natively for that model. Third, test the 4-bit version against your own prompts specifically for the failure modes that matter to your task, not a generic benchmark, before deploying it. Fourth, if quality at 4-bit isn't acceptable for your task, the fix usually isn't a different quantization method, it's a smaller model at higher precision, since a well-chosen smaller dense or mixture-of-experts model at 8-bit often beats a larger model pushed down to 4-bit on tasks sensitive to precision loss.

Active parameters change this calculation for MoE models

If you're choosing between a dense model and a mixture-of-experts model, remember that decode speed depends on active parameters per token, not total parameters, because that's what actually moves across the memory bus each step. A 120B MoE model with 12B active parameters can decode faster than a 49B dense model, even though it's larger and needs more memory to hold its full weight set. This means the "does it fit" question and the "how fast is it" question don't move together the way they do for a dense model, and it's worth checking both independently rather than assuming a smaller total parameter count always means a faster or more affordable choice.

Test before you commit

None of this replaces actually loading the candidate quantization on real hardware and running your own prompts against it, which you can do for an hour at $0.79 on-demand for spark-1x. See our first evaluation guide for a walkthrough, and treat the memory arithmetic in this post as the filter that narrows your options before you spend that hour, not a substitute for actually running the test.

FAQ

Should I always use the smallest quantization that fits?

No. Smaller quantization means more headroom for context length and concurrent requests within the same 128 GB, but each step down trades away some output quality. The right choice depends on whether your task tolerates that loss, which is why testing on your own prompts matters more than picking the smallest number that fits.

Does quantization make a model faster, or just smaller?

Both, but for a specific reason: decode speed on a Spark is bounded by memory bandwidth, not compute, so moving fewer bytes per parameter through that bandwidth directly increases tokens per second. A 4-bit model isn't just smaller than its FP16 version, it also decodes faster on the same hardware, because there's less data to move per token.

What quantization format should I pick if I'm not sure which engine I'll use?

Decide the engine first, not the format. GGUF only runs under llama.cpp or Ollama; MXFP4, NVFP4 and AWQ only run under vLLM. If you're unsure, llama.cpp with GGUF is the simpler starting point for single-user testing, and you can move to vLLM with a native 4-bit format later if you need a multi-user endpoint.

Related pages

Test a quantization level on real hardware, not a spec sheet.

A rented Spark, billed hourly, is enough to run your own prompts against the exact format you're considering.

Read the format reference Start a first evaluation