LLM and GPU terms explained
Short, self-contained definitions for the terms that show up constantly around LLM inference and GPU hardware. Where GPUwerk has already measured a number for a term, this page links to it instead of restating it.
Prefill
Prefill is the phase where the model reads and processes your entire input prompt before producing the first output token. It is compute-bound: the GPU is doing matrix multiplication over every token in the prompt at once, so prefill speed tracks FLOPS rather than memory bandwidth. A long prompt takes proportionally longer to prefill, and the relationship is not linear either, because attention cost grows quadratically with context length. GPUwerk's own measured prefill numbers, across several models on a DGX Spark, are on the benchmarks page.
Decode
Decode is the phase that generates output tokens one at a time after prefill finishes. Each step reads the model's active weights out of memory once, so decode is bandwidth-bound rather than compute-bound: the ceiling on decode speed is set by how fast the GPU can move weight bytes from memory, not by how many FLOPS it has. This is why a chip with modest compute but fast memory can out-decode a chip with more compute and slower memory. GPUwerk's benchmarks page works through the exact arithmetic and measured numbers for a DGX Spark.
Tokens per second
Tokens per second (tok/s) is the standard unit for both prefill and decode speed. A token is roughly three-quarters of a word in English, so 10 tok/s decode is in the neighborhood of comfortable reading speed for one person, and figures well below that feel sluggish in a chat interface. The same unit is also used for aggregate throughput when many requests are served at once, where it means total tokens produced across every request per second, not the speed any single user experiences. GPUwerk's benchmarks page reports both kinds and is explicit about which is which.
KV cache
The KV cache stores the attention keys and values computed for every token already processed in a request, so the model does not recompute them from scratch on each new token. It grows with context length and with the number of requests being served concurrently, and it lives in the same memory pool as the model's weights. A model with a small memory footprint leaves more room for KV cache, which is one reason a smaller or more efficient model can serve far more concurrent users than a larger one on identical hardware. GPUwerk's vLLM doc reports measured KV cache token budgets for specific models on a Spark.
Unified memory
Unified memory means the CPU and GPU share one physical memory pool rather than each having its own separate bank, as a discrete GPU with dedicated VRAM does. The practical effect is capacity: a machine with unified memory can dedicate nearly all of it to a single large model, rather than being capped by whatever VRAM a discrete card happens to carry. The DGX Spark's 128 GB of unified memory is the reason it fits 70B-class and larger models that would need multiple discrete GPUs otherwise, detailed on the DGX Spark hardware page.
Quantization
Quantization stores each model parameter in fewer bits than the format it was trained in, trading some precision for a smaller memory footprint and, on hardware with native support for the format, faster execution. A 4-bit format stores each parameter in roughly half a byte against 2 bytes at FP16, which is usually the difference between a model fitting on one GPU and not. GPUwerk's quantization page covers the format-by-format detail, including which engine each one runs on.
MXFP4, NVFP4, AWQ, GGUF, briefly
MXFP4 and NVFP4 are both 4-bit floating point formats with native hardware support on Blackwell chips like the Spark's GB10; NVFP4 is NVIDIA's own scheme, used by its Nemotron model family. AWQ is a 4-bit integer format from the open-source community, not tied to any vendor's silicon, and has the broadest model coverage of the four when a model has no native FP4 build. GGUF is a different thing entirely: it is llama.cpp's own container and quantization format, running under llama.cpp or Ollama rather than vLLM. The full comparison, with which GPUwerk-measured model uses which format, is on the quantization page.
Mixture-of-experts vs dense
A dense model activates every one of its parameters on every token it processes. A mixture-of-experts (MoE) model instead routes each token through a small subset of specialized sub-networks, so only a fraction of its total parameters are read from memory per token, even though the full set is stored. Because decode speed is set by how many weight bytes get read per token, a MoE model with a small active-parameter count can decode far faster than a dense model of similar total size. GPUwerk's own measurements show this directly: a 30B-class MoE model and a 70B dense model on the same DGX Spark hardware, with the MoE model roughly 13 times faster single-stream, reported on the benchmarks page.
Memory bandwidth
Memory bandwidth is how fast data can move between memory and the processor, usually measured in gigabytes or terabytes per second. It is the direct ceiling on decode speed, since every generated token means reading the active weights out of memory once: tokens per second is approximately bandwidth divided by the bytes of active weight read per token. GPUwerk's benchmarks page works through this formula with the DGX Spark's own 273 GB/s figure and checks it against measured results.
Active parameters vs total parameters
Total parameters is the full parameter count a model ships with, the number in its name. Active parameters is how many of those get used, and read from memory, on a given token. For a dense model the two numbers are the same. For a mixture-of-experts model they are not: a "120B" MoE model might activate only 12B parameters per token, and that smaller number, not the headline figure, is what predicts decode speed. This distinction is the single most useful thing to check before assuming a model's parameter count tells you how fast it will run; GPUwerk's benchmarks page and quantization page both key their speed claims to active count rather than total.