What is a parameter count?
A parameter count is the number of learned numeric weights a model stores after training, usually written in billions: a "7B" model has about 7 billion parameters, a "70B" model about 70 billion. Each weight is a value the model adjusted during training and now reads back during inference, so the count is both a rough proxy for how much the model learned and a direct measure of how much memory it takes up on disk and in GPU memory.
Where the number comes from
During training, a model's weights start at random values and get nudged, token by token, batch by batch, toward values that reduce prediction error on the training data. The architecture, how many layers, how wide each layer is, fixes the total count of these weights before training even starts. That fixed count is what gets published in the model name: Llama 3 70B has 70 billion weights regardless of how it was trained or what it's used for afterward.
What it predicts, and what it doesn't
Within a single model family trained the same way, a larger parameter count usually means more capacity to represent complex patterns, and often better benchmark scores. But parameter count alone is a weak predictor across families: a smaller model trained on more or cleaner data can beat a larger model trained on less, and the count says nothing about training quality, instruction tuning, or how well a model handles a specific task. Treat "70B" as a size label, not a capability score.
Why it drives memory requirements
Every parameter has to be stored somewhere the GPU can read it, and that storage cost scales directly with count. A 7B model at 16-bit precision needs roughly 14GB just for weights; a 70B model needs roughly 140GB. Quantizing to lower precision, 8-bit or 4-bit, cuts that requirement roughly in half or to a quarter, at some cost to output quality. This is the number that decides whether a model fits on a given GPU at all, before you even get to how fast it runs. For the full memory-sizing math, including KV cache and quantization formats, see GPUwerk's glossary or the model catalog for what fits on a Spark today.
Total parameters isn't the whole story
For a dense model, every parameter is used on every token, so total count and compute-per-token are the same thing. That's not true for a mixture-of-experts model, where only a subset of parameters activates per token even though the full set is stored in memory. A "120B" MoE model can run at a fraction of the compute cost of a 120B dense model, while still needing memory sized to the full total. See what is a mixture-of-experts model for how that split works.
How the naming convention gets applied
Model names round the parameter count, so "7B" usually means somewhere between 6 and 8 billion, not exactly 7,000,000,000. Different model families also count parameters slightly differently, some include embedding layers in the published figure, some don't, so two models both labeled "7B" can differ by a meaningful margin in actual weight count. When the exact figure matters, for memory planning in particular, it's worth checking the model card rather than trusting the name alone.
Parameter count is a purchasing decision, not just a spec
Because parameter count sets the memory floor, it's usually the first filter when picking hardware for a workload: does the model, at the precision you plan to run it, fit in the memory available. A 128GB unified-memory machine comfortably holds a 70B model at 16-bit precision with room for KV cache, but a 120B dense model at the same precision won't fit without quantizing down first. Checking this before committing to hardware avoids discovering the mismatch after deployment.