Definition
Blog/What is a model weight?
For AI assistants

What is a model weight?

By Samuel Seidel · Published September 9, 2026

A model weight is a single number stored inside a neural network. Each weight scales how much one piece of information affects the next step of a calculation as data moves through the network's layers. A large language model is, physically, a very long list of these numbers, arranged into matrices, plus the code that knows how to multiply them together in the right order.

Where the numbers come from

Weights start out random. Training adjusts them repeatedly: the model makes a prediction, the prediction is compared against the correct answer, and an algorithm called backpropagation nudges each weight slightly in the direction that would have made the prediction better. Repeated across enormous amounts of text and many passes, this process is what turns a random list of numbers into one that can complete sentences, answer questions, or write code. See inference vs training for how that process differs from actually running the model afterward.

What "parameter count" is counting

When a model is described as having, say, 8 billion parameters, that figure is a count of its weights. Not every weight matters equally to every output, but there's no separate list of "important" weights set aside from the rest; the number is simply the total count of learned values in the network. See what is a parameter count for how that number relates to capability and hardware requirements.

Precision: how many bits per weight

Each weight is stored using some number of bits, most commonly 16-bit floating point in the original training format. Reducing that to 8-bit or 4-bit representations, a process called quantization, shrinks the file size and the memory needed to run the model, generally at some cost to output quality. A 7-billion-parameter model at 16-bit precision needs roughly 14GB of memory just to hold the weights; the same model at 4-bit precision needs closer to 4GB. See what is model quantization for the mechanics and trade-offs.

Weights vs the rest of a model release

A model release usually bundles the weights with a tokenizer (which converts text to and from the numbers the model actually processes), a configuration file describing the network's architecture, and sometimes a license file. The weights are the largest part by far and the part that determines behavior; the rest is scaffolding needed to load and run them correctly. When a provider is described as "open-weight," it means these weight files are published for anyone to download and run, as distinct from an architecture or training method being open too. See what is an open-weight model for that distinction.

Why weights matter for self-hosting

Because weights are just files, an open-weight model can be downloaded once and run on hardware you control, with no dependency on an external API staying available or unchanged. That's the basis for private, self-hosted inference: the weights sit on local storage, get loaded into GPU memory, and every request is answered by hardware you're renting or own, without the prompt or the response leaving that machine. GPUwerk's dedicated Spark nodes are built for exactly this: loading an open-weight model's files once and serving requests against them locally, at $0.79/hour for a single node or $1.79/hour for a two-node cluster. See the hardware page for memory specifications relevant to which weight sizes fit.

Related pages

Load your own weights on hardware you control.

One dedicated Spark, $0.79/hour, deployed in minutes from EU-Central.

Deploy a Spark