What is a model checkpoint?
A model checkpoint is a saved snapshot of a neural network at one point in time: its weights, and usually additional state needed to resume training, such as the optimizer's internal values and which training step it stopped at. Checkpoints exist because training can take days or weeks, and nobody wants to redo that work from scratch after a crash, or lose the ability to go back to an earlier, better-performing version.
What's actually inside one
At minimum, a checkpoint contains the model's weights, the numeric parameters that define what the network has learned so far. A checkpoint meant for resuming training also stores the optimizer state (momentum and variance terms used by algorithms like Adam), the current training step or epoch count, and sometimes the random number generator's state, so training can pick back up as if it had never stopped. A checkpoint meant only for inference, by contrast, usually strips all of that away and keeps just the weights, since none of the training-resumption data is needed to run the model.
Why runs save more than one
Long training runs save checkpoints at intervals, sometimes every few hundred or few thousand steps, rather than waiting until the very end. That protects against hardware failures and bad training dynamics: if loss starts climbing or a machine crashes, the run can restart from the last good checkpoint instead of from zero. It also lets a team evaluate several checkpoints against a held-out test set and pick whichever one actually performs best, since the final step of training isn't guaranteed to be the best one.
Checkpoint vs weights vs release
These terms get used loosely and it's worth being precise. Weights are the numbers themselves. A checkpoint is a file (or directory) that bundles those weights with some amount of extra state, saved at a specific point. A public model release is usually a single checkpoint that the creator picked, cleaned of training-only data, and published under a license. So every release is a checkpoint, every checkpoint contains weights, but a checkpoint isn't automatically a release, and weights alone aren't automatically a full checkpoint if they've been stripped of the surrounding state.
Fine-tuning starts from a checkpoint
When a team fine-tunes a model rather than training from scratch, they load an existing checkpoint (usually a released base model) and continue training on new data, saving new checkpoints along the way. That's cheaper than training from random initialization because the base checkpoint already encodes general language ability; fine-tuning only has to adjust it toward the new task. See what is fine-tuning for how that process works and where it fits against retrieval-based approaches.
Where GPUwerk fits
GPUwerk doesn't run training; it rents dedicated DGX Spark hardware for inference and fine-tuning workloads that customers control directly. A checkpoint, whether it's a public release downloaded from a model hub or one you've fine-tuned yourself, is the file you load onto that hardware and serve requests against. Because each Spark is dedicated rather than shared, whatever checkpoint you load and whatever data you send it stays on that node, at $0.79/hour for a single node or $1.79/hour for a two-node cluster. See what is GPU rental for how that hourly model compares to owning the hardware outright.