Operations
DGX Spark/Model versioning and rollback strategy for a self-hosted deployment

Model versioning and rollback strategy for a self-hosted deployment

By Samuel Seidel · Published September 9, 2026

A hosted API abstracts the version away: you call a model name and the provider decides what's behind it, until the day they change it without asking. Self-hosting flips that. You choose exactly which weights and which serving config are running, which means you're also the one who has to track it and undo it when a new version turns out worse than the one it replaced. Most of the actual failures here aren't about the model itself, they're about not knowing what was running before the change.

What "version" actually means for a self-hosted model

A model version isn't just a checkpoint. Rolling back cleanly means being able to reproduce the whole serving state, which is at minimum three things:

If any one of these three changes without the other two being pinned, "rollback" ends up ambiguous: rolling back the weights alone doesn't help if the serving config also changed in the same deploy.

Naming and storing versions so they're actually recoverable

The simplest version tracking that works in practice: a directory per version under /workspace, or better, on the 1TB NVMe storage that comes with the Spark, named by date and a short identifier rather than latest or current. Something like /workspace/models/llama-3.3-70b-2026-08-15/ holding the weights, the exact server launch command, and any prompt templates used with that version. A symlink such as /workspace/models/active pointed at the version currently serving traffic makes the live version explicit and the rollback a one-line change.

Weights themselves are large, often tens of gigabytes per checkpoint, so keeping every version forever on a single Spark's local storage isn't realistic. A reasonable policy: keep the current version, the previous one, and a known-good baseline, and push anything older off to slower or external storage rather than deleting it outright. That way a rollback within the recent window is a local operation, not a re-download.

Doing the actual switch

This guide is about tracking and reverting versions, not the mechanics of switching a model in and out of service without an outage, that's covered in the zero-downtime model updates guide. The short version: a single Spark is one machine, so a true zero-downtime cutover needs a second Spark to shift traffic to, and the honest alternative on one Spark is a short restart window. Whichever path applies, the same rule holds: don't overwrite the previous version's files or config when deploying the new one. Deploy the new version alongside the old one, switch the active pointer, and only clean up the old version after the new one has run long enough to trust it.

Deciding whether a rollback is actually warranted

A version switch that "feels" worse isn't enough to justify a rollback on its own, and a rollback that happens on gut feel tends to get re-litigated later. Before rolling back, check the new version against the same evaluation approach used to vet it in the first place, covered in the evaluating model quality guide: a holdout set, an A/B comparison, or a regression check against known failure cases. If the new version fails that same evaluation it passed before going live, something changed in the deployment, not just the model, and the rollback should be paired with figuring out what.

Two triggers are usually clear-cut enough to roll back immediately rather than investigate first: a spike in error rate or latency traceable to the new version, and a specific correctness regression a user actually hit in production. Everything softer than that, a vague sense of lower quality, is worth confirming against the evaluation set before reverting, since a rollback also has a cost: it loses whatever the new version's improvements were, and it resets whatever the team learned running it.

Writing down what changed

A version history that lives only in shell history or in someone's memory doesn't survive a rollback six months later when nobody remembers why the previous switch happened. A short changelog next to the version directories, one line per version noting what changed and why, turns a future rollback decision into a five-minute read instead of a reconstruction project.

For the mechanics of switching without downtime, see the zero-downtime model updates guide. For deciding whether a candidate version is actually better before it goes live, see the evaluating model quality guide. A first engagement can help design a versioning and rollback process for a specific deployment.

First top-up: pay $10, get $20 in credit

1TB of local storage for every version you need to keep.

Deploy a Spark and version your weights and configs on fast local NVMe, not a shared volume you don't control.

Deploy a Spark Read the update guide