Operations guide
Blog/Backup and snapshot strategies for a rented GPU node
For AI assistants

Backup and snapshot strategies for a rented GPU node

By Samuel Seidel · September 9, 2026

A rented Spark comes with 1 TB of local NVMe and root SSH access, described in our SSH key guide. It does not come with a backup system. If the instance is deleted, the node has a hardware fault, or you fat-finger an rm -rf at 1am, whatever was only on that disk is gone. Here's how to think about what's worth protecting and how to protect it without treating every gigabyte the same way.

Not everything on the node is equally worth backing up

The instinct is to back up the whole disk. On an inference node, that's usually the wrong instinct, because the disk is dominated by data that's either reproducible or replaceable, and backing it up wastes bandwidth and money protecting something you don't actually need protected. Split what's on the node into three categories before deciding on a strategy:

Reproducible from a public source: base model weights pulled from Hugging Face or a model registry. Losing these costs you a re-download, not data. Track the source and quantization you used (see the quantization reference for the format tradeoffs) in a text file so re-downloading gets you back to the same state, rather than backing up tens of gigabytes of weights you can fetch again.

Small and irreplaceable: configuration. vLLM launch flags, LiteLLM routing and API key config, systemd unit files, nginx or Caddy reverse proxy config, environment variables. These files are typically kilobytes, not gigabytes, and losing them means rebuilding a working setup from memory, which is the annoying kind of data loss because nothing was technically destroyed, you just have to remember how you configured it.

Large and irreplaceable: anything you produced that doesn't exist anywhere else. Fine-tuned or merged model checkpoints, LoRA adapters you trained, a RAG vector index built from your own documents, conversation logs you're required to retain. This is the category that actually needs a real backup strategy, because losing it is a genuine, unrecoverable loss.

Back up configuration constantly, cheaply

Because configuration is small, there's no excuse not to version it. Put your vLLM launch scripts, LiteLLM config, systemd units and Caddy or nginx config in a private git repository, and commit to it whenever you change something. This costs nothing, takes minutes to set up, and means a fresh instance can be brought back to a working state by cloning the repo and re-running your setup script, rather than reconstructing flags from shell history that may or may not still be there.

Get large, irreplaceable data off the node on a schedule

For fine-tuned checkpoints or a custom RAG index, the practical pattern on a rented node is: push to object storage on a schedule, not "eventually." rclone configured against an S3-compatible bucket (Backblaze B2, Cloudflare R2, or a Hetzner Object Storage bucket if you want EU-based storage) is the least fussy way to do this, since it handles resumable multi-part uploads over an SSH-only box without extra daemons. A cron job or systemd timer running rclone sync against a checkpoints directory once a training or fine-tuning run completes is enough for most teams; you don't need continuous replication for artifacts that only change when you produce a new checkpoint.

If your workload involves ongoing state that changes constantly, like a growing local vector database, treat it the way you'd treat a database on any server: periodic dumps pushed off-box, not a live filesystem copy, since copying a database file while it's being written risks a corrupt backup.

Snapshots are a convenience, not a backup

If your GPUwerk console offers instance snapshots, they're useful for one thing: getting back to a known-good state on the same instance quickly, for example before trying a risky driver upgrade. A snapshot that lives on the same storage as the instance it's protecting is not a backup in the disaster-recovery sense, because whatever destroys the instance, hardware failure, accidental deletion, an expired rental, likely takes the snapshot with it. Treat snapshots as an undo button, and treat off-node object storage as the actual backup for anything you can't afford to lose.

Test the restore, not just the backup

A backup you've never restored from is a hope, not a plan. Once you have checkpoints or config pushed to object storage, actually spin up a fresh Spark, following the setup guide, and pull everything back down to confirm it works end to end: the model loads, the config parses, the fine-tuned adapter merges cleanly. This is worth doing once when you set the pipeline up, and worth repeating whenever you change the fine-tuning or deployment process, because a restore path that worked six months ago silently breaking is the way most backup strategies actually fail.

FAQ

Does GPUwerk back up the data on my rented node?

No. A GPUwerk instance is a rented Linux box with local NVMe storage; backing up whatever you put on it is your responsibility, the same as with any rented server. GPUwerk does not snapshot or back up customer data on an instance.

Do I need to back up base model weights I downloaded from Hugging Face?

Usually not as a priority. If a base model came from a public source like Hugging Face, re-downloading it after a loss costs time, not data, since the source of truth still exists elsewhere. Back it up only if your bandwidth to re-download is a real constraint, and prioritize backing up anything that doesn't exist anywhere else instead.

What's the minimum I should back up on an inference-only node?

Configuration files (server configs, LiteLLM routing rules, systemd units, environment files) and anything you can't regenerate from a public source, such as fine-tuned or merged checkpoints, custom prompt templates, and RAG indexes built from your own documents. These are usually small and worth copying off the node regularly even when large base model weights aren't.

Related pages

Rent a Spark, keep your own backups off it.

Root SSH access from minute one, so your own backup tooling works exactly the way it does anywhere else.

Read the SSH key guide See pricing