Backup and disaster recovery for your DGX Spark workload
GPUwerk's pricing page is direct about this: it keeps only a periodic recovery copy of /workspace, refreshed roughly every six hours, solely to recover from hardware failure, and that is not a backup service. That's not a limitation to work around, it's the actual architecture, and it means backup and recovery planning is entirely on you. This page covers what to back up, where to send it, and how often.
Why there's no platform backup
A Spark instance is one workspace on one node. GPUwerk's own recovery copy is a periodic snapshot for hardware failure, refreshed roughly every six hours, not a substitute for your own backup and not something you can restore from on request. The pricing page's callout is equally direct: "Protect your results... Keep your own backup." If your balance hits zero or you hit a monthly budget, the instance is stopped and the machine is released for someone else to rent; the workspace is saved off it and kept for 7 days so topping up and pressing start restores it, and deleted for good after that window, per the storage and persistence guide. A manual terminate skips straight to deleting the workspace, with no 7-day window and no copy kept.
What's actually worth backing up
Not everything in /workspace is equally valuable. Sort it into two piles:
- Irreplaceable, back it up: fine-tuned or LoRA weights you trained yourself, configuration files (LiteLLM's
config.yaml, systemd units, environment files), any dataset you curated or cleaned that isn't a straight copy of a public source, and conversation or usage logs if you need them for audit or debugging later. - Reproducible, skip it: base model downloads from Hugging Face or similar, since re-pulling them costs bandwidth and time but no data loss; installed packages and system dependencies, since these belong in a setup script or Dockerfile you keep in version control, not in the backup itself.
The practical test: if losing it means redoing a training run or a curation pass you can't easily repeat, back it up. If losing it means re-running a script, write the script down instead of backing up its output.
Where to send it
GPUwerk doesn't provide or manage a backup destination, so you're picking your own. Reasonable options, all reachable from a Spark's root SSH access:
- S3-compatible object storage (AWS S3, Backblaze B2, or a self-hosted MinIO instance elsewhere) via
aws s3 syncorrclone, which handles incremental syncs well for large model weight directories. - rsync or scp to a machine you control, simplest to set up, and fine for infrequent manual backups of a moderate amount of data.
- A git remote for config files, scripts and small datasets, since these benefit from version history in a way that large binary weight files don't.
Model weight files can run into tens of gigabytes, so factor egress and storage cost into the choice; GPUwerk's own network doesn't charge egress fees on outbound traffic per the pricing page, but your destination might.
How often
Match frequency to how much rework a loss would cost you, not to a fixed calendar schedule:
- A fine-tuning run in progress: checkpoint every few hours, or at each meaningful loss milestone, so a crash costs you an hour, not a day.
- A stable inference deployment: back up config after each change, and weights once, after training completes, not on a recurring schedule.
- Before any terminate: a final sync is the one backup you should treat as non-optional, since terminate deletes the workspace immediately and there's no console-level undo.
Practical checklist
- Assume nothing outside your own backup destination survives a terminate or a zero-balance wipe.
- Back up fine-tuned weights, config and curated data; skip base downloads and packages, keep those reproducible instead.
- Pick an S3-compatible target or a remote you control, and automate the sync rather than relying on remembering to run it.
- Always sync before you terminate, since the delete on confirm is immediate.
- Keep auto top-up on if you want to avoid an unplanned zero-balance termination catching you without a recent backup.
See the storage and persistence guide for the full stop-versus-terminate mechanics, or a first engagement if you want help designing a backup pipeline around a specific workload.