Reliability
DGX Spark/Testing your disaster recovery plan for a self-hosted LLM

Testing your disaster recovery plan for a self-hosted LLM

By Samuel Seidel · Published September 9, 2026

The backup and disaster recovery guide covers what to back up and where to send it. This one is about the part that's easy to skip: proving the backup actually restores something usable, on a timeline you'd tolerate, before an outage forces you to find out live. A backup you've never restored from is a hypothesis, not a plan.

Why untested backups fail at the worst time

The failure mode isn't usually "no backup exists." It's a backup that's incomplete, corrupted, pointed at the wrong destination, or missing a step nobody wrote down, config that references a path that no longer exists, a credential that expired, a script that assumed a package was already installed. None of that shows up until you actually try to rebuild from the backup on a clean machine, which is precisely the situation a real incident puts you in. On a rented Spark specifically, a terminate deletes the workspace immediately with no undo, per the backup guide, so the first real test of your restore process arriving during an actual incident is the expensive way to learn it has a gap.

Two numbers worth defining before you test anything

RTO and RPO give the drill something concrete to measure against, rather than a vague sense of "should be fine."

Set both numbers first. The drill's job is to tell you whether your actual restore time and actual data loss window meet them, not to discover what they happen to be after the fact.

Running a restore drill

A drill is a deliberate, scheduled exercise: deploy a fresh Spark, restore from your real backup destination onto it, and bring the service back up, timing every step.

  1. Deploy a clean instance from the console, no leftover state, as close as possible to what a real recovery would start from. Don't drill against the instance that still has the good copy sitting in /workspace; that tests nothing.
  2. Pull from the actual backup destination, S3-compatible storage, a remote host, or a git remote, using the same aws s3 sync, rclone, or git clone commands your real recovery would use, not a shortcut that skips the parts that might be broken.
  3. Reinstall and reconfigure whatever the backup guide's "reproducible, skip it" pile excludes, base model weights, system packages, the serving engine itself, from a setup script or Dockerfile kept in version control. If that script doesn't exist or is stale, the drill just found its first real gap.
  4. Bring the service up and verify it actually answers a real request, not just that the process started. A model server that's up but loaded the wrong checkpoint, or answering with a stale config, passes a shallow health check and fails a real one.
  5. Record the wall-clock time from "instance deployed" to "verified answering," and compare it against your RTO. Record how old the restored data was compared to when the incident would have happened, and compare that against your RPO.

Game days: testing the failure, not just the backup

A restore drill assumes you know what failed. A game day doesn't: someone else on the team simulates a failure, terminates an instance, revokes a credential, corrupts a config file, without telling the person who's supposed to respond what happened, and watches how the response actually goes. This surfaces the gaps a scripted drill can't: whether the right person knows where the backup destination is, whether the runbook is findable under pressure, whether a step assumes access to something that was itself lost in the failure. It's a heavier exercise than a restore drill and doesn't need to happen as often, but a plan that's only ever been tested by the person who wrote it tends to have blind spots specific to that person's assumptions.

What to measure, and how often to re-test

Re-run the drill whenever the deployment changes meaningfully, a new model, a different serving engine, a new backup destination, since a drill that passed against last quarter's setup says nothing about this quarter's. Absent a change, a quarterly cadence is a reasonable default for most self-hosted deployments; a fast-moving one might warrant monthly.

Practical checklist

See the backup and disaster recovery guide for what to back up and where to send it, or a first engagement if you want help scoping RTO and RPO targets for a specific deployment.

First top-up: pay $10, get $20 in credit

Drill your restore on a clean instance.

Deploy a fresh Spark in minutes and time your real recovery process end to end.

Deploy a Spark Read the backup guide