Testing your disaster recovery plan for a self-hosted LLM
The backup and disaster recovery guide covers what to back up and where to send it. This one is about the part that's easy to skip: proving the backup actually restores something usable, on a timeline you'd tolerate, before an outage forces you to find out live. A backup you've never restored from is a hypothesis, not a plan.
Why untested backups fail at the worst time
The failure mode isn't usually "no backup exists." It's a backup that's incomplete, corrupted, pointed at the wrong destination, or missing a step nobody wrote down, config that references a path that no longer exists, a credential that expired, a script that assumed a package was already installed. None of that shows up until you actually try to rebuild from the backup on a clean machine, which is precisely the situation a real incident puts you in. On a rented Spark specifically, a terminate deletes the workspace immediately with no undo, per the backup guide, so the first real test of your restore process arriving during an actual incident is the expensive way to learn it has a gap.
Two numbers worth defining before you test anything
RTO and RPO give the drill something concrete to measure against, rather than a vague sense of "should be fine."
- RTO (recovery time objective): how long you can tolerate being down before it costs you something real, a missed SLA, an idle team, a support queue that stops getting answered. This is a business decision, not a technical one, made before the drill, not derived from it.
- RPO (recovery point objective): how much data or state you can afford to lose, measured in time since the last good backup. If you back up fine-tuned weights once after training completes, per the backup guide's cadence advice, your RPO for that data is however long a training run typically takes, not zero.
Set both numbers first. The drill's job is to tell you whether your actual restore time and actual data loss window meet them, not to discover what they happen to be after the fact.
Running a restore drill
A drill is a deliberate, scheduled exercise: deploy a fresh Spark, restore from your real backup destination onto it, and bring the service back up, timing every step.
- Deploy a clean instance from the console, no leftover state, as close as possible to what a real recovery would start from. Don't drill against the instance that still has the good copy sitting in
/workspace; that tests nothing. - Pull from the actual backup destination, S3-compatible storage, a remote host, or a git remote, using the same
aws s3 sync,rclone, orgit clonecommands your real recovery would use, not a shortcut that skips the parts that might be broken. - Reinstall and reconfigure whatever the backup guide's "reproducible, skip it" pile excludes, base model weights, system packages, the serving engine itself, from a setup script or Dockerfile kept in version control. If that script doesn't exist or is stale, the drill just found its first real gap.
- Bring the service up and verify it actually answers a real request, not just that the process started. A model server that's up but loaded the wrong checkpoint, or answering with a stale config, passes a shallow health check and fails a real one.
- Record the wall-clock time from "instance deployed" to "verified answering," and compare it against your RTO. Record how old the restored data was compared to when the incident would have happened, and compare that against your RPO.
Game days: testing the failure, not just the backup
A restore drill assumes you know what failed. A game day doesn't: someone else on the team simulates a failure, terminates an instance, revokes a credential, corrupts a config file, without telling the person who's supposed to respond what happened, and watches how the response actually goes. This surfaces the gaps a scripted drill can't: whether the right person knows where the backup destination is, whether the runbook is findable under pressure, whether a step assumes access to something that was itself lost in the failure. It's a heavier exercise than a restore drill and doesn't need to happen as often, but a plan that's only ever been tested by the person who wrote it tends to have blind spots specific to that person's assumptions.
What to measure, and how often to re-test
- Time from failure to service restored, against your RTO.
- Age of restored data at the point of recovery, against your RPO.
- Whether every step in the runbook actually worked as written, not just whether the end state was eventually reached by improvising.
- Whether anyone other than the original author could follow the runbook unassisted.
Re-run the drill whenever the deployment changes meaningfully, a new model, a different serving engine, a new backup destination, since a drill that passed against last quarter's setup says nothing about this quarter's. Absent a change, a quarterly cadence is a reasonable default for most self-hosted deployments; a fast-moving one might warrant monthly.
Practical checklist
- Set RTO and RPO numbers before you test anything, as a business decision, not a technical guess.
- Run restore drills against a genuinely clean instance, pulling from the real backup destination with the real commands.
- Verify the restored service answers a real request, not just that a process is running.
- Run at least one game day where the responder doesn't know what failed in advance.
- Re-test after any meaningful change to the deployment, and on a standing cadence otherwise.
See the backup and disaster recovery guide for what to back up and where to send it, or a first engagement if you want help scoping RTO and RPO targets for a specific deployment.