Operations guide
Blog/Testing failover for a single-node deployment
For AI assistants

Testing failover for a single-node deployment

By Samuel Seidel · September 9, 2026

"Failover" usually implies a second system standing by to take over when the first one fails. A single dedicated Spark doesn't have that, and no amount of configuration on the one node changes it. What's worth testing on a single-node deployment is different and narrower: whether the service comes back correctly on its own after a crash, and how long that actually takes when you measure it instead of assume it.

Name the failure modes separately

A process crash, an OOM kill, an unhandled exception in the engine, is recoverable by a restart in minutes if the restart is configured correctly. A full node failure, hardware fault, network loss to the node, is not something a single node can recover from by itself; the fix is provisioning a replacement, which takes considerably longer. Conflating these two categories is the most common way single-node resilience planning goes wrong: the easy case gets tested and treated as proof the hard case is covered too, when it isn't.

Test the restart path deliberately

Kill the inference engine process, or the container it runs in, and time how long it takes for the service to be answering requests correctly again, rather than merely running. A container that reports as up before the model has finished loading into GPU memory will fail requests during that window while looking healthy at the process level, so the test needs to check for correct responses instead of process status alone. Configure `--restart unless-stopped` or the systemd equivalent so this recovery happens automatically rather than requiring someone to notice and intervene.

Verify state actually persists

If the deployment depends on anything written to local disk, model weights, a vector index, LiteLLM's request log, confirm those survive a restart and a reboot, not only a graceful stop. This overlaps with backup practice covered in backup and snapshot strategies, but the specific thing to test here is narrower: does the service come back into the same state it was in, or does a restart quietly lose something.

Decide what "acceptable downtime" means before an incident, not during one

A single-node deployment will have some downtime during a hardware fault; the honest planning question is how much, and what the team does while it's happening. Write that answer down, expected recovery time for a process crash versus a hardware failure, and what the fallback plan is for the harder case, as part of the runbook, per writing a runbook for your inference service, rather than discovering the answer live during the first real outage.

When single-node recovery isn't enough

If the cost of downtime is high enough that a restart-and-wait recovery time isn't acceptable, the honest next step is adding a second node, which turns this into a scaling and routing problem covered in scaling to a fleet, not a single-node configuration problem. No amount of tuning restart policies on one Spark produces the availability characteristics of two nodes behind a load balancer; that requires actually having two nodes.

Rehearse the replacement-node path

For teams that accept single-node risk but want to shorten hardware-failure recovery time, the useful preparation isn't a standby node, it's a tested, scripted path to standing up a replacement: weights staged somewhere reachable, a deployment script that doesn't depend on manual steps, and DNS or reverse proxy config that can be repointed quickly. Run through that sequence against a fresh node at least once before you need it for real, the same principle as disaster recovery testing more broadly, so the first time it's needed isn't also the first time it's tried.

FAQ

Is single-node failover even a real thing?

Not in the sense of a second node automatically taking over when the first fails; that requires a second node, which is a cluster, covered separately in scaling to a fleet. What single-node failover testing covers is narrower: whether the service running on that one node comes back correctly after a process crash, a container restart, or a reboot, without someone having to manually reconstruct its state.

What's the honest recovery time for a single node that goes down hard?

There isn't a universal number, it depends on whether the failure is a software crash that a restart fixes in minutes or a hardware fault that requires provisioning a replacement node and reloading weights, which can take considerably longer. State both cases explicitly to whoever depends on the service rather than quoting a single recovery time figure that only covers the easy case.

Should I keep a second Spark on standby just in case?

That depends on how much an outage costs you and is a cost-versus-risk decision each team has to make for itself, not a default recommendation. A cheaper middle ground many teams land on is keeping model weights and a deployment script ready to go, so a replacement node can be brought up quickly without idling a second Spark the whole time, rather than paying for permanent standby capacity.

Related pages

Spin up a replacement node in minutes, not days.

On-demand Sparks at $0.79/hour mean your recovery plan doesn't require a second node sitting idle.

Read the scaling guide Read the backup guide