Writing a runbook for your self-hosted inference service
Most self-hosted LLM setups start as one person's SSH session: start the engine, point the app at it, done. That's fine until the person who ran that session is asleep, on a call, or gone, and something breaks. A runbook is the document that lets someone else, or the same person at 3am with less patience than usual, fix the problem without reconstructing the whole system from first principles.
Start with what actually goes wrong
A useful runbook is organized around symptoms, not around the architecture diagram. "The endpoint returns 502" and "generation is slow but the endpoint responds" are different entries with different diagnostic steps, and a reader in the middle of an incident is searching for their symptom, not reading the document end to end. Build the list from what has actually happened, an engine crash on OOM per troubleshooting OOM errors, a certificate expiring, a container failing to restart, rather than a hypothetical inventory of everything that could theoretically fail.
Each entry needs a way to confirm the diagnosis
"Check if the GPU is out of memory" is not actionable on its own. "Run `nvidia-smi` and check whether memory used is near the card's total" is. Every runbook step should name the exact command, its expected output for the healthy case, and what a specific abnormal output means, so the person following it isn't guessing whether what they're looking at is the problem or a red herring.
Include the boring facts that are easy to forget under pressure
Where the systemd unit or Docker Compose file lives, which port each service binds to, what the health check endpoint returns when things are fine, where logs are written and how far back they retain. None of this is interesting to write down when everything is working, which is exactly why it's usually missing when it's needed. A short "topology" section at the top of the runbook, listing every process and port involved, saves more time during an incident than any individual fix instruction.
Write the restart procedure precisely
"Restart the service" undersells how much can go wrong in a restart on a GPU node: a process that doesn't release GPU memory cleanly, a container that comes back before the model has finished loading and starts failing health checks, a reverse proxy that needs restarting separately if its upstream changed. Write the actual sequence, in order, including how to confirm each step completed before moving to the next one, rather than trusting that "restart" is self-explanatory.
Note what a step can't fix
Some failures on a single dedicated node don't have a quick remediation, hardware failure being the obvious one. The runbook entry for that case isn't a fix, it's the decision tree: who to contact, what the recovery time estimate is, and what the fallback is while the node is down, which ties directly into failover planning for a single-node deployment. Writing "this can't be fixed quickly, here's what to do instead" is more useful than an entry that implies every problem has a five-minute solution.
Test it before you need it
A runbook step that was never executed, only written, is a guess dressed up as a procedure. Walk through each entry against a real or deliberately broken instance of the stack at least once, the same discipline covered for full disaster scenarios in disaster recovery testing, and fix the steps that don't actually work as written before trusting the document during a real incident.
FAQ
How is a runbook different from documentation?
Documentation explains how a system works and is read to build understanding. A runbook is written to be followed under time pressure by someone who may not have that understanding yet, or who has it but is too rushed to reconstruct it from memory. The test for a runbook step is whether it can be executed correctly without first understanding why it works.
Who should be able to follow the runbook?
Anyone with SSH access to the node and basic command-line familiarity, not only the person who built the stack. If a step assumes knowledge that only the original operator has, such as which config file controls a setting or what a specific error message means, that knowledge belongs in the runbook itself, not in the runbook writer's head.
How often should a runbook be updated?
After every incident it was used in, whether or not it worked as written, and after any change to the stack it describes, an engine upgrade, a new reverse proxy config, a different model. A runbook that wasn't updated after the underlying system changed is worse than no runbook, since it will be followed with confidence and produce the wrong result.