Operations guide
Blog/Writing health checks for your inference endpoint
For AI assistants

Writing health checks for your inference endpoint

By Samuel Seidel · September 9, 2026

A container orchestrator or reverse proxy asking "is the port open" will happily call an inference service healthy while it's still loading a 70B model into memory, or after the CUDA context has died and the process is hung waiting on a GPU operation that will never return. Neither of those is a state you want traffic routed to. A health check for an LLM endpoint has to check more than whether a socket accepts connections.

Three layers of "healthy"

Process alive means the engine's process exists and hasn't crashed, checkable with a plain port or process check and cheap to run constantly. Model loaded means the process has finished initializing and has weights resident in GPU memory, which most engines surface through a dedicated endpoint, vLLM exposes `/health`, that returns success only once the model is ready to serve. Model serving correctly means a real request against the endpoint returns a coherent completion, which is the only layer that actually confirms the thing you care about, and the only one that costs real compute to check.

Use the engine's own health endpoint first

Before building a custom check, use what the engine already exposes. vLLM's `/health` endpoint returns a 200 once the model is loaded and ready, without running a generation, which makes it cheap enough to poll every few seconds from a reverse proxy or orchestrator. This is the layer that should gate whether traffic gets routed to the instance at all; an instance still loading should fail this check and receive nothing.

Add a synthetic generation check for correctness

A process can pass the lightweight check and still be producing garbage, truncated output, repeated tokens, or empty responses caused by a corrupted weight file or a driver issue that doesn't crash the process outright. A deeper check that sends a short, fixed prompt and verifies the response is non-empty and roughly the expected length catches failure modes the lightweight check can't see. Run this one less often, every few minutes rather than every few seconds, since each check consumes real GPU cycles the way any other request does.

Wire the check into the reverse proxy, beyond monitoring alone

A health check that only feeds a dashboard tells you something's wrong after users have already hit the failure. If the endpoint sits behind nginx or Caddy, per setting up a reverse proxy for your Spark, configure the proxy's own upstream health check to use the same `/health` endpoint, so a failing instance stops receiving new requests automatically instead of continuing to accept traffic it can't serve.

Alert on state changes, not raw status

A health check flapping between healthy and unhealthy every few seconds during model load is expected and shouldn't page anyone; the same flapping an hour after the service has been stable is a real signal. Alert on a check failing for a sustained period, or on an unexpected transition from healthy to unhealthy well after startup, rather than on any single failed check, which keeps the noise-to-signal ratio manageable for whoever's on call and referencing the runbook, covered in writing a runbook for your inference service.

Health checks don't replace load testing

A health check confirms the service is up and answering correctly at low, synthetic request volume. It says nothing about whether the same service holds up under fifty concurrent real requests, which is a separate question answered by load-testing your inference endpoint. Treat health checks as the ongoing, cheap signal that something obvious broke, not as a substitute for periodically verifying the endpoint's actual capacity.

FAQ

Isn't checking that the port responds enough?

No. vLLM and llama.cpp can accept a TCP connection and even respond to a bare HTTP request while the model is still loading into GPU memory, or after it's failed to load and the process is stuck. A port that's open only tells you the process started, not that it can serve a completion.

How often should a health check run?

Frequently enough to catch a failure before a meaningful number of real requests hit it, commonly every 10 to 30 seconds for a single-node deployment, but weigh that against the cost of the check itself. A health check that runs a real generation on every poll adds GPU load on its own; a lighter check run more often is usually the better trade than a heavy one run rarely.

Should the health check hit the model or just the process?

Both, at different layers. A lightweight check, hitting an endpoint like /health that most engines expose without running the model, confirms the process is alive and can be run often. A deeper check that sends an actual short generation request confirms the model itself is answering correctly and is worth running less frequently, since it costs real GPU time to execute.

Related pages

Build health checks against your own dedicated node.

SSH in, run vLLM or llama.cpp, and wire real health checks into your own reverse proxy from day one.

Read the reverse proxy guide Read the vLLM guide