Operations guide
Blog/Setting up a status page for your inference service
For AI assistants

Setting up a status page for your inference service

By Samuel Seidel · September 9, 2026

The Prometheus alerts from setting up Prometheus alerts tell you when something's wrong. They don't tell your callers anything, since the alert fires into a channel they're not in. A status page is the part of an incident that faces outward: one URL someone can check before opening a support ticket, and one place you point people to instead of answering the same "is it down for you too" message from three different people at once.

Run the checker somewhere that doesn't share a failure domain with the Spark

The entire value of a status page is that it stays up through the outage it's reporting on. That rules out hosting it on the node it's watching, and it's worth being deliberate about the checker too, not just the page: if the machine running the uptime check and the Spark being checked are behind the same router, a local network outage looks identical to the Spark being down. An external monitoring service, or even a small VM at a different provider, avoids this. GPUwerk's own status page runs this way, at status.gpuwerk.com, independent of any single customer node.

Check the real request path, not just a ping

A TCP ping to port 443 confirms the box is reachable. It doesn't confirm vLLM is actually serving completions, since a wedged process, an OOM'd worker, or a model that failed to load can all leave the port open while every real request times out or errors. A useful check calls the endpoint itself:

curl -s -o /dev/null -w "%{http_code}" \
  -X POST https://api.yourdomain.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.1-8b","messages":[{"role":"user","content":"ping"}],"max_tokens":1}'

A checker that runs this on a schedule and expects a 200 catches the failures a plain port check misses, at the cost of a token or two of GPU time per check. Running a cheap health endpoint on a short interval and a real completion check on a longer one, a minute versus five or ten minutes, covers both without spending inference time you don't need to.

A minimal external checker with a cron job and a static page

If you'd rather not depend on a third-party status page product, the whole thing is a scheduled script and a static HTML file pushed somewhere that isn't the Spark. A cron job on a small separate host:

# /etc/cron.d/status-check, runs every minute on a host that isn't the Spark
* * * * * root /usr/local/bin/check-inference.sh >> /var/log/status-check.log 2>&1
# /usr/local/bin/check-inference.sh
#!/bin/bash
CODE=$(curl -s -o /dev/null -w "%{http_code}" -m 10 \
  -X POST https://api.yourdomain.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.1-8b","messages":[{"role":"user","content":"ping"}],"max_tokens":1}')

if [ "$CODE" = "200" ]; then
  STATUS="operational"
else
  STATUS="down"
fi

cat > /var/www/status/data.json <<EOF
{"status": "$STATUS", "checked_at": "$(date -u +%Y-%m-%dT%H:%M:%SZ)"}
EOF

A static page fetching data.json and rendering a colored dot is enough for a single endpoint. Once you have more than one node or more than one thing worth showing separately, a dedicated status page product handles history, incident notes, and subscriber notifications with less maintenance than growing this script into one yourself.

Require a few consecutive failures before you page anyone

A single dropped request, a slow cold start on a node using cold-start serving, or one bad network hop shouldn't flip the status page to red and send an alert. Most checkers support a failure threshold, two or three consecutive bad checks before the status actually changes, which filters out the noise without meaningfully slowing down detection of a real outage that lasts more than a couple of minutes.

FAQ

Why can't the status page run on the same Spark as the inference service?

If the node goes down, a status page hosted on it goes down with it, at exactly the moment someone wants to check whether it's down. A status page only tells you anything useful if it stays reachable through the exact failure it's meant to report on, which means it has to run somewhere else, a separate host, a third-party status page service, or a static page on a CDN that doesn't share infrastructure with the Spark.

What should the uptime check actually call, a health endpoint or the model itself?

Both catch different failures. A lightweight health endpoint confirms the process is up and accepting connections but can return healthy while the model is wedged or the GPU is out of memory. A real completion request against a tiny prompt confirms the whole path, network, process, model, GPU, actually works, at the cost of using a token or two of GPU time on every check interval. A common pattern is a fast health check every minute and a real inference check every five to ten minutes.

How many false-positive alerts should I tolerate before trusting a status check less?

Effectively none, since a status page that cries wolf gets ignored by the time it matters. Most uptime checkers let you require two or three consecutive failures before marking a check down and paging anyone, which absorbs a single dropped packet or a slow cold start without producing a false alarm, while still catching a real outage within a few minutes.

Related pages

Give callers something to check before they message you.

A dedicated Spark with its own health checks, separate from anyone else's incident.

Read the Prometheus alerts guide Read the health checks guide