Setting up Prometheus alerts for a Spark inference node
nvidia-smi and nvtop are fine for watching a node while you're looking at it. They tell you nothing about the three hours you weren't. Prometheus alerting closes that gap: metrics scraped continuously, and rules that page someone the moment a threshold is crossed rather than whenever a human next happens to check.
Get GPU metrics into Prometheus first
This builds on the monitoring setup in monitoring and observability. NVIDIA's DCGM exporter is the standard way to expose GPU metrics, utilization, memory used, temperature, in a format Prometheus can scrape on an interval.
# run the DCGM exporter as a container alongside your inference service docker run -d --gpus all --rm -p 9400:9400 \ --name dcgm-exporter \ nvcr.io/nvidia/k8s/dcgm-exporter:3.3.9-3.6.1-ubuntu22.04 # confirm it's exposing metrics curl -s localhost:9400/metrics | grep DCGM_FI_DEV_GPU_UTIL
Point Prometheus at it, and at your inference engine's own metrics endpoint if it exposes one, vLLM does out of the box on /metrics.
# /etc/prometheus/prometheus.yml
scrape_configs:
- job_name: 'dcgm'
static_configs:
- targets: ['localhost:9400']
scrape_interval: 15s
- job_name: 'vllm'
static_configs:
- targets: ['localhost:8000']
metrics_path: /metrics
scrape_interval: 15s
Three alerts worth having on day one
It's tempting to alert on everything the exporters expose. Start narrower: the endpoint being down, GPU memory heading toward exhaustion, and request latency crossing a threshold that means the service feels broken to a caller. Each of these maps to a real incident category rather than a metric that happened to be available.
# /etc/prometheus/rules/inference.yml
groups:
- name: inference-node
rules:
- alert: EndpointDown
expr: up{job="vllm"} == 0
for: 1m
labels: { severity: critical }
annotations:
summary: "Inference endpoint {{ $labels.instance }} is unreachable"
- alert: GPUMemoryNearLimit
expr: DCGM_FI_DEV_FB_USED / DCGM_FI_DEV_FB_TOTAL > 0.92
for: 5m
labels: { severity: warning }
annotations:
summary: "GPU memory on {{ $labels.instance }} above 92% for 5 minutes"
- alert: RequestLatencyHigh
expr: histogram_quantile(0.95, rate(vllm_request_latency_seconds_bucket[5m])) > 30
for: 5m
labels: { severity: warning }
annotations:
summary: "p95 request latency on {{ $labels.instance }} above 30s"
The GPUMemoryNearLimit rule is the one that earns its keep fastest: it's the leading indicator for the OOM crashes covered in troubleshooting OOM errors, and catching it at 92 percent gives you a chance to act, restart with tighter batching, hold off deploying a second model, before the process crashes on its own.
Use for: to stop noise from becoming the norm
Every rule above has a for: duration, meaning the condition has to hold for that whole window before it fires. Without it, a brief spike, a burst of long-context requests, a garbage-collection pause, fires an alert every time, and the fastest way to make people ignore alerts is to send one for things that resolve on their own in ten seconds. Tune the duration to your traffic pattern: five minutes is a reasonable default for a low-traffic internal endpoint, shorter if you're serving external users who feel every outage immediately.
Route severity, don't page for everything
Split alerts by severity label and route them differently: critical to whatever actually wakes someone up, warning to a channel someone checks during working hours. GPU memory climbing toward a limit is worth knowing about; it isn't worth a 3am page on its own, since the actual failure, the OOM, will fire its own critical alert if it happens. Save the paging tier for things that are already broken, not things that might be soon.
Verify the alert actually fires before you trust it
An alert rule that's never triggered is unverified, the same way an untested runbook step is a guess. Force a condition, load a model that pushes memory past 92 percent deliberately, or stop the service and confirm EndpointDown actually reaches wherever you've routed it, at least once after setting the rules up. This pairs with the health check work in writing health checks for your inference endpoint: the alert is only as good as the signal it's watching.
FAQ
Do I need DCGM exporter or does nvidia-smi metrics work for Prometheus?
NVIDIA's DCGM exporter is the standard choice: it runs as a long-lived process exposing a /metrics endpoint Prometheus can scrape on an interval, which is what alerting needs. nvidia-smi is built for a human reading a terminal, not for continuous scraping, and doesn't expose a Prometheus-format endpoint on its own.
What's the minimum set of alerts worth setting up first?
Endpoint down (the health check fails), GPU memory near the card's limit (the leading indicator of an imminent OOM), and request latency crossing a threshold that would mean the service feels broken to a caller. Those three catch most real incidents; add more once you understand your own failure patterns rather than alerting on everything Prometheus can measure.
How do I avoid alert fatigue on a single-node setup?
Use a `for:` duration on every rule so a brief spike doesn't page anyone, and route warnings differently from critical alerts, a warning to a channel you check occasionally, a critical alert to something that actually wakes you up. An alert that pages for a condition that resolves itself in ten seconds trains you to ignore alerts, which defeats the purpose of having them.