Setting up Grafana dashboards for a Spark inference node
The alert rules from setting up Prometheus alerts tell you when something's wrong. They don't show you what was happening in the ten minutes before it went wrong, or whether a metric has been drifting for days. That's what a dashboard is for, and since the DCGM exporter and vLLM metrics endpoint are already scraped, wiring them into Grafana is mostly a matter of a datasource and a panel layout.
Point Grafana at the Prometheus you already have
If Prometheus is scraping the DCGM exporter and vLLM's /metrics endpoint as set up in the alerting guide, Grafana needs nothing more than the same address as a datasource. Provision it from a file so it's reproducible rather than a one-time UI click.
# /etc/grafana/provisioning/datasources/prometheus.yml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://localhost:9090
isDefault: true
Provision the dashboard from JSON instead of building it by hand
A dashboard defined as a file lives in your repo next to the alert rules, gets reviewed the same way, and survives a Grafana reinstall. Point the provisioning config at a directory, then drop dashboard JSON files there.
# /etc/grafana/provisioning/dashboards/inference.yml
apiVersion: 1
providers:
- name: inference-node
folder: Inference
type: file
options:
path: /etc/grafana/dashboards
A minimal dashboard covering the four panels worth having first, GPU utilization, GPU memory, request rate, p95 latency:
# /etc/grafana/dashboards/inference-node.json
{
"title": "Inference node",
"panels": [
{
"title": "GPU utilization",
"type": "timeseries",
"targets": [{"expr": "DCGM_FI_DEV_GPU_UTIL"}],
"gridPos": {"x": 0, "y": 0, "w": 12, "h": 8}
},
{
"title": "GPU memory used / total",
"type": "timeseries",
"targets": [{"expr": "DCGM_FI_DEV_FB_USED / DCGM_FI_DEV_FB_TOTAL"}],
"gridPos": {"x": 12, "y": 0, "w": 12, "h": 8}
},
{
"title": "Request rate",
"type": "timeseries",
"targets": [{"expr": "rate(vllm_request_success_total[5m])"}],
"gridPos": {"x": 0, "y": 8, "w": 12, "h": 8}
},
{
"title": "p95 latency",
"type": "timeseries",
"targets": [{"expr": "histogram_quantile(0.95, rate(vllm_request_latency_seconds_bucket[5m]))"}],
"gridPos": {"x": 12, "y": 8, "w": 12, "h": 8}
}
]
}
These four expressions are the same metrics the alert rules already watch, DCGM's utilization and memory series, vLLM's latency histogram, just rendered as a trend instead of a threshold. That overlap is deliberate: a panel and an alert rule pointed at the same query means an alert firing sends you straight to a dashboard that already explains the shape of the problem.
Add a panel for the thing you actually got paged about
Once the base four are running, the next panel worth adding is whatever caused your last real incident. If an OOM was the last thing that woke you up, a panel tracking KV cache usage alongside GPU memory would have shown the trend building. Build the dashboard outward from your own failure history rather than trying to chart everything the exporters expose up front.
Keep the dashboard and the alert rules in the same place
Storing both the Prometheus rule file and the dashboard JSON in the same repo directory means a change to one prompts you to check the other, add an alert, add the panel that explains it; retire an alert, retire the panel that only existed to support it. Treating them as separate systems that happen to share a datasource is how a dashboard quietly drifts out of sync with what actually pages you.
FAQ
Do I need Grafana if I already have Prometheus alert rules?
Alert rules answer a yes-or-no question: has a threshold been crossed. They don't show you the shape of a metric over time, whether GPU memory has been climbing steadily for three days or jumped in the last ten minutes, which matters for diagnosing why an alert fired. A dashboard is where that investigation happens, not a replacement for alerting.
Can I provision a Grafana dashboard from a file instead of building it in the UI?
Yes, and it's the better approach for anything you want to reproduce or keep in version control. Grafana's provisioning system reads datasource and dashboard JSON from files on disk at startup, so the whole dashboard lives in your repo alongside the Prometheus rules it's built to complement, rather than only existing as UI state.
What's the minimum useful dashboard for a single inference node?
GPU utilization and memory used over time, request rate, and p95 latency. Those four panels cover capacity (is the GPU the bottleneck), headroom (how close to OOM), and caller experience (is it slow), which is most of what you'd actually look at during an incident or a capacity review.