Docs analysis
DGX Spark/Monitoring and observability for a self-hosted LLM

Monitoring and observability for a self-hosted LLM

By Samuel Seidel · Published September 9, 2026

A rented Spark gives you root over SSH and a console with instance status and billing. It doesn't give you dashboards for GPU load or request latency inside your own container, because that's application territory, not host territory. This page covers what's worth watching and which open-source tools do it well on a single machine.

What GPUwerk's console shows, and where it stops

The console and API tell you whether an instance is running, stopped or terminated, and what it's costing you per minute at the $0.79/hour on-demand rate. That's operational visibility at the platform level, not observability inside your workload. There's no built-in dashboard for GPU utilization, memory pressure, or inference latency, and no alerting on those signals. If you want that, you're setting it up yourself inside the container you have root on, the same as you would on hardware you owned outright.

What actually matters to watch

For a single machine serving an LLM, four things tell you most of what you need:

Start with nvidia-smi and nvtop

For a quick check, nvidia-smi gives you a snapshot of GPU utilization, memory used, and running processes, and ships with the driver stack already on the machine. nvidia-smi dmon streams that as a rolling log if you want to watch it change over a few minutes. nvtop is a fuller terminal UI on top of the same data, closer to htop for the GPU, and is a small install (apt install nvtop on most Debian-based images) if it isn't already there. Neither keeps history past what's on screen, which is fine for debugging a slow response right now and not enough if you want to know what happened overnight.

Prometheus and Grafana for anything you want to keep

Once you want a graph over time, an alert when memory crosses 90%, or a dashboard you can glance at instead of SSHing in, Prometheus and Grafana are the standard open-source pairing and work fine on a single node. The rough shape: run NVIDIA's DCGM exporter as a container alongside your workload to expose GPU metrics on a scrape endpoint, point a Prometheus instance at it to store the time series, and put Grafana on top for dashboards and alert rules. This is a real setup task, not a flag you flip, and worth doing once request volume or uptime expectations make "I'll check nvidia-smi if something feels slow" insufficient.

If your serving layer is LiteLLM, it exposes its own request-level metrics (spend, latency, error counts per key) through the same Prometheus format, so one Grafana instance can cover both the GPU and the gateway sitting in front of it.

Application-level logging

Whatever inference engine you run, vLLM, llama.cpp, TGI, log at minimum request duration, token counts, and any error or timeout. Plain structured logs (JSON lines to a file under /workspace so they survive a stop) are enough at small scale; ship them to something like Loki once you have more than one process to correlate across.

Practical checklist

None of this is something GPUwerk operates on your behalf; it's advice for running your own stack well on a single dedicated machine. See the LiteLLM setup guide if you're putting a gateway in front of your model, or a first engagement if you want help designing the observability layer for a specific deployment.

First top-up: pay $10, get $20 in credit

Root over SSH, your monitoring stack.

Deploy a dedicated Spark and instrument it exactly the way you would your own hardware.

Deploy a Spark Read the LiteLLM guide