Monitoring and observability for a self-hosted LLM
A rented Spark gives you root over SSH and a console with instance status and billing. It doesn't give you dashboards for GPU load or request latency inside your own container, because that's application territory, not host territory. This page covers what's worth watching and which open-source tools do it well on a single machine.
What GPUwerk's console shows, and where it stops
The console and API tell you whether an instance is running, stopped or terminated, and what it's costing you per minute at the $0.79/hour on-demand rate. That's operational visibility at the platform level, not observability inside your workload. There's no built-in dashboard for GPU utilization, memory pressure, or inference latency, and no alerting on those signals. If you want that, you're setting it up yourself inside the container you have root on, the same as you would on hardware you owned outright.
What actually matters to watch
For a single machine serving an LLM, four things tell you most of what you need:
- GPU utilization and memory, since a Spark's 128 GB is unified between CPU and GPU, and running out of it is the most common failure mode for a model that's too large or a batch size that's too ambitious.
- Request latency, specifically time to first token and total generation time, since these are what a user actually experiences and won't show up in raw GPU numbers.
- Throughput, tokens per second under real concurrent load, which is different from the single-request numbers in a benchmark.
- Error rates, timeouts, out-of-memory kills, and failed requests at whatever serving layer sits in front of the model.
Start with nvidia-smi and nvtop
For a quick check, nvidia-smi gives you a snapshot of GPU utilization, memory used, and running processes, and ships with the driver stack already on the machine. nvidia-smi dmon streams that as a rolling log if you want to watch it change over a few minutes. nvtop is a fuller terminal UI on top of the same data, closer to htop for the GPU, and is a small install (apt install nvtop on most Debian-based images) if it isn't already there. Neither keeps history past what's on screen, which is fine for debugging a slow response right now and not enough if you want to know what happened overnight.
Prometheus and Grafana for anything you want to keep
Once you want a graph over time, an alert when memory crosses 90%, or a dashboard you can glance at instead of SSHing in, Prometheus and Grafana are the standard open-source pairing and work fine on a single node. The rough shape: run NVIDIA's DCGM exporter as a container alongside your workload to expose GPU metrics on a scrape endpoint, point a Prometheus instance at it to store the time series, and put Grafana on top for dashboards and alert rules. This is a real setup task, not a flag you flip, and worth doing once request volume or uptime expectations make "I'll check nvidia-smi if something feels slow" insufficient.
If your serving layer is LiteLLM, it exposes its own request-level metrics (spend, latency, error counts per key) through the same Prometheus format, so one Grafana instance can cover both the GPU and the gateway sitting in front of it.
Application-level logging
Whatever inference engine you run, vLLM, llama.cpp, TGI, log at minimum request duration, token counts, and any error or timeout. Plain structured logs (JSON lines to a file under /workspace so they survive a stop) are enough at small scale; ship them to something like Loki once you have more than one process to correlate across.
Practical checklist
- Use
nvidia-smiornvtopfor a quick look, no setup required. - Set up DCGM exporter, Prometheus and Grafana once you want history and alerts, not before.
- Watch GPU memory specifically, since the 128 GB unified pool is the resource most likely to run out on a self-hosted deployment.
- Log request latency and errors at the serving layer, since GPU metrics alone won't show what users actually feel.
- Keep dashboard config and exporter setup in
/workspaceor a git repo, since anything installed elsewhere on the container doesn't survive a rebuild.
None of this is something GPUwerk operates on your behalf; it's advice for running your own stack well on a single dedicated machine. See the LiteLLM setup guide if you're putting a gateway in front of your model, or a first engagement if you want help designing the observability layer for a specific deployment.