Operations guide
Blog/Logging and log retention for a self-hosted LLM API
For AI assistants

Logging and log retention for a self-hosted LLM API

By Samuel Seidel · September 9, 2026

Logging for a typical HTTP API is a solved problem: log the request line, status code, latency, maybe headers, keep it for a few weeks, done. An LLM API breaks that pattern because the request body and response body can each be substantial text, and unlike a URL path or a status code, that text is often the exact thing a customer told you not to send anywhere, whether that's a contract, a patient note, or source code. Logging strategy for a self-hosted endpoint needs to treat content separately from everything else.

Split metadata logging from content logging

Metadata, timestamp, model name, token counts, latency, HTTP status, requesting API key, is low-risk to log broadly and keep for a long time. It's what you need for capacity planning, billing reconciliation, and spotting a client that's misbehaving. Content, the actual prompt and completion text, is a different category: useful for debugging quality problems and for some compliance obligations, but expensive to secure properly and the reason self-hosting was probably chosen in the first place. Treating these as one undifferentiated "request log" tends to end with either content logged somewhere it shouldn't be, or metadata deleted too aggressively because it got bundled in with a short retention policy meant for the sensitive part.

If you're running LiteLLM as a proxy in front of vLLM or llama.cpp, its request logging already separates these reasonably well: it can emit structured logs with token counts, latency, and model routing decisions, with content logging as a separate, explicit opt-in rather than the default.

Decide whether you need content logs at all

The honest first question is whether you need prompt and completion content in logs, rather than what to do with it once it's there. Debugging a quality regression, tuning a system prompt, or investigating a specific user complaint are real reasons to want it. "Might be useful someday" is not a reason, it's a default that turns your log storage into an unbounded copy of everything anyone has ever sent the model. If you do need content logs, scope them: log at the application layer for specific flows under investigation rather than turning on full-content logging for the whole API permanently.

Where logs should live

On a rented Spark, logs from the inference engine and any reverse proxy in front of it (see our guide on setting up TLS for your inference endpoint) land on local disk by default. That's fine for debugging in the moment, but local disk logs don't survive a node being redeployed, and they're not centralized if you're running more than one node, a scenario covered in scaling to a fleet. Shipping logs off the node to a log aggregator you control, whether that's a self-hosted stack like Loki or a managed service, keeps them queryable and durable independent of any one node's lifecycle. If content is being logged, the aggregator needs the same access controls as the inference endpoint itself; shipping sensitive content off one server just to leave it unprotected on another isn't progress.

A retention schedule that matches the risk

A reasonable starting point, adjustable to your actual compliance obligations: metadata (latency, token counts, status codes, request IDs) retained for months, since it's cheap to store and useful for trend analysis; access and authentication logs, who connected with which API key and when, retained per whatever audit requirement applies, discussed further in audit logging for compliance; content logs, if you're keeping them at all, on the shortest retention period that still serves their purpose, often measured in days rather than months. Set retention as an automated deletion job, not a manual cleanup task someone is supposed to remember. A `logrotate` config or a scheduled job that purges log files past a given age is a few lines and removes the dependency on anyone remembering.

Redaction is not a substitute for short retention

It's tempting to log full content and plan to redact sensitive fields later with a regex pass over the log files. This works poorly in practice: prompt content doesn't follow a fixed schema the way a form submission does, so pattern-based redaction misses whatever it wasn't specifically written to catch. Treat unredacted content logs as sensitive for their entire retention window rather than assuming a cleanup pass will make them safe to keep longer.

What this looks like end to end

A workable setup on a Spark: LiteLLM or your reverse proxy logs metadata for every request to a local file, shipped nightly to a central log store with a months-long retention. Content logging is off by default and turned on per debugging session, writing to a separate log stream with a retention measured in days and deleted automatically. Authentication events, covered in our guide on API key rotation, get their own longer-retained log independent of both. Three streams, three retention periods, matched to what each one is actually for.

FAQ

Should I log full prompts and completions by default?

Only if you have a specific reason to, such as debugging quality issues or a compliance requirement to retain them. Logging content by default means every log line is now sensitive data with the same handling requirements as the underlying request, which is a bigger commitment than most teams intend to make when they turn on request logging.

Does GPUwerk log the contents of my requests?

GPUwerk provides the node; what runs on it, including any inference server or proxy logging, is configured by the customer. GPUwerk's own infrastructure logging covers access to the node itself, not the contents of requests made to whatever API you run on it.

How long should I keep LLM request logs?

There's no single correct answer; it depends on what you're keeping logs for. Metadata used for capacity planning and billing is often useful indefinitely at low storage cost. Logs containing prompt or completion content should have an explicit retention period tied to why you're keeping them, since indefinite retention of content logs is the main way a logging setup turns into a data liability.

Related pages

Run your own logging stack on dedicated hardware.

Full root access on a Spark means logs never leave your control unless you decide to ship them.

Read the LiteLLM guide Read the audit logging guide