Audit logging for a compliance-sensitive self-hosted LLM deployment
If a compliance program requires a record of who asked a model what and what it answered, that record has to come from inside your own container. GPUwerk gives you a dedicated Spark and root SSH into it, not a managed platform with built-in audit trails, so the logging itself is work you set up and own. This guide covers the practical side: what to capture, where to put it, and how to keep it from becoming a liability of its own.
Where the responsibility line sits
Worth stating plainly before anything else: GPUwerk provisions the machine and the network path to it. What runs inside your container, including any audit logging, is yours to configure and yours to secure. We do not inspect, log, or retain the contents of requests sent to a model running in your instance, and we have no visibility into whether audit logging is set up at all. If an auditor asks for evidence of request logging, that evidence has to be produced from your own instance, not from us. This isn't a gap we plan to fill; a self-hosted deployment is attractive precisely because the operator controls this layer end to end, and that control comes with the responsibility attached to it.
What to actually capture
An audit log that tries to capture everything usually ends up too large to query and too sensitive to keep. Start narrower and add fields only when a specific requirement calls for them.
- Who: the authenticated identity or API key that made the call, not just an IP address. This depends on having per-user or per-key access set up in the first place, covered in the team and multi-user access guide.
- When: a timestamp with timezone, at request time and again at response time, so latency and any timeout behavior are reconstructable later.
- What: the model name and version invoked, the endpoint called, and either the full prompt and response or, if the content itself is too sensitive to store, a hash of both plus enough metadata (token counts, whether the call succeeded) to reconstruct the shape of what happened.
- Outcome: HTTP status, error type if any, and whether a safety filter or content policy fired, if one is running in front of the model.
Whether prompt and response text itself belongs in the log is a decision to make deliberately, not by default. Full-content logging is far more useful for an actual audit, but it also means the log becomes the largest concentration of sensitive data on the box. If regulatory or contractual language requires content retention, store it; if it doesn't, hashing plus metadata is usually the safer default.
Getting the logs out of the request path
If you're running a gateway like LiteLLM in front of the model, per the LiteLLM setup doc, its request logging can be pointed at a structured sink rather than a local file: a syslog target, an S3-compatible bucket, or a database it writes to over the network. The general pattern for whatever gateway or reverse proxy is in front of your model:
- Log at the gateway or proxy layer, not inside application code for every service that happens to call the model. One logging point is easier to audit than several.
- Write logs off the instance as they're generated, or on a short interval, rather than accumulating them only on local disk. A workspace on a stopped or terminated Spark is deleted with it, and terminate deletes it permanently, so logs held only in
/workspacedisappear with the instance. - Set retention centrally, at the sink, so a compliance-mandated retention period (say, one year of logs) is enforced by policy rather than by remembering to run a cleanup script.
Access control on the log itself
An audit log that anyone with SSH access can read or edit undermines its own purpose; the whole point is a record that's trustworthy after the fact. Two practical steps matter more than the storage technology chosen:
- Write access restricted to the logging process, ideally running as a non-root user with a narrow file permission set, so a compromised application account can't silently edit its own history.
- Append-only storage where the sink supports it. Object storage with versioning enabled, or a database with an insert-only role for the logging writer, both make after-the-fact tampering detectable even if the writer's credentials are compromised.
This overlaps directly with instance-level hardening in general, covered in the security hardening guide: running services under a non-root user and separating credentials by purpose reduces the blast radius of any one compromised process, audit logging included.
What compliance frameworks actually tend to ask for
The specific requirement varies by framework and by which auditor is reading it, and GPUwerk's legal documents haven't been reviewed by outside counsel, so treat this as a starting checklist rather than legal guidance: confirm the actual requirement with whoever owns the compliance program before building to it.
That said, a few things come up often enough to plan for from the start: a documented retention period rather than an indefinite one, a way to produce logs for a specific user or time window on request rather than only a full export, and a written description of what the logging captures and what it deliberately omits. Building the description alongside the logging, not after an auditor asks for it, saves a scramble later.
For the network and instance side of a compliance-sensitive deployment, see the security hardening guide. For getting per-user identity into the logs in the first place, see the team and multi-user access guide. A first engagement can help scope logging requirements against a specific framework before deployment.