Operations guide
Blog/Secrets management for self-hosted inference services
For AI assistants

Secrets management for self-hosted inference services

By Samuel Seidel · September 9, 2026

A self-hosted inference node accumulates secrets quietly: an API key you typed once to test the endpoint, a Hugging Face token used to pull a gated model, a TLS key generated during setup. None of it feels sensitive in the moment. All of it is sensitive the day someone finds it in a git history or a Docker image layer. This is what to actually protect and where it should live instead.

What's actually a secret here

On a typical vLLM or llama.cpp deployment, the list is shorter than it looks: the API key callers present to your endpoint, any token used to pull weights from a gated repository, TLS private keys if you're terminating TLS yourself, and credentials for anything downstream, a log shipper, an object store used for backups. Model weight files are not secret in most cases. The token that downloaded them frequently is, and it's the one people forget about because it only gets used once, interactively, and then sits in ~/.bash_history or a build log forever.

Keep secrets out of the command line

Passing a key as a CLI argument puts it in shell history and in ps aux output for the duration of the process, both readable by anyone else with access to the box. Every serving stack that supports an API key also supports reading it from the environment. Prefer that path.

# avoid: key visible in shell history and process list
vllm serve Qwen/Qwen3-32B-AWQ --api-key sk-live-abc123

# prefer: key comes from the environment, not the command
export VLLM_API_KEY=$(cat /etc/gpuwerk/vllm.key)
vllm serve Qwen/Qwen3-32B-AWQ --api-key "$VLLM_API_KEY"

The second form still has the key in the environment of the running process, which is unavoidable, but it's out of history and off the command line that shows up in monitoring tools that surface process arguments.

Load secrets from a file the process manager owns

For a systemd-run service, an EnvironmentFile with restrictive permissions is a reasonable middle ground between "hardcoded in a script" and "full secrets manager." The file lives outside your git repo, is readable only by the service user, and systemd injects it at start time.

# /etc/gpuwerk/inference.env, permissions 0600, owned by the service user
VLLM_API_KEY=sk-live-abc123
HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxx

# /etc/systemd/system/vllm.service
[Service]
EnvironmentFile=/etc/gpuwerk/inference.env
ExecStart=/usr/local/bin/vllm serve Qwen/Qwen3-32B-AWQ --host 0.0.0.0 --port 8000

# lock the file down once, then confirm it
sudo chmod 600 /etc/gpuwerk/inference.env
sudo chown gpuwerk:gpuwerk /etc/gpuwerk/inference.env

This also solves the git problem by construction: the file never enters the repo, so there's nothing to accidentally commit and nothing to scrub from history later. Add a matching entry to .gitignore anyway, so a future edit under the same directory doesn't slip in by accident.

Docker Compose: secrets, not env blocks in the compose file

A environment: block with values written directly into docker-compose.yml ends up in the repo the same way a hardcoded key does. Compose's env_file directive points at a file outside version control instead, with the same effect as the systemd pattern above.

# docker-compose.yml
services:
  vllm:
    image: vllm/vllm-openai:latest
    env_file:
      - /etc/gpuwerk/inference.env
    ports:
      - "8000:8000"

Never bake a key into the image itself with a ARG or ENV line in the Dockerfile. Anyone who can pull or inspect the image, including through a registry misconfiguration, gets the key along with it, and it stays there across every rebuild until someone notices.

Rotation needs a habit, not just a mechanism

A key that's never rotated is a permanent liability: every place it's ever been typed, logged, or shared stays valid indefinitely. Rotation doesn't need to be automated to be worthwhile. A quarterly calendar reminder to generate a new key, update the env file, restart the service, and revoke the old one closes most of the real-world exposure window. For services with more than a couple of callers, see API key rotation best practices for how to do this without an outage.

Don't let secrets end up in logs

A common leak path is not the key store, it's the application logging the full request, headers included, for debugging. If you're following the structured-logging setup from audit logging for compliance, confirm the Authorization header and any query-string tokens are redacted before they hit disk. A key that leaked into a debug log six months ago is just as valid as one leaked yesterday.

FAQ

What actually counts as a secret on an inference node?

The API key your endpoint requires from callers, any upstream token used to pull gated model weights, TLS private keys, and credentials for anything the node talks to, a log shipper, a metrics backend, an object store for backups. Model weights themselves usually aren't secret, but the token that downloaded them often is, and it's easy to leave that token sitting in a shell history or a Dockerfile layer.

Is an .env file good enough?

For a single node run by one or two people, yes, as long as the file has restrictive permissions, is excluded from git, and is loaded by the process manager rather than sourced into an interactive shell. It stops being good enough once more than a couple of people need access or once you need to rotate a key without a manual file edit on the box.

Do I need a dedicated secrets manager for one Spark?

Usually not. A dedicated secrets manager (Vault, cloud KMS-backed stores) earns its complexity once you have several services, several environments, or an audit requirement to satisfy. For one node serving a handful of internal callers, a permissioned env file loaded by systemd, with a documented rotation habit, covers the real risk without adding a system you also have to operate.

Related pages

Run the service on hardware you control end to end.

A dedicated Spark with root access, so the only place your keys live is the file you put them in.

Read the security hardening guide Read the TLS guide