Security
DGX Spark/API key rotation best practices for a private LLM deployment

API key rotation best practices for a private LLM deployment

By Samuel Seidel · Published September 9, 2026

A private inference endpoint usually accumulates keys faster than anyone plans for: one for the console, one per teammate, one per script, one for whatever internal tool started calling the model last month. Nobody rotates a key that isn't tracked, so the actual work here is less about the rotation mechanics and more about knowing what keys exist and what each one can reach.

Scope before you rotate

A key with narrow, specific scope is one you can rotate without a scramble; a key with broad, undocumented access is one nobody wants to touch because nobody's sure what breaks. GPUwerk's own console keys carry an explicit scope, for example an inference-only key that can call the model endpoint and nothing else in the account, per the vLLM verification steps. Apply the same discipline to keys you issue yourself:

LiteLLM's key management, if that's your gateway

If you're running LiteLLM in front of your model, per-key controls are built in rather than something you bolt on. Keys are issued through its key management API or admin UI, each with its own rpm_limit and tpm_limit, per the rate limiting guide, and the same issuance flow is where you set an expiration or a scoped model list. Rotation with LiteLLM is: issue a new key with the same scope, update the consumer to use it, confirm traffic has moved, then revoke the old one. Revocation takes effect immediately, so a key you're unsure is still in use is safer to leave alive one more cycle than to guess wrong and break something silently.

A rotation schedule that matches actual risk

Calendar-based rotation for its own sake is a compliance checkbox more than a security measure; what actually matters is rotating on a cadence that matches how exposed each key is.

Whatever cadence you pick, write it down. A rotation policy that lives only in someone's memory doesn't survive that person being on vacation when a key needs to turn over.

Rotating without an outage

The failure mode to avoid is revoking the old key before every consumer has switched to the new one. A safe rotation is sequential, not simultaneous:

  1. Issue the new key with identical scope to the one it's replacing.
  2. Update the consumer's configuration or environment variable to the new key, and confirm it's actually being used, not just deployed, by checking the gateway's usage or spend logs for that key.
  3. Leave the old key active for a short overlap window, long enough to confirm nothing is still calling with it.
  4. Revoke the old key, and confirm the endpoint still answers using only the new one.

For a key that's shared across several independent consumers, scripts, CI, and a person, this is exactly the argument for not sharing it in the first place: a single shared key turns step 2 into a coordination problem across everyone at once, instead of one script at a time.

Practical checklist

See the security hardening guide for the instance-level side of this, or the team access guide for setting up per-teammate keys in the first place. A first engagement can help design a key policy for a specific team's tooling.

First top-up: pay $10, get $20 in credit

Scoped keys from the start.

Deploy a dedicated Spark and issue narrow, trackable API keys for every teammate and script.

Deploy a Spark Read the security guide