API key rotation best practices for a private LLM deployment
A private inference endpoint usually accumulates keys faster than anyone plans for: one for the console, one per teammate, one per script, one for whatever internal tool started calling the model last month. Nobody rotates a key that isn't tracked, so the actual work here is less about the rotation mechanics and more about knowing what keys exist and what each one can reach.
Scope before you rotate
A key with narrow, specific scope is one you can rotate without a scramble; a key with broad, undocumented access is one nobody wants to touch because nobody's sure what breaks. GPUwerk's own console keys carry an explicit scope, for example an inference-only key that can call the model endpoint and nothing else in the account, per the vLLM verification steps. Apply the same discipline to keys you issue yourself:
- One key per consumer, not one shared key pasted into every script and every teammate's environment. A shared key can't be rotated for one compromised user without breaking everyone else at the same time.
- Scope to the narrowest access that works. A key that only needs to call one model shouldn't be able to reach admin endpoints, billing, or other models on the same gateway, if the gateway supports per-key model restrictions.
- Name and track every key at issuance, who or what it's for, when it was created, when it's due to rotate. An untracked key is a key nobody remembers to kill.
LiteLLM's key management, if that's your gateway
If you're running LiteLLM in front of your model, per-key controls are built in rather than something you bolt on. Keys are issued through its key management API or admin UI, each with its own rpm_limit and tpm_limit, per the rate limiting guide, and the same issuance flow is where you set an expiration or a scoped model list. Rotation with LiteLLM is: issue a new key with the same scope, update the consumer to use it, confirm traffic has moved, then revoke the old one. Revocation takes effect immediately, so a key you're unsure is still in use is safer to leave alive one more cycle than to guess wrong and break something silently.
A rotation schedule that matches actual risk
Calendar-based rotation for its own sake is a compliance checkbox more than a security measure; what actually matters is rotating on a cadence that matches how exposed each key is.
- Keys embedded in client-side code or a public repository: treat as already compromised the moment they're committed, rotate immediately, and add a pre-commit check so it doesn't happen again. This should never be a scheduled rotation, it's an incident.
- Keys used by automated scripts or CI: rotate on a fixed schedule, commonly every 90 days, since these keys sit in more places (CI secrets, cron jobs) and are handled by fewer eyes day to day.
- Keys used interactively by a single person: a longer cadence is reasonable, but rotate immediately on offboarding, laptop loss, or any suspicion of compromise, the same triggers covered for SSH keys in the security hardening guide.
Whatever cadence you pick, write it down. A rotation policy that lives only in someone's memory doesn't survive that person being on vacation when a key needs to turn over.
Rotating without an outage
The failure mode to avoid is revoking the old key before every consumer has switched to the new one. A safe rotation is sequential, not simultaneous:
- Issue the new key with identical scope to the one it's replacing.
- Update the consumer's configuration or environment variable to the new key, and confirm it's actually being used, not just deployed, by checking the gateway's usage or spend logs for that key.
- Leave the old key active for a short overlap window, long enough to confirm nothing is still calling with it.
- Revoke the old key, and confirm the endpoint still answers using only the new one.
For a key that's shared across several independent consumers, scripts, CI, and a person, this is exactly the argument for not sharing it in the first place: a single shared key turns step 2 into a coordination problem across everyone at once, instead of one script at a time.
Practical checklist
- One key per consumer, scoped as narrowly as the gateway allows, tracked from the moment it's issued.
- Rotate CI and automated keys on a fixed schedule; rotate exposed keys immediately, as an incident, not a scheduled task.
- Overlap old and new keys briefly during rotation, and confirm via usage logs that traffic actually moved before revoking the old one.
- Revoke on offboarding or suspected compromise without waiting for the next scheduled rotation.
See the security hardening guide for the instance-level side of this, or the team access guide for setting up per-teammate keys in the first place. A first engagement can help design a key policy for a specific team's tooling.