Setting up SSO and access control for a team LLM deployment
SSH keys and API tokens work fine for a small team, but they don't scale to "anyone in the company with the right badge can use the internal model." Past a certain team size, what's actually wanted is single sign-on: a teammate authenticates with the identity provider the company already uses, and access to the model endpoint follows from that, no separate credential to issue, track, or revoke. That layer doesn't ship with a rented Spark. It's something you put in front of your own endpoint, and this guide covers the general shape of doing that.
Where this fits relative to SSH and API keys
These aren't competing approaches so much as different layers for different consumers. Per the team and multi-user access guide, SSH access with per-user keys is right for people who need a shell on the box itself, and scoped API keys are right for scripts and services calling the model directly. SSO in front of a web-facing endpoint, a chat UI, an internal tool, is for the broader group of people who just need to use the model and shouldn't need SSH access or a personal API key to do it. A team commonly ends up with all three: SSH for the few people administering the instance, API keys for automated consumers, and SSO for everyone else.
The general shape of the setup
The common pattern is an authenticating reverse proxy sitting between the public internet and the model endpoint, sometimes called an OAuth2 proxy pattern: a lightweight proxy that redirects an unauthenticated request to the identity provider's login page, verifies the resulting token, and only then forwards the request to whatever's actually serving the model, LiteLLM, vLLM, or a chat interface layered on top of either. The proxy is the thing doing the auth check; the model server behind it doesn't need to know anything about SSO at all.
Several open-source options exist for the proxy layer itself, and which one fits best depends on the identity provider already in use, Google Workspace, Microsoft Entra ID, Okta, or a self-hosted alternative, so this guide stays general rather than naming one as a GPUwerk recommendation. Whichever is chosen, the setup follows roughly the same steps:
- Register an OAuth application with the identity provider, which gives the proxy a client ID and secret to authenticate itself, distinct from any individual user's credentials.
- Point the proxy at the model endpoint as its upstream, so authenticated traffic is forwarded there and nothing else is exposed directly.
- Put TLS in front of the whole thing. An SSO flow that redirects over plain HTTP exposes tokens in transit, so this needs a certificate, not the plain-HTTP setup that might be fine for a quick internal test.
- Restrict the login to your organization's domain or group, not just "anyone with an account with this identity provider." Most proxies support restricting by email domain or by group membership pulled from the provider.
Access control beyond just logging in
Authentication answers who someone is; it doesn't answer what they should be allowed to do once they're in. For a single shared model endpoint, that's less of a concern; everyone authenticated gets the same access. Once there's more than one model, or an admin function like changing serving config, alongside the model API itself, authorization needs its own layer: group membership pulled from the identity provider mapped to specific routes or models, so a general user reaches the chat interface and only whoever's in an admin group reaches configuration endpoints.
Rate limiting per authenticated user is worth setting up alongside this rather than after the fact, particularly if the deployment is shared across a team of any size. If the endpoint sits behind LiteLLM, its per-key rate limits, covered in the rate limiting guide, can be issued per SSO-authenticated identity so one person's heavy usage doesn't starve everyone else on the same Spark.
What SSO doesn't replace
SSO controls who reaches the endpoint through the front door. It doesn't touch SSH access to the underlying machine, which needs its own key management regardless, per the team access guide, and it doesn't replace API keys for machine-to-machine calls that aren't a person clicking through a login flow, covered in the API key rotation guide. Offboarding still needs to cover all three: removing someone from the identity provider's group revokes their SSO access, but any SSH key or API key they were separately issued has to be revoked on its own.
For SSH-based team access as an alternative or complement to SSO, see the team and multi-user access guide. For managing the API keys that sit alongside an SSO layer, see the API key rotation best practices guide. A first engagement can help design an access control setup around a specific identity provider.