Docs analysis
DGX Spark/Rate limiting and quotas for an internal LLM deployment

Rate limiting and quotas for an internal LLM deployment

By Samuel Seidel · Published September 9, 2026

A raw DGX Spark has no concept of a user, so it has no concept of a per-user limit either. If one person or one script can saturate the GPU and make everyone else's requests slow, that's an application-layer problem you solve yourself, and there's a straightforward way to do it if your serving stack includes a gateway.

Why the platform doesn't do this for you

GPUwerk's billing meters the instance itself, per-minute, at $0.79/hour on-demand, not requests or tokens flowing through whatever you're running inside the container. There's no per-user request throttling at the platform level, since the platform has no visibility into what's happening inside your root-owned container, only whether it's running. Anything resembling a quota, a request limit, or a fair-use cap on a shared deployment has to live in your own serving layer.

The straightforward path: LiteLLM

If your serving stack already runs LiteLLM as a gateway in front of your model, rate limiting is close to free, since it's built into how you issue keys. Two settings do most of the work, set when a virtual key is created:

Issuing a scoped key with limits looks like a single API call to LiteLLM's key endpoint, specifying the models that key can reach and its rpm_limit. Spend and usage per key, team or model is queryable from LiteLLM's /spend/report endpoint, or visible in its admin UI, so you can see who's actually hitting a limit before deciding whether to raise or lower it.

On a single shared machine, the meaningful constraint usually isn't cost, it's contention: one person's heavy job degrading everyone else's latency. Rate limits per key are the practical lever for that, more than a hard dollar budget.

If you're not running LiteLLM

Without an LLM-aware gateway, you're limited to what a general reverse proxy understands, which is connections and requests, not tokens. Nginx's limit_req module or Caddy's rate-limit plugin can cap requests per second per client IP or API key in front of whatever inference server you're running, which stops a client from hammering the endpoint even without visibility into token counts. This is coarser than LiteLLM's per-key token accounting, since a single large request still costs far more GPU time than a small one, and a request-count limit doesn't see that difference. It's a reasonable fallback if adding a full gateway isn't worth it for your setup, but if you expect real multi-user load, putting LiteLLM in front is worth the extra piece.

Some inference engines expose their own concurrency or queue-depth settings (a max batch size, a request queue limit) that act as a coarse throttle even without a gateway in front. Check your specific engine's documentation, since this varies and isn't something to assume is there.

Practical checklist

See the LiteLLM setup guide for the full key-issuance walkthrough, or a first engagement if you want help sizing limits for a specific team's usage.

First top-up: pay $10, get $20 in credit

One machine, keys that keep it fair.

Deploy a dedicated Spark and put LiteLLM in front to give every teammate their own limit.

Deploy a Spark Read the LiteLLM guide