Rate limiting and quotas for an internal LLM deployment
A raw DGX Spark has no concept of a user, so it has no concept of a per-user limit either. If one person or one script can saturate the GPU and make everyone else's requests slow, that's an application-layer problem you solve yourself, and there's a straightforward way to do it if your serving stack includes a gateway.
Why the platform doesn't do this for you
GPUwerk's billing meters the instance itself, per-minute, at $0.79/hour on-demand, not requests or tokens flowing through whatever you're running inside the container. There's no per-user request throttling at the platform level, since the platform has no visibility into what's happening inside your root-owned container, only whether it's running. Anything resembling a quota, a request limit, or a fair-use cap on a shared deployment has to live in your own serving layer.
The straightforward path: LiteLLM
If your serving stack already runs LiteLLM as a gateway in front of your model, rate limiting is close to free, since it's built into how you issue keys. Two settings do most of the work, set when a virtual key is created:
rpm_limitandtpm_limitcap requests and tokens per minute for that specific key. A lowrpm_limiton an automated or batch key keeps a runaway loop from starving interactive users, while a human's key can stay generous.max_budgetwithbudget_durationis a spend cap over a time window, useful mainly for keys that can reach a paid external provider, less relevant if everything routes to your own Spark where there's no per-token bill.
Issuing a scoped key with limits looks like a single API call to LiteLLM's key endpoint, specifying the models that key can reach and its rpm_limit. Spend and usage per key, team or model is queryable from LiteLLM's /spend/report endpoint, or visible in its admin UI, so you can see who's actually hitting a limit before deciding whether to raise or lower it.
On a single shared machine, the meaningful constraint usually isn't cost, it's contention: one person's heavy job degrading everyone else's latency. Rate limits per key are the practical lever for that, more than a hard dollar budget.
If you're not running LiteLLM
Without an LLM-aware gateway, you're limited to what a general reverse proxy understands, which is connections and requests, not tokens. Nginx's limit_req module or Caddy's rate-limit plugin can cap requests per second per client IP or API key in front of whatever inference server you're running, which stops a client from hammering the endpoint even without visibility into token counts. This is coarser than LiteLLM's per-key token accounting, since a single large request still costs far more GPU time than a small one, and a request-count limit doesn't see that difference. It's a reasonable fallback if adding a full gateway isn't worth it for your setup, but if you expect real multi-user load, putting LiteLLM in front is worth the extra piece.
Some inference engines expose their own concurrency or queue-depth settings (a max batch size, a request queue limit) that act as a coarse throttle even without a gateway in front. Check your specific engine's documentation, since this varies and isn't something to assume is there.
Practical checklist
- Don't expect the platform to rate-limit anything inside your container; it only meters instance uptime.
- If you're running LiteLLM, set
rpm_limitandtpm_limitper key at issuance, generous for humans, tight for automated callers. - Use LiteLLM's
/spend/reportor admin UI to see actual usage before tightening limits further. - Without LiteLLM, fall back to a reverse proxy's request-rate limiting, understanding it can't see token counts the way a gateway can.
- Treat contention, not cost, as the real reason to rate-limit on a single dedicated machine.
See the LiteLLM setup guide for the full key-issuance walkthrough, or a first engagement if you want help sizing limits for a specific team's usage.