Setting up API usage metering and per-tenant billing
One model, one API key, one team: nobody needs to know who used what. That stops being true the moment a second team starts calling the same endpoint, or you start running inference for more than one customer off the same Spark. At that point "who used how much" is a question you need an actual answer to, not a guess from eyeballing logs.
GPUwerk meters the Spark, not your callers
The GPUwerk bill is for the hardware, an hourly rate on a dedicated Spark, and it has no visibility into who's calling the inference endpoint running on it. If ten different clients share one API key against one vLLM instance, GPUwerk's billing can't tell them apart, and neither can vLLM by default, it just serves whatever request shape it's given. Per-tenant metering is something you add in front of the endpoint, not a dial in the console.
A gateway with per-key spend tracking is the fastest path
Rather than building a metering pipeline from scratch, a LiteLLM proxy in front of the Spark, covered in full in setting up LiteLLM as your AI gateway, already tracks spend and usage per virtual key once you're issuing one key per tenant. Mint a key per customer or team with metadata tagging who it belongs to:
curl http://localhost:4000/key/generate \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"models": ["spark-qwen"],
"metadata": {"tenant": "acme-corp", "plan": "pro"},
"max_budget": 200,
"budget_duration": "30d"
}'
Every request against that key gets attributed automatically, and the spend and usage per key, team, or model is queryable without writing your own aggregation:
curl "http://localhost:4000/spend/report?start_date=2026-09-01&end_date=2026-09-30" \ -H "Authorization: Bearer $LITELLM_MASTER_KEY"
This is the same mechanism the LiteLLM doc uses for team budgets and rate limits, applied here to tenants instead of internal teams. If you're already running LiteLLM for the per-user keys and fallback routing it covers, per-tenant metering is largely a matter of how you assign keys, not new infrastructure.
Turning token counts into a bill
Every OpenAI-compatible response, GPUwerk's endpoint included, returns a usage object with prompt and completion token counts. That's the raw unit most usage-based billing is built on:
{
"id": "chatcmpl-8f2a1c",
"choices": [...],
"usage": {
"prompt_tokens": 412,
"completion_tokens": 128,
"total_tokens": 540
}
}
If you're not routing through a gateway that aggregates this for you, logging the usage object alongside the tenant identifier on every request, into a table rather than a flat log file, gives you the raw data to sum into a monthly bill:
INSERT INTO usage_events (tenant_id, model, prompt_tokens, completion_tokens, created_at)
VALUES ('acme-corp', 'llama-3.1-8b', 412, 128, now());
-- monthly rollup for billing
SELECT tenant_id,
sum(prompt_tokens + completion_tokens) AS total_tokens,
sum(prompt_tokens + completion_tokens) * 0.0000015 AS estimated_cost
FROM usage_events
WHERE created_at >= date_trunc('month', now())
GROUP BY tenant_id;
Pick a per-token rate that reflects your actual cost, hourly Spark cost divided by a realistic tokens-per-hour throughput figure for the model in question, rather than copying a rate from a hosted API provider whose cost structure is nothing like a dedicated GPU you're paying for by the hour regardless of load.
Cap spend before it becomes a surprise, not after
A monthly report tells you what happened. A budget cap on the key itself stops it from happening in the first place. With max_budget and budget_duration set on a virtual key, the gateway rejects further requests once the tenant crosses their cap, rather than you finding out at the end of the month that one tenant's usage was ten times everyone else's. Pair this with a low rpm_limit on any automated or batch-style tenant, since a runaway loop hitting the endpoint in a tight retry cycle is the most common way a budget gets blown through quickly.
Separate metering from rate limiting
It's tempting to conflate the two since they're both per-key controls in the same config block, but they solve different problems. Metering answers "how much did this tenant use," for billing or capacity planning. Rate limiting, covered alongside queuing in configuring request queuing and backpressure, answers "how fast can this tenant hit the endpoint right now," to protect the Spark from being overwhelmed by one caller. A tenant can be well within budget and still need to be rate limited if they're sending requests faster than the node can serve everyone fairly.
FAQ
Does GPUwerk provide per-tenant billing or usage metering?
No. GPUwerk bills for the Spark itself, by the hour. If you're running that Spark for multiple internal teams or external customers and need to know who used how much, that's metering you set up yourself, typically with a gateway like LiteLLM sitting in front of the inference endpoint, not something the GPUwerk console tracks per end user.
Is metering tokens enough, or do I need to time GPU usage too?
Token counts are the standard unit and the one most billing models are built around, since they map directly to what a caller consumed regardless of how long the request took. Wall-clock GPU time matters more when tenants have very different latency profiles, a tenant sending long prompts that take much longer per token isn't fairly charged the same per-token rate as one sending short ones, so a metering setup that also logs latency lets you catch that case even if the bill is still token-based.
How do I stop one tenant from running up an unexpectedly large bill?
Set a hard budget cap per tenant key rather than relying on catching it after the fact in a usage report. A gateway that supports per-key max_budget with a budget_duration rejects requests once a key crosses its cap, which turns an unexpected overage into a blocked request the tenant notices immediately instead of a bill you have to explain at the end of the month.