On-premise LLM cost vs cloud API cost: how to actually compare them
These two pricing models don't compare directly, and that's the whole problem with most vendor comparisons. A cloud API charges per token: no traffic, no bill. Dedicated infrastructure charges per hour, whether the GPU is generating tokens or sitting idle. The question isn't which is cheaper in the abstract, it's where your own volume and utilization put you relative to the crossover point. This page gives you the method, not a number that happens to flatter one side.
Two pricing models, one comparison that actually works
A per-token API bills you for exactly what you generate. Idle time costs nothing, because there's no dedicated hardware sitting there waiting for your next request; you're sharing a provider's fleet with everyone else calling the same endpoint. A dedicated machine works the opposite way: you pay a flat rate for the hours it's reserved to you, and that rate doesn't change whether you send it one request an hour or run it flat out. The cost per token you actually generate on dedicated hardware falls as utilization rises, and rises toward infinity as utilization falls toward zero.
That's the entire shape of the tradeoff. Low, spiky volume favors metered billing. Steady, high volume favors flat-rate hardware. The crossover between them is a number you calculate from your own traffic, not one either pricing model can hand you directly.
The formula
Four steps, the same ones used in GPUwerk's own comparison against Azure OpenAI:
- Find your dedicated hardware's sustained throughput. Not the single-stream number, the aggregate tokens/second your model reaches at the concurrency your real traffic runs at. Batch-1 and high-concurrency throughput on the same machine can differ by more than an order of magnitude, so use a number measured at a concurrency close to your own.
- Convert to tokens per hour, then to cost per token. Tokens/second × 3,600 gives tokens/hour. Divide the hourly rate by that figure and you have a cost-per-token number for the dedicated machine, assuming it stays busy at that concurrency the entire hour.
- Compare that figure to the API's per-token price. Most API pricing is split between input and output tokens at different rates, so match the ratio of input to output tokens in your own workload rather than comparing a single blended number.
- Multiply both by your actual monthly volume, then check the utilization assumption. The dedicated-hardware number only holds if the machine is actually busy at that concurrency for the hours you're counting. If your real traffic arrives in bursts with long gaps, the effective cost per token on dedicated hardware is higher than step 2 suggested, because you're still paying for the idle hours.
Worked example, using GPUwerk's own numbers
Take a workload generating 500 million output tokens a month, in line with the worked example on our Azure OpenAI comparison page. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests on a single DGX Spark, per Dendro Logic's concurrency benchmark cited on our benchmarks page. That's 862.8 × 3,600 = 3,106,080 tokens/hour if the node stays saturated at that concurrency continuously. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At GPUwerk's published $0.79/hour on-demand rate, 161 × $0.79 = $127.19 before tax.
That figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back, with no idle time and enough real traffic to keep the queue full. Real traffic arrives in bursts, not a constant stream at exactly your saturation concurrency, so a machine billed by the hour that sits half-idle waiting for requests can lose ground fast, you're paying for every hour on the clock whether the GPU is generating tokens or not. A per-token API doesn't care about idle time because there is no idle time to bill; you pay only for tokens actually generated. Put your chosen API's per-token price for 500 million tokens next to $127.19 and you have the real comparison for that specific volume; the crossover point for your own traffic is a utilization question, not a fixed answer either pricing model can give you on its own.
What tips the answer one way or the other
- Volume and steadiness push toward dedicated hardware. A continuous internal assistant, a batch document pipeline, or an always-on agent fleet keeps utilization high, which is exactly the condition the worked example above assumes.
- Low or spiky volume pushes toward a metered API. If your traffic is a handful of requests an hour with long silent gaps, you're paying for a mostly idle machine on the dedicated side, and the per-token API's "no traffic, no bill" property wins by default.
- Model choice changes the throughput number, not the method. A mixture-of-experts model like gpt-oss-120b reaches much higher aggregate throughput at high concurrency than a dense model of similar size, because a MoE model activates only a fraction of its parameters per token. Rerun step 1 with your own model's measured throughput rather than reusing the number above.
- Data location and who can read your prompts are separate questions from cost. They often decide the choice before cost does; see private LLM hosting for that side of the comparison.
Run it on your own numbers
The fastest way to get a real answer is to measure your own model's throughput rather than borrow someone else's benchmark. Rent a DGX Spark by the hour, run your actual workload at your actual concurrency, and plug the measured tokens/second into step 1 above. That replaces the Dendro Logic figure in this page's worked example with a number that describes your traffic, not ours.