What does it actually cost to self-host an LLM?
Most self-hosting cost guides stop at the hardware price. That's the easy part to look up and the least representative of what running a model actually costs over a year. This one covers the hardware or rental line, the ops line most people leave out, and, since we sell the hardware, an honest section on when the answer is: don't.
The hardware or rental line
You have two ways to get compute: buy it once, or rent it by the hour. Both put a specific model in your control; they differ in where the cost lands and how it moves with your usage.
Renting. GPUwerk's own published rate for a dedicated DGX Spark is $0.79/hour on-demand, billed per minute, with no separate egress fee. A stopped instance that holds its reservation runs at 75% of that rate. Renting turns hardware into an operating expense that scales with hours used rather than a lump sum up front, and it's the way to find out whether a model fits and performs before committing to anything larger.
Buying. The EU street price range we've published for a DGX Spark-class machine runs roughly €3,800 to €4,800, spanning the ASUS Ascent GX10 board at the low end to the NVIDIA Founders Edition at the top; see the full breakdown in our before-you-buy guide. Whether buying beats renting depends entirely on how many hours a year the machine runs; our rent-vs-buy break-even page works that arithmetic in detail, with a calculator for your own hours.
The ops line people forget
A model server doesn't run itself. Someone has to install it, keep it patched, watch for it falling over, and be the person who gets paged when it does, at whatever hour that happens to be. None of that shows up on a hardware invoice or a rental bill, and it's easy to leave entirely out of a cost comparison because it doesn't arrive as a line item, it arrives as somebody's time.
Concretely, that ongoing work includes: keeping the model server and its dependencies current as the software stack moves (an area our own buying guide flags as still rough on newer hardware combinations), monitoring for degraded performance or outright failures, managing who has access and how, deciding what happens when the machine is at capacity, and owning backups, since a self-hosted setup typically has no managed backup layer behind it unless you build one. If nobody on the team already does this kind of infrastructure work, the real cost of self-hosting includes either training someone to do it or paying for it externally.
That's the gap a scoped engagement is meant to close rather than leave to guesswork. See first engagement for what a fixed-scope rollout, including the operational handoff, actually covers.
When self-hosting is not worth it
We sell the hardware and the rental, and it's still not the right answer for every workload. Skip self-hosting, or at least start with a metered API, when:
- Your volume is low and spiky. A flat-rate machine costs the same whether it's generating tokens or sitting idle. If your real traffic is a handful of requests scattered through the day rather than a steady stream, you end up paying for a lot of idle time that a per-token API simply wouldn't bill you for. See our crossover-point comparison for the exact arithmetic on where that line sits for your own numbers.
- Nobody on the team can own the ops work. If keeping a model server running competes with every other priority the team has, that cost shows up as downtime and slow fixes rather than a bill, which is often worse.
- You need a specific closed frontier model by name. Self-hosting only runs open-weight models. If your evaluation depends on a particular closed model, no self-hosted setup gets you there; see our Azure OpenAI comparison for that tradeoff in detail.
- You haven't tested the model on the hardware yet. Committing to a purchase before confirming a model fits and performs at an acceptable speed is the single most common way people overspend on this. Rent an hour first.
Self-hosting tends to make sense on the other side of each of those: steady volume, someone able to own operations, an open-weight model that already meets the quality bar, and a workload you've actually tested. If that's where you are, start with a rental to confirm the fit rather than buying outright.