Private LLM hosting vs Hetzner GPU servers
We rent DGX Sparks, so weigh that against everything below. Hetzner is a German cloud and dedicated-server provider with a long reputation, deserved or not, for value pricing in the EU market; they've added GPU instances alongside their existing server lineup. Hetzner's GPU servers are general-purpose, discrete-GPU machines you configure and manage like any of their other dedicated or cloud servers. A DGX Spark is a single machine built around 128GB of unified memory shared between CPU and GPU, aimed specifically at running LLM inference. The comparison is architecture and focus more than price alone.
Side by side
| Hetzner GPU servers | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| What it is | General-purpose GPU cloud and dedicated servers, one product line inside Hetzner's much larger server and cloud catalogue. | A single-purpose GPU host: one dedicated DGX Spark, rented by the hour, built around unified memory for inference. |
| Region | Germany and Finland primarily, per Hetzner's own current data center list; check their site for availability by product. | EU-Central (Prague), exclusively. One location, no region selection to get wrong. |
| Memory architecture | Conventional discrete GPU cards with separate VRAM and system RAM; exact configuration varies by server type, check Hetzner's current specs. | 128GB unified memory shared between CPU and GPU, NVIDIA's DGX Spark architecture, suited to large quantized models without splitting memory across pools. |
| Who can see your data | Governed by Hetzner's own terms and data processing addendum; check their legal pages for the current text. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." |
| Model choice | Whatever you deploy, unrestricted by the provider. Discrete-GPU servers with more raw VRAM can suit models tuned for that memory layout. | Whatever you deploy too, bounded by 128GB unified memory: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and similar fit; the two-node cluster extends this for larger workloads. |
| Speed and concurrency | Depends on the GPU card and server tier rented; Hetzner publishes server specs but tok/s for a given model is yours to benchmark. | A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page. |
| Pricing model | Per hour for cloud GPU instances, or monthly for dedicated GPU servers, varying by card and tier, published on Hetzner's own pricing page. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of workload, per pricing. |
| Worked cost example | See the full arithmetic below the table. | |
| Contracts and DPA | Governed by Hetzner's own customer agreement and data processing terms; both companies are EU entities, so check the specifics against your own requirements. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| What you operate yourself | Everything on the server: OS, model server, monitoring, backups. Hetzner manages the physical hardware and network underneath. | The same: root SSH access to your own container, the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:
Hetzner. GPU server pricing runs hourly for cloud instances or monthly for dedicated servers, varying by card, published on Hetzner's own pricing page. A fair comparison needs a specific server type and its measured tok/s for the model you'd run, both of which depend on your choice, so GPUwerk did not invent a blended figure here. Pull the current rate for the server you'd actually use from Hetzner's site and benchmark your model on it for a real number.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That Spark figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back. Hetzner's reputation is for value pricing on raw compute, so a well-chosen GPU server there could come out cheaper per hour, if it delivers a comparable tok/s for your model on its discrete-GPU layout. That's exactly the kind of comparison neither of us can settle in the abstract, benchmark your own model on both before deciding.
Migration path
Both are bare GPU instances, so migration is mostly re-deploying your own stack. A model served through vLLM on a Spark and the same setup on a Hetzner server expose an identical API shape if you run the same serving software on both, the OpenAI-compatible endpoint comes from vLLM, not from the underlying hardware provider. Put LiteLLM in front of either for request logging and key management. What changes is the memory layout underneath: a model tuned for a discrete-GPU card's VRAM may behave differently on the Spark's unified 128GB pool, and vice versa, so re-test throughput after moving either direction.
When Hetzner is the right choice
- You want a discrete GPU card with a specific VRAM footprint, or need raw compute for training rather than inference.
- You're already running dedicated or cloud servers on Hetzner and want the GPU on the same account and tooling.
- Value pricing on general compute matters more than an inference-specific architecture, worth checking Hetzner's current rates directly.
When a dedicated Spark is the right choice
- Your workload is LLM inference and a large unified-memory pool fits how you'd quantize and serve the model.
- You want a single, fixed hourly rate and no GPU-card decision to make first.
- Your model fits comfortably in 128GB and you'd rather not configure a general-purpose server for a specific job.
FAQ
Is Hetzner a good alternative to GPUwerk for LLM hosting?
It can be, depending on what you're running. Hetzner has a long-standing reputation in the EU for value dedicated and cloud servers, and has added GPU instances to that lineup; check their own site for current GPU specs and availability. Those are general-purpose GPU servers, not built specifically around unified memory for inference. GPUwerk's DGX Spark is one machine with 128GB unified memory, aimed specifically at running LLM inference workloads at a fixed rate. If your workload is inference on models that fit in 128GB, the Spark's architecture is purpose-built for that; if you need a different GPU class or footprint, compare against Hetzner's current specs directly.
What's the difference between a DGX Spark and a typical Hetzner GPU server?
The DGX Spark uses NVIDIA's unified memory architecture, 128GB shared between CPU and GPU, which is well suited to running large quantized LLMs without splitting memory across separate VRAM and system RAM pools. Hetzner's GPU servers, per their own published specs, use conventional discrete GPU cards; check their site for exact memory configurations, which vary by server type.
Is a DGX Spark cheaper than a Hetzner GPU server?
It depends on the server type you'd compare it against. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time. Hetzner publishes its server pricing on its own site; pull the current rate for the GPU server you'd actually run before comparing directly. Hetzner has historically priced dedicated servers competitively, so a direct comparison is worth doing against their current numbers rather than assuming either side wins.
Why choose a DGX Spark over a Hetzner GPU server?
Mainly the memory architecture and the fact that it's purpose-built for inference rather than a general GPU server you configure yourself. If you specifically need a large unified-memory pool for a big quantized model and want a flat hourly rate with no server-spec decision, the Spark fits. If you need a different GPU class, more raw VRAM on a discrete card, or you're already comfortable managing Hetzner's dedicated-server tooling, their catalogue may serve you as well or better.
GPUwerk did not find a single Hetzner GPU server price that fairly represents every card in their catalogue; check Hetzner's own pricing page for the server you'd run. Spark throughput figures come from Dendro Logic's concurrency benchmark, detailed on our benchmarks page.