Private AI/vs Exoscale
Comparison

Private LLM hosting vs Exoscale

By Samuel Seidel · Published September 9, 2026 · 7 min read

We rent DGX Sparks, so weigh that against everything below. Exoscale is a Swiss-owned cloud provider with data centers in Switzerland and elsewhere in Europe, offering general-purpose compute, storage, managed databases, and networking, aimed largely at teams that want a European alternative to the large US hyperscalers. GPU instances, where offered, sit alongside that broader catalogue rather than being the whole product. A DGX Spark is a single-purpose machine, one dedicated node with 128GB unified memory, built specifically for LLM inference.

Side by side

ExoscaleDedicated DGX Spark (GPUwerk)
What it is A general-purpose Swiss and European cloud platform: compute, storage, managed databases, and networking, with GPU instances as one part of a wider catalogue. A single-purpose GPU host: one dedicated DGX Spark, rented by the hour, built around unified memory for inference.
Region Switzerland, Germany, Austria, and Bulgaria, per Exoscale's own current data center list; check their site for GPU availability by region. EU-Central (Prague), exclusively. One location, no region selection to get wrong.
Memory architecture Conventional cloud VM architecture; exact GPU and memory configuration, where offered, varies by instance type, check Exoscale's current specs. 128GB unified memory shared between CPU and GPU, NVIDIA's DGX Spark architecture, suited to large quantized models without splitting memory across pools.
Who can see your data Governed by Exoscale's own terms and data processing addendum; a Swiss and EU-based provider with its own published legal pages, worth checking directly. Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance."
Model choice Whatever you deploy, unrestricted by the provider, within whatever GPU and memory specs their instance types offer. Whatever you deploy too, bounded by 128GB unified memory: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and similar fit.
Speed and concurrency Depends on the instance type rented, if GPU instances are currently offered; Exoscale publishes specs but tok/s for a given model is yours to benchmark. A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page.
Pricing model Per hour for compute instances generally, varying by instance type and resources, published on Exoscale's own pricing page. Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of workload, per pricing.
Worked cost example See the full arithmetic below the table.
Contracts and DPA Governed by Exoscale's own customer agreement and data processing terms; both companies are EU or EU-adjacent entities, so check the specifics against your own requirements. A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads.
What you operate yourself Everything on the instance: OS, model server, monitoring, backups. Exoscale manages the physical hardware and network underneath. The same: OS, model server, any RAG or agent layer, monitoring, backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours.

The worked cost example

Take a workload that runs inference continuously across a full month. Two ways to serve it:

Exoscale. Compute instance pricing runs hourly, varying by instance type and resources, published on Exoscale's own pricing page. A fair comparison needs a specific instance type and its measured tok/s for the model you'd run, both of which depend on your choice and on whether a suitable GPU instance is currently offered, so GPUwerk did not invent a blended figure here. Pull the current rate for the instance you'd actually use from Exoscale's site and benchmark your model on it for a real number.

Dedicated Spark. At $0.79/hour, running continuously for a 730-hour month costs 730 × $0.79 = $576.70. If you hold the reservation instead of running it, the held rate drops that to 730 × $0.59 = $430.70.

A general-purpose cloud instance and a purpose-built inference machine aren't priced against the same yardstick. Exoscale's value is a broader platform you'd already be paying for if you run other workloads there; the Spark's value is one flat number for exactly one job. Weigh the platform consolidation against the specialization before deciding.

Migration path

Both are instances you configure yourself, so migration is mostly re-deploying your own stack. A model served through vLLM on a Spark and the same setup on an Exoscale instance expose an identical API shape if you run the same serving software on both, the OpenAI-compatible endpoint comes from vLLM, not from the underlying cloud provider. Put LiteLLM in front of either for request logging and key management. What changes is the memory layout underneath: a model tuned for a discrete-GPU instance's VRAM may behave differently on the Spark's unified 128GB pool, and vice versa, so re-test throughput after moving either direction.

When Exoscale is the right choice

When a dedicated Spark is the right choice

FAQ

Is Exoscale a good alternative to GPUwerk for LLM hosting?

Exoscale is a Swiss-owned cloud provider with data centers in Switzerland and elsewhere in Europe, known primarily for general-purpose compute, storage, and managed services rather than GPU inference specifically. If GPU instances are part of their current lineup, check their site for the models and regions on offer. GPUwerk's DGX Spark is narrower and built specifically for LLM inference, with 128GB unified memory. If you're already running other workloads on Exoscale and want GPU compute on the same account, that convenience is worth weighing against a purpose-built inference machine elsewhere.

What's the difference between a DGX Spark and an Exoscale GPU instance?

Exoscale's compute instances, GPU-equipped tiers included where offered, use conventional cloud VM architecture; check their current lineup for specs and available GPU models. A DGX Spark uses NVIDIA's unified memory architecture, 128GB shared between CPU and GPU, purpose-built for LLM inference rather than general cloud compute.

Is a DGX Spark cheaper than Exoscale?

It depends on the instance type you'd compare it against, and Exoscale publishes its own current pricing on its site. A dedicated Spark costs $0.79/hour flat, one known number regardless of workload. Pull the current rate for the specific GPU instance you'd actually use, if one is on offer, and benchmark your model on it before comparing directly.

Why choose a DGX Spark over Exoscale?

Mainly the memory architecture and the fact that it's purpose-built for inference rather than a general cloud instance. If your workload is inference on a model that fits 128GB unified memory and you want a flat hourly rate with no instance-type decision, the Spark fits. If you're consolidating broader infrastructure, including storage, networking, and managed services, on one Swiss or EU cloud provider, Exoscale's wider catalogue may serve you better.

GPUwerk did not find a single Exoscale GPU instance price that fairly represents every available configuration; check Exoscale's own pricing page for the instance you'd run. Spark throughput figures come from Dendro Logic's concurrency benchmark, detailed on our benchmarks page.

Related pages

First top-up: pay $10, get $20 in credit

Run your own numbers before you commit either way.

Deploy a dedicated DGX Spark in EU-Central and test your actual prompts against an open model before comparing quotes.

Deploy a Spark Read the benchmarks