Private LLM hosting vs Scaleway
We rent DGX Sparks, so weigh that against everything below. Scaleway is a French cloud provider and, like GPUwerk, operates from the EU, so this isn't a US-vs-EU comparison the way some of our other pages are. The real difference is scope. Scaleway sells a full cloud catalogue: GPU Instances, object storage, managed Kubernetes, databases, and more, so its GPU offering sits inside a much bigger product line. GPUwerk sells one thing, a dedicated DGX Spark by the hour, with nothing else to configure around it.
Side by side
| Scaleway GPU Instances | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| What it is | A GPU compute product inside Scaleway's broader cloud platform, alongside storage, managed databases, Kubernetes, and more. | A single-purpose GPU host: one dedicated DGX Spark, rented by the hour, nothing else attached. |
| Region | Scaleway operates data centers in France and elsewhere in the EU; check their own region list for which products are available where. | EU-Central (Prague), exclusively. One location, no region selection to get wrong. |
| Hardware choice | Multiple GPU instance types across NVIDIA cards, per Scaleway's own current catalogue. | One machine: the DGX Spark, 128GB unified memory, with an optional two-node cluster for larger workloads. |
| Who can see your data | Governed by Scaleway's own terms and data processing addendum; check their legal pages for the current text. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." |
| Model choice | Whatever you deploy, unrestricted by the provider. Larger GPU instance types can run bigger models or serve higher concurrency than a single Spark. | Whatever you deploy too, but bounded by 128GB unified memory: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and similar fit; models needing substantially more memory don't, without the two-node cluster. |
| Speed and concurrency | Depends on the GPU type and count rented; Scaleway publishes instance specs but tok/s for a given model is yours to benchmark. | A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page. |
| Pricing model | Per hour or per minute, varying by GPU type, published on Scaleway's own pricing page. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of workload, per pricing. |
| Worked cost example | See the full arithmetic below the table. | |
| Contracts and DPA | Governed by Scaleway's own customer agreement and data processing terms; both companies are EU entities, so check the specifics against your own requirements. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| What you operate yourself | Everything on the instance: OS, model server, monitoring, backups. Scaleway manages the physical hardware and network underneath, and offers additional managed products if you want to build on top of them. | The same: root SSH access to your own container, the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:
Scaleway. GPU Instance pricing runs per hour and varies by card, published on Scaleway's own pricing page. A fair comparison needs a specific instance type and its measured tok/s for the model you'd run, both of which depend on your choice, so GPUwerk did not invent a blended figure here. Pull the current rate for the instance you'd actually use from Scaleway's site and benchmark your model on it for a real number.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That Spark figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back. A larger Scaleway GPU instance could plausibly clear the same 500 million tokens in fewer wall-clock hours at a higher hourly rate, so whether it's cheaper overall depends on the instance's price-per-token-per-hour relative to a Spark's, not just its sticker price. Run both through your own model and traffic pattern; a raw GPU-hour comparison without a specific instance type and a specific tok/s figure isn't one either provider can hand you.
Migration path
Both are bare GPU instances, so migration is mostly re-deploying your own stack. A model served through vLLM on a Spark and the same setup on a Scaleway instance expose an identical API shape if you run the same serving software on both, the OpenAI-compatible endpoint comes from vLLM, not from the underlying hardware provider. Put LiteLLM in front of either for request logging and key management. What changes is the hardware spec underneath: a model tuned to a larger Scaleway instance's memory and bandwidth may need re-quantizing or a smaller variant to fit a single Spark's 128GB, and vice versa moving the other way.
When Scaleway is the right choice
- You already run other infrastructure on Scaleway, object storage, managed Kubernetes, databases, and want the GPU on the same platform and billing.
- You need a GPU type or memory footprint beyond a single Spark's 128GB, or want the flexibility to switch instance types between workloads.
- You want a broader menu of regions or products than a single dedicated machine offers.
When a dedicated Spark is the right choice
- You want GPU compute and nothing else to configure, no instance-type catalogue to read through first.
- A single, fixed hourly rate matters more than platform breadth.
- Your model fits comfortably in 128GB and you'd rather not pay for headroom you won't use.
FAQ
Is Scaleway a good alternative to GPUwerk?
It depends what you're optimizing for. Both are EU-headquartered infrastructure providers (Scaleway is French, GPUwerk operates from EU-Central in Prague), so data residency concerns are similar in kind. Scaleway is a full general-purpose cloud, GPU Instances are one product line among object storage, Kubernetes, managed databases, and more. GPUwerk sells one thing: a dedicated DGX Spark by the hour. If you need only GPU compute for LLM inference and want to avoid picking an instance type, GPUwerk is simpler. If you need the GPU alongside other Scaleway infrastructure you already run, staying on one platform may be worth more than the simplicity.
Does Scaleway offer dedicated hardware or is it shared/virtualized?
Check Scaleway's own GPU Instances documentation for the current instance types and whether a given tier is dedicated or shared. GPUwerk's DGX Spark is dedicated hardware for the duration of your reservation, no other tenant runs on the same box while it's yours.
Is a DGX Spark cheaper than Scaleway GPU Instances?
It depends on the GPU type you'd compare it against. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time. Scaleway publishes its GPU Instance rates on its own pricing page; pull the current rate for the instance you'd actually run before comparing directly.
Why pick a single dedicated Spark over a cloud GPU instance from a general-purpose provider?
Mainly simplicity. Scaleway's catalogue spans many products and GPU instance types, which is useful if you need that range, but it also means comparing specs and prices before you can start. A Spark is one machine, one rate, and it's what GPUwerk operates end to end, so there's less to configure before you're running a model.
GPUwerk did not find a single Scaleway GPU Instance price that fairly represents every card in their catalogue; check Scaleway's own pricing page for the instance you'd run. Spark throughput figures come from Dendro Logic's concurrency benchmark, detailed on our benchmarks page.