Private LLM hosting vs Fly.io GPUs
We rent DGX Sparks, so weigh that against everything below. Fly.io built its platform around running small application instances close to users in many regions, with fast boot times and a deploy workflow centered on its own Machines API. GPU-backed Machines are one option inside that platform, useful mainly if inference needs to sit next to an application already running on Fly.io. A DGX Spark takes a narrower approach: one dedicated machine, operated by GPUwerk directly, built specifically for LLM inference rather than as an option inside an edge-app platform.
Side by side
| Fly.io GPU Machines | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| What it is | A GPU-backed Machine type inside a platform built for edge-deployed application instances, sold alongside Fly.io's standard CPU Machines. | A single-purpose GPU host: one dedicated DGX Spark, rented by the hour, built around unified memory for inference. |
| Who operates the hardware | Fly.io, across whichever regions it makes GPU capacity available in. | GPUwerk directly, in a data center we control in Prague. |
| GPU model and specs | Varies by region and tier; check Fly.io's current GPU pricing page for what's on offer, since lineup and availability change. | NVIDIA DGX Spark: Grace Blackwell, 128GB unified memory, the same specification on every node. |
| Region | Many global regions for standard Machines; GPU capacity is available in a smaller subset that changes over time, so check current availability for the EU specifically. | EU-Central (Prague), exclusively. One location, no region selection to get wrong. |
| Who can see your data | Governed by Fly.io's own terms and data processing agreement; review those directly for the specifics of GPU Machine handling. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." |
| Pricing model | Per-second billing by machine size and GPU tier, published on Fly.io's own pricing page and subject to change; check it directly for a current number. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of workload, per pricing. |
| Contracts and DPA | Fly.io's standard terms and DPA, applying across its whole platform; check current terms for GPU-specific clauses. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| What you operate yourself | OS image and model server inside a Fly.io Machine, plus Fly-specific deploy and scaling configuration if you're using it. | OS, model server, any RAG or agent layer, monitoring, backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that runs inference continuously across a full month. Two ways to serve it:
Fly.io GPU Machines. Billed per second by machine size and GPU tier, published on Fly.io's own pricing page and changing over time. A fair dollar comparison needs the specific machine size you'd actually run, so check Fly.io's current GPU pricing for that number rather than trusting a figure repeated secondhand here.
Dedicated Spark. At $0.79/hour, running continuously for a 730-hour month costs 730 × $0.79 = $576.70, flat, at one known rate from one known operator. If you hold the reservation instead of running it, the held rate drops that to 730 × $0.59 = $430.70.
Fly.io's per-second billing can work in your favor for bursty, short-lived inference jobs that spin up and down, since you only pay for the seconds a Machine runs. For a workload that's on continuously, that same fine-grained billing doesn't inherently beat a flat hourly rate, it depends on the actual per-second rate for the GPU tier you'd need. Check Fly.io's current numbers for the configuration you'd actually run before deciding either way.
Migration path
Both are Linux GPU environments, so migration is mostly re-deploying your own model server. Serve a model through vLLM on a Spark and inside a Fly.io GPU Machine, and both expose an OpenAI-compatible endpoint from vLLM itself, not from the underlying platform. Put LiteLLM in front of either for request logging and key management. If your application is already built around Fly.io's deploy tooling and fly.toml configuration, expect to keep that wiring and just point the inference call at whichever backend you're testing.
When Fly.io is the right choice
- Your application already runs on Fly.io and you want inference colocated with it under the same deploy workflow.
- Your inference load is bursty and short-lived, where per-second billing on a Machine that scales to zero fits the traffic pattern.
- Committing to a single EU-hosted operator matters less than staying inside a platform you already use.
When a dedicated Spark is the right choice
- You want a machine built specifically for LLM inference rather than a GPU option bolted onto an edge-app platform.
- Your traffic is steady enough that a flat hourly rate is simpler to reason about than per-second billing.
- EU data residency and a standard GDPR DPA are requirements, not nice-to-haves.
FAQ
Does Fly.io offer GPU machines for LLM hosting?
Yes, Fly.io added GPU-backed Machines to its platform, which is otherwise built around running small, fast-booting app instances close to users at the edge. Which GPU models and regions carry GPU capacity changes over time, so check Fly.io's current GPU pricing page directly before comparing.
Is Fly.io a good alternative to a DGX Spark for running an LLM?
It depends on what you need. Fly.io's GPU Machines fit naturally if your application already runs on Fly.io and you want inference colocated with the rest of your app's edge deployment. A dedicated DGX Spark is a single machine built specifically for LLM inference, with 128GB of unified memory, operated by GPUwerk directly in EU-Central. If edge deployment alongside your app matters more than a machine sized for the model itself, Fly.io is worth checking; otherwise weigh the two on hardware fit.
Is a Spark cheaper than Fly.io GPU Machines?
It depends on the specific GPU tier and how you bill Fly.io Machines, since Fly.io prices GPU-backed instances on its own pricing page and that changes over time. A dedicated Spark costs $0.79/hour flat, one number regardless of workload. Check Fly.io's current GPU pricing for the tier you'd actually need before comparing.
Does Fly.io run GPU machines in the EU?
Fly.io operates machines across many regions worldwide, but GPU capacity by region varies and changes; check their current site for which regions carry GPU machines today. A GPUwerk Spark is fixed in EU-Central, Prague, with no region selection because there is only the one machine.
GPUwerk's own figures on this page ($0.79/hour, $0.59/hour, $576.70/month) come from our published pricing. Fly.io's GPU pricing and availability are not something GPUwerk can verify or restate; check Fly.io's own site directly.