Private AI/vs Modal
Comparison

Private LLM hosting vs Modal

By Samuel Seidel · Published September 9, 2026 · 9 min read

We rent DGX Sparks, so weigh that against everything below. Modal is a serverless compute platform, billed per second, with functions and containers that spin up on demand and scale to zero when idle, popular well beyond inference for batch jobs, data pipelines, and training runs. A Spark is the opposite shape: one machine, always on, billed by the hour whether it's busy or not. For a bursty workload, Modal's model can genuinely cost less than an always-on Spark; for a steady one, the fixed rate on a Spark can be simpler and cheaper to reason about.

Side by side

ModalDedicated DGX Spark (GPUwerk)
What it is A serverless compute platform: functions and containers run on demand, scale to zero when idle, and are billed per second of actual compute used, per Modal's documentation. Used for inference, batch processing, and training. One dedicated physical machine, rented by the hour, always on for as long as you keep it running. No autoscaling, no scale-to-zero.
Where data is processed Whichever infrastructure region Modal provisions your workload to; check Modal's own documentation for current region options. EU-Central (Prague), on one dedicated machine, no other region to configure.
Who can see your data Governed by Modal's own terms of service and data processing agreement; check Modal's current legal pages for the specifics that apply to your account. Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance."
Workload shape Built for bursty, event-driven, or scheduled work: batch inference jobs, training runs, data pipelines, and request-driven inference that spins containers up and down as traffic arrives. Built for continuous serving: a model loaded once and kept warm, answering requests as they arrive without a cold-start penalty because nothing ever spins down.
Model choice Whatever you package into a Modal function or container, open-weight, fine-tuned, or custom, unrestricted by the platform. Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else you can fit in 128 GB, run directly since you have root access.
Speed and concurrency Depends on the GPU type you request in your Modal function and how many containers scale up under load; Modal does not publish a single tokens/second figure that applies across models and GPU choices. A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page.
Pricing model Per second of compute actually used, no charge while scaled to zero. See Modal's own pricing page for current per-second GPU rates. Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of traffic, per pricing.
Worked cost example See the full arithmetic below the table.
Contracts and DPA Governed by Modal's own terms of service and data processing agreement. A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads.
What you operate yourself Your function code and container image; Modal manages provisioning, scheduling, and scaling the compute underneath it. Everything, root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours.

The worked cost example

Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:

Modal. Billing is per second of GPU compute actually consumed, with containers scaling to zero between requests, per Modal's own pricing page. For a workload that's genuinely continuous, requests arriving constantly enough to keep a container warm the whole month, the per-second rate approximates a per-hour rate, and Modal's own pricing page has the current figure for the GPU type you'd use. GPUwerk did not attempt to reconstruct that number here since it depends on the specific GPU and container configuration you'd choose.

Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.

That $127.19 figure assumes the Spark is saturated at 256 concurrent requests for the full 161 hours, back to back, with zero idle time, which is a best case a real traffic pattern rarely hits. This is precisely where Modal's model changes the arithmetic: real traffic arrives in bursts, and a Spark billed by the hour keeps charging $0.79 whether it's generating tokens or sitting idle between requests, while Modal charges close to nothing during the gaps. The less continuous your traffic, the more that per-second billing favors Modal; the closer your traffic is to constant, all-day saturation, the more a flat hourly Spark rate tends to win, since Modal is also charging for GPU time under the hood, just metered more finely.

Migration path

A model wrapped in a Modal function commonly exposes an HTTP endpoint you define yourself, and a Spark serving a model through vLLM exposes an OpenAI-compatible endpoint by default; matching the two takes deliberate API design on the Modal side, it isn't automatic. Put LiteLLM in front of either for a consistent gateway and key management. What doesn't move is Modal's function-based scheduling, its scale-to-zero behavior, and its handling of batch and training jobs as first-class workloads, a Spark has none of that; you'd run training or batch work as ordinary processes on the machine instead.

When Modal is the right choice

When a dedicated Spark is the right choice

FAQ

Is Modal only for inference, or does it do more?

More. Modal is a general serverless compute platform popular for the full range of ML workloads: batch data processing, model training and fine-tuning jobs, scheduled pipelines, and inference, all billed per second of actual compute used, per Modal's own documentation. A Spark is inference-shaped by comparison, one machine you keep running, well suited to serving a model continuously but not built around Modal's job-scheduling and ephemeral-container model for batch and training workloads.

How does Modal's pricing actually work?

Per-second billing for the compute resources (CPU, GPU, memory) a function or container actually consumes while running, with containers spinning up on demand and down to zero when idle, per Modal's own pricing page. There's no fixed hourly commitment: a job that runs for 90 seconds is billed for roughly 90 seconds of compute, not a full hour.

Does Modal have EU infrastructure?

Check Modal's own documentation for current region and infrastructure options, availability changes over time. GPUwerk's Spark fleet runs exclusively from EU-Central (Prague), a fixed location rather than a region choice among several.

Is a DGX Spark cheaper than Modal?

It depends almost entirely on how continuous your traffic is. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time, back to back with no idle time. Modal's per-second billing means a workload with real idle gaps, most production traffic, pays only for active compute, so the honest comparison depends on your specific request pattern, not the sticker rate on either side. Check Modal's own pricing page for current per-second GPU rates.

Can I run training jobs on a Spark the way I can on Modal?

You can run training or fine-tuning code directly on a Spark since you have root access and 128GB unified memory to work with, but there's no job scheduler, no automatic container spin-up per job, and no built-in queueing the way Modal provides. A Spark is one machine running whatever you start on it; Modal is built to fan a training or batch job out across many ephemeral containers and bill only for what each one uses.

GPUwerk did not find a single flat per-token or blended per-hour Modal price that holds across GPU types and container configurations; Modal prices per-second usage on its own pricing page, so that row above reflects that structure rather than a guess. Spark throughput figures come from Dendro Logic's concurrency benchmark, detailed on our benchmarks page.

Related pages

First top-up: pay $10, get $20 in credit

Run your own numbers before you commit either way.

Deploy a dedicated DGX Spark in EU-Central and test your actual prompts against an open model before comparing quotes.

Deploy a Spark Read the benchmarks