Private LLM hosting vs Databricks Mosaic AI
We rent DGX Sparks, so weigh that against everything below. Databricks Mosaic AI and a Spark aren't really the same shape of product. Mosaic AI is a training, fine-tuning, and serving layer built into the Databricks lakehouse, meant for teams whose data pipelines, governance, and ML workflows already run there. A Spark is one dedicated machine you rent by the hour, with no platform underneath it. If your data already lives in Databricks, Mosaic AI keeps model work in the same place your pipelines run. If it doesn't, or you don't want the platform commitment, a Spark is a shorter path to a private model endpoint.
Side by side
| Databricks Mosaic AI | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| What it is | A set of ML tools, training, fine-tuning, vector search, and model serving, built into the Databricks lakehouse platform, governed by Unity Catalog and run on Databricks-managed compute. | One dedicated physical machine, rented by the hour, with root access. No platform, no catalogue, no workspace to adopt. |
| Prerequisite | An existing Databricks account and workspace. Mosaic AI is not a standalone product; it runs inside your Databricks environment. | None. Sign up and deploy directly against a Spark, no workspace or platform migration required. |
| Where data is processed | Whichever cloud region your Databricks workspace is deployed in (AWS, Azure, or GCP), configured through Databricks' own account and region settings. | EU-Central (Prague), on one dedicated machine, no other region to configure. |
| Who can see your data | Governed by your Databricks account's own access controls and Unity Catalog permissions, and by whichever Databricks terms and data processing addendum apply to your contract. Check those documents for your specific arrangement. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." |
| Model choice | Open-weight and fine-tuned models through Mosaic AI Model Serving, plus access to external model APIs through the same gateway, per Databricks' documentation. | Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else you can fit in 128 GB. No external model gateway, no closed frontier models. |
| Speed and concurrency | Databricks does not publish a general tokens/second figure for Mosaic AI Model Serving; throughput depends on the endpoint's provisioned scaling configuration and the model served. | A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page. |
| Pricing model | Consumption-based Databricks Units (DBUs) that vary by workload type, instance size, and commitment level, on top of the underlying cloud compute cost. See Databricks' own pricing page for figures specific to your account and cloud. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of model or how many requests you send, per pricing. |
| Worked cost example | See the full arithmetic below the table. | |
| Contracts and DPA | Governed by your Databricks Master Cloud Services Agreement and Data Processing Addendum. Enterprise terms are typical at volume. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| What you operate yourself | Prompts, pipelines, and model configuration inside the workspace; Databricks manages the underlying compute, scaling, and platform services. | Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:
Databricks Mosaic AI. Model serving is billed in Databricks Units that vary by endpoint type, provisioned throughput, and your account's commitment tier, on top of the underlying cloud instance cost. Databricks does not publish a flat per-token price that applies across accounts and clouds, so GPUwerk did not attempt to reconstruct one here. Run your own model and volume through Databricks' pricing calculator on their site for a number specific to your workspace.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That Spark figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back, with no idle time. A team already paying for Databricks compute anyway, because their pipelines run there regardless, may find the marginal cost of adding Mosaic AI serving lower than standing up a separate Spark just for inference. A team with no existing Databricks footprint is comparing a new platform subscription against a single hourly-billed machine, and the fixed $0.79/hour is the easier number to reason about without a calculator.
Migration path
Mosaic AI Model Serving exposes an OpenAI-compatible endpoint shape for many model types, per Databricks' documentation, and a Spark serving a model through vLLM exposes the same shape, so client code that only calls a chat completions endpoint moves with a base URL and API key change. What doesn't move is anything tied to Unity Catalog governance, Databricks' pipeline orchestration, or MLflow tracking wired into the workspace; none of that has an equivalent on a Spark, you would either keep those workflows in Databricks and only move inference, or build your own tracking and access control on top of the Spark.
When Databricks Mosaic AI is the right choice
- Your data pipelines, feature engineering, and training already run in Databricks, and you want model serving governed by the same Unity Catalog permissions as the rest of your data.
- You need autoscaling endpoints that Databricks operates for you, rather than a fixed machine you scale yourself.
- Your workload mixes training, fine-tuning, and serving in one pipeline, and keeping all three in one platform is worth more than the flexibility of separate tools.
When a dedicated Spark is the right choice
- You don't already run Databricks and don't want to adopt a lakehouse platform just to serve one or two models.
- You want a single, fixed hourly rate you can put in a spreadsheet without a pricing calculator.
- You want the shortest possible answer to "who can access our data": one dedicated machine, root access under your own SSH keys, and an operator who states in writing that it does not read your instance's content.
FAQ
Is Databricks Mosaic AI the same kind of product as a dedicated Spark?
No. Mosaic AI is a set of tools inside the Databricks lakehouse platform for training, fine-tuning, and serving models, tied to your existing Databricks workspace, Unity Catalog governance, and compute. A DGX Spark is a single dedicated machine you rent by the hour and operate yourself. If you already run your data pipelines in Databricks, Mosaic AI keeps model work in the same place; if you don't, adopting it means adopting the platform underneath it too.
Do I need an existing Databricks account to use Mosaic AI?
Yes. Mosaic AI features are part of a Databricks workspace, so you need a Databricks account and workspace set up first, with billing running through your Databricks contract (typically consumption-based Databricks Units, per Databricks' own pricing page). A Spark has no such prerequisite: you sign up and deploy directly.
Can I run open-weight models on either platform?
Yes on both. Mosaic AI Model Serving hosts open-weight and fine-tuned models behind an autoscaling endpoint inside your workspace, per Databricks' documentation. A Spark runs the same class of open-weight models, gpt-oss-120b, Llama 3.3, Qwen3-Coder, DeepSeek, on a single machine you control directly rather than behind a managed serving layer.
Is a DGX Spark cheaper than Databricks Mosaic AI?
It depends on your volume and whether you already pay for Databricks compute. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time. Databricks prices Mosaic AI serving through consumption-based Databricks Units that vary by workload and commitment level; check Databricks' own pricing page for a number specific to your account before comparing.
What do I lose by leaving the Databricks platform for a Spark?
Unity Catalog governance, the notebook and pipeline environment, MLflow experiment tracking wired into the same workspace, and autoscaling model-serving endpoints that Databricks operates for you. A Spark gives you root access to a single machine and none of that platform layer; you would run your own tracking, scaling, and access control on top of it.
GPUwerk did not find a published flat per-token price for Databricks Mosaic AI Model Serving that holds across accounts, clouds, and workload types; Databricks prices this through consumption-based Databricks Units on its own pricing page, so that row above reflects that structure rather than a guess. Spark throughput figures come from Dendro Logic's concurrency benchmark, detailed on our benchmarks page.