Private LLM hosting vs Google Vertex AI
We rent DGX Sparks, so weigh that against everything below. Vertex AI is Google Cloud's managed AI platform: Gemini for closed frontier models, and Model Garden for a broader catalogue that includes open-weight models too. If your task needs Gemini by name, this page cannot get you there, a Spark only runs open-weight models. If an open model already meets your quality bar and your reason for looking at Vertex AI is where the data goes, the answer is more layered than "pick the EU region," and it's worth being precise about that before you decide.
Side by side
| Google Vertex AI | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| Where data is processed | Google Cloud operates EU regions, and Vertex AI deployments can be configured to process and store data there. Google itself, however, is a US-headquartered company, and the US CLOUD Act asserts jurisdiction over data controlled by a US company regardless of where the servers sit. An EU region is a real residency control; it is not the same as removing US legal exposure entirely. | EU-Central (Prague), on one dedicated machine, operated by a Czech company. Not a complete answer to every jurisdictional question either, but a structurally different starting point than a US-headquartered cloud provider's EU region. |
| Who can see your prompts | Governed by Google Cloud's own data governance and generative AI service terms. GPUwerk did not independently verify the current specifics of those terms for this page; read Google's own documentation for what applies to your account and model. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." No abuse-monitoring or content-scanning layer exists at all. |
| Model choice | Gemini, Google's closed frontier model family, plus Model Garden, a catalogue spanning Google models, third-party models, and a number of open-weight models hosted on Google's infrastructure. | Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, Google's own open Gemma family, and anything else that fits in 128 GB. Gemini is not available on a Spark under any arrangement; an open-weight model that also appears in Model Garden may be. |
| Speed and concurrency | Google does not publish a general tokens/second figure for Vertex AI; usage limits are typically expressed as quota per model rather than throughput. | A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page. |
| Pricing model | Per token for Gemini and most Model Garden models, varying by model and context length. See cloud.google.com/vertex-ai/generative-ai/pricing for current rates; GPUwerk did not fetch a specific figure for this page. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of model. USD, tax extra, per pricing. |
| Worked cost example | See the full arithmetic below the table. | |
| Contracts and DPA | Governed by Google Cloud's own data processing addendum and service terms for Vertex AI. Enterprise agreements are typical at volume. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| What you operate yourself | Nothing at the infrastructure layer; Google manages capacity, scaling, and model serving. You manage prompts, application code, and your own use of quota. | Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:
Vertex AI. Pricing is per token, varying by model, Gemini tier, and context length, presented on Google's pricing page through model-specific tables. GPUwerk did not fetch and verify an exact current per-token price for a specific model at the time of writing. Run your own model and volume through cloud.google.com/vertex-ai/generative-ai/pricing for a figure you can trust.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That Spark figure only holds under the assumption stated: the node busy at 256 concurrent requests for the full 161 hours, back to back, with no idle time. Real traffic arrives in bursts, so a Spark billed by the hour that sits half-idle waiting for requests can lose ground against a per-token price with no idle charge. Put a current Vertex AI number for your chosen model next to $127.19 for your own volume, and you have the real comparison; which one wins is a utilisation question specific to your traffic, not something either provider can answer for you in the abstract.
Migration path
Where Vertex AI serves a model through an OpenAI-compatible endpoint, or where you're using an open-weight model from Model Garden, a Spark serving the same model through vLLM exposes a matching chat completions shape, so an application that only calls that endpoint moves with a base URL and key change. Put LiteLLM in front of the Spark and you get a drop-in OpenAI-compatible gateway with request logging, key management, and rate limits, useful for normalising between Gemini's native API shape and a Spark's OpenAI-compatible endpoint if you're running both. Gemini-specific features, grounding, built-in search, and Vertex's evaluation and pipeline tooling have no Spark equivalent because there's no managed platform layer underneath it.
When Vertex AI is the right choice
- Your task needs Gemini specifically, or a Model Garden feature like built-in grounding or evaluation pipelines, and no open-weight model paired with a Spark matches it.
- Your token volume is spiky or low, so paying only for tokens generated beats paying for a machine that sits idle between requests.
- You're already standardised on Google Cloud and want AI deployment inside the same billing, IAM, and networking boundary.
When a dedicated Spark is the right choice
- An open-weight model already meets your quality bar, so Gemini's advantage doesn't apply to your task.
- Where the data sits and who is legally able to compel access to it matters to you beyond picking a region, and you want a non-US operator as the starting point.
- Your traffic is steady enough to keep the machine busy for most of the hours you're paying for.
FAQ
Can I run Gemini on a DGX Spark?
No. Gemini is a closed model family and runs only on Google's own infrastructure. A Spark runs open-weight models: gpt-oss-120b, Llama 3.3, Qwen3-Coder, DeepSeek, and similar, including Google's own open-weight Gemma family if it fits your memory budget. Vertex AI's Model Garden also lists a number of open-weight models alongside Gemini; if you're using one of those through Model Garden, the same model may run directly on a Spark instead.
Does hosting in a Google Cloud EU region avoid US legal exposure?
Not entirely. Google Cloud operates EU regions and can store and process data there, but Google is a US-headquartered company, and the US CLOUD Act asserts jurisdiction over data controlled by US companies regardless of where the servers physically sit. Choosing an EU region is a real data-residency control, but it does not by itself remove that legal exposure. GPUwerk is a Czech company and PRINT IT! SE, its operator, is not a US entity, which is a structurally different starting point, though it does not make GPUwerk immune to every legal process either, no operator can claim that.
What models does Vertex AI offer?
Vertex AI's primary offering is Google's own Gemini model family, plus Model Garden, a catalogue that includes both Google models and a range of third-party and open-weight models hosted on Google's infrastructure. Check Google Cloud's own current Vertex AI documentation for exactly which models are available in Model Garden at any given time, that catalogue changes.
Is a DGX Spark cheaper than Vertex AI?
It depends on which model you'd use on Vertex AI and your token volume. GPUwerk did not fetch a specific current Vertex AI price for this page, since pricing varies by model and Google's pricing pages present it through per-model tables; check cloud.google.com/vertex-ai/pricing directly. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time.
Who can see my prompts on Vertex AI?
Google publishes its own data governance and generative AI terms describing how customer data submitted to Vertex AI is handled and whether it's used for model improvement. GPUwerk did not independently verify those specifics for this page, read Google Cloud's own current documentation for the terms that apply to your account. On a Spark, GPUwerk's DPA states GPUwerk hosts the machine but does not access, read, copy, index, or analyse the content of your instance.
Can I switch between Vertex AI and a Spark without rewriting my application?
For models served through Vertex AI's OpenAI-compatible endpoint (where Google offers one) or through Model Garden's open-weight models, client code that only targets a chat completions shape ports to a Spark with a base URL and key change. Gemini's native API shape and Vertex-specific platform features, grounding, built-in tools, evaluation pipelines, do not port and have no Spark equivalent, you'd build any of that yourself or normalise the difference through a gateway like LiteLLM.
GPUwerk did not fetch a current per-token Vertex AI price for this page, prices vary by model and change often; check cloud.google.com/vertex-ai/generative-ai/pricing directly. The data-handling claim above reflects the existence of Google's own published terms, not independent verification of their current content by GPUwerk. Google does not publish a general tokens/second figure for Vertex AI anywhere GPUwerk could find; that row above reflects that gap rather than a guess.