Private LLM hosting vs Azure OpenAI
We rent DGX Sparks, so weigh that against everything below. If you need a specific closed model by name, GPT-4o, GPT-5.1, o3, this page cannot get you there, a Spark only runs open-weight models. If your workload can run on an open model and your reason for looking at Azure OpenAI is where the data goes and who can read it, a dedicated Spark in EU-Central gives you a shorter answer to both questions than any per-token API can, because the machine is yours alone and nothing about it is shared.
Side by side
| Azure OpenAI | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| Where data is processed | The Azure geography you pick, unless you use a Global or Data Zone deployment. In a Data Zone deployment created in an EU member nation, Microsoft states prompts and responses "may be processed in that or any other European Union Member Nation." Microsoft, "Data, privacy, and security for Foundry Models sold by Azure," fetched 7 September 2026. | EU-Central (Prague), on one dedicated machine. No other region, no cross-border processing option to configure because there's only the one machine. |
| Who can see your prompts | Not other customers, not OpenAI, and not used to train models, per Microsoft's own statement. For accounts without modified abuse monitoring, a sample of flagged prompts is reviewed, first automatically, then by "authorized Microsoft employees" via "point wise queries using request IDs, Secure Access Workstations (SAWs), and Just-In-Time (JIT) request approval," located in the EEA for EEA deployments. Microsoft, fetched 7 September 2026. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." There is no abuse-monitoring layer at all, because GPUwerk operates infrastructure, not a model API. |
| Model choice | Closed frontier models: GPT-5.1, GPT-5.4, o3, and the rest of the current OpenAI lineup, plus a growing catalogue of other vendors' models in Foundry. | Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else you can fit in 128 GB. GPT-class closed models are not available on a Spark, full stop, that is not a current limitation, it is the shape of the product. |
| Speed and concurrency | Microsoft does not publish tokens/second; Azure OpenAI pricing and SLAs are stated in tokens and requests per minute (TPM/RPM) quota, not throughput. | A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). On our own fleet, Qwen3-Coder 30B-A3B AWQ decodes at 80.9 tok/s single-stream, and Llama 3.3 70B AWQ decodes at 6.0 tok/s single-stream but reaches 485 tok/s aggregate at 128 concurrent before its performance collapses past that point. Raw logs and methodology on our benchmarks page. |
| Pricing model | Per token, billed separately for input and output, varying by model and deployment type (Global, Data Zone, Regional). | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of model or how many requests you send. USD, tax extra, per pricing. |
| Worked cost example | See the full arithmetic below the table. | |
| Contracts and DPA | Governed by the Microsoft Products and Services Data Protection Addendum. Enterprise agreements and negotiated terms are typical at volume. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| Lock-in and portability | Both expose an OpenAI-compatible chat completions shape, so basic client code ports either direction. Prompts tuned to a specific closed model's behaviour, and platform features like the Assistants API or stored completions, do not move with you. | Open weights mean you can move the exact model file to another Spark, another GPU box, or your own hardware; the model itself is never Azure's or GPUwerk's to withhold. You do lose Azure's managed platform layer (content filters, quota management, stored completions) and have to build any of that yourself. |
| What you operate yourself | Nothing at the infrastructure layer; Microsoft manages capacity, scaling, and model serving. You manage prompts, application code, and your own use of quota. | Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, content filtering, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:
Azure OpenAI. Microsoft's pricing page lists per-token input and output prices for each model and deployment type, but the figures load through a client-side calculator rather than plain page text, so GPUwerk could not fetch and verify an exact per-token price for a current GPT-5.1 deployment at the time of writing. Run your own model and volume through azure.microsoft.com/pricing/details/cognitive-services/openai-service for a number you can trust, since Azure OpenAI pricing is per-token, billed separately for input and output and varying by model and deployment type as noted in the table above.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That Spark figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back, with no idle time and enough real traffic to keep the queue full. Real traffic arrives in bursts, not a constant stream at exactly your saturation concurrency, so a Spark billed by the hour that sits half-idle waiting for requests can lose ground fast, you're paying $0.79 for every hour on the clock whether the GPU is generating tokens or not. Azure's per-token price does not care about idle time because there is no idle time to bill, you pay only for tokens you actually generate. Put the Azure number from the calculator above next to $127.19 for your own volume and you have the real comparison; the crossover point is a utilisation question specific to your traffic pattern, not a fixed answer either provider can give you.
Migration path
Both sides speak the same wire protocol at the client level. Azure OpenAI exposes an OpenAI-compatible chat completions API, and a Spark serving a model through vLLM exposes the same shape, so an application that only calls that endpoint moves with a base URL and API key change. Put LiteLLM in front of the Spark and you get a drop-in OpenAI-compatible gateway with request logging, key management, and rate limits, the pieces Azure's managed layer gives you for free and that you have to stand up yourself on a dedicated machine. What does not move is anything tuned to one closed model's specific behaviour, output formatting, or context window, and platform-only features like Azure's Assistants API or stored completions have no equivalent on a Spark because there is no managed platform layer underneath it.
When Azure OpenAI is the right choice
- Your evaluation requires a specific closed frontier model by name, GPT-5.1, o3, or a future release, and no open-weight model has matched it on your task yet.
- Your token volume is spiky or low, so paying only for tokens generated beats paying for a machine that sits idle between requests.
- You want Microsoft managing capacity, scaling, and platform features like the Assistants API or built-in content filtering, rather than building and operating that layer yourself.
When a dedicated Spark is the right choice
- An open-weight model, gpt-oss-120b, Llama 3.3, Qwen3-Coder, or similar, already meets your quality bar, so the closed-model advantage doesn't apply to your task.
- Your traffic is steady enough to keep the machine busy for most of the hours you're paying for, batch document processing, an internal assistant used all day, or an agent fleet running continuously.
- You want the shortest possible answer to "who can access our data": one dedicated machine, root access under your own SSH keys, and an operator who states in writing that it does not read your instance's content.
FAQ
Can I run GPT-4o or GPT-5-class models on a DGX Spark?
No. GPT-4o, GPT-5.1, and the rest of OpenAI's closed model family run only on OpenAI's and Microsoft's infrastructure, not on hardware you rent or own. A Spark runs open-weight models: gpt-oss-120b, Llama 3.3, Qwen3-Coder, DeepSeek, and similar. If your evaluation requires a specific closed frontier model by name, Azure OpenAI is the only place to get it.
Is Azure OpenAI GDPR compliant if I use the EU Data Zone?
Microsoft's own documentation states that a Data Zone deployment created in a Foundry resource located in an EU member nation processes prompts and responses in that or any other EU member nation, and that data stored at rest stays in the customer-designated geography. That is a data-residency control you configure, not an automatic GDPR compliance statement; Microsoft's Products and Services Data Protection Addendum is the governing document, and you remain the controller responsible for the lawfulness of your own processing, the same as with a GPUwerk Spark.
Who can see my prompts on Azure OpenAI?
Per Microsoft's data, privacy, and security documentation, your prompts and completions are not available to other customers or to OpenAI, and are not used to train the underlying models. Microsoft does log and evaluate prompts for harmful content, and for standard (non-modified) abuse monitoring, a sample of flagged prompts can be stored and read by authorized Microsoft employees, located in the European Economic Area for EEA deployments, through point wise queries using request IDs and Secure Access Workstations. On a Spark, GPUwerk's published terms state GPUwerk hosts the machine but does not access, read, copy, index, or analyse the content of your instance, and there is no automated content scanning at all.
What happens to my data on a Spark when I stop or terminate the instance?
Terminating deletes your container and its /workspace volume from the node; a node is sanitised before it is offered to another customer. GPUwerk keeps only a periodic recovery copy of /workspace for hardware failures, not a live backup: while running or stopped, the workspace itself is on the node. Stopping keeps the workspace in place at 75% of the running hourly rate. Keep your own backup regardless.
Is a DGX Spark cheaper than Azure OpenAI?
It depends entirely on token volume and how continuously you can keep a Spark busy. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, but that Spark figure assumes the node sustains 256 concurrent requests the entire time. A lightly loaded Spark, paid for by the hour whether it is busy or not, can end up costing more per generated token than Azure's per-token price.
Can I switch between Azure OpenAI and a Spark without rewriting my application?
Mostly. Both expose an OpenAI-compatible chat completions API, so client code that only calls that endpoint shape ports with a base URL and API key change. What does not port is anything tuned to a specific model's behaviour, output format, or context window, and Azure-specific features like the Assistants API, stored completions, or content-filter configuration have no Spark equivalent because there is no managed platform layer, you would build any of that yourself.
Every Microsoft figure and quote on this page comes from pages GPUwerk fetched directly on 7 September 2026: learn.microsoft.com/.../openai/data-privacy for data location and abuse monitoring, and azure.microsoft.com/pricing/details/cognitive-services/openai-service for pricing. Microsoft does not publish a tokens/second figure for Azure OpenAI models anywhere GPUwerk could find; that row above reflects that gap rather than a guess.