Private AI/vs Azure AI Foundry
Comparison

Private LLM hosting vs Azure AI Foundry

By Samuel Seidel · Published September 9, 2026 · 9 min read

We rent DGX Sparks, so weigh that against everything below. Azure AI Foundry is Microsoft's platform for building and deploying AI applications, model catalogue, evaluation, and agent tooling, layered across models from OpenAI and other vendors. Microsoft has reorganised this part of Azure more than once, so we're deliberately keeping the Foundry-specific claims here general rather than guessing at product details we haven't verified against Microsoft's current documentation. If your interest is specifically GPT-model access through Azure, read our Azure OpenAI comparison instead, it goes deeper on that exact question.

Side by side

Azure AI FoundryDedicated DGX Spark (GPUwerk)
Where data is processed Depends on the model deployment and region you configure within Foundry. Microsoft's documentation for the specific model family in question, not a platform-wide statement, is the source to check, since options and defaults vary by which model and deployment type you pick. EU-Central (Prague), on one dedicated machine. No deployment configuration to get wrong, because there's only the one machine and one region.
Who can see your prompts Governed by Microsoft's platform-wide data, privacy, and security documentation, with specifics that can differ by model vendor within the catalogue. See our Azure OpenAI comparison for the exact abuse-monitoring language Microsoft publishes for its OpenAI-model deployments; GPUwerk did not verify whether that same language applies unchanged to every other model in Foundry's catalogue. Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." No abuse-monitoring or content-scanning layer exists at all.
Model choice A catalogue spanning closed frontier models (GPT-5.1 and similar) and a growing set of open-weight models from third-party vendors, plus Microsoft's own model families, all behind one platform. Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else that fits in 128 GB. If a model in Foundry's catalogue has open weights, the same model may run directly on a Spark; closed models in the catalogue do not.
Speed and concurrency Throughput depends on the specific model and deployment tier chosen; Foundry pricing and quota are generally expressed in tokens and requests per minute rather than a published tok/s figure. A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page.
Pricing model Per token or per deployment unit, varying by model vendor and tier within Foundry. GPUwerk did not fetch a current figure for this page; check learn.microsoft.com's Foundry pricing documentation for your specific model. Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of model. USD, tax extra, per pricing.
Worked cost example See the full arithmetic below the table.
Contracts and DPA Governed by the Microsoft Products and Services Data Protection Addendum, the same document that covers Azure OpenAI. Enterprise agreements are typical at volume. A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads.
What you operate yourself Nothing at the infrastructure layer; Microsoft manages capacity and serving. You manage prompts, application code, and evaluation/agent pipelines you build on top of Foundry's tooling. Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours.

The worked cost example

Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:

Azure AI Foundry. Pricing depends on which model and deployment tier you choose from the catalogue, and Foundry's own pricing pages present that through model-specific tables that GPUwerk did not fetch and verify for this page. Run your own model and volume through Microsoft's current Foundry pricing documentation for a number you can trust.

Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.

That Spark figure only holds under the assumption stated: the node busy at 256 concurrent requests for the full 161 hours, back to back. Real traffic arrives in bursts, so a Spark billed by the hour that sits half-idle can lose ground fast against a per-token price with no idle charge. Put a current Foundry number for your chosen model next to $127.19 and you have the real comparison for your own volume.

Migration path

For any Foundry model served through an OpenAI-compatible chat completions endpoint, a Spark serving a model through vLLM exposes the same shape, so an application that only calls that endpoint moves with a base URL and API key change. Put LiteLLM in front of the Spark and you get a drop-in OpenAI-compatible gateway with request logging, key management, and rate limits, useful for routing between a Foundry-hosted closed model and a Spark-hosted open model behind one interface. Foundry's platform-specific tooling, evaluation pipelines, agent orchestration, model routing built into the console, has no Spark equivalent, you'd stand up an alternative yourself or accept you're comparing infrastructure rather than a platform.

When Azure AI Foundry is the right choice

When a dedicated Spark is the right choice

FAQ

What is Azure AI Foundry?

Azure AI Foundry is Microsoft's platform for building and deploying AI applications on Azure, covering model deployment, evaluation, and agent tooling across a catalogue of models from OpenAI and other vendors. Microsoft has restructured and renamed this part of Azure more than once; for the specifics of any single model family it hosts, its own current documentation is the source of truth, not this page.

How is this different from your Azure OpenAI comparison?

Azure OpenAI is the OpenAI model deployment inside Microsoft's Azure platform; Foundry is the broader platform surface around it, model catalogue, evaluation tooling, and agent building blocks, that has absorbed and rebranded what used to be called Azure AI Studio. If your interest is specifically GPT-model access, data handling, and pricing, our dedicated Azure OpenAI comparison covers that in more depth and is what this page defers to for those specifics.

Can I run Foundry's models on a DGX Spark?

Closed models in Foundry's catalogue, GPT-5.1 and similar, run only on Microsoft's infrastructure. Foundry's catalogue also lists a number of open-weight models; if a model you'd use through Foundry has open weights, you may be able to run that same model directly on a Spark instead, at flat hourly cost rather than per token. Check the specific model's licence and weight availability before assuming portability.

Who controls where my data goes on Azure AI Foundry?

Data handling depends on the specific model deployment and region you configure within Foundry, and Microsoft's own documentation for the model family in question is the governing source, not a general platform-wide guarantee. A dedicated Spark sidesteps that configuration question entirely: one machine, one region, EU-Central, no deployment options to get wrong.

Is a DGX Spark cheaper than Azure AI Foundry?

It depends on which model and deployment type you'd use in Foundry, since pricing varies by model vendor and tier, and on how continuously you can keep a Spark busy. GPUwerk did not fetch a specific current Foundry price for this page; check Microsoft's own Foundry pricing documentation for your own model. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time.

Can I switch between Azure AI Foundry and a Spark without rewriting my application?

For models Foundry serves through an OpenAI-compatible chat completions endpoint, yes, mostly: client code that only calls that endpoint shape ports with a base URL and key change. Foundry's platform-specific tooling, evaluation pipelines, agent orchestration, model routing, has no equivalent on a Spark because there's no managed platform layer, you'd build any of that yourself or run it through a gateway like LiteLLM.

GPUwerk did not fetch or verify current Azure AI Foundry pricing, regional deployment options, or catalogue-wide data-handling terms for this page, Microsoft has changed this product's shape and name more than once and its own current documentation is the source to check. Where this page's claims overlap with Azure OpenAI specifically, see the sourced quotes on our Azure OpenAI comparison instead of treating this page as a source for that narrower product.

Related pages

First top-up: pay $10, get $20 in credit

Run your own numbers before you commit either way.

Deploy a dedicated DGX Spark in EU-Central and test your actual prompts against an open model before comparing quotes.

Deploy a Spark Read the benchmarks