Private AI/vs OpenAI API
Comparison

Private LLM hosting vs the OpenAI API

By Samuel Seidel · Published September 9, 2026 · 10 min read

We rent DGX Sparks, so weigh that against everything below. If your task needs a specific closed frontier model by name, GPT-5.1 or whatever OpenAI ships next, this page cannot get you there, a Spark only runs open-weight models. That's not a current gap we're working to close, it's the shape of the product. If an open model already meets your quality bar and your reason for looking elsewhere is where the data goes, a dedicated Spark in EU-Central gives you a single machine, assigned to you alone, instead of a shared API endpoint.

Side by side

OpenAI APIDedicated DGX Spark (GPUwerk)
Where data is processed OpenAI's own infrastructure, operated from the US. OpenAI publishes trust and compliance documentation describing its data handling and any regional options; GPUwerk did not independently verify current processing-location specifics for this page, check OpenAI's own documentation for what applies to your account type. EU-Central (Prague), on one dedicated machine. No shared endpoint, no other tenant, no region to pick because there's only the one machine.
Who can see your prompts OpenAI states in its own published API data usage policy that data sent through the API is not used to train its models by default, and describes retention and access controls in that same policy. That is OpenAI's statement about its own systems; verify current details at openai.com before relying on it. Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." There is no abuse-monitoring or logging layer at all, because GPUwerk operates infrastructure, not a model API.
Model choice Closed frontier models: GPT-5.1 and the rest of the current OpenAI lineup, plus reasoning models, image and audio endpoints, and platform features like the Assistants API and built-in tools. Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else you can fit in 128 GB. GPT-class closed models are not available on a Spark under any arrangement.
Speed and concurrency OpenAI does not publish a tokens/second figure for the API; usage limits are stated as tokens and requests per minute (TPM/RPM) tiers rather than throughput. A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Qwen3-Coder 30B-A3B AWQ decodes at 80.9 tok/s single-stream on our own fleet. Raw logs and methodology on our benchmarks page.
Pricing model Per token, billed separately for input and output, varying by model. See platform.openai.com/docs/pricing for current rates; GPUwerk did not fetch a specific figure for this page since prices shift with each model release. Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of model or how many requests you send. USD, tax extra, per pricing.
Worked cost example See the full arithmetic below the table.
Contracts and DPA Governed by OpenAI's own business terms and data processing addendum for API customers. Enterprise agreements are available at volume. A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads.
Lock-in and portability The API's chat completions shape is widely copied, so basic client code ports to other OpenAI-compatible endpoints, including a Spark's, with a base URL and key change. Prompts tuned to a specific closed model's behaviour, and platform-only features like the Assistants API, do not move with you. Open weights mean you can move the model file itself to another Spark, another GPU box, or your own hardware; the model is never OpenAI's or GPUwerk's to withhold. You lose OpenAI's managed platform layer and have to build any of that yourself.
What you operate yourself Nothing at the infrastructure layer; OpenAI manages capacity, scaling, and model serving. You manage prompts, application code, and your own use of rate limits. Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours.

The worked cost example

Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:

OpenAI API. Pricing is per token, billed separately for input and output, and varies by model. GPUwerk did not fetch and verify a current per-token price for a specific model at the time of writing, prices change with each release and a stale number here would be worse than none. Run your own model and volume through platform.openai.com/docs/pricing for a figure you can trust.

Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.

That Spark figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back, with no idle time. Real traffic arrives in bursts, so a Spark billed by the hour that sits half-idle waiting for requests can lose ground fast, you're paying $0.79 for every hour on the clock whether the GPU is generating tokens or not. OpenAI's per-token price doesn't have that problem, you pay only for tokens you actually generate, but you're also paying a per-token rate for a closed model rather than an hourly rate for a dedicated machine. Put a current OpenAI number next to $127.19 for your own volume and you have the real comparison; which one wins depends on your traffic pattern and whether the task in front of you needs a closed model at all.

Migration path

Both sides speak the same wire protocol at the client level. The OpenAI API's chat completions shape is what most tooling targets, and a Spark serving a model through vLLM exposes the same shape, so an application that only calls that endpoint moves with a base URL and API key change. Put LiteLLM in front of the Spark and you get a drop-in OpenAI-compatible gateway with request logging, key management, and rate limits, useful if you want to run both a closed model through OpenAI and an open model on a Spark behind one interface, routing by task. What doesn't move is anything tuned to one closed model's specific behaviour or output formatting, and OpenAI-only features like the Assistants API or built-in web search have no Spark equivalent because there's no managed platform layer underneath it.

When the OpenAI API is the right choice

When a dedicated Spark is the right choice

FAQ

Can I get GPT-5-class models on a DGX Spark?

No. GPT-5.1 and the rest of OpenAI's closed model family run only on OpenAI's own infrastructure, not on hardware you rent or own. A Spark runs open-weight models: gpt-oss-120b, Llama 3.3, Qwen3-Coder, DeepSeek, and similar. If your task requires a specific closed frontier model by name, the OpenAI API is the only place to get it, GPUwerk cannot offer that model on a Spark under any arrangement.

Does OpenAI train on my API data?

OpenAI's own published API data usage policy states that data submitted through the API is not used to train its models by default. That is OpenAI's policy statement, not something GPUwerk can verify independently; read it directly at openai.com's data usage page before relying on it for a compliance decision. On a Spark, GPUwerk's published DPA states GPUwerk hosts the machine but does not access, read, copy, index, or analyse the content of your instance, which is a structural fact about the product rather than a training policy.

Where does the OpenAI API process my data?

OpenAI is a US company and its standard API endpoints are US-operated; check OpenAI's own trust and compliance documentation for current details on any regional processing options. A dedicated Spark runs in one place: EU-Central, on a single machine assigned to you alone, so there's no processing-location question to ask about a shared multi-tenant system.

Is a DGX Spark cheaper than the OpenAI API?

It depends on token volume and how continuously you can keep a Spark busy. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time. A lightly loaded Spark, billed by the hour whether it's busy or not, can end up costing more per generated token than a per-token API.

Can I switch between the OpenAI API and a Spark without rewriting my application?

Mostly. A Spark served through vLLM exposes an OpenAI-compatible chat completions API, so client code that only calls that endpoint shape ports with a base URL and API key change. What doesn't port is anything tuned to a specific closed model's behaviour or output format, and OpenAI-only platform features like the Assistants API have no Spark equivalent because there's no managed platform layer, you'd build any of that yourself.

How does this compare to Azure OpenAI or Anthropic's API?

The core trade, closed frontier models on someone else's shared infrastructure versus open-weight models on a dedicated machine, is the same shape across OpenAI, Azure OpenAI, and Anthropic. See our separate comparisons against Azure OpenAI and the Anthropic API for the vendor-specific details on data location and prompt handling.

GPUwerk did not fetch a current per-token OpenAI API price for this page; prices vary by model and change often, so check platform.openai.com/docs/pricing directly for your own model and volume. The data-usage claim above reflects OpenAI's own published policy, not independent verification by GPUwerk. OpenAI does not publish a tokens/second figure for the API anywhere GPUwerk could find; that row above reflects that gap rather than a guess.

Related pages

First top-up: pay $10, get $20 in credit

Run your own numbers before you commit either way.

Deploy a dedicated DGX Spark in EU-Central and test your actual prompts against an open model before comparing quotes.

Deploy a Spark Read the benchmarks