Private LLM hosting vs the Mistral API
We rent DGX Sparks, so weigh that against everything below. Mistral is the one API in this series that doesn't need a data-residency argument made against it: it's a French company, headquartered in the EU, and its API already runs on EU infrastructure. If you picked Mistral because you wanted an EU vendor, you already made a reasonable call. The question worth asking next is narrower: Mistral's API is still shared multi-tenant infrastructure that a third party operates, billed per token, and a dedicated Spark is a machine only you touch. Whether that narrower difference is worth anything to you depends on what you're building.
Side by side
| Mistral API | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| Company and jurisdiction | Mistral AI, headquartered in Paris, France, an EU company under EU jurisdiction and GDPR directly. | PRINT IT! SE, operating from EU-Central (Czech Republic), no US parent. |
| Where data is processed | Mistral's own infrastructure; GPUwerk did not fetch Mistral's current documentation on which specific regions or data centers back the API, check Mistral's own trust and security pages for the current answer. | EU-Central (Prague), on one dedicated machine, no region choice needed because there's only the one machine. |
| Who can see your prompts | Mistral, as the operator of the API, has some level of access to run the service. GPUwerk did not fetch Mistral's current data-usage and training policy for this page; check Mistral's own privacy policy and API terms for what is retained, for how long, and whether staff can access it. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." No abuse-monitoring layer, because GPUwerk operates infrastructure, not a model API. |
| Model choice | Mistral's own model family: Mistral Large, Mistral Small, Codestral, and others as Mistral releases them, plus any Mistral open-weight models the API exposes. | Open-weight models you choose and run yourself: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, Mistral's own open-weight releases if you want them, and anything else that fits in 128 GB. You are not limited to one vendor's model family, but you also have to pick and manage the model instead of calling an endpoint. |
| Speed and concurrency | Mistral's API pages describe rate limits in requests and tokens per minute by tier; GPUwerk did not fetch a current tokens/second figure for this page. | Measured, not estimated: gpt-oss-120b decodes at 33.5 tok/s single-stream and reaches 862.8 tok/s aggregate at 256 concurrent requests; Qwen3-Coder 30B-A3B AWQ decodes at 80.9 tok/s single-stream. Full methodology on our benchmarks page. |
| Pricing model | Per token, priced separately for input and output, varying by model. Check mistral.ai/pricing for current rates, GPUwerk did not fetch a current figure to avoid quoting a stale price. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop. One rate regardless of model or request volume, per pricing. |
| Contracts and DPA | Mistral publishes its own data processing terms for API customers; GPUwerk has not reviewed them for this page. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| What you operate yourself | Nothing at the infrastructure layer; Mistral manages capacity and model serving. You manage prompts and application code. | Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
Where the honest difference actually is
Mistral being an EU company matters for a specific, narrow reason: it puts the vendor under GDPR by default, with no adequacy decision, no Standard Contractual Clauses, and no argument about whether a US CLOUD Act request could reach your data through a US parent. That's a genuine advantage over a US-headquartered API, and it's the reason a lot of EU companies default to Mistral when they need a hosted model and don't want to build infrastructure. Nothing on this page disputes that.
What doesn't change is the shape of the trust relationship. Calling Mistral's API means Mistral's servers process your prompt, on infrastructure Mistral operates, alongside every other customer's traffic. That's true of any API, EU-based or not. Whatever Mistral's retention and access policy says, and GPUwerk hasn't fetched the current version to quote it here, there is still a third party in the loop by the nature of the product. A dedicated Spark removes that party entirely: the machine runs your container and nothing else's, and GPUwerk's published terms commit to not looking at what's on it.
If your bar is "an EU vendor under GDPR," Mistral clears it. If your bar is "no third party sees the prompt at all," only a machine nobody else touches clears that, whether it's a Spark you rent from GPUwerk or hardware you run yourself.
The worked cost example
Take a workload generating 500 million output tokens a month, the kind of volume a busy internal tool or document pipeline running continuously might produce.
Mistral API. GPUwerk did not fetch current per-token pricing for a specific Mistral model at the time of writing, and pricing varies by model tier. Run your own volume and model choice through mistral.ai/pricing for a number you can trust.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per the concurrency benchmark on our benchmarks page. That's 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency throughout. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That figure assumes constant saturation at 256 concurrent requests for the full 161 hours. Real traffic is bursty, and a Spark billed hourly whether busy or idle can end up costing more per generated token than a per-token API once utilisation drops. Put Mistral's per-token number for your model and volume next to $127.19 and you have the real comparison for your traffic pattern.
Migration path
Mistral's API and a Spark serving a model through vLLM both expose an OpenAI-compatible chat completions shape, so client code that only calls that endpoint moves with a base URL and API key change. What doesn't move is anything tuned to Mistral's specific model behaviour or output formatting. If you want the same model family on both sides, Mistral publishes open-weight releases that run on a Spark too, in which case the model itself is identical and only the operator changes.
When the Mistral API is the right choice
- You want an EU-headquartered vendor under GDPR by default and don't want to operate infrastructure yourself.
- Your token volume is spiky or low, so paying only for tokens generated beats paying for a machine that sits idle between requests.
- You want Mistral's own model family specifically, with Mistral managing capacity and scaling.
When a dedicated Spark is the right choice
- You want to remove the third party from the loop entirely, not just move it to a friendlier jurisdiction.
- Your traffic is steady enough to keep the machine busy for most of the hours you're paying for.
- You want the model itself, open weights, so you can move it to another machine or verify exactly what it is, rather than calling an endpoint whose weights you never see.
FAQ
Is Mistral's API GDPR compliant?
Mistral AI is a French company headquartered in the EU, which puts it under GDPR directly rather than through a US-adequacy mechanism. That is a real difference from US-headquartered providers. It does not by itself mean nobody at Mistral can see your prompts: check Mistral's own data processing agreement and privacy policy for what its staff and any sub-processors can access, and under what conditions.
Does Mistral train on my API prompts?
GPUwerk did not fetch Mistral's current data-usage terms for this page and will not state a training policy we have not verified. Check Mistral's own privacy policy and API terms for the current answer, and whether it differs between the free tier and paid API usage.
Is a DGX Spark more private than the Mistral API?
On the specific question of who else can read a given prompt, yes: a Spark is one dedicated machine under your own SSH keys, and GPUwerk's terms state GPUwerk does not access, read, copy, index, or analyse instance content. Mistral's API is still shared multi-tenant infrastructure that a third party, Mistral, operates and necessarily has some level of access to in order to run it, whatever their stated retention and training policy is.
Is a Spark cheaper than the Mistral API?
It depends on token volume and how continuously you can keep a Spark busy. See the worked example on this page; run your own volume against Mistral's published per-token rates for a direct comparison.
GPUwerk did not fetch Mistral's current pricing, data-usage policy, or infrastructure documentation for this page; every Mistral-specific claim above is either general public knowledge (headquarters, jurisdiction) or explicitly hedged. Check mistral.ai directly for current terms. Spark figures come from our benchmarks page and our pricing page.