What is a private LLM?
A private LLM is a large language model deployment where inference runs on infrastructure the deploying company controls or has exclusive access to, rather than on a shared multi-tenant API where requests from many customers are processed on the same pool of hardware. The distinction is about who else's workload touches the machine your prompt runs on, not about which model you're running.
The core distinction: exclusive access, not just contractual promises
Most AI vendors run multi-tenant infrastructure: one fleet of GPUs, shared across every customer sending requests to a given endpoint. Isolation between customers is enforced in software, at the request level, not at the hardware level. That's how OpenAI, Anthropic, and most cloud AI APIs work, and it's a reasonable model for most uses. A private LLM deployment is different in one specific way: the GPU (or GPUs) processing your requests are not processing anyone else's requests at the same time. That can mean hardware you own and run in your own data center, or a dedicated machine you rent from a provider, where the whole unit is allocated to you rather than time-sliced or shared.
Why "private" gets used loosely
Vendors sometimes label a tier of their shared API as "private," meaning something narrower: your inputs and outputs are excluded from training data, retained for a shorter period, or covered by a stricter data processing agreement. Those are real and useful guarantees, but they don't change where the computation happens. The model still runs on shared hardware, alongside other customers' traffic, under infrastructure the vendor operates and you never see. Calling that a "private" deployment conflates a data-handling policy with a physical isolation guarantee, and the two aren't the same claim. If a vendor's own documentation doesn't say the hardware is dedicated to you, assume it isn't.
What actually changes with hardware-level isolation
On a genuinely private deployment, there's no other tenant's process running on the same GPU at the same moment, so there's no shared queue determining how fast your requests get served, and no theoretical path for another customer's workload to affect yours through resource contention (see our explainer on dedicated vs shared GPUs for the isolation mechanics in more depth). It also usually means you can see and control what's actually running: which model, which quantization, which inference engine, and what logging (if any) happens on the box. On a shared API, those are implementation details the vendor doesn't expose.
What it doesn't automatically guarantee
Private, in this sense, is a statement about hardware exclusivity, not automatically about data sovereignty or jurisdiction. A dedicated machine can still be operated by a company subject to a foreign government's legal process, or physically located somewhere your data protection rules don't intend it to be. Those are separate questions worth checking independently; see our explainer on data sovereignty vs data residency for that distinction.
How to tell which kind you're actually being offered
The reliable signal isn't the word "private" itself; it's whether the vendor's own documentation states that the hardware serving your requests is not shared with other customers, and ideally names the specific mechanism (a dedicated instance, a reserved node, hardware you provision yourself). If a pricing page or FAQ describes "private" only in terms of data retention or training exclusion, without mentioning hardware allocation, assume the underlying compute is still multi-tenant unless told otherwise. It's a reasonable question to put directly to a vendor's sales or support team, and a vendor with a genuine hardware-isolation offering will usually answer it plainly, because it's a selling point rather than something to obscure.
Where GPUwerk fits
A GPUwerk Spark is a dedicated machine: one DGX Spark with 128GB of unified memory, allocated entirely to one customer, starting at $0.79/hour from EU-Central. Nobody else's inference runs on it while it's yours. More detail on the setup is on the private LLM hosting page.