Dedicated vs shared GPU: what's the difference
A shared GPU serves multiple tenants' workloads concurrently, either by time-slicing compute between them or by partitioning the GPU's resources across tenants at once. A dedicated GPU is allocated entirely to one tenant for the duration of their session, with no other workload running on it at the same time. The difference shows up as performance variance and, for AI inference specifically, as a question of memory isolation between tenants.
How sharing actually works
Cloud GPU providers share a physical GPU across customers in a few ways: time-slicing, where each tenant's process gets a scheduled turn on the same hardware; MIG (Multi-Instance GPU) partitioning on NVIDIA hardware that splits one physical GPU into several fixed hardware-isolated instances; or software-level multiplexing, where a hypervisor or driver layer schedules access. Each approach lets a provider sell fractional capacity, which is cheaper per unit than dedicating a whole GPU, but it also means your workload's performance depends partly on what other tenants are doing at the same moment.
Noisy-neighbor variance
The practical symptom of sharing is variance: the same request can take noticeably different amounts of time depending on what else is running on the GPU when it arrives. For interactive workloads, like a chat interface where users expect a steady response rate, this shows up as inconsistent latency that has nothing to do with your own traffic pattern. A dedicated GPU removes this specific problem, because no other tenant's process ever competes for the same compute or memory bandwidth; whatever variance remains is caused by your own workload, which is much easier to reason about and tune for.
The memory isolation question
For AI inference specifically, sharing raises a sharper question than latency: could another tenant's process, running on the same physical GPU, ever read data left in memory by your workload. Modern virtualization and isolation mechanisms, MIG partitioning in particular, are designed to prevent exactly this by giving each partition its own dedicated slice of memory, not just scheduled compute time. Whether a given implementation of that isolation is correctly configured and holds up in practice depends on the specific vendor's setup, and that's not something we've independently audited for any provider other than our own infrastructure, so treat any specific vendor's isolation claim as something to verify with their own security documentation rather than take on faith. A dedicated GPU sidesteps the question entirely: there's no other tenant on the hardware, so there's nothing to isolate against.
When shared is the right trade-off
Shared GPUs aren't a worse choice in every case. For bursty, low-volume workloads, or ones without strict latency or data-isolation requirements, paying for fractional shared capacity is usually cheaper than reserving a whole dedicated GPU that sits idle between requests. Dedicated hardware earns its cost when performance consistency matters, when the workload runs close to continuously, or when a specific data-isolation posture is a requirement rather than a nice-to-have.
How to tell which one you're actually getting
Pricing structure is usually the clearest signal. Per-token or per-request pricing on a shared model implies fractional, time-sliced access to hardware you never fully occupy. Per-hour or per-instance pricing for a named unit of hardware, billed whether or not you're actively sending requests, implies the whole unit is reserved for you. It's also worth checking whether a provider's documentation names the isolation mechanism explicitly (MIG, a dedicated instance type, bare-metal allocation) rather than just using the word "dedicated" in marketing copy without backing it up.
Where GPUwerk fits
A GPUwerk Spark is a dedicated DGX Spark: one whole unit with 128GB of unified memory, allocated to a single customer for as long as it's rented, starting at $0.79/hour from EU-Central. No time-slicing, no partitioning, no other tenant on the same silicon. Details on the DGX Spark hardware page.