Best GPU for running open-source LLMs privately: how to think about the choice
There's no single best GPU for this, and any list that claims otherwise skipped the part where your workload determines the answer. What matters is two numbers, memory capacity and memory bandwidth, and one question, are you the only person using it. Get those three things straight and the hardware choice mostly picks itself.
Capacity and bandwidth pull in opposite directions
A model has to fit entirely in memory before it can run at all, so capacity decides which models are even on the table. Bandwidth decides how fast the ones that fit actually generate text, because token generation reads the active weights out of memory once per token. The arithmetic is simple enough to do before you buy anything: tokens per second is roughly memory bandwidth divided by the bytes of active parameters read per token. We walk through this in more detail, with real measured numbers rather than theoretical ones, in our benchmarks page.
The tension is that the hardware with the most memory bandwidth (a high-end datacenter GPU) usually ships with less total memory than you'd want for the largest useful models, and the hardware with the most memory capacity per dollar tends to have comparatively modest bandwidth. There's no configuration that gives you both without paying datacenter prices, so the real decision is which constraint bites harder for your specific use case: do you need to fit a big model at all, or do you need an already-fitting model to feel fast under one user typing.
Single-user chat and shared serving are different engineering problems
If you're the only person using the model, latency per response is what you feel, and that's a bandwidth-bound, mostly-idle workload: the GPU sits waiting for you to read the last answer and type the next question, and when it does work, it's working on one request. That workload wants low-latency single-stream throughput more than raw compute headroom.
If the model is serving a team, a support queue, or a batch pipeline, the picture changes completely. Multiple requests can share the same forward pass through the model, which is why serving frameworks batch requests together, and aggregate throughput across many concurrent users climbs well past what a single stream achieves, because the cost of reading weights out of memory gets amortized across everyone's requests at once rather than paid once per person. A machine that feels merely adequate for one user in a chat window can serve a much larger number of people at once, because the bottleneck that limits a single conversation isn't the bottleneck that limits a shared endpoint.
This is the single most common hardware-sizing mistake we see: someone benchmarks a card with themselves as the only user, decides it's too slow, and rules it out for a team deployment where the real number that matters, aggregate throughput under concurrency, would have told a different story.
Where each class of hardware actually fits
A DGX Spark's 128GB of unified memory (see our models guide for what fits at that capacity) is built around the capacity side of the tradeoff: it holds 100–120B-class models whole, at a memory bandwidth around 273GB/s that's modest by datacenter standards but is enough, paired with a mixture-of-experts model's small active-parameter count, to feel genuinely responsive. It's a strong fit for a team that wants a large model resident locally without renting datacenter-card time.
A consumer GPU like an RTX 5090 sits on the opposite end: excellent bandwidth and raw compute for its price, but VRAM measured in tens of gigabytes rather than over a hundred, which rules out the largest open models outright regardless of how fast the card could theoretically run them. It's the right call when the model you actually need is small enough to fit, and you want the fastest single-stream experience money can buy at that size. We compare the two directly, with real numbers, in DGX Spark vs RTX 5090.
A datacenter card like an H100 gives up neither capacity nor bandwidth, which is exactly why it costs what it does. It's the right answer when you need both the largest models and datacenter-grade throughput under heavy concurrency, and the budget matches that requirement. Most teams evaluating private LLM hosting don't actually need this tier; see DGX Spark vs H100 for where the line falls in practice.
A Mac Studio with a lot of unified memory is a legitimate alternative for the same reason a Spark is: high capacity, and enough bandwidth to run MoE models reasonably. It tends to appeal to individual developers already inside the Apple ecosystem more than teams standing up a shared inference endpoint, partly because the software stack for serving to multiple concurrent users is less mature there than on CUDA. See DGX Spark vs Mac Studio for the direct comparison.
How to actually decide
Start from the model, not the hardware. Figure out the smallest model that does your task well enough, because that determines the capacity floor. Then figure out whether you're serving one person or many, because that determines whether bandwidth-per-user or aggregate-throughput-under-concurrency is the number to optimize. Only after both of those are pinned down does it make sense to compare specific cards, at which point the comparisons above will tell you more than a generic "best GPU for LLMs" ranking ever could, because those rankings are answering a question that doesn't have one answer.