Decision guide
Blog/vLLM vs llama.cpp: choosing an inference engine
For AI assistants

vLLM vs llama.cpp: choosing an inference engine

By Samuel Seidel · September 9, 2026

GPUwerk's vLLM docs and llama.cpp docs each cover installing and running that engine in full. This page answers the question that comes before either of those: which one should you actually install. The short answer is that it comes down to how many people are going to hit the box at once, and everything else follows from that.

The one number that decides most of this

How many concurrent users you're serving is the single biggest factor in this choice, and GPUwerk's own measurements on a DGX Spark make the tradeoff concrete. gpt-oss-120b at MXFP4 decodes at 60.5 tok/s for a single user under llama.cpp, against 33.5 tok/s under vLLM for the same single user, per GPUwerk's measurement. For one person typing into a chat window, llama.cpp is meaningfully faster on identical hardware and the identical model.

That ranking flips under load. The same gpt-oss-120b under vLLM scales to 862.8 tok/s aggregate once 256 requests share the box, because vLLM's continuous batching lets concurrent requests share forward passes instead of queueing behind each other one at a time. llama.cpp wasn't built around that batching model, so its advantage is concentrated in the single-stream case and doesn't carry over once you're serving a crowd. Neither number is wrong; they're answering different questions, and the question your deployment actually asks is what decides which engine wins.

Choose llama.cpp for single-user and simpler setups

llama.cpp is the better fit when you're serving one person at a time: a personal coding assistant, a single researcher's chat interface, a local tool nobody else is hitting concurrently. Its setup is comparatively simple, it supports a wide range of GGUF quantization levels so you can trade memory for speed with fine granularity, and it can run hybrid across CPU and GPU when a model doesn't fit entirely in accelerator memory, which vLLM doesn't offer in the same way. If you're prototyping, evaluating a model before committing to a production deployment, or building something genuinely single-user, start here.

Choose vLLM once more than one person needs the box

vLLM is the right engine once concurrency is a real requirement rather than a hypothetical one. Continuous batching is the whole reason it exists in this comparison: it services many requests in overlapping forward passes rather than one at a time, so aggregate throughput keeps climbing with load in a way llama.cpp's architecture doesn't match. Qwen3-Coder 30B-A3B illustrates this well, an MoE model that decodes at 80.9 tok/s for a single user under vLLM and keeps scaling to 3,235 tok/s aggregate at 256 concurrent requests, per GPUwerk's measurement. If you're building anything with more than one simultaneous user, an internal team tool, a customer-facing endpoint, an API other services call, vLLM's batching is the feature that makes the deployment viable at all, not just faster.

The setup cost is real, though: GPUwerk's vLLM doc notes one sharp edge that costs people an afternoon on GB10 hardware specifically, worth reading before you start rather than after you're stuck.

Format flexibility is a second, smaller factor

Beyond concurrency, the two engines differ in how many quantization formats they'll load and how forgiving they are about hardware they run on. llama.cpp supports a wide span of GGUF quant levels and can split a model across CPU and GPU memory when the whole thing doesn't fit in accelerator memory alone, which matters if you're working with a model close to the edge of what a single Spark's 128 GB can hold. vLLM is narrower on format support, generally AWQ, MXFP4, or NVFP4, but that narrower surface is also what its continuous batching engine is built and optimized around, so the tradeoff isn't a flaw so much as the cost of the concurrency machinery. If you're already choosing vLLM for its batching, the format constraint is one you accept as part of that choice rather than a separate downside to weigh.

This is also why the two engines aren't a drop-in swap for each other. A GGUF checkpoint tuned for llama.cpp doesn't load into vLLM, and an AWQ or NVFP4 checkpoint built for vLLM doesn't load into llama.cpp, so switching engines partway through a project means requantizing, not just changing a command-line flag.

A short framework

Ask: will more than one person or process be sending requests to this deployment at overlapping times? If no, llama.cpp gives you the faster single-user experience and the simpler path to a working setup, plus CPU/GPU hybrid flexibility if memory is tight. If yes, vLLM's continuous batching is what makes serving a crowd practical, and the aggregate throughput number, not the single-stream one, is what you should be optimizing for.

It's fine to start on llama.cpp to validate a model works the way you want, then move the production deployment to vLLM once concurrency becomes real. The two engines use different quantization format families, GGUF against AWQ, MXFP4, or NVFP4, so budget time to requantize rather than expecting the same checkpoint to carry over; GPUwerk's quantization decision guide covers picking the right format once you've settled on the engine.

Related pages

Test both engines on the same hardware.

Deploy a Spark hourly and measure single-user and concurrent throughput for your own model before committing.

Deploy a Spark vLLM docs