llama.cpp vs vLLM vs SGLang on a DGX Spark: flags and startup
The decision-guide version of this question covers the tradeoff at a high level: llama.cpp for one fast user, vLLM once several people share the box. This page is the hands-on companion, the actual commands, flags, and startup output for all three engines on GB10, including SGLang, so you can go from "which engine" to a running server without re-deriving the syntax from each project's own docs.
llama.cpp: build once, serve GGUF
llama.cpp needs building from source on GB10; there's no prebuilt wheel that targets sm_121 correctly. Following NVIDIA's own DGX Spark playbook:
git clone https://github.com/ggml-org/llama.cpp ~/llama.cpp && cd ~/llama.cpp cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON -DGGML_CURL=ON -DGGML_RPC=ON \ -DCMAKE_CUDA_ARCHITECTURES=121a-real cmake --build build --config Release --target llama-server -j
The architecture flag is the detail people copy wrong: GB10 is Blackwell at sm_121, not the datacenter Blackwell parts, so a generic 90 (Hopper) or 120 (desktop Blackwell) from an unrelated guide will build clean and then underperform or miss the GPU path entirely. Use 121a-real. The build takes 5 to 10 minutes; llama-server lands in build/bin/.
Serving a GGUF checkpoint with -hf pulls it from Hugging Face on first run and caches it under ~/.cache/huggingface/hub:
./build/bin/llama-server \ -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL \ --host 0.0.0.0 \ --port 30000
By default llama-server tries to fit the model's full context window with room for 4 concurrent requests, backing off automatically if that doesn't fit; set --ctx-size explicitly if you want a firm ceiling. Multi-modal checkpoints need --mmproj pointed at the matching projector unless -hf fetched it for you. See the llama.cpp docs page for the full walkthrough, including speculative decoding with an MTP draft head via --spec-type draft-mtp.
vLLM: container, not pip install
pip install vllm in a fresh venv on GB10 can look like it succeeded and still leave you with a server that won't start, or one that silently runs on CPU. The aarch64 PyPI resolution has historically pulled CPU-only PyTorch, and even with CUDA-enabled PyTorch in place, FlashInfer's JIT compilation for sm_121 needs CUDA 12.9 or newer. Use the CUDA 13 container instead, per vLLM's own DGX Spark writeup:
docker run -d --name vllm \
--gpus all --ipc=host --restart unless-stopped \
-p 8000:8000 \
-e HF_TOKEN="$HF_TOKEN" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:cu130-nightly \
vllm serve Qwen/Qwen3.8-27B-FP8 \
--gpu-memory-utilization 0.85 \
--max-model-len 32768 \
--max-num-seqs 4
--ipc=host gives the container the host's shared memory segment, without which larger models fail on worker startup. Pin the cu130-nightly tag to a specific release once you've validated a build; it moves, and a restart weeks later can silently pick up a different one. --gpu-memory-utilization is a fraction of total memory, not free memory, so a value that worked yesterday can fail today if something else is resident, per the vLLM memory-math guide. --max-num-seqs around 4 favors single-user latency; raise it once you're serving a team and throughput across users matters more than any one person's speed.
SGLang: the third option, less charted on this hardware
SGLang targets the same territory as vLLM, high-throughput serving with continuous batching and its own RadixAttention prefix-caching scheme, and it's worth knowing about even though GPUwerk's own published benchmarks don't yet cover it on GB10 the way they do vLLM and llama.cpp. Its Docker path looks structurally similar to vLLM's:
docker run -d --name sglang \
--gpus all --ipc=host --restart unless-stopped \
-p 30000:30000 \
-e HF_TOKEN="$HF_TOKEN" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path Qwen/Qwen3.8-27B-FP8 \
--host 0.0.0.0 --port 30000 \
--mem-fraction-static 0.85
--mem-fraction-static is SGLang's equivalent of vLLM's --gpu-memory-utilization, and the same caution applies: check what's already resident before trusting a fraction that worked on an empty box. Blackwell and sm_121 support in SGLang's release cadence moves independently of vLLM's, so confirm the image tag you pull actually targets GB10 correctly before relying on it; if in doubt, verify with a real completion request the way the checks below show, rather than trusting a clean startup log alone.
Where SGLang earns a look over vLLM: RadixAttention's prefix caching can help workloads with heavy shared-prefix traffic, like a coding agent replaying the same system prompt across many requests, more than vLLM's own prefix caching does in some published community comparisons. That's a workload-shape question, not a hardware one, and it's worth testing against your own traffic pattern rather than assuming either engine wins in general.
Verifying any of the three the same way
All three speak the same /v1/chat/completions shape, so the same curl check works regardless of which you picked:
# confirm it's loaded and answering
curl http://localhost:<port>/v1/models
curl http://localhost:<port>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<served-model-name>",
"messages": [{"role": "user", "content": "Explain KV cache in two sentences."}]
}'
Anything built against an OpenAI-compatible client, a coding agent, Open WebUI, or a LiteLLM gateway, works unchanged against whichever port you picked. That portability is exactly why the choice between engines is safe to revisit later: swapping the backend behind a gateway doesn't touch anything downstream of it.
Which one, in practice
- One interactive user, or a coding agent on your own machine: llama.cpp. Community benchmarks on the llama.cpp GitHub discussions measured gpt-oss-120b MXFP4 at roughly 60.5 tokens/second decode under llama.cpp against roughly 33.5 tokens/second under vLLM on the same hardware, an 80% single-stream gain for the cost of a source build. See the benchmarks page for more reference points.
- A team sharing one endpoint, or you want the most mature memory-and-quantization tooling: vLLM. Its continuous batching is built for many concurrent users, and its DGX Spark install path is the most documented of the three.
- Heavy shared-prefix workloads, and you're willing to validate the GB10 build yourself: SGLang is worth a bench-off against vLLM on your actual traffic, but treat it as the less-charted option on this specific hardware until you've measured it.
See the decision guide for the throughput-vs-latency framework behind this choice, or a first engagement if you want help benchmarking a specific model across engines before committing.