Serving
DGX Spark/llama.cpp vs vLLM vs SGLang on a DGX Spark: flags and startup

llama.cpp vs vLLM vs SGLang on a DGX Spark: flags and startup

By Samuel Seidel · Published September 9, 2026

The decision-guide version of this question covers the tradeoff at a high level: llama.cpp for one fast user, vLLM once several people share the box. This page is the hands-on companion, the actual commands, flags, and startup output for all three engines on GB10, including SGLang, so you can go from "which engine" to a running server without re-deriving the syntax from each project's own docs.

llama.cpp: build once, serve GGUF

llama.cpp needs building from source on GB10; there's no prebuilt wheel that targets sm_121 correctly. Following NVIDIA's own DGX Spark playbook:

git clone https://github.com/ggml-org/llama.cpp ~/llama.cpp && cd ~/llama.cpp
cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON -DGGML_CURL=ON -DGGML_RPC=ON \
  -DCMAKE_CUDA_ARCHITECTURES=121a-real
cmake --build build --config Release --target llama-server -j

The architecture flag is the detail people copy wrong: GB10 is Blackwell at sm_121, not the datacenter Blackwell parts, so a generic 90 (Hopper) or 120 (desktop Blackwell) from an unrelated guide will build clean and then underperform or miss the GPU path entirely. Use 121a-real. The build takes 5 to 10 minutes; llama-server lands in build/bin/.

Serving a GGUF checkpoint with -hf pulls it from Hugging Face on first run and caches it under ~/.cache/huggingface/hub:

./build/bin/llama-server \
  -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL \
  --host 0.0.0.0 \
  --port 30000

By default llama-server tries to fit the model's full context window with room for 4 concurrent requests, backing off automatically if that doesn't fit; set --ctx-size explicitly if you want a firm ceiling. Multi-modal checkpoints need --mmproj pointed at the matching projector unless -hf fetched it for you. See the llama.cpp docs page for the full walkthrough, including speculative decoding with an MTP draft head via --spec-type draft-mtp.

vLLM: container, not pip install

pip install vllm in a fresh venv on GB10 can look like it succeeded and still leave you with a server that won't start, or one that silently runs on CPU. The aarch64 PyPI resolution has historically pulled CPU-only PyTorch, and even with CUDA-enabled PyTorch in place, FlashInfer's JIT compilation for sm_121 needs CUDA 12.9 or newer. Use the CUDA 13 container instead, per vLLM's own DGX Spark writeup:

docker run -d --name vllm \
  --gpus all --ipc=host --restart unless-stopped \
  -p 8000:8000 \
  -e HF_TOKEN="$HF_TOKEN" \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:cu130-nightly \
  vllm serve Qwen/Qwen3.8-27B-FP8 \
    --gpu-memory-utilization 0.85 \
    --max-model-len 32768 \
    --max-num-seqs 4

--ipc=host gives the container the host's shared memory segment, without which larger models fail on worker startup. Pin the cu130-nightly tag to a specific release once you've validated a build; it moves, and a restart weeks later can silently pick up a different one. --gpu-memory-utilization is a fraction of total memory, not free memory, so a value that worked yesterday can fail today if something else is resident, per the vLLM memory-math guide. --max-num-seqs around 4 favors single-user latency; raise it once you're serving a team and throughput across users matters more than any one person's speed.

SGLang: the third option, less charted on this hardware

SGLang targets the same territory as vLLM, high-throughput serving with continuous batching and its own RadixAttention prefix-caching scheme, and it's worth knowing about even though GPUwerk's own published benchmarks don't yet cover it on GB10 the way they do vLLM and llama.cpp. Its Docker path looks structurally similar to vLLM's:

docker run -d --name sglang \
  --gpus all --ipc=host --restart unless-stopped \
  -p 30000:30000 \
  -e HF_TOKEN="$HF_TOKEN" \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  lmsysorg/sglang:latest \
  python3 -m sglang.launch_server \
    --model-path Qwen/Qwen3.8-27B-FP8 \
    --host 0.0.0.0 --port 30000 \
    --mem-fraction-static 0.85

--mem-fraction-static is SGLang's equivalent of vLLM's --gpu-memory-utilization, and the same caution applies: check what's already resident before trusting a fraction that worked on an empty box. Blackwell and sm_121 support in SGLang's release cadence moves independently of vLLM's, so confirm the image tag you pull actually targets GB10 correctly before relying on it; if in doubt, verify with a real completion request the way the checks below show, rather than trusting a clean startup log alone.

Where SGLang earns a look over vLLM: RadixAttention's prefix caching can help workloads with heavy shared-prefix traffic, like a coding agent replaying the same system prompt across many requests, more than vLLM's own prefix caching does in some published community comparisons. That's a workload-shape question, not a hardware one, and it's worth testing against your own traffic pattern rather than assuming either engine wins in general.

Verifying any of the three the same way

All three speak the same /v1/chat/completions shape, so the same curl check works regardless of which you picked:

# confirm it's loaded and answering
curl http://localhost:<port>/v1/models

curl http://localhost:<port>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<served-model-name>",
    "messages": [{"role": "user", "content": "Explain KV cache in two sentences."}]
  }'

Anything built against an OpenAI-compatible client, a coding agent, Open WebUI, or a LiteLLM gateway, works unchanged against whichever port you picked. That portability is exactly why the choice between engines is safe to revisit later: swapping the backend behind a gateway doesn't touch anything downstream of it.

Which one, in practice

See the decision guide for the throughput-vs-latency framework behind this choice, or a first engagement if you want help benchmarking a specific model across engines before committing.

First top-up: pay $10, get $20 in credit

Test all three engines on the same hardware.

A dedicated Spark, root access, and no shared tenants to skew the numbers.

Deploy a Spark Read the vLLM guide