Docs/Serving
Serving

llama.cpp on a DGX Spark: build, serve, when to pick it

By Samuel Seidel · Updated September 4, 2026 · 5 min read

If you're serving one user at a time, llama.cpp beats vLLM on this hardware, not by a little. This walks through building it for GB10, serving a model through its OpenAI-compatible endpoint, and the benchmark numbers behind that claim.

On this page
  1. Why llama.cpp over vLLM here
  2. Build with CUDA for GB10
  3. Serve a model
  4. Test the endpoint
  5. Speculative decoding with MTP
  6. Context length and next steps

Why llama.cpp over vLLM here

vLLM's continuous batching is built to serve many concurrent users well; llama.cpp is built to serve one user fast. On the DGX Spark's 273 GB/s of memory bandwidth, that difference shows up directly in single-stream decode. User eugr's community benchmarks on the llama.cpp GitHub discussions measured gpt-oss-120b MXFP4 at roughly 60.5 tokens/second decode and 1,956 tokens/second prefill under llama.cpp, against roughly 33.5 tokens/second decode under vLLM on the same model and hardware. Full figures and more reference points are on the benchmarks page.

That's roughly an 80% single-stream speedup for no cost beyond building from source. The trade you're making: llama.cpp's batching is weaker once several people share the box, which is where vLLM takes over.

The rule of thumb: one interactive user, or a coding assistant on your own machine, run llama.cpp. A team sharing one endpoint, run vLLM.

Build with CUDA for GB10

Steps below follow NVIDIA's own DGX Spark playbook for llama.cpp. Install the build dependencies:

sudo apt update
sudo apt install -y git clang cmake libcurl4-openssl-dev libssl-dev

git clone https://github.com/ggml-org/llama.cpp ~/llama.cpp
cd ~/llama.cpp

Configure CMake with CUDA and GB10's sm_121 architecture, then build just the server target:

cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON -DGGML_CURL=ON -DGGML_RPC=ON \
  -DCMAKE_CUDA_ARCHITECTURES=121a-real
cmake --build build --config Release --target llama-server -j

The build takes 5 to 10 minutes. llama-server lands in build/bin/ when it's done.

The architecture flag is the part people get wrong. GB10 is Blackwell at compute capability sm_121, not the datacenter Blackwell parts. A generic -DCMAKE_CUDA_ARCHITECTURES=90 (Hopper) or 120 (desktop Blackwell) copied from an unrelated guide will build without error and then either fail to use the GPU correctly or underperform. Use 121a-real.

Serve a model

llama.cpp loads GGUF checkpoints. Pass -hf and it pulls the model from Hugging Face into ~/.cache/huggingface/hub the first time and reuses it after. NVIDIA's playbook example is Qwen3.6-35B-A3B, which fits comfortably in 128 GB with room for a long context:

cd ~/llama.cpp/build
./bin/llama-server \
  -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL \
  --host 0.0.0.0 \
  --port 30000

By default llama-server tries to fit the model's full context window with room for 4 concurrent requests, adjusting automatically if it doesn't fit. Swap in any GGUF you like with --model /path/to/file.gguf instead of -hf; multi-modal checkpoints need --mmproj pointed at the matching projector file unless you used -hf, which fetches it automatically.

Test the endpoint

Large GGUFs take a minute or more to load. Wait for the health check before sending requests:

# wait for the server to be ready
timeout 900 bash -c 'until curl -sf http://127.0.0.1:30000/health > /dev/null 2>&1; do sleep 5; done'

# a real completion
curl -X POST http://127.0.0.1:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL",
    "messages": [{"role": "user", "content": "Explain KV cache in two sentences."}],
    "max_tokens": 200
  }'

The endpoint speaks the same /v1/chat/completions shape as vLLM, so anything built against an OpenAI client works unchanged: point Open WebUI, Continue.dev, or a LiteLLM gateway at http://<spark-host>:30000/v1.

Speculative decoding with MTP

Models that ship an MTP (multi-token prediction) draft head, like the Qwen3.6-35B-A3B-MTP checkpoint above, can use it for speculative decoding, which drafts several tokens ahead and verifies them in one pass:

./bin/llama-server \
  -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL \
  --host 0.0.0.0 \
  --port 30000 \
  --chat-template-kwargs '{"preserve_thinking": true}' \
  --spec-type draft-mtp \
  --spec-draft-n-max 3

Check the model card before turning this on: --spec-type draft-mtp only works with a checkpoint that actually ships MTP weights, and enabling it on one that doesn't will fail to start rather than silently fall back. The preserve_thinking flag above is unrelated to MTP; it keeps prior reasoning blocks in the conversation history, which matters for agentic workflows on Qwen's thinking models.

Context length and next steps

llama.cpp allocates the model's maximum supported context by default if memory allows, or set it explicitly with --ctx-size. For coding agents or anything agentic, NVIDIA's own guidance is a floor of 32,768 tokens and a preference for 100,000 or more, since tool calls and file contents eat context fast.

Remember that "if memory allows" is checked against the Spark's unified 128 GB, which the OS and SSH also live in; leave a few gigabytes of headroom rather than letting --ctx-size default to the model's absolute maximum on a box you're also relying on to stay reachable. Watch free -g and keep "available" comfortably above zero.

Once it's running, treat it like any other OpenAI-compatible endpoint: put it behind LiteLLM if more than one script or teammate will call it, or point a coding agent straight at it per the vLLM guide's integration notes, which apply here unchanged.

See alsovLLM: when you need to serve a team, not one user See alsoFull benchmark numbers across engines

Try it on real hardware, not a spec sheet.

Deploy a dedicated DGX Spark in EU-Central and build llama.cpp against the actual GB10 GPU.

Deploy a Spark