Cost
DGX Spark/Prompt caching to cut inference cost

Prompt caching to cut inference cost

By Samuel Seidel · Published September 9, 2026

Every request to a self-hosted model pays for two phases: prefill, where the engine processes your input tokens and builds a KV cache entry for each of them, and decode, where it generates output one token at a time. If the same block of text opens request after request, a fixed system prompt, a long set of tool definitions, a document pasted into every question, reprocessing it from scratch each time is wasted prefill work. Prompt caching skips that work by reusing the KV cache from a previous request when the new one shares the same prefix.

What actually gets cached

The cache key is a prefix match on tokens, not a semantic match on meaning. If request A starts with the exact same sequence of tokens as request B, up to some point, the engine can reuse the attention keys and values it already computed for that shared portion and only run prefill on what comes after. Change one word in the system prompt, reorder the tool list, or vary a timestamp embedded near the top of the prompt, and the match breaks at that point, forward.

This is why prompt structure matters more than prompt caching itself. Put everything static first: system instructions, tool schemas, few-shot examples, a fixed knowledge-base excerpt. Put everything that changes per request last: the user's actual message, retrieved context that varies by query, a timestamp. A prompt built the other way around, variable content first, gets no benefit even if the engine supports caching, because nothing after token one matches between requests.

Where the savings show up

On a rented Spark billed by the hour, prompt caching doesn't reduce your bill directly, you're not paying per token the way you would on a hosted API. What it buys is throughput: less GPU time spent on prefill means more requests served per hour on the same box, or lower latency on each individual request since the cached portion doesn't wait in the compute queue. If you're deciding whether a workload fits on one Spark or needs the two-node cluster, prompt caching is one of the levers that can keep you on the smaller, cheaper option longer.

The effect is largest for workloads with a long, repeated static prefix and a short variable tail: a coding assistant with a large system prompt and short user turns, a support bot answering against the same policy document over and over, a batch job applying one instruction template to many inputs. It's smallest for workloads where every request is mostly unique text, there just isn't much prefix to reuse.

Enabling it in vLLM

vLLM supports automatic prefix caching, and it's worth turning on for almost any deployment with repeated prompt structure. Start the server with the flag enabled:

vllm serve /path/to/model \
  --enable-prefix-caching \
  --gpu-memory-utilization 0.85

With this on, vLLM keeps recently computed KV cache blocks around and matches new requests against them automatically, no client-side change required beyond structuring your prompts with static content first. The cache lives in the same memory pool as everything else on the box, so it competes with weights and per-request KV cache for space; see the memory management guide if you're running more than one model and need to budget for this alongside them. Cache blocks are evicted under memory pressure using an LRU policy, so a cache hit isn't guaranteed on a busy box, but it costs nothing to have enabled when it doesn't hit.

Check the hit rate rather than assuming caching is working. vLLM exposes prefix cache statistics through its Prometheus metrics endpoint (vllm:gpu_prefix_cache_hit_rate in recent versions); if it's near zero on a workload you expected to benefit, the likely cause is prompt structure, variable content bleeding into what should be a stable prefix, rather than a configuration problem.

Enabling it in llama.cpp

llama.cpp's server also reuses KV cache across requests when the new prompt shares a prefix with the previous one, and this is on by default in recent builds via slot-based context reuse; no flag is required for the single-conversation case. Where it needs more attention is concurrency: llama.cpp allocates a fixed number of parallel slots (--parallel N), each with its own KV cache, so caching is per-slot rather than shared across every in-flight request the way it can be with vLLM's block-level approach. If your workload rotates between a handful of fixed prompt templates, pinning related requests to the same slot (or keeping concurrency low enough that they land there naturally) keeps the reuse working; a high-concurrency workload with many distinct prefixes will thrash slots and see less benefit.

A worked example

Say a support tool sends every request with a 1,200-token system prompt covering tone, policies and three tool definitions, followed by a short user question averaging 40 tokens. Without caching, every request pays prefill on all 1,240 tokens. With prefix caching hitting on the shared 1,200-token block, only the 40-token tail needs fresh prefill, plus whatever the model generates in decode. Prefill is typically cheaper per token than decode, but at high request volume that 1,200-token block adds up, and skipping it repeatedly is where the aggregate saving comes from. The exact speedup depends on your model size, context length and concurrency, run your own measurement against your actual prompt rather than assuming a fixed multiplier.

Practical checklist

See cost forecasting for AI infrastructure for how caching factors into a broader hourly-cost model, or the glossary entry on KV cache for the underlying mechanism.

First top-up: pay $10, get $20 in credit

Cut prefill, not corners.

Deploy a dedicated Spark and turn on prefix caching in vLLM or llama.cpp from the console.

Deploy a Spark Read the vLLM guide