Benchmarking your own workload on a rented Spark
A vendor's tokens-per-second number was measured on a benchmark suite, not your prompts. Our own benchmarks page is upfront about this: we publish specific numbers for specific models, quantizations and engines because a single "tok/s" figure without those three details is close to meaningless. Here's how to get a number that actually describes your workload, on hardware you can rent by the hour to do it.
Decide what you're actually measuring
"Fast" means different things depending on what you're building. A chat interface cares about time to first token and steady-state decode speed for a single user. A batch summarization job cares about total throughput across many concurrent requests, and barely notices latency on any one of them. A coding agent that reads a large repository into context cares about prefill speed, since it processes long prompts before generating short diffs. Pick the metric that matches your product before you start measuring, because optimizing for the wrong one will lead you to the wrong model or the wrong quantization.
The two numbers worth separating are prefill (tokens per second processing the input, compute-bound) and decode (tokens per second generating output, bandwidth-bound on a Spark's 273 GB/s of memory bandwidth). Most public benchmarks report decode only, at short prompt lengths, which understates cost for anyone running long-context workloads.
Reproduce your actual prompt shape, not a synthetic one
Grab 10 to 20 real inputs from your intended use case: actual document lengths, actual system prompt size, actual expected output length. If you're building a RAG pipeline that stuffs 8,000 tokens of retrieved context into every call, benchmark with 8,000-token prompts, not the 200-token prompts a generic benchmark script defaults to. Prefill time scales with prompt length in a way decode time doesn't, and a model that looks fast on short synthetic prompts can feel slow in your actual product once real context length is added.
Both vLLM and llama.cpp ship their own load-testing tools: vllm bench serve replays a prompt set against a running server and reports prefill, decode and end-to-end latency percentiles; llama.cpp's llama-bench does the same for GGUF models. Point either one at your own prompt file rather than the tool's default dataset.
Match concurrency to your real traffic pattern
A single-user throughput number and a ten-concurrent-user throughput number on the same hardware can differ substantially, because concurrent requests share the same memory bandwidth. If your deployment is one engineer testing a model interactively, benchmark single-stream. If it's an internal tool with a dozen simultaneous users, benchmark at that concurrency level, because the single-stream number will overstate what any one user actually experiences once the node is under real load. vLLM's guide covers the batching and continuous-batching settings that determine how a Spark handles concurrent requests in practice.
Test the quantization you'll actually run in production
A model benchmarked at FP16 and a model benchmarked at 4-bit (MXFP4, NVFP4, AWQ or GGUF-quantized) are, for throughput purposes, close to different models. Our quantization reference explains the formats and which engine each belongs to. If you're planning to run a model at 4-bit in production to fit it inside 128 GB, benchmark it at 4-bit. A benchmark run at full precision, on a different piece of hardware than you'll deploy on, tells you almost nothing about the number you'll actually get.
A minimal benchmarking checklist
Rent a Spark for the hour, at $0.79 for spark-1x or $1.79 for spark-2x on the on-demand rate. Load the exact model, quantization and engine you intend to run in production, following the vLLM guide or the llama.cpp guide. Assemble 10 to 20 prompts that match your real prompt length and content. Run them at the concurrency level your product actually expects, recording prefill tok/s, decode tok/s, and time to first token separately. Repeat the run two or three times, since the first request after a cold start is often slower due to CUDA graph compilation and KV-cache warmup, and you want the steady-state number, not the cold one.
None of this needs to take more than an hour, and it's the only way to know your real number rather than someone else's. We don't publish workload-specific numbers for you because there isn't a way to do that honestly: a number measured on our test prompts is still someone else's benchmark from your point of view.
What to do with the result
If decode speed comes back slower than expected, the two usual causes are a model too large for the active-parameter math to favor a Spark's bandwidth (dense models read every parameter per token; mixture-of-experts models read only the active subset, which is why a 120B MoE model can outrun a 49B dense one on the same hardware), or a quantization mismatch with the engine you chose. If prefill is the bottleneck instead, the fix is usually a shorter context window or prompt caching rather than a different model. The benchmarks page has worked examples of this active-versus-total-parameter math if you want to reason about it before you rent anything.
FAQ
Why don't published benchmark numbers transfer to my deployment?
A published number is tied to a specific engine, quantization, batch size and prompt length. Decode speed on a Spark is bandwidth-bound, so changing any of those inputs, especially prompt length and concurrency, changes the number you get, sometimes by a large margin.
What's the difference between prefill and decode speed, and why does it matter for benchmarking?
Prefill is processing the input prompt, which is compute-bound and reported in tokens per second across the whole prompt. Decode is generating output one token at a time, which is bandwidth-bound and what most people mean by tok/s. A workload with long inputs and short outputs, like summarization, is dominated by prefill time; a workload with short inputs and long outputs, like a coding agent, is dominated by decode. Benchmark whichever one your workload actually spends its time in.
How long does it take to benchmark a model properly on a rented node?
Loading a model, running a representative set of your own prompts at a few concurrency levels, and recording the numbers usually takes under an hour on a Spark, billed at the on-demand hourly rate. That's enough to catch the throughput and memory issues that matter before committing to a longer rental or hardware purchase.