Model quality
DGX Spark/Evaluating model quality before switching models

Evaluating model quality before switching models

By Samuel Seidel · Published September 9, 2026

A newer model, a smaller quantization, or a cheaper checkpoint can look like a clear upgrade on a public leaderboard and still be worse on your actual traffic. Public benchmarks measure general capability; they don't know what your prompts look like, what your users expect, or which failure modes cost you the most. This page is about building a check against your own workload before a swap goes live, not after.

Build a holdout set from your own traffic

A holdout set is a fixed collection of real prompts, ideally pulled from actual usage logs, that neither model has been tuned against, held aside specifically for comparing candidates. It needs to look like what you actually serve, not a generic benchmark:

Strip anything sensitive from the set before it's used for repeated evaluation runs, the same way you'd handle any other stored copy of user data.

A/B approaches: running both models against real traffic

A holdout set catches obvious regressions before deployment; an A/B comparison catches what only shows up under real, live conditions. With two models running behind a gateway like LiteLLM, split a portion of live traffic to the candidate model and compare outcomes against the incumbent over the same window:

A/B testing needs live traffic and a gateway that can route a percentage split, so it's a heavier setup than a holdout run; use it once the holdout set has already ruled out an obvious regression, not as the first check.

Regression checks: what not to lose in the swap

Improvement on average is not the same as no regressions. A candidate model can raise your headline quality metric while getting meaningfully worse on a specific case your current model handles well, and an average masks that. Run both models against the same holdout set and diff the results per-example, not just the aggregate score:

Where quantization changes the answer

Switching models often means switching quantization too, a 4-bit version of a new model against an FP8 version of the old one, and it's worth evaluating those as separate variables where you can. A quality drop after a swap might be the new model itself, or it might be that you compared it at a lower precision than what you were running before. If you have the memory budget, run the same holdout set against the candidate at more than one quantization level before concluding the model itself is the problem.

Rolling it out once evaluation clears

Passing evaluation is the gate before a swap, not the swap itself. Once a candidate has cleared the holdout set and, if you ran one, an A/B window without unacceptable regressions, the actual cutover is a separate operational question: how you move traffic without an outage. The zero-downtime model updates guide covers the restart-window and second-Spark cutover options for the mechanics of that step.

Practical checklist

See the zero-downtime model updates guide for the cutover mechanics, or a first engagement if you want help designing an evaluation set for a specific workload.

First top-up: pay $10, get $20 in credit

Test the candidate on real hardware, side by side.

Deploy a second Spark and run your holdout set against both models before you switch.

Deploy a Spark Read the cutover guide