Evaluating model quality before switching models
A newer model, a smaller quantization, or a cheaper checkpoint can look like a clear upgrade on a public leaderboard and still be worse on your actual traffic. Public benchmarks measure general capability; they don't know what your prompts look like, what your users expect, or which failure modes cost you the most. This page is about building a check against your own workload before a swap goes live, not after.
Build a holdout set from your own traffic
A holdout set is a fixed collection of real prompts, ideally pulled from actual usage logs, that neither model has been tuned against, held aside specifically for comparing candidates. It needs to look like what you actually serve, not a generic benchmark:
- Sample from real requests, not synthetic examples you write to sound representative. Actual usage surfaces edge cases and phrasing quirks a hand-written test set won't.
- Weight it toward what matters most, not an even spread. If 80% of your traffic is one task type, a holdout set that's evenly split across five task types will hide a regression in the task you actually depend on.
- Include known-hard cases deliberately, prompts that have previously tripped up a model, ambiguous instructions, long context, edge-case formatting. These are where quality differences between models actually show up; easy prompts tend to look fine on almost anything.
- Keep it fixed across comparisons. A holdout set that changes between evaluating candidate A and candidate B isn't comparing the models, it's comparing the models against different tests.
Strip anything sensitive from the set before it's used for repeated evaluation runs, the same way you'd handle any other stored copy of user data.
A/B approaches: running both models against real traffic
A holdout set catches obvious regressions before deployment; an A/B comparison catches what only shows up under real, live conditions. With two models running behind a gateway like LiteLLM, split a portion of live traffic to the candidate model and compare outcomes against the incumbent over the same window:
- Start with a small split, a low percentage of traffic to the candidate, so a real regression affects a small fraction of users rather than everyone at once.
- Compare outcomes that matter to your product, not just raw model output. Response acceptance rate, retry rate, explicit user feedback if you collect it, or downstream task success are more informative than eyeballing sample completions.
- Run it long enough to see the tail, not just the average case. A model that's fine on typical requests and bad on the 5% that are long, ambiguous, or adversarial won't show a problem in an average-latency or average-quality metric.
A/B testing needs live traffic and a gateway that can route a percentage split, so it's a heavier setup than a holdout run; use it once the holdout set has already ruled out an obvious regression, not as the first check.
Regression checks: what not to lose in the swap
Improvement on average is not the same as no regressions. A candidate model can raise your headline quality metric while getting meaningfully worse on a specific case your current model handles well, and an average masks that. Run both models against the same holdout set and diff the results per-example, not just the aggregate score:
- Flag any example where the candidate's answer is clearly worse than the incumbent's, even if the aggregate score improved.
- Pay particular attention to instruction-following and formatting regressions, since these often break downstream parsing even when the substance of the answer is fine.
- Decide in advance how many regressions, and of what severity, are acceptable for the switch to proceed. Without that threshold set beforehand, it's easy to rationalize a bad result after the fact because the average looks good.
Where quantization changes the answer
Switching models often means switching quantization too, a 4-bit version of a new model against an FP8 version of the old one, and it's worth evaluating those as separate variables where you can. A quality drop after a swap might be the new model itself, or it might be that you compared it at a lower precision than what you were running before. If you have the memory budget, run the same holdout set against the candidate at more than one quantization level before concluding the model itself is the problem.
Rolling it out once evaluation clears
Passing evaluation is the gate before a swap, not the swap itself. Once a candidate has cleared the holdout set and, if you ran one, an A/B window without unacceptable regressions, the actual cutover is a separate operational question: how you move traffic without an outage. The zero-downtime model updates guide covers the restart-window and second-Spark cutover options for the mechanics of that step.
Practical checklist
- Build a holdout set from real traffic, weighted toward what you actually serve and including known-hard cases.
- Run a small-percentage A/B split on live traffic once the holdout set clears, comparing outcomes that matter to your product.
- Diff results per-example, not just on aggregate score, and set an acceptable-regression threshold before you look at results.
- Evaluate quantization separately from model choice when both are changing at once.
- Treat evaluation as the gate before rollout, then use a deliberate cutover plan for the swap itself.
See the zero-downtime model updates guide for the cutover mechanics, or a first engagement if you want help designing an evaluation set for a specific workload.