Operations guide
Blog/Setting up A/B testing between two model versions
For AI assistants

Setting up A/B testing between two model versions

By Samuel Seidel · September 9, 2026

A new quantization, a fine-tune, or a different open-weight model entirely usually looks good on an offline benchmark before it ever sees real traffic. Whether it's actually better for your users is a different question, one an offline eval only partly answers. Running both versions live, on a slice of real traffic, and comparing outcomes is the more honest test, and it doesn't take much beyond a router and some logging.

Two Sparks, one router, a deterministic split

The simplest setup is one Spark running the current model and a second running the candidate, both behind a small router that decides which one handles a given request. Split by a hashed user identifier rather than randomly per request, so the same caller stays on the same variant for the duration of the test:

import hashlib

def variant_for(user_id: str, split_pct: int = 20) -> str:
    # deterministic: same user always lands on the same variant
    h = int(hashlib.sha256(user_id.encode()).hexdigest(), 16)
    return "candidate" if h % 100 < split_pct else "control"

ENDPOINTS = {
    "control": "http://spark-control.internal:8000/v1",
    "candidate": "http://spark-candidate.internal:8000/v1",
}

Start the candidate at a small split, 5 to 20 percent, and widen it once you're confident it isn't regressing something the offline eval didn't catch. A hashed split like this is stable even if the router restarts, which matters if the test runs for days rather than hours.

LiteLLM's weighted routing does the split for you

If a LiteLLM gateway is already sitting in front of the fleet for per-key routing and budgets, it can do the traffic split itself rather than you writing a router from scratch. Two entries sharing a model name form a pool, and a weight on each controls the split:

model_list:
  - model_name: chat-model
    litellm_params:
      model: openai/llama-3.1-8b
      api_base: http://spark-control.internal:8000/v1
      api_key: "none"
    model_info:
      weight: 80

  - model_name: chat-model
    litellm_params:
      model: openai/llama-3.1-8b-ft-v2
      api_base: http://spark-candidate.internal:8000/v1
      api_key: "none"
    model_info:
      weight: 20

This gets you the split without a per-user hash, which is a reasonable tradeoff for a stateless task where a given caller not sticking to one variant across requests doesn't matter. Where session consistency does matter, a router with the hashed-split logic above is still the better fit.

Log enough to actually compare the two afterward

The split is the easy part. What decides whether the test tells you anything is what gets recorded per request: which variant served it, the prompt, the response, latency, token counts, and whatever quality signal you have, an explicit thumbs up or down, a downstream conversion, or a regenerate click as an implicit negative signal.

{
  "request_id": "req_8f2a1c",
  "user_id": "u_5521",
  "variant": "candidate",
  "model": "llama-3.1-8b-ft-v2",
  "latency_ms": 842,
  "prompt_tokens": 312,
  "completion_tokens": 96,
  "user_feedback": null,
  "regenerated": false,
  "timestamp": "2026-09-09T14:02:11Z"
}

Write this to a structured log or a table you can query, since the analysis at the end of the test is a group-by on variant. A plain text log line doesn't support that. Without user_feedback or an equivalent quality signal, all you can compare is latency and token usage, which tells you the candidate is faster or cheaper but says nothing about whether it's actually better.

Decide the sample size and the stopping rule before you start

Checking the results every hour and stopping the moment the candidate looks ahead is how a real difference of zero turns into a false positive. Pick a minimum sample size per variant, a few hundred requests is a reasonable floor for a rough comparison, and a fixed running duration before the first look, and write both down before the test starts. A basic significance check on the win rate, a two-proportion z-test is enough for most cases, keeps the eventual call from being a guess dressed up as data.

# rough two-proportion z-test, not a substitute for a real stats library
import math

def z_test(wins_a, n_a, wins_b, n_b):
    p_a, p_b = wins_a / n_a, wins_b / n_b
    p_pool = (wins_a + wins_b) / (n_a + n_b)
    se = math.sqrt(p_pool * (1 - p_pool) * (1 / n_a + 1 / n_b))
    return (p_b - p_a) / se if se > 0 else 0.0

Roll forward or roll back the same way you'd deploy anything else

Once the test resolves, promoting the winner is the same operation as any other model swap, ideally through the same reviewed pipeline covered in setting up a CI pipeline for model updates and blue-green deployments for model updates, rather than a manual cutover done in the moment the test ends. The A/B split gave you the evidence; the rollout mechanics that follow don't need to be improvised.

FAQ

Does GPUwerk provide A/B testing or traffic splitting as a feature?

No. A GPUwerk console deployment gives you an OpenAI-compatible endpoint per Spark, with no built-in traffic splitter between two model versions. Running two Sparks, one per model version, and splitting requests at your own routing layer or with a proxy like LiteLLM is how you get an A/B test on this infrastructure, not a feature you enable in the console.

How much traffic needs to go to each variant before the results mean anything?

There's no universal number, it depends on how large a difference you're trying to detect and how much your evaluation metric varies request to request. As a practical floor, a few hundred comparable requests per variant is usually the minimum before a difference in win rate or an average quality score is more than noise, and a subtle difference needs meaningfully more than that. Run a basic significance check, a two-proportion z-test on win rate is enough, before drawing a conclusion from a small sample.

Should the split be by request or by user?

By user, wherever the traffic has a stable user or session identifier. Splitting by individual request means the same user can hit both model versions across a session, which is confusing for them and muddies any comparison that depends on consistent behavior across a conversation. Hash a stable ID, user ID, API key, session token, into the split decision so a given caller stays on one variant for the duration of the test.

Related pages

Run the candidate on its own Spark.

A second dedicated instance for the test, no shared GPU memory to fight the control for.

Deploy a Spark Read the LiteLLM doc