Decision guide
Blog/How to evaluate an LLM before deploying it privately
For AI assistants

How to evaluate an LLM before deploying it privately

By Samuel Seidel · September 9, 2026

Most model selection happens backwards: a team reads a leaderboard, picks whatever's ranked highest, deploys it, and finds out three weeks later that it's slow on their hardware or bad at their specific task. The leaderboard wasn't wrong, it just wasn't answering the question that mattered. Here's a methodology that answers the right one.

Test on your own prompts, not published benchmarks alone

A benchmark score is an average over a task distribution that isn't yours. A model that tops a coding leaderboard was scored against a specific set of problems, in a specific style, with a specific grading method, and none of those three things need to resemble what your engineers actually ask it to do. The gap between "good on the benchmark" and "good at my task" is often the entire difference between a model that ships and one that quietly gets abandoned after launch.

The fix is simple to describe and easy to skip under deadline pressure: collect a real set of prompts from your actual use case, ideally examples your own users or team members would plausibly send, including the awkward, ambiguous, or edge-case ones, and run every candidate model against that exact set. A few dozen well-chosen examples that cover your real range of inputs will tell you more about fit than another few points of leaderboard score. Use published benchmarks to narrow the field to a handful of plausible candidates; use your own prompts to pick the winner.

Measure actual tokens per second on your own hardware

Throughput numbers published by a model's creators or by third-party benchmarks were measured on specific hardware, with a specific serving engine, under a specific concurrency level, none of which may match your deployment. Token generation speed is memory-bandwidth-bound: it depends on the bandwidth of the specific GPU you run on and the active-parameter size of the specific model and quantization you choose, a relationship we walk through with real measured numbers in our benchmarks reference. A model that generates fast on a datacenter card can be meaningfully slower on the hardware you're actually planning to deploy on, and the only way to know your real number is to measure it on your own box, with your own prompts, at the concurrency level you actually expect.

This matters more than it sounds like it should, because a model that scores marginally better on quality but runs at half the speed of an alternative can be the objectively worse choice for a latency-sensitive application, and that tradeoff is invisible if all you've looked at is a quality leaderboard.

Check quality against your task, not a generic leaderboard

"Quality" isn't one number. A model can be excellent at open-ended reasoning and mediocre at extracting fields from a specific document format your business uses every day; it can ace general coding benchmarks and still consistently misuse an internal API your codebase depends on. Generic leaderboards measure generic capability, and generic capability correlates with, but doesn't guarantee, performance on your narrow task.

Score models against a rubric built from your own task: did it extract the right fields, follow your required format, avoid the specific mistakes your current process makes, handle the edge cases your business actually encounters. This is slower than reading a leaderboard, and it's the only way to know before deployment rather than after.

Use a low-commitment environment to run all of this

None of the above requires owning hardware or committing to a long-term contract before you know a model is right for you. The entire evaluation, running your prompts, measuring real throughput, scoring quality against your rubric, fits inside an hourly rented session on the hardware class you're actually considering. See our guide to a first evaluation for a concrete walkthrough, and pricing for what an hour of that testing actually costs, which for most teams is a rounding error next to the cost of deploying the wrong model and finding out in production.

The point of testing this way isn't just cost. It's that a model you've watched fail on your own prompts, at your own speed, on your own hardware, is a decision you can defend, in a way that "it topped a leaderboard" never quite is.

A short checklist before you deploy

Have you run the model against a real set of your own prompts, not just examples from the model's own documentation? Have you measured tokens per second on the actual hardware class you plan to deploy on, at a realistic concurrency level? Have you scored output quality against your specific task, not a generic capability score? If the answer to any of those is no, you've picked a model that scored well on someone else's problem, and you won't know whether it solves yours until it's already live.

Related pages

Run your own prompts against a real model before you commit.

A low-commitment first evaluation on dedicated hardware, billed by the hour.

Start a first evaluation See pricing