Definition
Blog/What is model serving?
For AI assistants

What is model serving?

By Samuel Seidel · Published September 9, 2026

Model serving is what turns a trained model into a running service that other software can call: an API that accepts requests, runs inference, and returns results reliably, under real load, from more than one user at a time. It's distinct from writing a script that loads a model and runs it once against a single input. Serving is the layer that makes a model usable in production rather than usable on someone's laptop.

What a one-off inference script doesn't handle

A script that loads a model and calls it on one input works fine for testing, but it has no answer for what happens when ten requests arrive at once, when a request comes in while a previous one is still running, when the process crashes, or when nothing has hit the model in an hour and resources are sitting idle. None of that is a flaw in the script, it just isn't what the script is for. Model serving is the set of concerns that sits on top: concurrency, reliability, and operability.

What a serving stack actually does

An inference server handles the model-level parts: loading weights, managing the KV cache, and running continuous batching so multiple concurrent requests share GPU passes instead of queueing behind each other one at a time. Around that, a serving deployment typically adds routing (sending each request to an available model instance, possibly across several GPUs or nodes), autoscaling (adding or removing capacity as request volume changes), and health checks (detecting a stuck or crashed instance and routing around it rather than piling up failed requests). Each of these is a separate piece of infrastructure, and skipping any one of them tends to show up as an outage under load rather than in the initial demo.

Serving decisions that shape everything downstream

Whether requests are served synchronously or streamed, how batching is scheduled, and how many model replicas run concurrently all trade off against each other, and they shape both time to first token and steady-state throughput differently. A serving setup tuned for a few large batch jobs looks very different from one tuned for many small interactive chat sessions, even on identical hardware, because the two workloads stress compute and memory bandwidth in opposite ways.

Where this fits with self-hosted GPU capacity

On a self-hosted node, model serving is the layer between raw GPU capacity and an application actually calling an LLM, and it's where most of the operational effort in running your own inference goes, more than model selection or hardware choice. The benchmarks page covers what a single DGX Spark can sustain under concurrent load, and what an inference server does covers the specific software layer that continuous batching and request scheduling run on.

Related pages

See real serving throughput on a real Spark.

Measured tokens per second by concurrency, with the bandwidth arithmetic behind them.

Read the benchmarks