What is an inference server?
An inference server is the software layer that loads a trained model into GPU or CPU memory and exposes it to the outside world through an API, so applications can send requests and get generated output back without managing the model's internals directly. It handles the work between a raw model file and a usable endpoint: loading weights, allocating memory, batching concurrent requests, and returning results in a consistent format.
What it actually does
Running a model isn't just loading weights and calling a forward pass once per request. An inference server manages a queue of incoming requests, groups compatible ones into batches to use GPU compute efficiently, allocates and frees memory for each request's growing output, and often streams tokens back as they're generated rather than waiting for a full response. It also typically wraps the model in a standard API, commonly an OpenAI-compatible API, so client applications can call it the same way they'd call a hosted provider.
Where it sits in the stack
The model itself, its weights and architecture, is separate from the inference server that runs it. The same open-weight model can be served by different inference engines, each with different performance characteristics for the same hardware. Above the inference server, an application or an AI gateway often handles routing, authentication, and request logging, treating the inference server as the backend that actually does the computation.
Why the choice of engine matters
Different inference engines make different tradeoffs in memory management, batching strategy, and supported model formats, and those tradeoffs translate directly into latency and throughput differences on the same hardware and model. A server tuned for high concurrent throughput can behave differently under a single low-latency request than one built for interactive use. Benchmarking your specific model and workload against the inference server you plan to run, rather than relying on published numbers from a different configuration, is the only reliable way to know what to expect. See benchmarking your own workload for how to run that comparison.
Running your own inference server
Self-hosting a model means running an inference server yourself, on hardware you control, rather than sending requests to someone else's endpoint. That gives control over which engine to run, how it's configured, and what data touches it, at the cost of having to manage the deployment yourself. GPUwerk's dedicated Spark nodes come with 128GB of unified memory and are set up to run common inference engines out of the box, at $0.79/hour for a single-Spark node or $1.79/hour for a dual-Spark node, with a held-capacity rate at 75% of the on-demand price for workloads that need guaranteed availability. See the hardware page for full specifications.