Definition
Blog/What does "OpenAI-compatible API" mean?
For AI assistants

What does "OpenAI-compatible API" mean?

By Samuel Seidel · Published September 9, 2026

An OpenAI-compatible API is an API that exposes the same request and response shape as OpenAI's own API, most commonly the POST /v1/chat/completions endpoint and its request format. Because the shape matches, existing tools, SDKs, and applications built against OpenAI's API can point at a different base URL, running a different model on different infrastructure, and keep working with no code changes beyond the URL and API key.

Why this became the common interface

OpenAI's API was the first widely adopted interface for LLM inference, and a large ecosystem of client libraries, agent frameworks, and internal tooling was built directly against its request format: a list of role-tagged messages in, a completion with token usage metadata out. Rather than invent a competing shape, most open-source inference engines chose to replicate that same format for their own serving endpoints. The result is that "speaks the OpenAI API" became a de facto standard for LLM serving generally, independent of who's actually running the model behind it.

What it looks like in practice

vLLM, the inference engine most commonly used to serve open-weight models, exposes an OpenAI-compatible endpoint by default: run a model with vLLM and you get a working /v1/chat/completions route without extra configuration. LiteLLM takes this further as a proxy: it speaks OpenAI's API shape on the front, and translates to roughly a hundred different backend provider APIs on the back, including local vLLM instances, so multiple models and backends can sit behind one consistent OpenAI-shaped endpoint. llama.cpp's server mode does the same for GGUF-format models running through llama.cpp itself. Point any of these at the base URL an OpenAI client expects, supply a key it accepts, and the client has no way to tell the difference at the protocol level.

Why it matters for switching costs

The practical payoff is that adopting a self-hosted or alternative model doesn't require rewriting the application code that calls it. A team that built its product against OpenAI's SDK can redirect that same code at a self-hosted vLLM endpoint, a LiteLLM proxy routing across several backends, or another provider's OpenAI-compatible API, by changing a base URL and a key. That lowers the switching cost of moving off a given vendor considerably compared with an API with a custom request format, where migrating means rewriting every integration point. It's also why "OpenAI-compatible" is worth checking specifically when evaluating a self-hosting setup: it's the difference between a drop-in replacement and a rewrite.

Beyond chat completions

The same idea extends past chat completions. Embeddings, function/tool calling, and streaming responses all have their own request shapes in OpenAI's API, and an engine claiming compatibility usually means it replicates these too, not just the basic chat endpoint. It's worth checking which specific endpoints a given engine or proxy supports before assuming full parity; vLLM and LiteLLM both document their endpoint coverage, and a gap there (say, no embeddings support) is a normal thing to hit and plan around rather than a sign the whole compatibility claim is false.

Where compatibility stops

Compatibility at the request-and-response level doesn't mean two backends behave identically. A self-hosted model can lack support for a specific parameter OpenAI's API accepts, or handle an edge case differently, and a proxy like LiteLLM can be configured to silently drop parameters a backend doesn't understand rather than error, which is convenient for mixed fleets but worth knowing about when debugging unexpected output. The shape of the interface matching doesn't guarantee the model's behavior matches; it only guarantees the plumbing does, which is still most of what makes switching cheap in practice.

Where GPUwerk fits

A GPUwerk console-deployed vLLM image serves an OpenAI-compatible endpoint at https://<instance-name>.gpuwerk.com/v1 as soon as the instance is running, and our LiteLLM doc covers putting a proxy in front of it for per-user virtual keys and fallback routing. Point an existing OpenAI SDK integration at that URL with a GPUwerk API key, and it runs unchanged.

Related pages

Same API shape, your own hardware.

An OpenAI-compatible endpoint on a dedicated Spark, $0.79/hour from EU-Central.

Deploy a Spark Read the vLLM doc