Operations guide
Blog/Configuring CORS for your inference API
For AI assistants

Configuring CORS for your inference API

By Samuel Seidel · September 9, 2026

The first sign your inference endpoint has a CORS problem is usually a working curl request next to a browser console full of red text, same URL, same API key, same body. vLLM's OpenAI-compatible server doesn't set permissive CORS headers by default, so a page that calls it directly from client-side JavaScript gets blocked before your API key check ever runs. This is a browser-side restriction, not a server bug, and it's fixed at whichever layer terminates the request.

Where CORS actually gets enforced

A browser sends a preflight OPTIONS request before certain cross-origin calls, checks the response headers, and only lets the page's JavaScript read the real response if Access-Control-Allow-Origin matches. Nothing about this involves your API key or your server's own logic; it's the browser deciding whether to hand the response to the calling script. If you're terminating TLS at a reverse proxy, as described in setting up TLS for your inference endpoint, that's the natural place to add the headers, since it already sits in front of every request and doesn't need the inference engine itself to know about CORS at all.

A CORS block in the nginx config fronting vLLM

Adding the headers to the same server block that proxies to vLLM keeps the CORS policy in one file, next to the TLS config:

# /etc/nginx/conf.d/inference.conf
server {
    listen 443 ssl;
    server_name api.yourdomain.com;

    # reflect a fixed list of allowed origins
    set $cors_origin "";
    if ($http_origin = "https://app.yourdomain.com") {
        set $cors_origin $http_origin;
    }
    if ($http_origin = "https://staging.yourdomain.com") {
        set $cors_origin $http_origin;
    }

    add_header Access-Control-Allow-Origin $cors_origin always;
    add_header Access-Control-Allow-Methods "GET, POST, OPTIONS" always;
    add_header Access-Control-Allow-Headers "Authorization, Content-Type" always;

    if ($request_method = OPTIONS) {
        add_header Access-Control-Allow-Origin $cors_origin always;
        add_header Access-Control-Allow-Methods "GET, POST, OPTIONS" always;
        add_header Access-Control-Allow-Headers "Authorization, Content-Type" always;
        add_header Content-Length 0;
        return 204;
    }

    location / {
        proxy_pass http://localhost:8000;
        proxy_buffering off;
    }
}

Listing exact origins rather than a wildcard is the right default for a browser front end that authenticates with a session cookie, since a wildcard origin combined with credentialed requests is exactly the combination browsers refuse to allow. If the front end sends the API key as a header instead of relying on a cookie, a wildcard is safe and simpler, covered below.

If vLLM is running standalone, its flag does the same thing

vLLM's server accepts CORS configuration directly if you'd rather not add it at the proxy layer, or you're testing without one yet:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --port 8000 \
  --allowed-origins '["https://app.yourdomain.com"]'

This is the same header logic nginx would otherwise add, and the two shouldn't both be configuring it, since duplicate Access-Control-Allow-Origin headers from proxy and origin server are a common cause of CORS still failing after you thought you fixed it. Pick one layer, and if you're running a reverse proxy at all, do it there rather than in the inference engine.

When a wildcard origin is the right call

A public-ish demo, an internal tool with no login, or any endpoint whose only access control is an API key supplied explicitly in the request rather than a cookie, doesn't gain anything from restricting origins. The key is the thing standing between a stranger's page and your inference spend, and CORS can't add meaningfully more than that without the credentialed-cookie case, which a wildcard already blocks by browser rule.

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --port 8000 \
  --allowed-origins '["*"]'

Keep the rate limiting from API rate limiting in place regardless. CORS controls who can read a response in a browser; it says nothing about how many requests a caller with a valid key can send.

FAQ

Why does my API key check not stop unwanted browser requests?

It does stop them, but not before the browser has already sent a preflight OPTIONS request and, for a simple request, the actual call. CORS is enforced by the browser reading response headers, not by your server refusing the connection; a missing or wrong Access-Control-Allow-Origin header just means the browser hides the response from the calling page's JavaScript. Your API key check still runs and still rejects unauthenticated requests, CORS is a separate, earlier layer about which web pages are allowed to read the response at all.

Is a wildcard Access-Control-Allow-Origin ever fine for an inference API?

Yes, specifically when the endpoint requires an API key or bearer token and doesn't rely on cookies for authentication. A wildcard origin means any web page's JavaScript can read the response if it also has the key, but it can't get the key from a cookie the browser sends automatically, since wildcard origins are incompatible with credentialed requests. If the only credential is a header the calling code has to supply explicitly, a wildcard origin doesn't hand out access to anyone who doesn't already have the key.

Does CORS protect the endpoint from non-browser clients?

No. CORS is a browser-enforced restriction on JavaScript running in a web page; it has no effect on curl, a Python script, a mobile app, or a server calling your API directly, all of which never send an Origin header the way a browser does and aren't subject to the same-origin policy at all. CORS headers control what a browser page can do, not who can reach the endpoint over the network.

Related pages

Build your front end against a dedicated endpoint.

A Spark that's yours alone, so the CORS policy you write only ever governs your own traffic.

Read the TLS guide Read the rate limiting guide