Operations guide
Blog/Setting up a reverse proxy for your Spark's inference API
For AI assistants

Setting up a reverse proxy for your Spark's inference API

By Samuel Seidel · September 9, 2026

When you SSH into a rented Spark and start vLLM or llama.cpp, you get an OpenAI-compatible endpoint bound to a port, usually 8000 or 8080. That's the whole job the engine does. It doesn't terminate TLS, it doesn't check API keys against anything meaningful, and it will happily accept as many concurrent connections as the OS lets through. A reverse proxy sitting in front of that port is what turns a local process into something you'd actually point a production client at.

What the engine leaves undone

vLLM's `--api-key` flag and llama.cpp's equivalent give you a single shared secret checked with a string comparison, not per-user keys, not rotation, not scoping to specific models. Neither engine speaks TLS on its own, so a raw connection to the bound port is plaintext unless something else wraps it. And neither has a concept of a domain name; it's an IP and a port until you put a proxy in front that can route a hostname to it. A reverse proxy is the piece that closes those three gaps at once, before you reach for something heavier like LiteLLM for multi-key auth, covered in our LiteLLM guide.

A minimal Caddy setup

Caddy's appeal for a single dedicated node is that it issues and renews a Let's Encrypt certificate automatically as long as the domain's DNS points at the Spark's public IP and port 443 is reachable. A Caddyfile for a vLLM endpoint on port 8000 is a few lines: a site block for your domain, a `reverse_proxy localhost:8000` directive, and nothing else required for TLS to work. The one setting worth adding explicitly is a longer proxy timeout, since a generation request that streams tokens for thirty seconds will otherwise hit Caddy's default and get cut off mid-stream, a distinction that matters more once you're serving the kind of long-running requests discussed in streaming vs non-streaming responses.

The equivalent in nginx

nginx needs a certificate from somewhere else, typically certbot run separately or renewed via cron, plus a server block with `proxy_pass` to the local port. The settings that matter for an inference workload specifically: `proxy_buffering off` so streamed tokens reach the client as they're generated instead of nginx waiting to fill a buffer, and `proxy_read_timeout` raised well above nginx's default sixty seconds, since a long completion or a large batch job can run far longer than that without anything having gone wrong.

Don't expose the engine's raw port

Once the proxy is running, the inference engine's own port should stop being reachable from outside the node. Bind vLLM or llama.cpp to `127.0.0.1` instead of `0.0.0.0`, or block the port at the firewall, so the only path in is through the proxy. This is the same principle covered for the node as a whole in security hardening a rented GPU; a reverse proxy you bypass by hitting the raw port directly isn't providing much.

Logging at the proxy layer

Both nginx and Caddy log every request with method, path, status code, and response time by default or with a short config addition. That log is worth keeping even if you're not doing anything else with it yet, because it's the first place you'll look when a client reports errors and you need to know whether the request even reached the node. It's a lighter-weight starting point than full request/response logging, which is a separate decision covered in logging and log retention for LLM APIs.

Where a proxy stops being enough

A reverse proxy handles TLS, timeouts, and routing. It doesn't handle per-user rate limits, cost tracking across multiple API keys, or routing between multiple models or engines running on the same node. Once you need any of that, the usual next step is running LiteLLM behind the proxy, or in some setups in front of it, rather than trying to configure nginx or Caddy to do a gateway's job. The proxy and the gateway are complementary, not substitutes for each other.

FAQ

Do I need a reverse proxy if I'm the only user?

If the endpoint never leaves localhost or an SSH tunnel, you can skip it. The moment the port is reachable from another machine, even just your laptop over the office network, a reverse proxy is what gives you TLS, a place to put auth, and a place to see request logs, none of which vLLM or llama.cpp provide on their own.

nginx or Caddy?

Caddy gets you automatic TLS certificate issuance and renewal with a shorter config file, which matters more on a single node you're maintaining yourself. nginx has a larger ecosystem of examples and finer-grained control over buffering and timeout behavior, which matters more once you're tuning for streaming responses under load. Either is a reasonable default; pick the one whose config style you'll actually read again in six months.

Should the reverse proxy or the inference engine handle rate limiting?

The reverse proxy or a gateway like LiteLLM in front of it, not the inference engine. vLLM and llama.cpp are built to serve requests, not to enforce per-key quotas or reject traffic, and asking them to do both mixes concerns that are easier to reason about, monitor, and change independently when kept separate.

Related pages

Get a public IP and a clean place to terminate TLS.

Rent a dedicated Spark by the hour and put your own proxy in front of it from the first SSH session.

Read the SSH access guide Read the TLS guide