Setting up TLS for your inference endpoint
vLLM and llama.cpp both start up serving plain HTTP by default. That's fine for a local test on the same box, but the moment a client outside the node talks to your API, whether that's a teammate's laptop, a production app, or a mobile client, the request and response bodies, including whatever's in the prompt, travel unencrypted unless you put TLS in front of the server yourself. GPUwerk gives you SSH access to a dedicated node; setting up TLS on it is the same job it would be on any rented Linux box.
Terminate TLS at a reverse proxy, not in the inference server
Neither vLLM nor llama.cpp's server mode is built to handle certificate management, renewal, or the handshake overhead of terminating TLS directly, and adding that responsibility to the inference process is more moving parts than necessary. The standard approach is to run the inference server bound to localhost or an internal interface, and put nginx, Caddy, or a similar reverse proxy in front of it listening on 443 with the certificate. The proxy terminates TLS and forwards plain HTTP to the inference server over the loopback interface, which never leaves the box.
Caddy is worth considering specifically because it handles Let's Encrypt certificate issuance and renewal automatically with a few lines of config, no separate certbot cron job to maintain. nginx needs certbot or a similar ACME client set up alongside it, but is more likely to already be familiar to anyone who's run a reverse proxy before. Either works; pick based on what you or your team already knows.
A minimal Caddy config for an OpenAI-compatible endpoint
A Caddyfile pointing at a vLLM server running on port 8000 locally is short:
api.yourdomain.com {
reverse_proxy localhost:8000
}
Caddy requests and renews a certificate from Let's Encrypt automatically the first time it starts, provided the domain's DNS A record points at the node's public IP and port 80 is reachable for the ACME HTTP-01 challenge. If you're running LiteLLM as a proxy in front of multiple backend engines, as described in our LiteLLM guide, point the reverse_proxy directive at LiteLLM's port instead of the inference engine directly, so TLS termination sits in front of the whole stack rather than one engine.
Don't let the proxy buffer streaming responses
Both engines support streaming completions, where tokens arrive incrementally rather than all at once at the end. A reverse proxy that buffers the full response before forwarding it defeats this, turning a stream into what looks to the client like a single slow response. nginx buffers by default; the fix is proxy_buffering off; on the relevant location block. Caddy doesn't buffer by default for reverse_proxy, which is one reason it needs less tuning for this specific case. Whichever proxy you use, test with a streaming client request and confirm tokens appear progressively rather than in one burst, since this is easy to configure wrong and not obviously broken until someone notices output arriving all at once.
Certificate renewal and expiry
Let's Encrypt certificates last 90 days. Both Caddy and certbot renew automatically well before expiry as long as the process keeps running and the DNS record hasn't changed, but a node that gets rebooted or redeployed and comes back up without the proxy configured to start on boot will eventually serve an expired certificate to clients, who will see connection errors that have nothing to do with the model or the inference engine. Enable the proxy as a systemd service (systemctl enable) rather than starting it manually in a terminal session, so it survives reboots and restarts automatically if it crashes.
Internal traffic and mutual TLS
If the endpoint is only ever called by services you also control, on the same private network or over a VPN, plain HTTP between those services is a defensible choice and TLS overhead buys you little. Once the endpoint is reachable from the public internet or from a client you don't fully control, TLS stops being optional. For endpoints that need to authenticate the caller as well as encrypt traffic, mutual TLS (client certificates checked by the proxy) is available in both nginx and Caddy, but it's more operational overhead than most self-hosted setups need; API key checking at the LiteLLM or application layer, covered in our API key rotation guide, handles authentication for most teams without the certificate distribution problem mTLS introduces.
Checking the setup
After the proxy is up, confirm the certificate chain is valid and complete with openssl s_client -connect api.yourdomain.com:443 -servername api.yourdomain.com, and check that plain HTTP on port 80 either redirects to HTTPS or is closed rather than serving the API unencrypted as a fallback. A common oversight is leaving the inference engine's port open on the node's firewall alongside the proxy's port, which lets anyone bypass TLS entirely by connecting to the unencrypted port directly. Firewall rules should only allow the proxy's port from outside the node.
FAQ
Does GPUwerk provide TLS for my inference endpoint?
No. GPUwerk gives you SSH access to a dedicated Spark node and the network path to it; what you run on that node, including whether it's exposed over HTTPS or plain HTTP, is entirely up to you. TLS termination is your responsibility, same as it would be on any rented server.
Does a TLS reverse proxy add noticeable latency to streaming responses?
The TLS handshake adds a fixed cost at connection setup, not per token, so once a streaming connection is established the proxy just passes bytes through as they arrive. Reused connections with keep-alive skip the handshake cost on subsequent requests. The bigger risk to streaming is proxy-level response buffering, which can be disabled in nginx and most proxies without touching TLS at all.
Can I use a self-signed certificate instead of Let's Encrypt?
You can, and it works fine for testing or for clients you control where you can install the certificate as trusted. It's a poor fit for anything a browser or a third-party client will connect to, since every connection will either show a security warning or require disabling certificate verification, which defeats much of the point of using TLS.