Setting up a VPN for your Spark
An inference endpoint doesn't need to be reachable from the entire internet to be useful. If every caller is a known machine, your own laptop, a handful of application servers, a VPN puts the node's API on a private network instead, so TLS and an API key aren't the only thing standing between the endpoint and anyone who finds the IP.
Why this is worth doing even with TLS and an API key already in place
TLS and an API key, covered in setting up TLS for your inference endpoint and secrets management for inference services, protect the traffic and authenticate the caller. Neither hides the endpoint from being discovered and probed in the first place. A VPN removes the node's inference port from the public internet entirely: it's only reachable from machines already inside the tunnel, which shrinks the attack surface before authentication is even in play.
WireGuard server config on the Spark
WireGuard is a good fit here: a small kernel module, one config file per side, no separate daemon to babysit. Generate a key pair on the Spark and set up the interface.
# generate the server keypair wg genkey | tee server-private.key | wg pubkey > server-public.key # /etc/wireguard/wg0.conf on the Spark [Interface] Address = 10.66.0.1/24 ListenPort = 51820 PrivateKey = <contents of server-private.key> # one [Peer] block per client allowed to connect [Peer] PublicKey = <client public key> AllowedIPs = 10.66.0.2/32
sudo wg-quick up wg0 sudo systemctl enable wg-quick@wg0
Client config and bind the inference service to the tunnel
On the calling machine, the mirror config points back at the Spark's public IP as the endpoint:
# /etc/wireguard/wg0.conf on the client
[Interface]
Address = 10.66.0.2/24
PrivateKey = <client private key>
[Peer]
PublicKey = <server public key>
Endpoint = 203.0.113.10:51820
AllowedIPs = 10.66.0.1/32
PersistentKeepalive = 25
The part that actually closes off public access is binding the serving process to the tunnel's private address rather than 0.0.0.0. If you're running the reverse proxy pattern from setting up a reverse proxy for your Spark, bind nginx's listen directive to 10.66.0.1 instead of the public interface, and the inference port is unreachable from anywhere outside the VPN, full stop, independent of any firewall rule.
# /etc/nginx/conf.d/inference.conf server { listen 10.66.0.1:443 ssl; # ... rest of the TLS and proxy config unchanged }
Add peers as callers change
Each new caller, a teammate's laptop, a new application server, gets its own key pair and its own [Peer] block, which also means you can revoke one caller by deleting a block and reloading, without touching anyone else's access or rotating a shared key. This maps cleanly onto the key-rotation habit in API key rotation best practices: the API key still identifies the caller at the application layer, the WireGuard key controls whether they can reach the network at all.
sudo wg set wg0 peer <public key> allowed-ips 10.66.0.3/32 sudo wg-quick save wg0
FAQ
Do I still need TLS and an API key if traffic goes over a VPN?
Yes. A VPN restricts who can reach the endpoint at the network layer; it doesn't authenticate individual callers or encrypt anything the VPN tunnel itself isn't already covering internally. Keep the API key so you know which caller made a request, and keep TLS on the endpoint if multiple people share the VPN and you don't want them reading each other's traffic.
WireGuard or Tailscale for a single Spark?
Tailscale is WireGuard underneath with a managed control plane, less config, a hosted coordination server, easier multi-device setup. Plain WireGuard is one config file, no third-party coordination server, and nothing running that you didn't configure yourself. For a single node with a handful of known callers, either works; plain WireGuard is the better fit if you'd rather not depend on an external service to establish connections.
Does routing inference traffic through a VPN add noticeable latency?
WireGuard's overhead is small, typically well under a millisecond of added latency on a local or well-peered connection, dwarfed by LLM generation times that run from tens of milliseconds to tens of seconds. It's not the bottleneck; the model's own decode speed is.