Security hardening a rented GPU instance
Renting a Spark gets you a machine with root access and a public network path. What runs on it, and how exposed it is, is entirely your configuration from that point on. This is the in-instance half of security: the steps you take once you are logged in as root. It is deliberately narrow, GPUwerk's own infrastructure-level posture, physical security, network isolation between tenants, and so on, is a separate layer that this page does not speak for, and the site does not publish a specific security certification to cite here.
SSH: the door you actually control
Every instance ships with password authentication disabled, per GPUwerk's SSH key guide, so a key pair is the only way in by default, which already removes the most common brute-force target. From there, the hardening is on you:
- Use a passphrase-protected key, loaded into an agent rather than typed on every connection. A private key with no passphrase is a plaintext credential; if your laptop is compromised, so is every instance that key opens.
- Do not paste the same key into services you do not trust equally. If a key is deployed to a rented Spark and also to unrelated infrastructure, a compromise of the weaker target compromises both.
- Rotate the key if you ever suspect it leaked, for example from a laptop theft or an accidental commit. Deploying a fresh key to an existing instance, or using the console's SSH reset flow described in the SSH guide, replaces
authorized_keyswithout touching your workspace. - Treat host key warnings as real. A changed host key fingerprint on an address you have connected to before is the expected result of an instance being terminated and its port recycled, per the SSH guide, but it is also what a genuine machine-in-the-middle looks like. Don't clear the warning reflexively on a host you have not just redeployed.
Network exposure: expose only what needs exposing
Every instance gets a public HTTPS endpoint routed to port 8888 in the container, and SSH on its own hostname at the standard port 22. Anything else you bind to a port is your call, and the default should be not public.
- Tunnel, don't publish, for anything you are not intentionally sharing. An admin panel, a Jupyter server, a metrics dashboard, or a vLLM endpoint you want for yourself belongs behind an SSH tunnel (
ssh -N -L 8080:localhost:8080 spark, as covered in the SSH guide), not bound to a public interface with no auth in front of it. - If a service must be public, put authentication in front of it. An inference endpoint meant for external callers needs an API key or token check at minimum; a model server with no auth in front, reachable at a public HTTPS address, is an open invitation to anyone who finds the hostname.
- Bind to localhost by default, and only widen a service's bind address when you have a specific reason to. Most serving frameworks default to
0.0.0.0, which listens on every interface; check the config rather than assume the default is safe. - Firewall what you can within the container. A container-level firewall (iptables, ufw where available) limiting which ports actually answer is a second layer behind "just don't publish it," useful for catching a service you forgot was listening.
Don't run services as root just because you land there
SSH puts you in as root, and it's tempting to run everything that way since nothing stops you. Don't. A service running as root that gets compromised, through a dependency vulnerability, a malicious model artifact, or an exposed debug endpoint, gives an attacker root on the box immediately, rather than the more limited blast radius of a service account.
- Create a dedicated non-root user for your serving process (
useradd -m -s /bin/bash svc && su - svc) and run the model server, API layer, and any long-running process under it. - Keep root for setup and package installation, not for the process that stays up and faces the network.
- If you're using containers inside the instance (nested containers, or a process manager), don't drop the privilege boundary just because the outer layer is already a rented tenant container.
Secrets: don't leave them in shell history or plaintext env files
API keys, model repository tokens, and any credentials for services you call out to need the same discipline you'd use anywhere else.
- Don't pass secrets as command-line arguments; they land in shell history and in
psoutput visible to any other process on the box. Use environment variables set from a file, or your framework's secret-loading mechanism. - Keep a secrets file (
.envor similar) outside version control and readable only by the user that needs it (chmod 600). - Because
/workspaceis what survives a stop and gets deleted on terminate, per GPUwerk's SSH guide, don't assume a secret written there is more durable than that: terminate the instance and it's gone with everything else, including GPUwerk's own recovery copy of it.
Ephemeral vs persistent: pick deliberately
A stopped instance keeps its workspace, SSH port, and hostname exactly as they were, billed at 75% of the running rate, per GPUwerk's pricing page. A terminated instance is wiped for the next tenant and your workspace is deleted with no copy kept. That split is also a security decision, not just a billing one:
- Treat a long-lived stopped instance like any other long-lived server: it accumulates config drift, unpatched packages, and forgotten open ports the longer it sits there, stopped or running. Patch and review it periodically rather than assuming a paused instance is a frozen, safe instance.
- Treat a terminate-and-redeploy cycle as your patching mechanism for anything where you'd rather start clean than maintain drift. A freshly deployed instance starts from the current tenant image with no accumulated state.
- Never assume state on the instance is backed up. GPUwerk states plainly on its pricing FAQ that it keeps only a periodic recovery copy of /workspace for hardware failure, not a backup service. If losing the instance would lose something you can't regenerate, that thing needs to live somewhere else too, not just in
/workspace.
What this page is not claiming
This is in-instance hardening: what you configure once you have root on a rented machine. GPUwerk's own infrastructure security, isolation between tenants at the platform level, and physical access controls are a separate posture that this page does not describe or vouch for beyond what is already published on GPUwerk's own pages. No specific security certification is claimed here; if that matters for your compliance requirements, ask directly rather than infer one from this guide.
A first engagement can help scope a hardening checklist against your specific compliance requirements before you deploy production traffic.