Operations guide
Blog/Multi-user access on a shared GPUwerk node
For AI assistants

Multi-user access on a shared GPUwerk node

By Samuel Seidel · September 9, 2026

When you rent a Spark, you get one thing: root SSH access to a Linux box, described in the SSH key guide. If your team is three engineers instead of one, that single root login is a problem you have to solve yourself, the same way you would on any rented server. None of this is GPUwerk-specific. It's ordinary Linux administration, applied to a machine that happens to have a GPU in it.

Start by not sharing the root key

The instinct on a first rental is to hand the same private key to everyone on the team, or paste the root password into a group chat. Skip that. Create a real user account per person, add each person's own public key to that account's authorized_keys, and keep root logins limited to whoever actually needs to run system-level commands, like installing drivers or managing disk partitions. This gets you two things a shared key can't: an audit trail in last and shell history that says who did what, and the ability to revoke one person's access without rotating a key everyone else also uses.

Create the accounts as normal Linux users, add them to a shared group (say, gpuwerk), and put anything they'll need to read or write, model weight caches, shared datasets, into a directory owned by that group with the setgid bit set so new files inherit group ownership automatically. A minimal setup looks like:

useradd -m -G gpuwerk alice, then drop alice's public key into /home/alice/.ssh/authorized_keys, set ownership and chmod 700/600 on the directory and file. Repeat per teammate.

Decide who runs the inference server, and how others reach it

The failure mode we see most often isn't a permissions bug, it's three people each starting their own vLLM or llama.cpp process, each trying to load a model into the same pool of unified memory. A Spark has 128 GB total, shared between every process on the box: two 70B-class models loaded at once, even quantized, can exhaust it fast, and the second process to start usually just fails to allocate rather than politely queuing.

The cleaner pattern is one inference server, run by one designated account or a system service, that the rest of the team reaches over the network rather than each running their own copy. LiteLLM is built for exactly this: it sits in front of vLLM or an Ollama/llama.cpp backend, gives each teammate their own API key, and lets you set per-key rate limits or budgets so one person's batch job doesn't starve everyone else's interactive session. If your team just needs a shared chat interface rather than API access, Open WebUI supports its own user accounts on top of a single backend, which is often simpler for a small team than managing Linux accounts at all.

Isolate workloads with systemd or containers, not hope

Once more than one person can start processes on the box, you want some way to stop a runaway job from taking the whole node down. Two practical options, neither requiring GPU-level isolation the Spark's single-GPU design doesn't offer:

Run each person's long-lived service as a systemd unit with resource limits set through cgroups, MemoryMax for RAM and, if you're disciplined about it, a wrapper script that checks GPU memory before launch. Or run each workload in its own container with the NVIDIA container runtime, which at least gives you a clean filesystem and dependency boundary per user even though the GPU memory pool underneath is still shared. Containers are covered in more depth in the vLLM setup guide and in the coding agents doc, which walks through running an agent's tool-execution sandbox alongside a model server on the same box.

Neither approach gives you hardware-level GPU partitioning. A single DGX Spark has one GPU; there's no MIG-style split the way there is on a datacenter A100. What you're managing is contention for memory and bandwidth, not hard isolation, and the tools above manage that contention rather than eliminate it.

Set expectations about noisy neighbors on the same node

If your team shares one rented Spark, agree on ground rules before someone's 2am batch job blocks someone else's demo. A simple convention that works for small teams: pick a time window or a Slack-style "claiming" message before anyone loads a second large model, and default to the shared LiteLLM endpoint for anything that isn't explicitly a benchmark or a one-off experiment. If contention becomes a regular problem, it's usually a sign the team has outgrown one node, at which point renting a second Spark at the same $0.79/hour on-demand rate for spark-1x, or moving to spark-2x at $1.79/hour, is cheaper than the time lost to scheduling conflicts.

Keep the audit trail when someone leaves

Because access is per-account, offboarding is a one-line fix: remove the departing person's key from their account's authorized_keys, or delete the account with userdel, and check last -a for anything they had left running. This is the whole reason to avoid a shared root key in the first place. With one shared credential, someone leaving the team means rotating a key everyone else depends on and hoping nobody wrote it down somewhere. With per-person accounts, it's a five-second change with no blast radius.

FAQ

Does GPUwerk provide separate logins for each team member?

No. A GPUwerk instance ships with one root SSH login, the same as any rented Linux box. Splitting that into per-person accounts is standard Linux system administration you do yourself once you're in, not a feature GPUwerk configures for you.

Can two people run inference on the same Spark at once?

Yes, as long as the models they're running fit together inside 128 GB of unified memory and they aren't both trying to saturate the same 273 GB/s of bandwidth for latency-sensitive work at the same time. A shared LiteLLM or vLLM endpoint that both people call is usually a better fit than two separate model processes fighting over the same memory pool.

What happens if one user's process crashes the GPU driver?

A hard driver crash typically takes down every GPU process on the node, regardless of which user started it, because the driver and CUDA context are shared kernel-level resources, not something Linux user permissions can isolate. The mitigation is process-level, not permission-level: keep individual jobs bounded so they fail on their own rather than taking the driver with them.

Related pages

Give your team a shared endpoint, not five copies of the same model.

Rent a Spark, set up LiteLLM once, and hand out API keys instead of root passwords.

Set up LiteLLM See pricing