What is a GPU node?
A GPU node is a single machine, physical or virtual, with one or more GPUs attached, treated as one unit of compute. In a rental context, one GPU node is typically what you provision and get billed for: a discrete machine with its own CPU, memory, storage, and network connection, built around the GPU or GPUs that do the actual computation.
Node as the unit of allocation
"Node" is cluster terminology: a group of machines working together is a cluster, and each individual machine in it is a node. The word carries over into GPU rental because it describes the same thing, one discrete, addressable machine, whether or not it's actually part of a larger cluster. When a provider advertises a GPU node, they mean one machine you get exclusive or shared access to, with a specific GPU model, memory capacity, and CPU/RAM/storage attached to it.
Dedicated vs shared nodes
A dedicated node is allocated entirely to one renter; nothing else runs on it while it's yours. A shared node splits its GPU or GPUs, by time-slicing, memory partitioning, or virtualization, across multiple renters or workloads at once. The distinction matters for two reasons: performance predictability (a dedicated node's full capacity is available to you at all times, a shared node's isn't) and isolation (a dedicated node has no other tenant's code or data on the same GPU memory, which matters for anything handling sensitive information). See dedicated vs shared GPUs for more on the isolation mechanics.
Single-node vs multi-node
Some workloads fit entirely within one node's GPU memory and compute; others need more than one node working together, connected over a network, to handle a model too large for a single machine or to parallelize training across more GPUs than one node has. Multi-node setups add networking overhead and coordination complexity that single-node setups don't need to deal with, so it's worth checking whether a workload actually requires multiple nodes before provisioning them. Most LLM inference for models that fit in one node's memory runs fine single-node.
What defines a node's capacity for LLM inference
For serving language models specifically, a node's GPU memory sets the ceiling on how large a model (and how much context) it can hold, and memory bandwidth sets how fast it can generate tokens. CPU and system RAM matter less for the inference itself, though they matter for data preprocessing and general system overhead. A node's specification sheet is usually the fastest way to tell whether it can handle a given model before renting it.
Where GPUwerk fits
GPUwerk's unit of rental is a dedicated DGX Spark node: one machine with 128GB of unified GPU memory, allocated entirely to you, at $0.79/hour. For workloads needing more, a two-node cluster is available at $1.79/hour, connected for larger models or higher throughput than a single Spark can provide. See pricing for current rates and node specifications.