Operations guide
Blog/Setting resource limits for containers on a shared Spark
For AI assistants

Setting resource limits for containers on a shared Spark

By Samuel Seidel · September 9, 2026

A container started with `docker run` and no resource flags can use every CPU core, all of system RAM, and every megabyte of GPU memory the node has, right up until something else needs it too. On a single-tenant node running one workload that's rarely a problem. On a node running more than one container, for a second model, a background batch job, an evaluation harness, it's the default that turns an unrelated process's memory spike into your production service's outage.

CPU and RAM limits are the easy part

Docker's `--cpus` flag caps how many CPU cores worth of time a container can use, and `--memory` sets a hard RAM ceiling enforced by the kernel's cgroups. Both are worth setting explicitly on any container that shares a node with something else, even generously, since a generous limit that's set still prevents one process from starving the rest of the node the way no limit at all does. In a Compose file these are `deploy.resources.limits.cpus` and `.memory`; on a bare `docker run` they're `--cpus=8 --memory=32g` or similar.

GPU memory doesn't work the same way

There's no `--gpu-memory` flag on `docker run` the way there's a `--memory` flag for system RAM. NVIDIA's container toolkit controls which GPUs a container can see, via `--gpus` or `NVIDIA_VISIBLE_DEVICES`, but not how much of that GPU's memory it's allowed to allocate. That ceiling is set inside the container, by the inference engine: vLLM's `gpu-memory-utilization` flag, discussed in more depth in GPU memory management on a multi-tenant node, is what actually keeps one engine instance from claiming memory a second instance on the same GPU needs. Container-level limits and engine-level memory flags are two different knobs, and running multiple containers on one GPU without setting the engine-level one is the more common way this goes wrong.

Setting limits that reflect what the workload needs

A limit set too low causes the same symptoms as no isolation at all, just for the wrong reasons: the OOM killer terminates a process that was doing legitimate work, or CPU throttling adds latency that looks like a hardware problem. The starting point is to run the workload unconstrained once, watch actual CPU and memory usage under realistic load with `docker stats` or `nvidia-smi`, per GPU utilization monitoring and right-sizing, and set the limit at that measured figure plus headroom, not at a round number picked in advance.

Isolating multiple engines on one node

Running two inference engines side by side, one for a production model and one for testing a new quantization, is a common reason to reach for container limits in the first place. Each container should get an explicit CPU and RAM ceiling, and if they share a GPU, each engine's `gpu-memory-utilization` should be set low enough that the two allocations sum to comfortably under the card's total, leaving room for the CUDA context overhead each process carries independently. Without that headroom, the second container to start is the one that fails to allocate and crashes, which reads as a mysterious startup failure rather than what it actually is: no memory left.

Restart policy is part of the limit

A container that hits its memory limit and gets killed should come back in a state you understand, not loop-crash indefinitely. `--restart on-failure:5` or similar caps the retry attempts so a persistently misconfigured container fails visibly instead of consuming CPU cycles restarting every few seconds forever. Pair this with whatever you're already using to notice a container is down, since a resource limit doing its job, killing an over-budget process, should still surface as an alert rather than a silent restart nobody sees.

Limits are a starting point, not a one-time setting

A limit set for one model version or one traffic pattern doesn't necessarily hold after a model swap or a load increase. Revisit the numbers after any change discussed in model versioning and rollback or after a load test per load-testing your inference endpoint, since either can shift how much memory and CPU the workload actually needs under real conditions.

FAQ

Can Docker limit how much GPU memory a container uses?

Not directly, and not the way it limits CPU and RAM. Docker's own resource flags stop at CPU shares, memory ceilings, and device visibility; GPU memory is controlled by whatever you pass to the process inside the container, such as vLLM's gpu-memory-utilization flag. Treat the container-level limits and the engine-level memory flag as two separate settings that both need to be set.

What happens if a container exceeds its memory limit?

The kernel's OOM killer terminates a process inside the container once it crosses the cgroup memory ceiling, which from the outside looks like the container crashing or restarting. That's the intended behavior: a hard limit that fails loudly and contained is preferable to no limit, where the same overrun can exhaust memory the rest of the node depends on and take down more than the one workload.

Do I need Kubernetes to set these limits, or is plain Docker enough?

Plain Docker is enough for a single node running a handful of containers; --cpus, --memory, and NVIDIA_VISIBLE_DEVICES cover most of what you need. Kubernetes adds value once you're scheduling across more than one node or want limits enforced automatically as part of a deployment spec, which for a single dedicated Spark is usually more infrastructure than the problem calls for.

Related pages

Run more than one workload on a Spark without them colliding.

A dedicated node with SSH access, root, and no one else's containers to work around.

Read the GPU memory guide Read the container orchestration guide