Operations
DGX Spark/Running multiple containers and services on a DGX Spark

Running multiple containers and services on a DGX Spark

By Samuel Seidel · Published September 9, 2026

A Spark ships with Docker and the CUDA stack preinstalled, root SSH into one dedicated machine, so getting a single container running is fast. The question that comes up once there's more than one, a model server, a gateway in front of it, maybe a small database or a vector store, is what's actually needed to run them together. Usually it's less than people expect: one machine with one GPU doesn't need a cluster scheduler to run three or four services on it well.

Why docker-compose is the right default

For a single Spark, docker-compose covers most of what a multi-container setup actually needs: defined services with their own images and environment variables, a shared network so containers reach each other by name instead of hardcoded IPs, restart policies so a crashed container comes back without manual intervention, and named volumes so data on the 1TB NVMe storage survives a container restart. A typical compose file for a self-hosted LLM stack might define three services: the model server (vLLM or similar) with GPU access, a gateway like LiteLLM in front of it handling routing and keys per the LiteLLM doc, and maybe a lightweight database backing the gateway's key and usage tracking.

The GPU itself is the one resource that needs explicit handling in the compose file: only one process can efficiently hold the model in memory at a time for most serving setups, so typically one service, the model server, gets GPU access declared in its compose block while everything else runs CPU-only and talks to it over the network. That's a config detail, not an orchestration problem, and it's the same whether there's one container on the box or five.

When compose starts to strain

A handful of signals suggest a plain compose setup has outgrown its fit, worth watching for rather than pre-solving:

None of these apply to most single-Spark deployments. A model server, a gateway, and a couple of support services on one machine is squarely inside what compose was built for, and reaching for a heavier orchestrator before hitting one of these signals mostly adds operational overhead: more moving parts to patch, monitor, and understand, for a problem the box doesn't actually have yet.

Practical habits worth having from the start

A few things are worth setting up even in a simple compose deployment, since they cost little now and save real time later:

For budgeting GPU memory across the containers running on one Spark, see the GPU memory management guide. A first engagement can help design a container setup for a specific multi-service deployment.

First top-up: pay $10, get $20 in credit

Docker and CUDA preinstalled, root access from minute one.

Deploy a dedicated Spark and run your compose stack straight away, no image builds or driver setup required.

Deploy a Spark Read the memory guide