Running multiple containers and services on a DGX Spark
A Spark ships with Docker and the CUDA stack preinstalled, root SSH into one dedicated machine, so getting a single container running is fast. The question that comes up once there's more than one, a model server, a gateway in front of it, maybe a small database or a vector store, is what's actually needed to run them together. Usually it's less than people expect: one machine with one GPU doesn't need a cluster scheduler to run three or four services on it well.
Why docker-compose is the right default
For a single Spark, docker-compose covers most of what a multi-container setup actually needs: defined services with their own images and environment variables, a shared network so containers reach each other by name instead of hardcoded IPs, restart policies so a crashed container comes back without manual intervention, and named volumes so data on the 1TB NVMe storage survives a container restart. A typical compose file for a self-hosted LLM stack might define three services: the model server (vLLM or similar) with GPU access, a gateway like LiteLLM in front of it handling routing and keys per the LiteLLM doc, and maybe a lightweight database backing the gateway's key and usage tracking.
The GPU itself is the one resource that needs explicit handling in the compose file: only one process can efficiently hold the model in memory at a time for most serving setups, so typically one service, the model server, gets GPU access declared in its compose block while everything else runs CPU-only and talks to it over the network. That's a config detail, not an orchestration problem, and it's the same whether there's one container on the box or five.
When compose starts to strain
A handful of signals suggest a plain compose setup has outgrown its fit, worth watching for rather than pre-solving:
- Multiple models that need to share the GPU dynamically, loaded and unloaded based on demand rather than each pinned to a fixed container. That's a scheduling problem compose doesn't solve on its own; it needs either an application-level router that manages model loading itself, or manual scripting around
dockercommands. - Services that need to scale independently, for instance a gateway handling far more request volume than the model server behind it can serve, where running multiple gateway replicas would help but running multiple model-server replicas wouldn't fit in memory. Compose can run multiple replicas of a service, but load balancing across them and health-aware routing is thinner than a dedicated orchestrator provides.
- A move from one Spark to several, per the GPU memory management guide, where the math says splitting workloads across nodes makes more sense than budgeting everything onto one machine's 128GB of unified memory. Coordinating containers across multiple physical machines is exactly the problem Kubernetes and similar tools exist for; compose is scoped to a single host.
None of these apply to most single-Spark deployments. A model server, a gateway, and a couple of support services on one machine is squarely inside what compose was built for, and reaching for a heavier orchestrator before hitting one of these signals mostly adds operational overhead: more moving parts to patch, monitor, and understand, for a problem the box doesn't actually have yet.
Practical habits worth having from the start
A few things are worth setting up even in a simple compose deployment, since they cost little now and save real time later:
- Pin image versions rather than
latest, so a container restart doesn't silently pull a different version of a service than the one that was tested. - Keep the compose file itself in version control, even a single-file git repo on the Spark, so a change to service configuration has a history and a rollback path.
- Set explicit restart policies (
unless-stoppedis a reasonable default) so a container that crashes overnight comes back without someone SSHing in to restart it by hand. - Keep volumes for anything stateful outside the container's writable layer, the model server's cache directory, a database's data directory, so recreating a container to pick up a config change doesn't lose data in the process.
For budgeting GPU memory across the containers running on one Spark, see the GPU memory management guide. A first engagement can help design a container setup for a specific multi-service deployment.