Operations guide
Blog/Load balancing across multiple Sparks
For AI assistants

Load balancing across multiple Sparks

By Samuel Seidel · September 9, 2026

One Spark handles a lot of concurrent inference load before it needs help, but every model has a concurrency knee past which one node isn't enough. Once you're past yours, load balancing across two or more Sparks is a smaller change than it sounds, mostly configuration, not new infrastructure.

Know your knee before you add a node

Our benchmarks page measured a dense 70B model's aggregate throughput on one GB10 node peaking at 485 tok/s around concurrency 128, then falling to 300 tok/s at concurrency 256, with median time to first token going from 8 seconds to 40. A mixture-of-experts 30B model on the same hardware instead plateaus, still gaining a little aggregate throughput at concurrency 256 without collapsing. If you're running a dense model and your real traffic regularly pushes past its knee, load balancing across nodes fixes what no further tuning on one node will: see capacity planning for LLM workloads for the full arithmetic behind that call.

The simplest setup: identical nodes, least-connections routing

Run the same model on two or more Sparks and put a load balancer in front that routes each new request to whichever node currently has the fewest active connections. This matters more for LLM traffic than for typical web traffic, because request duration varies enormously with output length: a round-robin balancer can stack three long generations on one node while another sits idle. Least-connections avoids that by tracking actual load rather than assuming requests are roughly equal in cost.

# /etc/nginx/conf.d/spark-pool.conf
upstream spark_pool {
    least_conn;
    server 10.0.0.11:8000 max_fails=2 fail_timeout=10s;
    server 10.0.0.12:8000 max_fails=2 fail_timeout=10s;
}

server {
    listen 443 ssl;
    server_name inference.internal;

    location / {
        proxy_pass http://spark_pool;
        proxy_read_timeout 300s;   # generation can run long, don't time out a healthy stream
        proxy_buffering off;       # required for token streaming to reach the client live
    }
}

The proxy_buffering off line matters specifically for LLM traffic: nginx buffers responses by default, which would hold back streamed tokens until the buffer fills or the response completes, defeating the point of a streaming API.

HAProxy, if you want per-node health awareness

HAProxy's active health checks pull a node out of rotation automatically the moment it stops responding, which nginx's passive max_fails approach does more slowly. Worth the extra setup once you have more than two nodes or the load balancer itself is a single point of failure you're trying to shrink.

# /etc/haproxy/haproxy.cfg
backend spark_pool
    balance leastconn
    option httpchk GET /health
    server spark1 10.0.0.11:8000 check inter 5s fall 2 rise 2
    server spark2 10.0.0.12:8000 check inter 5s fall 2 rise 2

The /health path here should be a real health check, not just a port-open probe, the same distinction covered in writing health checks for your inference endpoint. A node that's up but stuck should be pulled from rotation just as fast as one that's fully down.

Sticky sessions are usually the wrong instinct

Unlike a stateful web app, a single inference request to a stateless serving endpoint doesn't need to land on the same node as a previous one, there's no server-side session to preserve between requests. Don't add session affinity unless you're doing something unusual like node-local prompt caching; it only reduces the load balancer's ability to spread load evenly.

Two nodes is a cluster, not just a pair

Once you're running more than one Spark behind a shared entry point, the operational questions multiply: how you roll out a model update without taking capacity offline, see blue-green deployments for model updates, and how far this scales before the coordination overhead outweighs adding more nodes, covered in scaling to a fleet. If you're deciding whether to go from one node to a small cluster at all, single Spark vs. cluster is the place to start that comparison.

FAQ

Round robin or least-connections for load balancing across Sparks?

Least-connections, almost always. LLM requests have wildly different durations depending on output length, so round robin can send a new request to a node already handling several long generations while a node that just finished sits idle. Least-connections routes to whichever node currently has the fewest active requests, which tracks actual load far better for this workload.

Does load balancing require the same model on every node?

Not necessarily, but it's the simplest setup and the one worth starting with. Routing different request types to different models on different nodes works too, but it means the load balancer needs to inspect the request, not just distribute it blindly, which is a meaningfully more complex configuration.

How do I know when I need a second node instead of just tuning the first?

When your measured per-request throughput at the concurrency you actually need has already dropped below what your use case tolerates, and you're running a dense model, where that drop is steep rather than gradual. Check current numbers against the concurrency knee described on the benchmarks page before assuming another tuning pass will fix it.

Related pages

Add a second node to a fleet you control.

Dedicated Sparks on the same network, ready to sit behind your own load balancer.

Read the two-node cluster guide Read the benchmarks page