Batch processing large inference jobs on a rented Spark
Bulk document processing, dataset labeling at scale, running a fixed prompt against a large pile of inputs, is a different workload from an interactive chat endpoint, and it deserves different settings. For the conceptual difference between batch and interactive inference, see the batch vs. interactive inference post; this page is the hands-on version: how to actually configure and run a large batch job on a Spark you're renting by the hour.
Why batch settings differ from serving settings
An interactive endpoint optimizes for latency on each individual request, one user is waiting on the response right now. A batch job optimizes for total throughput across the whole job, nobody's watching any single request, so the engine can queue many requests and process them together, maximizing GPU utilization at the cost of any one item's latency. The practical consequence: batch jobs should run with higher concurrency and larger effective batch sizes than you'd ever use for a chat endpoint, and the metric that matters shifts from time-to-first-token to total job completion time and tokens processed per hour.
Running a batch job with vLLM
vLLM's offline batch inference API (LLM.generate() on a list of prompts, rather than the OpenAI-compatible server) is the most efficient path for a large fixed job, since it skips the HTTP layer and lets vLLM's continuous batching schedule the whole list at once:
from vllm import LLM, SamplingParams
llm = LLM(model="/path/to/model", gpu_memory_utilization=0.9)
params = SamplingParams(temperature=0, max_tokens=512)
prompts = [build_prompt(row) for row in dataset]
outputs = llm.generate(prompts, params)
Set gpu_memory_utilization higher than you would for a mixed-use box, per the memory management guide, since a dedicated batch run doesn't need headroom for other resident models. If your job is large enough to not fit in memory as a single Python list of prompts, chunk it and call generate() per chunk rather than trying to hold everything at once; vLLM's internal batching handles concurrency within a chunk on its own, you don't need to manually throttle beyond that.
If you'd rather use the OpenAI-compatible HTTP server (useful if your pipeline already speaks that protocol, or if the job runs from a language other than Python), raise --max-num-seqs above the default and fire requests concurrently from the client, since the server-side scheduler needs enough in-flight requests to batch effectively; a client sending one request at a time defeats continuous batching regardless of server-side settings.
Sizing the job against the hourly rate
Renting a Spark at $0.79/hour (or $1.79/hour for the two-node cluster, per pricing) makes the cost model straightforward for a batch job: total cost is hours run × hourly rate, so the only variable worth optimizing is job duration, tokens processed per hour. Two changes typically move that number more than anything else: enabling prompt caching if your job shares a fixed instruction template across every item (a common case in labeling and extraction jobs), and choosing the smallest model that still hits the accuracy bar for the task, since a 27B model finishing a job in a third of the time a 70B model takes is very often the better economic choice even before accounting for the per-token cost difference.
Before committing a large job, run a timed sample of 50-100 representative items and extrapolate; job duration on real data rarely matches a back-of-envelope estimate based on average token counts, since real inputs vary in length far more than synthetic test prompts do.
When the job doesn't fit one node
A batch job that's simply large in item count, not per-item complexity, generally doesn't need the two-node cluster; running it longer on one node is usually cheaper than doubling the hourly rate, since batch jobs are throughput-bound rather than latency-bound and don't benefit from the cluster's networking the way a single very large model would. The cluster is the right call when the model itself is too large for one node's memory, not when the queue is simply long; see the memory management guide for that distinction.
Checkpointing and resumability
A job running for hours on a rented instance should write results incrementally rather than holding everything in memory until the end, both because a crash partway through shouldn't cost the whole job, and because it lets you monitor progress and catch a systematic formatting problem early rather than after the full run completes. Write completed items to disk (or to your target destination) as they finish, and track which input rows have been processed so a restarted job skips completed work instead of reprocessing it.
Practical checklist
- Use vLLM's offline batch API (
LLM.generate()) for large fixed jobs rather than the HTTP server, when your pipeline can call it directly. - Set memory utilization and concurrency higher than interactive-serving defaults; a dedicated batch run doesn't need headroom for other resident models.
- Enable prompt caching if the job shares a fixed template across items, and default to the smallest model that clears your accuracy bar.
- Time a representative sample before committing to a full run; real-data duration rarely matches a token-count estimate.
- Checkpoint incrementally so a crash or restart doesn't cost the whole job.
See batch vs. interactive inference for the conceptual framing, or private AI for data labeling and annotation for the labeling-specific case.