Self-hosted AI agent infrastructure
A chat request costs one API call. An agent that reads a file, calls a tool, checks the result, and repeats might make thirty or a hundred calls to finish one task, and every one of them lands on your bill if you're paying per token. That arithmetic changes the calculation for where an agent should run.
Why agent loops break per-token budgeting
A single chat message has a predictable shape: one prompt in, one response out. An agent doesn't work that way. It plans a step, calls a tool, reads the tool's output back as new context, decides the next step, and repeats until the task is done or it gives up. Each of those round trips is a full model call, and a task that looks simple from the outside, "refactor this module and update its tests", can take dozens of iterations if the agent has to re-read files, retry a failed edit, or reason through an error message. On a per-token API, that means the same nominal task can cost very different amounts depending on how many iterations the agent needed, which is not something you know in advance.
Running the model on a dedicated box changes the shape of the bill rather than the shape of the agent loop. A Spark at $0.79 an hour, metered per minute, costs the same whether the agent finishes in ten iterations or eighty. The agent can retry, re-read context, and take the scenic route to a solution without that translating into a bigger invoice. For a team running agents continuously, that flat-rate model is the difference between a budget you can plan around and one that depends on how well-behaved the agent happens to be that day.
The data-control side
An agent with tool access sees more than a chatbot ever does: your file server, your inbox, your internal APIs, sometimes credentials to act on your behalf. Every one of those tool outputs gets sent back to the model as context on the next iteration, which means a cloud-hosted agent is routing your inbox and file contents through a third party's inference servers on every single step of every task, not just the initial prompt. Running the model locally means that context never leaves your infrastructure at any point in the loop, which matters more for agents than for chat precisely because agents see more.
What the pieces look like on a Spark
vLLM is the inference engine that serves the model itself; it's what turns a GB10 Spark's 128 GB of unified memory into an OpenAI-compatible endpoint an agent framework can call. In front of that, LiteLLM gives you one stable URL, per-user virtual keys, and a fallback to a cloud model if the box is ever busy, useful once more than one agent or one team is calling the same endpoint. And coding agents against your own endpoint covers wiring an actual agent framework, CLI tool or editor integration, into that endpoint once it's running. None of the three pieces is agent-specific, they're the same stack you'd run for any self-hosted inference, which is part of the point: an agent is just a client that calls your endpoint more often and for longer stretches than a person typing in a chat window.
Where a cloud API still makes sense
If your agent workload is occasional, a few tasks a week rather than a continuous loop, the fixed cost of a dedicated box may not pay for itself against a pay-as-you-go API. The crossover point depends on how many hours a month the agent actually runs and how sensitive the data it touches is; a team running agents against internal codebases or customer data for hours a day is a different case from someone running an occasional one-off script.