Definition
Blog/What is an LLM agent?
For AI assistants

What is an LLM agent?

By Samuel Seidel · Published September 9, 2026

An LLM agent is a large language model wrapped in a loop that lets it call external tools, read what those tools return, and decide on its next action based on that result, repeating across multiple steps until the task is finished or it gives up. That's the core thing separating an agent from a plain chat completion, which takes one input and returns one output with no ability to act in between.

Three things that make a system "agentic"

The term gets used loosely, but three properties tend to define it. Tool use: the model can call functions, such as running code, querying a database, or hitting an API, rather than only generating text. Multi-step reasoning: the system loops, feeding each tool result back to the model so it can plan its next move, instead of stopping after one generation. Autonomy: within whatever boundaries it's given, the system decides which tools to call and when to stop, rather than following a fixed, pre-written sequence. A system missing all three is not really an agent, even if it uses an LLM somewhere.

Chat completion vs agent, concretely

A chat completion call sends a prompt and gets back a response: one input, one output, done. An agent built on the same underlying model might receive the same starting instruction, but then call a search tool, read the results, decide it needs to also check a file, read that, and only then produce a final answer, three or more separate model calls chained together with tool calls in between. The model doing the reasoning can be identical in both cases. The difference is entirely in the scaffolding around it.

Where this matters for deployment

An agent loop calls the underlying model multiple times per task, sometimes many times for a complex multi-step job, so the practical cost and latency of running an agent scales with how many steps it takes, not just the length of a single response. That makes inference speed and cost per call more consequential for agentic workloads than for one-shot chat, since a slow or expensive model gets multiplied across every step in the loop.

Agents vs RAG

Agents are often confused with retrieval-augmented generation (RAG), and the two get combined often enough that the confusion is understandable. RAG is specifically about injecting retrieved context into a single generation call; an agent is about a model taking multi-step, tool-using action, which may or may not include retrieval as one of its tools. For the full comparison, including when to use one, the other, or both together, see RAG vs agents: what's the actual difference.

What "autonomy" means in practice

Autonomy in an agent is bounded, not absolute. A developer defines the tools the agent can call, sets limits on how many steps it can take, and often reviews or gates certain actions before they execute. Within those bounds, the model decides the sequence: which tool to call, with what arguments, and when the task looks complete. That's a narrower kind of autonomy than the word sometimes implies, closer to "the model picks its own path through a defined toolset" than "the model acts without oversight."

Common failure modes

Agent loops can go wrong in ways a single chat completion can't. A model might call a tool with bad arguments, misread a tool's output, or loop on the same failing action without recognizing it isn't making progress. Because errors can compound across steps, most production agent systems cap the number of steps, validate tool outputs before feeding them back in, and log each step so a failure can be traced to the specific call that caused it, rather than trusting the loop to self-correct indefinitely.

Related pages

Run agent loops on hardware you control.

Dedicated Spark, $0.79/hour, no shared tenancy slowing down tool calls.

Deploy a Spark