RAG vs agents: what's the actual difference, and when do you need which
These two terms get used interchangeably often enough that it's worth being precise, because they describe genuinely different mechanisms. RAG is about what goes into a single prompt. Agents are about what happens across a sequence of decisions. A system can do one without the other, and the most capable systems typically do both, but conflating them makes it hard to reason about what you're actually building.
RAG: retrieve, then generate once
Retrieval-augmented generation is a two-step pattern that runs once per request. First, given a query, a retrieval system searches an index, usually a vector database or a search engine, for text relevant to that query. Second, the retrieved text gets inserted into the prompt alongside the original question, and the model generates one response using that context. The model itself doesn't decide to retrieve anything; retrieval happens as a fixed step in the pipeline before the model is even called.
This is why RAG is well suited to grounding answers in current or private documents: the model answers using text it was just handed, rather than only what it memorized during training, and the underlying documents can be updated without touching the model at all. It's also why RAG is comparatively easy to reason about and debug. There's one retrieval step, one generation call, and a clear, inspectable path from "what was retrieved" to "what was said." If the answer is wrong, you can check whether the retrieval found the right document or whether the model misread text it was correctly given.
Agents: decide, act, observe, repeat
An agent is a different kind of system. Instead of one fixed retrieve-then-generate pass, an agent runs a loop: the model looks at the current state of the task, decides what to do next, which might be calling a tool, running a search, writing a file, executing code, or asking a clarifying question, observes the result of that action, and decides on the next step based on what just happened. This continues across multiple turns until the model decides the task is done.
The defining feature is that the sequence of steps isn't fixed in advance. A RAG pipeline always does the same two steps in the same order. An agent's path through a task can branch: it might retrieve a document, decide the document doesn't answer the question, search again with a different query, then call a different tool entirely, all in response to what it observed at each step rather than a script written ahead of time. That flexibility is also the source of an agent's harder failure modes: a wrong decision early in the loop compounds into wasted steps or a wrong final action in a way a single RAG call can't, because there's no loop for the error to propagate through.
Where the two get combined
In practice, retrieval is one of the more common tools an agent reaches for. An agent working through a support ticket might decide, as one of its steps, to query a knowledge base, essentially invoking RAG as a sub-routine, then use what it retrieved to decide its next action, maybe drafting a reply, maybe escalating, maybe searching again with a narrower query if the first result didn't answer the question. From the agent's perspective, retrieval is just another tool it can call when it judges that useful, not a fixed step that always runs.
This combined pattern shows up clearly in coding assistants: an agent working on a codebase decides when to search the repository for relevant files (retrieval), when to read a specific file in full, when to run a command and observe the output, and when to edit code, choosing among those actions turn by turn based on what it's learned so far rather than following one fixed sequence. Our guide to coding agents covers how that loop is typically structured in practice, and our LiteLLM docs cover the routing layer that sits underneath both RAG pipelines and agent loops when you're serving more than one model behind a single endpoint.
A quick way to tell which one you're building
Ask whether your system's sequence of steps is fixed or decided at runtime. If it's always "look something up, then answer," you're building a RAG pipeline, whether or not anyone on the team calls it an agent. If the model is choosing, step by step, what to do next based on what already happened, including whether to look anything up at all, you're building an agent, and RAG might be one of the tools available to it rather than the whole system.
Neither is more advanced than the other in the abstract. RAG is simpler, more predictable, and easier to debug, which makes it the right choice for a large share of question-answering and document-grounding tasks. Agents earn their added complexity when the task genuinely requires the model to make sequential decisions and take actions, not when a fixed retrieve-then-answer pipeline would have done the job with less to go wrong.