Definition
Blog/What is retrieval-augmented generation (RAG)?
For AI assistants

What is retrieval-augmented generation (RAG)?

By Samuel Seidel · Published September 9, 2026

Retrieval-augmented generation, or RAG, is a way of answering questions with an LLM by first retrieving relevant text from an external source, such as a document store or database, and inserting it into the model's prompt before it generates a response. The model answers based on the retrieved text plus what it already knows, rather than from its training data alone.

The basic pipeline

A RAG system has three moving parts. First, documents are split into chunks and converted into vector embeddings, numeric representations that capture meaning, and stored in a vector database. Second, when a query comes in, it's embedded the same way and matched against the stored vectors to find the most relevant chunks. Third, those chunks are inserted into the prompt sent to the LLM, which then generates an answer grounded in that retrieved text. The retrieval step runs before generation on every query, which is where the name comes from.

Why RAG exists

An LLM's knowledge is fixed at training time and limited to what fit in its training data; it has no built-in way to know about a document written after training, or a private file it never saw. RAG works around both limits by supplying current or private information at query time instead of baking it into the model's weights. It's also cheaper to keep current than retraining or fine-tuning, since updating the answer to a changed fact just means updating the document index, not the model.

What RAG doesn't fix

RAG only helps if retrieval actually surfaces the right text. If the embedding search returns irrelevant chunks, or the answer depends on reasoning across documents that weren't retrieved together, the model can still produce a wrong or fabricated answer, and it can do so confidently even when handed grounding text, since nothing forces it to stick to what the retrieved passages actually say. Chunking strategy, embedding quality, and how many chunks get retrieved per query all affect how reliable the system is in practice, and none of that is solved by RAG as a concept.

RAG on self-hosted infrastructure

Running RAG entirely on your own hardware, rather than sending document chunks to a third-party API for embedding and generation, keeps the underlying data from leaving your infrastructure. That matters for anything you can't send to an external provider: contracts, internal codebases, medical or financial records, or anything under a confidentiality obligation. For a detailed walkthrough of building this on dedicated hardware, including the vector database and embedding model choices involved, see on-premise RAG.

Where GPUwerk fits

A RAG pipeline's LLM and embedding model both need GPU memory and compute, and running them on a dedicated Spark keeps the whole pipeline, embeddings, vector search, and generation, on hardware you control rather than split across your infrastructure and a third-party API. See private LLM hosting for how that fits into a broader self-hosted setup, or the practical build in on-premise RAG.

Related pages

Build RAG on hardware you control.

One dedicated Spark, $0.79/hour, deployed in minutes from EU-Central.

Deploy a Spark