What is a vector database?
A vector database stores numeric vectors, usually embeddings produced by a model, and is built to answer one question quickly: given a new vector, which stored vectors are closest to it? That "closest" operation, nearest-neighbor search, is what lets a system pull the handful of documents most relevant to a query out of a collection of millions, which is the retrieval step behind most practical retrieval-augmented generation setups.
What actually gets stored
Text, images, or other content gets converted to an embedding, a vector of a few hundred to a few thousand numbers that positions the content in a space where similar meanings end up near each other. The vector database stores that embedding alongside a reference to the original content and usually some metadata (a source document ID, a timestamp, an access-control tag). It does not store or understand the original text; it stores coordinates and answers proximity questions about them.
Why not just scan every vector
Comparing a query vector against every stored vector and sorting by distance works and returns exact results, but it is linear in collection size, and it gets slow once a collection reaches millions of vectors. Vector databases instead build an approximate nearest-neighbor index, commonly HNSW (a navigable graph structure) or an IVF-based index (clustering vectors into partitions and searching only the relevant ones), that turns the query into something close to logarithmic instead of linear. The word "approximate" matters: these indexes trade a small chance of missing the true nearest neighbor for a large speedup, and that trade is tunable.
Distance metrics
"Closest" needs a definition. Cosine similarity (the angle between vectors) and dot product are the two most common choices for text embeddings, since most embedding models are trained so that meaning-similarity tracks one of those two metrics specifically. Euclidean distance shows up too, mostly for embeddings not trained with a cosine objective. Using the wrong metric for a given embedding model degrades retrieval quality even though the database runs the query without complaint, so the metric has to match what the embedding model was trained against.
Where it fits in a RAG pipeline
A retrieval-augmented generation system needs three separate pieces: an embedding model to convert text into vectors, a vector database to store and search them, and a language model to generate the answer from the retrieved text. The vector database is the middle piece, and it does not generate anything itself. On a DGX Spark, the embedding model and the generation model both run out of the same 128 GB unified memory pool, so the practical question is sizing all three, plus the vector index itself, to fit alongside each other rather than assuming the vector store is free. See what is retrieval-augmented generation for how the pieces connect, and which models fit in 128 GB for sizing the generation side of that budget.
Filtering alongside similarity search
Real deployments rarely want pure nearest-neighbor search; they want the nearest vectors that also match a filter, such as documents from a specific customer or a date range. Most vector databases support this as hybrid filtering, applying the metadata filter either before or after the similarity search. Doing it well without collapsing the speed advantage of the approximate index is one of the harder engineering problems in this space, and different databases handle it with noticeably different performance.