What is an embedding?
An embedding is a list of numbers, called a vector, that represents a piece of content, usually text, in a way that captures its meaning rather than its exact wording. A model trained to produce embeddings places content with similar meaning close together in that numeric space and content with different meaning further apart, so comparing two embeddings tells you how similar the underlying content is, without either piece of text needing to share the same words.
Why numbers instead of keywords
Traditional keyword search matches literal words: searching "car" won't find a document that only says "automobile." Embeddings solve that by capturing meaning instead of matching text. A sentence about cars and a sentence about automobiles get embedded into vectors that end up close together in the numeric space, even though they share no words at all, because the embedding model learned during training that these concepts are related. Comparing vectors, typically by measuring the angle or distance between them, is how "similar meaning" gets turned into a number a computer can rank and sort.
How embeddings get produced
A specialized embedding model, usually a smaller neural network trained for exactly this purpose, reads a piece of text and outputs a fixed-length vector, commonly a few hundred to a couple thousand numbers. Every piece of content run through the same embedding model lands in the same numeric space, which is what makes comparisons between them meaningful. Embedding models are generally distinct from the larger generative language model that writes conversational responses; a deployment often runs both, side by side, for different jobs.
Where embeddings get used
The most common use in LLM systems is retrieval-augmented generation: a knowledge base of documents gets embedded once and stored in a vector database, and at query time the user's question is embedded too, then compared against the stored vectors to find the most relevant passages. Those passages get inserted into the model's context window so it can ground its answer in them, rather than relying only on what it memorized during training. See what is retrieval-augmented generation for the full pipeline. Embeddings also power semantic deduplication, recommendation, and clustering, anywhere "find things similar to this" is the underlying task.
Embedding vs weight
It's easy to conflate these since both are just numbers inside a model, but they're different things. A model weight is a learned parameter used internally to compute an output; embeddings are that output, a representation of a specific piece of content that gets stored, compared, and searched. A model has a fixed set of weights regardless of what you ask it; it produces a new embedding for every new piece of text you give it.
Where GPUwerk fits
Running an embedding model, and the retrieval pipeline built on top of it, entirely on infrastructure you control keeps the documents being embedded, and the vectors derived from them, off any third-party API. GPUwerk's dedicated Spark nodes can run an embedding model alongside a generative LLM on the same hardware, which is the common setup for self-hosted RAG, at $0.79/hour for a single node or $1.79/hour for a two-node cluster. See on-premise RAG for a fuller look at building that pipeline privately.