Running embeddings and vector search on a DGX Spark
This is the setup guide: how to get an embedding model and a vector database actually running on one Spark, alongside your generation model, with a memory budget that doesn't blow up when the corpus grows. For the concept-level case for self-hosting semantic search, or the broader RAG-versus-agents framing, see the semantic search and RAG vs. agents posts; this page assumes you've already decided to build a RAG pipeline and covers the mechanics.
What runs where
A RAG pipeline on one Spark typically has three components sharing the box's 128 GB of unified memory: your generation model, an embedding model, and a vector database. The embedding model only runs during ingestion and at query time to embed the user's question, it's not in the hot path for every generated token the way the main model is, so it doesn't need to stay loaded continuously if memory is tight. The vector database, by contrast, needs its index resident (or memory-mapped) whenever a query comes in, and its footprint scales with corpus size rather than model size.
Choosing and serving the embedding model
Embedding models are small relative to generation models, most useful open embedding models run in the 100M to 8B parameter range, so a strong one is a modest addition to the box's memory budget, typically well under 5 GB even for a larger multilingual model at a reasonable quantization. Serve it through the same engine you're already running: vLLM supports embedding models via its /v1/embeddings endpoint (launch with --task embed for a model that supports it), and llama.cpp's server exposes a similar embeddings endpoint for GGUF embedding models. Running it as a separate small process alongside your main generation model is usually simpler than trying to multiplex both onto one engine instance.
Model choice affects retrieval quality more than almost anything else in the pipeline. A default or small embedding model is often tuned for size over accuracy; for anything beyond a quick prototype, benchmark a couple of candidates against your own documents rather than assuming the first one you try is good enough, retrieval quality varies significantly by domain and language, and there's no single model that's best across all corpora.
Choosing a vector store
For a corpus in the low millions of chunks or fewer, an embedded vector database that runs in-process or as a lightweight local service (Chroma, Qdrant in embedded mode, or a Postgres instance with pgvector if you already run Postgres) avoids adding another network hop and another service to operate. These handle approximate nearest-neighbor search well within that range without needing a dedicated cluster. Move to a standalone vector database server, still self-hosted on the same box or a neighboring one, only once you need features an embedded store doesn't offer: multi-tenant isolation, hybrid search combining dense and sparse retrieval, or a corpus large enough that index build time becomes a real operational concern.
Memory budgeting
Three things draw from the same 128 GB pool, and none of them are optional to account for:
- The generation model's weights and KV cache, sized per the memory management guide if you're running it alongside anything else.
- The embedding model, small on its own, but still a fixed cost while it's loaded.
- The vector index, which scales with corpus size and embedding dimension. A rough estimate: number of chunks × embedding dimension × 4 bytes (for float32 vectors) plus index overhead, which varies by index type but commonly adds 20-50% on top of the raw vector data for graph-based indexes like HNSW. A million chunks at 1024 dimensions in float32 is roughly 4 GB of raw vectors before index overhead, small next to a 70B model's weights, but worth checking against your actual corpus size rather than assuming it's negligible.
Most RAG deployments on a single Spark are generation-model-bound, not vector-index-bound, the index rarely competes seriously for memory unless the corpus is very large or the embedding dimension is unusually high. Confirm this for your own numbers before assuming it.
A basic ingestion and query flow
The ingestion side: chunk your documents, embed each chunk through the embedding model's endpoint, and write the resulting vectors plus metadata into the vector store's index. The query side: embed the incoming question with the same model (embedding a query with a different model than the one that indexed the corpus produces meaningless similarity scores), retrieve the top-k nearest chunks from the vector store, and pass them as context into the generation model's prompt, structured with the retrieved context near the end of the prompt if you're also using prompt caching, since the retrieved content changes per query and shouldn't sit in the static, cacheable portion of the prompt.
Practical checklist
- Serve the embedding model through the same engine (vLLM or llama.cpp) you're already running, as a separate small process rather than multiplexed onto the generation model.
- Benchmark embedding model choice against your own documents; don't assume a default is good enough for anything beyond a prototype.
- Start with an embedded vector store for corpora in the low millions of chunks; move to a standalone server only when you need features it doesn't have.
- Budget vector index memory against corpus size and embedding dimension, not against a rule of thumb, and check it against the generation model's footprint rather than assuming it's negligible.
- Always embed queries with the same model that indexed the corpus; a mismatch silently degrades retrieval without an obvious error.
See why self-host semantic search and RAG vs. agents for the conceptual background, or the memory management guide for budgeting alongside other models on the same box.