Self-hosted semantic search and embeddings
Keyword search over your company's documents fails the moment someone searches for "cancellation terms" and the actual document says "termination for convenience". Semantic search, matching on meaning rather than exact words, fixes that. The usual way to build it is to send every document through a hosted embeddings API. For a public knowledge base that's fine. For contracts, HR files, or internal engineering docs, it means every sentence in your document store passes through a third party before you've even built the search index.
The whole stack, embeddings, vector database, and the LLM that turns retrieved chunks into an answer, can run on hardware you own. None of it requires a hosted API.
The three pieces
Semantic search over documents has three components, and it's worth understanding what each one does before wiring them together:
- An embeddings model turns a chunk of text into a vector, a list of numbers that captures its meaning, so that texts with similar meaning end up as nearby vectors. This runs once per document at index time, and again for every search query.
- A vector database stores those vectors and can quickly find the ones nearest to a query vector, i.e. the document chunks most relevant to what someone searched for.
- An LLM takes the retrieved chunks and the original question, and writes an actual answer instead of leaving the user to read through five document excerpts themselves. This is the "synthesis" step, and it's what makes the system feel like a search assistant rather than a search index.
That combination, retrieve relevant chunks, then have an LLM answer using them, is retrieval-augmented generation (RAG). Semantic search is the retrieval half of it.
Running the embeddings model locally
Open-weight embedding models have caught up to the hosted options for general document search. BGE-M3 (BAAI) and Nomic Embed are both strong, run comfortably on a single GPU, and are small enough that embedding a large document store takes hours, not days. Serve one through an inference server with an OpenAI-compatible embeddings endpoint:
vllm serve BAAI/bge-m3 --host 0.0.0.0 --port 8001 --task embed
Then embed every document, chunked into passages of a few hundred words each, since embedding a whole long document as one vector loses the fine-grained meaning a search query needs to match against:
# embed.py - chunk documents and store vectors in a local Qdrant instance
import requests, uuid
from qdrant_client import QdrantClient
from qdrant_client.models import PointStruct, VectorParams, Distance
qc = QdrantClient(url="http://localhost:6333")
qc.recreate_collection("docs", vectors_config=VectorParams(size=1024, distance=Distance.COSINE))
def embed(text):
r = requests.post("http://spark.internal:8001/v1/embeddings",
json={"model": "bge-m3", "input": text})
return r.json()["data"][0]["embedding"]
def index_chunk(text, source):
vec = embed(text)
qc.upsert("docs", [PointStruct(id=str(uuid.uuid4()), vector=vec,
payload={"text": text, "source": source})])
Qdrant is a good default vector database here: it's open source, runs as a single container, and handles collections of a few million chunks without tuning. For a document store under a few hundred thousand chunks, which covers most internal knowledge bases, it runs fine on the same box as the models.
Answering, not just retrieving
Retrieval alone gives you a ranked list of document excerpts, which is already more useful than keyword search, but the step that makes it feel like a real tool is having an LLM read the top results and write a direct answer with citations back to source. Serve a general instruction model alongside the embeddings model:
vllm serve Qwen/Qwen2.5-32B-Instruct --host 0.0.0.0 --port 8000 # search.py - embed the query, retrieve top chunks, synthesize an answer def search(query, k=5): query_vec = embed(query) hits = qc.search("docs", query_vector=query_vec, limit=k) context = "\n\n".join(f"[{h.payload['source']}]\n{h.payload['text']}" for h in hits) prompt = ( f"Answer the question using only the context below. Cite sources by name.\n\n" f"Context:\n{context}\n\nQuestion: {query}" ) r = requests.post("http://spark.internal:8000/v1/chat/completions", json={ "model": "Qwen2.5-32B-Instruct", "messages": [{"role": "user", "content": prompt}], }) return r.json()["choices"][0]["message"]["content"]
The "answer using only the context below" instruction matters: without it, a capable model will happily fill gaps with what it already knows rather than what your documents actually say, which defeats the point of building this over your own corpus in the first place.
Practical notes from actually running this
Chunk size is the setting that affects quality the most, and it's usually the first thing worth tuning. Too small and chunks lose context (a clause without the contract it's from); too large and the embedding blurs multiple ideas together, hurting retrieval precision. A few hundred words with modest overlap between consecutive chunks is a reasonable starting point.
Re-index on a schedule, not once. Documents change, and a search tool that answers from a stale contract version is worse than no tool at all. A nightly job that re-embeds anything modified since the last run keeps the index current without re-processing the whole store each time.
Keep access control in mind before this becomes company-wide: if the underlying documents have permissions (some people shouldn't see certain HR files or contracts), the search layer needs to respect that too, or it becomes the fastest way to leak something that was previously just hard to find.
Where the hardware fits
The embeddings model, vector database, and a 32B-class synthesis model all fit on a single Spark's 128GB of unified memory with room to spare, so this is genuinely a one-machine internal tool, not a project that needs a cluster. If your document store grows into the millions of chunks or you want a larger synthesis model for better answer quality, a two-node cluster at $1.79/hour adds the memory headroom without changing anything else about the setup.