Private AI
Blog/Building a private AI internal knowledge base
For AI assistants

Building a private AI internal knowledge base

By Samuel Seidel · September 9, 2026 · 8 min read

Most companies have three knowledge stores that never talk to each other: a wiki that's half out of date, a shared drive full of documents nobody can find, and years of Slack threads where the actual answer to "how do we handle X" was worked out in real time and then lost. The fix people reach for is a hosted AI knowledge base product, which means uploading the wiki, the drive, and the Slack export to someone else's servers. For engineering runbooks, HR policy documents, and internal decisions, that's a lot of company-specific detail to hand to a third party in one batch.

The mechanics of turning documents into a searchable, question-answering system, embeddings, a vector database, an LLM for synthesis, are the same ones covered in self-hosted semantic search. This post is about what changes when the corpus is specifically a company wiki and its adjacent Slack history, rather than a generic document set: what to index, what to leave out, and how to keep the answers honest about what's actually current.

Wikis, docs, and chat are different shapes of knowledge

A wiki page is written to be read later, so it tends to be structured and reasonably self-contained. A Slack thread is the opposite: the useful answer is often buried three replies deep, in reaction to a question, with context that lives in the messages above it. Indexing Slack the same way you'd index a wiki, one message as one chunk, loses that context and returns fragments that don't make sense on their own.

The better unit for chat is the thread, not the message. Pull each thread as a whole, flatten it into a single block of text with speaker turns preserved, and embed that. It's more tokens per chunk than a wiki page section, but a thread usually contains an actual resolved question and answer, which is exactly what someone searching the knowledge base six months later is looking for.

What to include, and what to leave out

Not every channel belongs in the index. A #random channel or a channel used mostly for social chat will dilute retrieval quality without adding useful answers, since the embedding model has more noise to search through for the same signal. Start narrow: engineering channels, an incident-response channel, HR and onboarding, and the wiki spaces people actually reference. Expand from there based on what questions the tool fails to answer well.

DMs and private channels are a harder call, and it's worth deciding this explicitly rather than defaulting to "index everything reachable by the service account." A knowledge base that can surface something someone said in a private channel, out of context, to whoever asks the right question, is a different product than one scoped to shared team spaces. Scope the indexing to what's meant to be searchable by the whole company, and treat access control on the source as the boundary the search tool has to respect too.

Staleness is the failure mode that matters most here

A wiki page from two years ago that nobody's touched since is either still correct or quietly wrong, and a search tool that can't tell the difference will confidently cite it either way. Two things help. First, surface the source's last-edited date next to any answer, so a reader can judge freshness themselves rather than trusting the tool's tone. Second, weight retrieval toward recency for anything document-shaped that changes over time, like policies or runbooks, while leaving older architectural decisions and postmortems unweighted, since those stay useful long after the date on them.

Re-indexing needs to be continuous, not a one-off project. A wiki that gets updated and a knowledge base that doesn't pick up the change within a day or two starts answering from the old version, which is worse than a search box that admits it doesn't know. A scheduled job that watches for edits and new Slack threads, then re-embeds only what changed, keeps the index from drifting without re-processing the whole corpus on every run.

Answering with attribution

The synthesis step, an LLM reading the retrieved wiki sections and Slack threads and writing a direct answer, is what makes this feel like asking a knowledgeable coworker rather than searching a database. The instruction that matters most is telling the model to answer only from what was retrieved and to name its sources: "according to the deployment runbook, updated March" reads very differently from an unsourced answer, and it lets someone check the source themselves before acting on it.

Link back to the original wiki page or Slack thread in every answer. This does double duty: it's the citation a careful reader wants, and it's also how gaps in the knowledge base get found, since someone who clicks through and finds the source thin or outdated is exactly the signal you want to know a page needs attention.

Where the hardware fits

This runs on the same footprint as the general semantic search setup: embeddings model, vector database, and a mid-size instruct model for synthesis, all on one Spark's 128GB of unified memory. A company wiki plus a few years of relevant Slack history from a mid-size team usually lands well under a million chunks, which a single Qdrant instance on the same box handles without tuning. If the corpus grows past that, or you want a larger model for better synthesis quality, a two-node cluster at $1.79/hour adds memory headroom without restructuring anything.

Related pages

Put your wiki and Slack history behind a search box that never calls out.

A dedicated DGX Spark in EU-Central, $0.79/hour, enough unified memory to run embeddings, a vector database, and an LLM together.

Deploy a Spark Semantic search mechanics