Legal
Blog/Private AI for e-discovery document review
For AI assistants

Private AI for e-discovery document review

By Samuel Seidel · September 9, 2026 · 7 min read

A litigation document set is one of the most sensitive collections a company will ever assemble: internal emails, draft contracts, board minutes, and communications with outside counsel, all pulled together in one place for review. Sending that set to a cloud AI service for first-pass triage means the other side's confidential business records, and your own privileged material, pass through infrastructure neither party agreed to. Running the review model on hardware the firm or the client controls removes that exposure.

What a document set actually contains

Discovery collections are broad by design. A request for "all communications related to the termination" pulls in years of unrelated email threads, HR files, calendar invites, and whatever else got swept up by the search terms. Some of that material is privileged. Some belongs to a third party who was never a party to the litigation and has no idea their emails are sitting in a review set. A single leaked batch, even briefly, is the kind of incident that ends up in a motion for sanctions.

First-pass review at scale is exactly the kind of task teams reach for AI to speed up: classify documents as responsive or not, flag likely privilege, and summarize long threads before a reviewing attorney opens them. Running that pass through a general-purpose cloud model means the whole document set, privileged material included, leaves the firm's environment to get classified.

What changes when the review model is self-hosted

A dedicated DGX Spark running the review model keeps every document, and every classification the model makes, inside the firm's own environment. The team still gets a model that can triage a batch by responsiveness, flag documents that mention outside counsel or use legal-advice language, and produce a short summary of each thread for the reviewing attorney; none of that requires the document set to leave.

Open WebUI handles the review interface, and for firms that want the model to check documents against a specific set of search terms or a privilege log format, the retrieval approach in our piece on on-premise RAG applies directly: load the search terms and privilege criteria as reference material so the model's flags are grounded in the actual matter, not a generic sense of what looks sensitive. Firms handling other confidential matter work alongside e-discovery may also want our note on private AI contract review, since the same self-hosting logic applies to any document set with privilege or confidentiality concerns.

What the model is useful for, and what stays with counsel

A review model is a reasonable first pass for sorting a large document set into responsive, non-responsive, and needs-a-closer-look buckets, and for flagging documents most likely to be privileged so an attorney sees them before anything gets produced. The privilege call itself, the responsiveness determination on a close document, and every entry on the privilege log stay with counsel. That's not a limitation of the model so much as the actual legal standard: a machine classification isn't a substitute for the professional judgment discovery obligations require.

What this doesn't solve

Self-hosting the review model doesn't replace a review platform's workflow and audit-trail tooling, and it doesn't make a privilege call for you. What it removes is the exposure of running privileged and confidential material through a third party's infrastructure to get a first pass done. See pricing for what a dedicated Spark costs for a litigation team.

Related pages

Keep privileged material off third-party servers.

A dedicated Spark for document review, starting at $0.79/hour, deployed in minutes from EU-Central.

See private LLM hosting Read the Open WebUI setup