Private AI for M&A due diligence document review
This is not legal or financial advice, and it doesn't replace outside counsel or a deal team's own judgment. It's about one narrow question: where the documents in a due diligence data room go when someone uses AI to help review them, and what changes when the model runs on hardware you control instead of a vendor's.
What sits in a data room before a deal is announced
A due diligence data room for an acquisition typically holds financial statements not yet public, customer contracts with pricing terms, cap tables, pending litigation summaries, employee compensation data, and often a first draft of the deal terms themselves. All of it is material non-public information under securities law in most jurisdictions, and all of it is sensitive independent of whether the deal ever closes. If the deal falls through, or leaks before signing, the damage isn't hypothetical: stock moves, competitors adjust, and counterparties who trusted the confidentiality of the process stop trusting it.
Deal teams increasingly want AI to help work through the volume: summarizing hundreds of contracts for change-of-control clauses, flagging inconsistencies between representations and underlying financials, or drafting first-pass diligence memos. That's a real efficiency gain on work that used to be junior-associate hours. The question is where the documents go to get that summary.
What changes when the model is on hardware you control
A general-purpose AI API is a shared endpoint. Whatever the vendor's contract says about retention or training, the documents still leave your infrastructure and cross a network boundary you don't control, and a signed NDA with the counterparty doesn't automatically cover what your own AI vendor does with the text you send it. Running the review model on a dedicated DGX Spark instead removes that hop: the deal documents go to a model in your own environment, and nowhere else. There's no vendor log to worry about, no plan-tier distinction between "used for training" and "not," because the text never left your control.
In practice this looks like Open WebUI pointed at the deal team's document set, or a retrieval setup over the data room contents so associates can ask questions against the actual filings instead of re-reading each one. The same architecture we describe for on-premise RAG applies directly here, the difference is the sensitivity of what's being indexed.
How a deal team actually uses this
A typical setup: the data room's contract set gets loaded into a retrieval index on the Spark, and associates ask it questions like "which customer contracts have change-of-control termination rights" or "flag any employment agreements with golden parachute clauses above a threshold." The model returns citations back to the specific document and page, which a human still opens and reads before it goes in a memo. That's the same pattern law firms use for internal document review tools, run on hardware the deal team controls instead of a firm-wide platform, which matters when the deal is one where even the fact of the negotiation is not yet public.
Because a Spark is billed per minute against a prepaid balance, a deal team can spin one up for the diligence window and terminate it once the deal signs or dies, rather than carrying a standing subscription to a document-review platform between deals. The instance and its workspace go away with it, which for MNPI is a feature: nothing about the deal lingers in a system after the reason for having it access is gone.
Where the line still sits with counsel
None of this changes who is responsible for the substance of the diligence. Deal counsel still reviews the actual representations, still runs the real risk assessment, and still decides what goes in the closing memo. An AI-assisted first pass through hundreds of contracts is useful for the same reason a well-organized data room is useful, it lets the lawyers spend their time on judgment calls instead of manually opening every file. It doesn't change the standard of care expected of them, and it shouldn't be treated as having done the legal analysis on its own.
What this doesn't solve
Self-hosting the model doesn't create attorney-client privilege, doesn't substitute for a lawyer's read of a representations and warranties clause, and doesn't change who on your team has access to the data room, that's still governed by your own access controls and the deal's confidentiality agreement. An AI summary of a contract is a starting point for a human reviewer, not a substitute for one, particularly on anything that will end up in a purchase agreement. What a dedicated model removes is one specific exposure: a third-party AI vendor sitting between your deal team and documents that, if they moved at all before signing, would move the stock.