Private AI
Blog/AI-assisted loan underwriting support without sending applicant financials to a third party
For AI assistants

AI-assisted loan underwriting support without sending applicant financials to a third party

By Samuel Seidel · September 9, 2026 · 7 min read

A loan file, whether it's a small-business term loan or a mortgage, is a bundle of bank statements, tax returns, pay stubs, existing debt schedules and a credit report. An underwriter reads all of it to build a picture of repayment capacity before a decision gets made. An LLM can speed up the reading part: pull the numbers into a consistent summary, flag inconsistencies between what an application states and what a bank statement shows, note missing documents. What it should never do is decide whether the loan gets approved.

Not financial or legal advice. This post describes an infrastructure setup, not underwriting guidance. Nothing here should be read as advice on credit policy, regulatory compliance for lending decisions, or fair-lending obligations, all of which vary by jurisdiction and product. Check with your own compliance and legal counsel before using AI anywhere near a credit decision.

Why applicant financial data doesn't belong on a third-party API

A loan application is dense with the kind of personal financial data that regulators, and applicants, expect to be handled carefully: account balances, income sources, existing debts, sometimes years of transaction history. Running that through a general-purpose AI API means a third party now has a copy of someone's financial life, subject to whatever retention and training policy applies to that vendor's plan tier, which is rarely something a lender can fully verify or control. For a regulated lender, that's a data-handling question on top of the underwriting one, and it complicates any answer to "where does applicant data go" that a regulator or auditor might ask.

Self-hosting the model on a dedicated instance changes that: the loan file goes to a model you run, on hardware you control, and it doesn't leave. No vendor log retention question, no ambiguity about whether a document became training data. That's the same point covered in more general terms in private AI for financial analysis, applied here specifically to the underwriting file rather than financial statements broadly.

What it's actually good for

The clearest win is document triage and summarization: turning a folder of PDFs into a structured summary of stated income, verified income from bank statements, existing monthly debt obligations, and any red flags like a large unexplained deposit or a recent account opening that doesn't match the applicant's stated history. A model can also do a first pass at reconciling numbers across documents, catching the case where a tax return and a pay stub imply different annual income, which is exactly the kind of cross-checking that's tedious for a person and mechanical for a model.

A second use is drafting the underwriter's write-up: a first-pass summary memo an underwriter edits and signs off on rather than writing from a blank page, which saves time on the repetitive parts of the file while leaving judgment where it belongs.

Where it stops being a drop-in replacement

The model does not make a credit decision, and it shouldn't be positioned as making one, even informally. It has no accountability for a wrong call, no ability to apply the judgment a human underwriter applies to a borderline file, and no visibility into factors a regulator would expect a human to weigh, like fair-lending considerations that go beyond what's in the numeric data. Any output that gets close to a decision, a risk score, an approve/deny recommendation with a stated confidence, needs to be treated as a draft for a licensed underwriter to review, not as the underwriting itself.

It's also not the tool for anything requiring specialized credit-risk modeling. Traditional underwriting models trained and validated on historical default data, tuned to a specific portfolio and regulatory framework, do a job a general-purpose LLM was never built for. Keep the LLM in the summarization-and-flagging role and the underwriting model, if you have one, doing the actual risk scoring.

A concrete example

A small-business lender runs three months of bank statements plus the application form through a prompt template that extracts: stated monthly revenue, deposit pattern consistency, any NSF or overdraft events, and existing debt payments visible in the statements. The output is a one-page summary an underwriter reviews in a few minutes instead of scrolling through fifty pages of statements manually. The underwriter still pulls the credit report, still makes the call, still signs the file. The model just got them to that point faster, on infrastructure that never sent the applicant's bank statements anywhere outside the company.

Where the hardware fits

A loan file's document set is usually a few hundred pages at most, well within a single model's context window without chunking. A DGX Spark's 128GB of unified memory runs a large instruct model for this kind of extraction and summarization comfortably. At $0.79/hour on-demand, or $0.59/hour to keep a Spark reserved between sessions, it's a low-cost way to keep applicant financial data off third-party infrastructure while still getting the speed benefit of AI-assisted review. See the pricing page for the full breakdown.

Related pages

Keep loan files off third-party AI infrastructure.

A dedicated DGX Spark in EU-Central, $0.79/hour, for underwriting support that stays on hardware you control.

Deploy a Spark Private LLM hosting