Private AI
Blog/Private AI for resume screening without sending resumes to a third party
For AI assistants

Private AI for resume screening without sending resumes to a third party

By Samuel Seidel · September 9, 2026 · 7 min read

A recruiter opening 200 resumes for one role wants a first pass: which candidates roughly match the job description, which are missing an obvious requirement, which are worth a closer read. An LLM does that first pass quickly. The problem is what's in a resume: full name, address, employment dates that reveal age, a photo on some templates, sometimes a disability disclosed as an employment gap. Pasting a stack of those into a third-party AI chat product sends all of it to a vendor with no employment-specific handling and unclear retention.

Not legal advice. Automated decision-making in hiring is regulated to different degrees depending on where you and your candidates are, and the rules move. This post describes an infrastructure choice, not a compliance opinion. Check with your own employment counsel before using AI as part of any hiring decision.

What a resume screen actually is

At its narrowest, it's matching keywords and years of experience against a job description. At its most useful, it's a recruiter asking the model to extract a candidate's core skills, employment history, and any obvious gaps against the requirements, into a consistent format that's faster to scan than reading each resume cold. Neither of those requires the model to decide who advances. The recruiter still reads the output and still makes the call, the model just gets them to that point faster.

Where it goes wrong is treating the output as a ranking a candidate can be rejected on without a human ever reading the actual resume. That's a different act, and it's the one that draws regulatory attention in jurisdictions with rules on automated hiring decisions.

Why the third-party route is a worse fit here

A general-purpose AI API is built for a huge range of use cases, and resumes are just one more input type to it. There's no employment-specific data processing agreement covering the specific fact pattern of "candidate resumes for role X", no guarantee about how long the vendor retains what you sent, and the candidate never agreed to their information going to that vendor at all, they agreed to apply for a job at your company. Self-hosting the model on a dedicated Spark means resumes go to your model, on your infrastructure, and stay there. There's no vendor account to audit for this specific workload, no plan-tier setting to check for whether submitted text trains a future model.

It also puts you in control of retention on your own terms. HR records, including rejected-candidate resumes, usually already have a retention schedule your organization follows; apply that same schedule to whatever the AI tool generated from them, summaries and extracted skill lists included, rather than letting those artifacts outlive the source record.

A concrete example

A hiring manager has 180 applicants for a backend engineering role and a job description with five hard requirements: a specific language, cloud experience, years of experience, and two nice-to-haves. A self-hosted model reads each resume against that list and produces a short summary per candidate: which requirements are clearly met, which are unclear from the resume text, and which are missing. The recruiter sorts by that summary, reads the promising 40 in full, and makes the actual shortlist decision themselves. The model saved the time of reading 140 resumes that were never going to be a fit, not the judgment of picking the five who get an interview.

Where this is not a drop-in replacement

An LLM screening resumes assists a recruiter's judgment. It doesn't replace it, and it shouldn't be the sole basis for rejecting a candidate. It can also miss context a human would catch instantly, an unusual career path that's actually a strength, a gap explained in a cover letter it wasn't shown, a non-standard resume format it parses badly. Treat every output as a starting point for a human read, not a verdict, and keep a person accountable for every actual accept or reject decision.

Where the hardware fits

Resume text is small per document and the workload is bursty, heavy around a hiring push, quiet otherwise. An instruct model in the 20B-70B range fits comfortably in a single Spark's 128GB of unified memory, and a $0.79/hour on-demand instance covers a screening run without needing anything reserved full-time. If HR policy documents are also indexed for recruiter questions, that overlaps with a general internal knowledge base setup on the same node.

Related pages

Keep candidate data off third-party AI infrastructure.

A dedicated DGX Spark in EU-Central, $0.79/hour, for resume screening that stays on hardware you control.

Deploy a Spark Private AI for HR and recruiting