Private AI for QA and software testing
A QA engineer asking a cloud AI tool to write test cases, or a developer pasting a diff in for review, is sending your source code, your API surface, and often your test data to a third party's servers. For most codebases that's a minor risk. For anything under an NDA, anything pre-acquisition, or anything a security team would flag if it showed up in a vendor's training pipeline, it's a real problem. Running the model yourself removes that question rather than managing it.
What QA teams actually use LLMs for
Three tasks come up repeatedly. Test case generation from a spec or a set of acceptance criteria, where the model reads a ticket and drafts edge cases a human might not think of first, like empty inputs, boundary values, and concurrent-access scenarios. Code review assistance, where the model reads a diff and flags things like an unhandled error path, a missing null check, or a change that looks inconsistent with the rest of the file. And bug triage, where the model reads an incoming bug report plus a stack trace and suggests a likely root cause or duplicate ticket, which speeds up the first ten minutes of investigation before an engineer even opens the code.
None of these need a frontier model. They're pattern-matching and summarization tasks over text you already have, which is exactly the kind of workload a mid-size open model handles well on a single GPU.
Why the codebase question matters
A cloud coding assistant's terms usually say your prompts and code snippets aren't used for training on paid tiers, and most vendors mean it. But "usually" and "most" are not the same as "never," and the actual guarantee depends on which product tier a given engineer is signed into that week, whether they're using a personal account instead of the company's enterprise agreement, and whether a browser extension is quietly routing text through a different endpoint than the one procurement approved. Self-hosting sidesteps the whole question: the code never leaves infrastructure your own team controls, so there's no vendor policy to audit and no tier mismatch to worry about.
This matters more for some codebases than others. A team building an internal admin tool has less at stake than one working on a security product, a trading system, or code under an active NDA with a customer. Match the effort to what's actually being protected rather than treating every repo the same way.
What quality actually looks like
Here's the honest tradeoff: a self-hosted 30-70B class model is noticeably behind GPT-5 or Claude on code review depth. It catches the obvious issues, an unclosed resource, a missing await, an off-by-one in a loop bound, but it misses subtler bugs a frontier model would flag, like a race condition that only shows up under specific ordering, or a security issue that requires reasoning across multiple files at once. If your team is used to a frontier model's review quality, a self-hosted model will feel like a downgrade on the hardest cases.
Where it holds up well is volume work: generating a first draft of test cases from a spec, summarizing a long bug report into a triage-ready one-liner, or doing a first pass over a diff to catch the mechanical issues before a human reviewer looks at the substantive ones. Treat it as a first-pass filter, not a replacement for a senior engineer's review on anything that matters.
Setup effort, honestly
Running an open coding model like Qwen2.5-Coder or a similar instruct model isn't a one-click install. Someone on the team needs to stand up an inference server, wire it into whatever IDE plugin or CI hook the team wants to use, and tune the prompt templates so the model's output format actually matches what your test framework or ticketing system expects. Budget a few days of setup for a small team, more if you want it integrated into CI so every PR gets an automatic first-pass review comment. It's not free, and a team expecting Copilot-level polish out of the box will be disappointed by the raw experience before the integration work is done.
Once it's running, maintenance is light: updating the model occasionally as better open weights ship, and revisiting the prompt templates when the team's conventions change.
A concrete workflow
A pattern that works well: wire the model into a pre-PR hook that reads the diff and posts a comment with a first-pass review before a human even looks at it, catching the mechanical stuff, unhandled errors, obvious style deviations, missing tests for new branches, so the human reviewer's time goes to logic and design questions instead. Separately, point the model at a ticket queue once a day to draft test cases for anything newly marked ready-for-QA, which a QA engineer then edits rather than writes from scratch. Neither task requires the model to be right every time; it needs to be right often enough to save more time than it costs to review its output.
Where the hardware fits
A 30-70B class coding model fits comfortably in a single Spark's 128GB of unified memory with room for a long context window, which matters for code review since a meaningful diff plus surrounding file context can run to several thousand tokens. At $0.79/hour in EU-Central, a small QA or platform team can run this continuously for a fraction of what a per-seat AI coding tool costs across a dozen engineers, while keeping the codebase off third-party infrastructure entirely.