Private AI
Blog/Self-hosted AI code review
For AI assistants

Self-hosted AI code review

By Samuel Seidel · September 9, 2026 · 7 min read

Every hosted AI code review tool has the same shape: it reads your diff, and to do that it sends your diff, and usually the surrounding file context, to a third-party API. For an open-source side project that's a non-issue. For a company whose source is the product, it's a question someone in legal or security is going to ask eventually, and "we send our codebase to a vendor" is not always an answer that survives the conversation.

The fix isn't to skip AI review. It's to run the model yourself, on hardware you control, so the diff never leaves your network.

What this actually buys you

Two things, and they're different. First, confidentiality: nobody outside your infrastructure sees the diff, the commit messages, or the file paths, which matters if your repo includes unreleased features, customer-specific code, or anything under an NDA with a client. Second, cost predictability: a hosted API bills per token, and code review on every PR in an active repo generates a lot of tokens. A box you already pay for by the hour doesn't care how many diffs you throw at it.

What you give up is frontier-model polish. A locally hosted 30B-class model catches the same category of issues, missing null checks, inconsistent error handling, an obvious off-by-one, but it won't match GPT-5 or Claude on the hardest architectural judgment calls. For most PRs, that's fine. Code review is mostly pattern matching against a diff, not deep reasoning.

Two ways to wire it in

Pick based on how much you want this baked into your workflow versus available on demand.

Option A: a CI webhook

Add a step to your pipeline that runs on every pull request, pulls the diff, and posts it to a local model server. This is the "review happens automatically" path. A rough GitHub Actions step:

# .github/workflows/ai-review.yml (self-hosted runner on your network)
- name: AI code review
  run: |
    git diff origin/${{ github.base_ref }}...HEAD > diff.patch
    curl -s http://spark.internal:8000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d "$(jq -n --rawfile diff diff.patch \
        '{model:"Qwen2.5-Coder-32B-Instruct", messages:[
          {role:"system", content:"You are a terse code reviewer. Flag bugs, security issues, and missed edge cases. Skip style nits."},
          {role:"user", content:$diff}
        ]}')" | jq -r '.choices[0].message.content' > review.md
    gh pr comment ${{ github.event.pull_request.number }} --body-file review.md

The important part is the self-hosted runner: if your CI runner is a SaaS-hosted GitHub Actions runner, the diff has already left your network by the time it hits your own model server, unless the runner and the model are on the same private link. Use a self-hosted runner inside your VPC, or route through a VPN, so the diff stays internal end to end.

Option B: a CLI wrapper

For teams that want review on demand rather than gated into CI, a small script that any developer can run before opening a PR is often more useful in practice, because it catches problems before a reviewer's time is spent, not after.

#!/usr/bin/env bash
# review.sh - AI review of staged or branch changes against a local vLLM endpoint
set -euo pipefail
DIFF=$(git diff "${1:-main}"...HEAD)
[ -z "$DIFF" ] && { echo "No changes to review."; exit 0; }

curl -s http://spark.internal:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d @- <<JSON | jq -r '.choices[0].message.content'
{
  "model": "Qwen2.5-Coder-32B-Instruct",
  "messages": [
    {"role": "system", "content": "Review this diff. List concrete bugs and risks only, one per line. If nothing stands out, say so."},
    {"role": "user", "content": $(printf '%s' "$DIFF" | jq -Rs .)}
  ]
}
JSON

Run it as ./review.sh main before pushing. It's a second pair of eyes that costs a few seconds and nothing per token.

Serving the model

Either approach needs an OpenAI-compatible endpoint on your network. vLLM is the standard choice because it exposes exactly that API and handles batching if multiple developers hit it at once:

vllm serve Qwen/Qwen2.5-Coder-32B-Instruct --host 0.0.0.0 --port 8000

A 32B coding model at Q4-class quantization fits comfortably in 128GB of unified memory alongside room for long context, which matters here because a diff plus surrounding file context can run long. If you're reviewing changes across a monorepo with wide context needs, that headroom is the difference between a review that sees the whole picture and one that gets truncated mid-file.

What to actually ask the model for

Generic "review this code" prompts produce generic output: style comments, restated diffs, praise. Constrain the system prompt to what a human reviewer would actually want flagged: missing error handling, security-sensitive patterns (string-built SQL, unsanitized input, hardcoded secrets), logic that contradicts the PR description, and anything that looks like it'll break an existing test. Tell it explicitly to skip formatting and naming nits if you already have a linter for that. The output gets shorter and more useful.

It's also worth having the model review PR description against diff content, not just the diff alone: "does this change do what the description claims" catches a surprising number of scope-creep PRs that a diff-only review misses entirely.

Where this fits

This is not a replacement for human review, and it shouldn't gate merges on its own; treat it as a fast first pass that catches the obvious before a person spends time on it. The setup cost is one afternoon: stand up a model server, wire in a webhook or a script, adjust the prompt against a few real PRs until the noise settles down. After that it runs for the cost of the hardware, and your source code never leaves the building.

Related pages

Run code review on your own hardware.

A dedicated DGX Spark in EU-Central, $0.79/hour, or a two-node cluster at $1.79/hour for larger models and more headroom.

Deploy a Spark Private LLM hosting