Operations guide
Blog/Setting up a CI pipeline for model updates
For AI assistants

Setting up a CI pipeline for model updates

By Samuel Seidel · September 9, 2026

A model swap done by hand, SSH in, stop the old process, pull the new weights, start it back up, works fine right up until it's done at the wrong hour by someone in a hurry, and the endpoint is down for longer than the swap itself should have taken. Blue-green deployments for model updates covers the swap pattern that avoids the downtime; this is about the pipeline that decides a new model version is safe to swap to in the first place, so the decision doesn't rest on someone eyeballing it.

What the pipeline should validate before anything ships

Two checks catch most of what actually goes wrong with a model update: the new version loads at all, and it produces a sane response to a basic prompt. A corrupted download, a quantization format the inference engine doesn't support, a config mismatch, all show up as a failure to load rather than a subtle quality problem, and catching them here means they never reach a production swap.

# validate-model.sh, run against a staging instance
#!/bin/bash
set -euo pipefail

MODEL_PATH="$1"
PORT=8001

vllm serve "$MODEL_PATH" --port $PORT &
SERVER_PID=$!

for i in $(seq 1 30); do
  curl -s "http://localhost:$PORT/health" && break
  sleep 2
done

RESPONSE=$(curl -s -X POST "http://localhost:$PORT/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{"model":"'"$MODEL_PATH"'","messages":[{"role":"user","content":"Say ok."}],"max_tokens":10}')

kill $SERVER_PID

echo "$RESPONSE" | jq -e '.choices[0].message.content | length > 0'

This is also the natural place to run the staging checks from setting up a staging environment for model testing, a fixed set of prompts with expected properties rather than just one smoke-test call, if a bad update has previously slipped through a load-only check.

A GitHub Actions workflow around it

The GPU-touching steps need a runner with GPU access, GitHub's own hosted runners don't have one, so this assumes a self-hosted runner registered on or with access to a Spark. The non-GPU steps, everything before the actual load test, can run anywhere.

# .github/workflows/model-update.yml
name: Validate and deploy model update

on:
  push:
    branches: [main]
    paths:
      - 'models/manifest.json'

jobs:
  validate:
    runs-on: self-hosted
    steps:
      - uses: actions/checkout@v4

      - name: Read target model from manifest
        id: manifest
        run: echo "model_path=$(jq -r '.current' models/manifest.json)" >> "$GITHUB_OUTPUT"

      - name: Load and smoke-test the model
        run: ./scripts/validate-model.sh "${{ steps.manifest.outputs.model_path }}"

  deploy:
    needs: validate
    runs-on: self-hosted
    steps:
      - uses: actions/checkout@v4

      - name: Hand off to blue-green swap
        run: ./scripts/blue-green-swap.sh "${{ needs.validate.outputs.model_path }}"

The manifest file, a single JSON file naming the current model version, is what a pull request actually changes to trigger an update. Reviewing a one-line diff to that file is a much smaller surface than reviewing a full deploy script each time, and it means the update history lives in git log rather than in someone's memory of what they ran last Tuesday.

Keep validation and the swap as separate jobs

The deploy job only runs if validate passes, and it calls out to the blue-green swap script rather than reimplementing the cutover logic inline. Splitting them means a validation failure just fails a GitHub check with a clear reason, model didn't load, response was empty, rather than leaving the deployment in a half-swapped state because a script that was doing both jobs at once died partway through the second one.

Alert on the pipeline the same way you'd alert on the endpoint

A failed validation run should notify someone the same way a production alert would, not just show up as a red X in GitHub that nobody checks until the next PR. Wiring the job's failure into the same channel as the Prometheus alerts from setting up Prometheus alerts keeps model update failures and production incidents visible in one place instead of two.

FAQ

What should a CI pipeline actually check before a model update ships?

At minimum, that the new model loads without error and responds to a basic completion request, which catches a corrupted download or an incompatible quantization format before it reaches production. Beyond that, a small fixed set of evaluation prompts with expected properties, response isn't empty, stays under a token limit, doesn't contain a known failure string, catches regressions that a load check alone would miss. It won't catch a subtle quality regression; that needs the kind of evaluation covered separately, but it catches the failures that would otherwise page someone at deploy time.

Should the pipeline swap the model directly, or hand off to a separate deploy step?

Hand off. The CI pipeline's job is deciding whether a model version is safe to ship, running the validation and blocking on failure; the actual swap is a separate, idempotent operation, ideally the blue-green pattern of standing up the new version alongside the old one and cutting over traffic only once it's confirmed healthy. Conflating the two means a flaky validation step can leave a deployment half-finished instead of just failing a check.

Does a model update pipeline need a self-hosted runner, or does GitHub's hosted runner work?

GitHub's hosted runners don't have a GPU, so any step that needs to actually load the model and run inference has to happen somewhere else, either a self-hosted runner registered on a machine with GPU access, or a step that SSHes into the Spark and runs the validation there rather than trying to do it on the runner itself. Steps that don't touch the model, linting a config file, checking a checksum against a manifest, can run on the hosted runner fine.

Related pages

Validate model updates on hardware that's only running your jobs.

A dedicated Spark for staging and validation, so a load test isn't queued behind someone else's workload.

Read the blue-green deployments guide Read the staging environment guide