Automating model downloads with checksum verification
Downloading a gguf or safetensors file by hand is fine the first time. It stops being fine once you're doing it on a schedule, pulling a new quant as it's published, refreshing a fine-tune nightly, and the download runs unattended with nobody watching to notice a truncated file or a hash that doesn't match. This is a systemd timer that fetches and checksums a model without a human in the loop, and fails loudly instead of quietly serving a broken file.
Download to a temp path, never straight into the serving directory
The core safety property is simple: the file the running inference service has open is never the file being written to. A download-in-progress or a failed one just leaves a stale temp file behind and changes nothing about what's currently serving traffic.
#!/usr/bin/env bash
# /usr/local/bin/fetch-model.sh
set -euo pipefail
MODEL_DIR=/srv/models
TMP_DIR=/srv/models/.download
REPO="Qwen/Qwen3-32B-GGUF"
FILE="qwen3-32b-q4_k_m.gguf"
mkdir -p "$TMP_DIR"
curl -fSL "https://huggingface.co/${REPO}/resolve/main/${FILE}" \
-o "${TMP_DIR}/${FILE}.part"
Verify against the checksum the publisher actually provides
Hugging Face exposes a SHA-256 for each repo file via its API. Fetch that value rather than trusting your own re-computed hash from a prior download, comparing a hash to itself only proves the file didn't change between two downloads, not that either one is correct.
# fetch the published checksum for this specific file EXPECTED=$(curl -fsSL "https://huggingface.co/api/models/${REPO}" \ | jq -r --arg f "$FILE" '.siblings[] | select(.rfilename==$f) | .lfs.sha256 // .sha') ACTUAL=$(sha256sum "${TMP_DIR}/${FILE}.part" | cut -d' ' -f1) if [ "$EXPECTED" != "$ACTUAL" ]; then echo "checksum mismatch for ${FILE}: expected ${EXPECTED}, got ${ACTUAL}" >&2 rm -f "${TMP_DIR}/${FILE}.part" exit 1 fi # only now does the verified file replace anything the service reads from mv "${TMP_DIR}/${FILE}.part" "${MODEL_DIR}/${FILE}" echo "verified and installed ${FILE}"
A mismatch exits non-zero and leaves the previous model file in MODEL_DIR untouched. That failure needs to actually reach someone; wiring the timer's failure state into the alert rules from setting up Prometheus alerts via a node exporter textfile collector, or a simpler webhook on OnFailure=, is what turns a silent cron failure into something you find out about the same day.
Run it on a systemd timer instead of cron
A timer unit gives you structured logging in journalctl, an OnFailure= hook, and dependency ordering, none of which cron provides on its own.
# /etc/systemd/system/fetch-model.service [Unit] Description=Fetch and verify latest model release OnFailure=notify-failure@%n.service [Service] Type=oneshot ExecStart=/usr/local/bin/fetch-model.sh User=gpuwerk # /etc/systemd/system/fetch-model.timer [Unit] Description=Nightly model refresh [Timer] OnCalendar=*-*-* 03:00:00 Persistent=true [Install] WantedBy=timers.target
sudo systemctl enable --now fetch-model.timer
Downloading a new model isn't the same as putting it into production
This script only handles the fetch and verify step. Swapping the running service to actually serve the new file is a separate decision that deserves its own gate, covered in blue-green deployments for model updates: verify first, stage the file, run it through whatever evaluation you already do before a swap, and only then point the serving process at it. Automating the download is about removing manual toil from a repetitive step, not about removing the human decision of when a new model actually goes live.
FAQ
Why verify a checksum if the download completed without an error?
A completed HTTP transfer confirms the connection didn't drop, not that every byte matches what the publisher uploaded. Silent corruption from a flaky network, a truncated write from running out of disk mid-download, or a partial retry that appended instead of overwriting can all produce a file that looks complete and loads without erroring, but generates subtly wrong output. A checksum catches all three before the file is ever loaded.
Where do I get the checksum to verify against?
Hugging Face repos expose a SHA-256 for each file via the API (the same value shown in the repo's file listing), which is the value to check against, not one you compute yourself from a first download and trust from then on. Recomputing your own hash and comparing it to itself only catches a corrupted re-download, not a wrong or tampered file to begin with.
Should the download script overwrite a model that's currently serving traffic?
No. Download to a temporary path, verify it there, and only move it into the serving directory after verification passes. This is what makes the automation safe to run unattended: a failed or corrupted download never touches the file the running service has open, and a bad fetch just leaves the old model in place untouched.