Automating certificate renewal on a Spark
Setting up TLS for your inference endpoint covers getting a certificate issued for the first time. The failure mode worth planning for happens 90 days later, when a Let's Encrypt certificate expires and every caller starts seeing TLS handshake errors that have nothing to do with the model or the GPU. Both Caddy and certbot can renew without a human involved; the only real work is confirming the automation is actually wired up, and adding a check that would catch it if it wasn't.
Caddy renews on its own, no timer to configure
If you're running Caddy as the reverse proxy in front of vLLM, as covered in the TLS guide, renewal is already handled: Caddy checks certificate expiry as part of its own background operation and renews automatically once the certificate is within its renewal window, with no separate cron job or systemd timer needed. The only thing to automate is keeping the Caddy process itself alive across reboots and crashes.
# /etc/systemd/system/caddy.service (if not already installed by the package)
[Unit]
Description=Caddy
After=network.target
[Service]
ExecStart=/usr/bin/caddy run --config /etc/caddy/Caddyfile
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
sudo systemctl enable --now caddy
certbot needs a scheduled renewal job, usually already installed
If nginx is fronting the endpoint instead, certbot handles issuance but needs something to actually invoke certbot renew on a schedule. The certbot package on Debian and Ubuntu installs a systemd timer for this automatically, worth confirming rather than assuming:
systemctl list-timers | grep certbot
If nothing shows up, or you're on a distribution or container setup that doesn't ship the timer, add it directly:
# /etc/systemd/system/certbot-renew.timer
[Unit]
Description=Run certbot renew twice daily
[Timer]
OnCalendar=*-*-* 00,12:00:00
RandomizedDelaySec=3600
Persistent=true
[Install]
WantedBy=timers.target
# /etc/systemd/system/certbot-renew.service
[Unit]
Description=certbot renew
[Service]
Type=oneshot
ExecStart=/usr/bin/certbot renew --quiet --deploy-hook "systemctl reload nginx"
sudo systemctl enable --now certbot-renew.timer
The --deploy-hook is the part easiest to skip and most likely to bite you: certbot renewing the certificate files on disk doesn't reload nginx to pick up the new ones on its own, so a renewal can succeed and the running process can keep serving the old, soon-to-expire certificate anyway until something restarts it.
Watch expiry from outside the renewal job, not just trust it worked
A renewal job that silently fails, a changed DNS record, a firewall change blocking the ACME HTTP-01 challenge, a typo in the deploy hook, tends to fail quietly into a log nobody's watching until the certificate is already expired. The fix is a check that doesn't depend on the renewal job telling the truth about itself: an external monitor, either an addition to the alert rules in setting up Prometheus alerts or a check on the status page from setting up a status page, that reads the certificate's actual expiry date and alerts well before it's close.
# check days until expiry, exits non-zero under 14 days
EXPIRY=$(echo | openssl s_client -connect api.yourdomain.com:443 -servername api.yourdomain.com 2>/dev/null \
| openssl x509 -noout -enddate | cut -d= -f2)
DAYS_LEFT=$(( ($(date -d "$EXPIRY" +%s) - $(date +%s)) / 86400 ))
[ "$DAYS_LEFT" -lt 14 ] && echo "certificate expires in $DAYS_LEFT days" && exit 1
exit 0
Running this on the same schedule as your other health checks means a broken renewal shows up as a warning two weeks out, not as callers reporting connection errors on the day it actually expires.
FAQ
Does Caddy really need no cron job or timer for renewal?
Correct, as long as the Caddy process itself stays running. Caddy checks certificate expiry as a background task of the running server and renews automatically well before the 90-day Let's Encrypt window closes, with no separate scheduled job to configure or forget. The thing to automate instead is making sure the Caddy process itself survives a reboot, which a systemd service with restart-on-failure handles.
Why does certbot need a systemd timer instead of just running once after install?
A certificate issued once is good for 90 days and then expires; certbot doesn't renew it on its own unless something invokes certbot renew on a recurring schedule. The Debian and Ubuntu certbot packages install a systemd timer for this automatically, but a manual install or a certbot run inside a container might not, so it's worth checking with systemctl list-timers rather than assuming it's there.
How would I know if automated renewal silently stopped working?
Not from the renewal job itself, since a failed renewal usually just logs an error nobody reads until the certificate actually expires and callers start seeing TLS errors. The reliable way to know is an external check, a Prometheus alert or an uptime checker, that watches certificate expiry directly rather than trusting the renewal job succeeded. Alerting at 14 days before expiry gives you a window to fix a broken renewal before it becomes an outage.