Setting up webhook notifications for job completion
A batch script that runs against a Spark for twenty minutes and then just exits leaves whoever kicked it off checking a log file, or worse, polling a status endpoint every few seconds until it's done. Neither scales past a couple of jobs. A webhook sent at the moment the job finishes, or fails, moves that information to whoever needs it without anyone watching a terminal.
There's no GPUwerk-side job queue to hook into
A deployed Spark runs an OpenAI-compatible inference endpoint, not a managed job scheduler, so there's no first-party "job completed" event to subscribe to. If "job" means a batch script that fires off a series of completions against an OpenAI-compatible endpoint and then does something with the output, the webhook has to be sent by that script itself, from your own code, once it reaches the end. This is a pattern you build on top of a Spark, not a feature the platform provides.
Sign the payload, don't just POST it
A webhook receiver is a public URL. Without verification, anyone who guesses or leaks that URL can send a fake completion notification with whatever body they want. Sign every outgoing payload with an HMAC over a shared secret, and check the signature before doing anything with the request:
# sender: after the batch job finishes
import hmac, hashlib, json, time, requests
payload = {
"job_id": "batch-2026-09-09-014",
"status": "completed",
"output_rows": 4820,
"finished_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
}
body = json.dumps(payload).encode()
signature = hmac.new(WEBHOOK_SECRET.encode(), body, hashlib.sha256).hexdigest()
requests.post(
WEBHOOK_URL,
data=body,
headers={
"Content-Type": "application/json",
"X-Signature-256": f"sha256={signature}",
},
timeout=10,
)
# receiver: verify before trusting the body from flask import Flask, request, abort import hmac, hashlib app = Flask(__name__) @app.route("/webhooks/job-complete", methods=["POST"]) def job_complete(): signature = request.headers.get("X-Signature-256", "") expected = "sha256=" + hmac.new( WEBHOOK_SECRET.encode(), request.data, hashlib.sha256 ).hexdigest() if not hmac.compare_digest(signature, expected): abort(401) # record the event durably before returning 200 save_job_event(request.json) return "", 200
Use hmac.compare_digest rather than == for the comparison. A plain string comparison short-circuits on the first mismatched byte, which leaks timing information an attacker can use to guess the signature one byte at a time.
Retry on the sending side, dedupe on the receiving side
A single unretried POST is one dropped connection away from a notification that never arrives. Retry with exponential backoff on anything other than a 2xx response, and give up after a bounded number of attempts rather than looping forever:
import time
def send_with_retry(url, body, headers, attempts=5):
for i in range(attempts):
try:
resp = requests.post(url, data=body, headers=headers, timeout=10)
if resp.status_code < 300:
return True
except requests.RequestException:
pass
time.sleep(min(2 ** i, 60))
return False
Retries mean the receiver can see the same job ID twice, so make the handler idempotent: check whether that job ID was already recorded before acting on it a second time. Otherwise a retried notification for a job that already triggered a downstream action, sending an email, kicking off a dependent job, triggers it again.
Return 200 only once the event is durably stored
Doing slow work, sending an email, calling a third API, inside the request handler before returning a response is how a receiver ends up timing out on the sender's side and triggering an unnecessary retry. Write the event to a queue or a database row first, return 200 immediately, and let a separate worker process it:
@app.route("/webhooks/job-complete", methods=["POST"])
def job_complete():
# verify signature as above, then:
if not already_recorded(request.json["job_id"]):
enqueue_event(request.json) # fast, durable write
return "", 200 # respond before doing anything slow
This also protects the sender's retry logic from your own downstream latency. A slow email provider shouldn't be the reason a batch script thinks its notification failed.
Where this fits with the rest of your pipeline
If the batch job itself is running against rate limits and retry logic on request timeouts and retries, the completion webhook is the last step in that same chain, one more place where a transient failure needs a retry rather than a silent drop. And if the notification needs to reach a channel rather than an HTTP endpoint, the same signed payload can be forwarded from the receiver into Slack, PagerDuty, or wherever the team already watches for alerts, instead of building a second delivery path for every destination.
FAQ
Does GPUwerk send webhooks for batch jobs?
GPUwerk doesn't run a batch job scheduler of its own, a deployed Spark just exposes an OpenAI-compatible inference endpoint. If a job means a long-running batch script calling that endpoint, the webhook has to be sent by the script itself, from your own code, at the point it finishes, rather than by any GPUwerk-side infrastructure.
Why sign a webhook payload instead of just trusting the sender?
An unauthenticated webhook receiver is a public endpoint that will act on whatever POST body it's sent, so anyone who finds the URL can forge a completion notification. Signing the payload with an HMAC and a shared secret, then verifying that signature before processing the request, is the difference between a webhook and an open door.
What happens if the receiver is down when the job tries to notify it?
Without retry logic on the sending side, nothing, the notification is lost and no one knows the job finished until someone checks manually. A sender that retries with backoff on a non-2xx response, and a receiver that returns 200 only after the notification is durably recorded, covers the common failure of the endpoint being briefly unreachable during a deploy or restart.