LibrePortal/containers/stalwart/scripts/stalwart_upgrade_hooks.sh
librelad 0eda3104c7 stalwart: report a missing admin console without failing the upgrade
A failed verify makes the engine abort and restore, and a restore cannot
put back a bundle that was never downloaded — it would roll a working
mail server back a version to fix a missing web page, then hit the same
empty GitHub fetch next time. So the console check now warns loudly and
returns 0; readiness stays the only gate.

Renamed to stalwart_upgrade_check_admin_ui so the name cannot be read as
part of the gate, and bounded its poll to a 60s grace window (capped by
the caller's deadline) — the upgrade result is already decided by then,
so there is no reason to hold the run open on a web asset. The unreach-
able-probe branch is advisory for the same reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:59:02 +01:00

127 lines
6.3 KiB
Bash

#!/bin/bash
# Stalwart upgrade verifier.
# ---------------------------------------------------------------------------
# Stalwart publishes Kubernetes-style probes on its admin port (8080):
# GET /healthz/live — the process is alive and not deadlocked
# GET /healthz/ready — dependencies initialised, config loaded, accepting traffic
# https://stalw.art/docs/http/overview/
#
# Readiness is the one that matters after a version move: it only answers 200
# once the storage backend is open and the config has loaded, which is exactly
# the window where a schema change would otherwise go unnoticed. Liveness alone
# would pass on a process that is up but unable to serve.
#
# Deliberately weaker than the Nextcloud verifier: Stalwart's probes confirm the
# server is serving, but do not report a version, so this asserts readiness
# rather than "running exactly $expected". Stated plainly instead of implied —
# the ladder is only ever as strong as the check underneath it, and pretending
# otherwise is how a half-migrated app advances a rung.
#
# Readiness is still the only thing that GATES the upgrade. Behind it sits a
# second, advisory check. Since v0.16 the WebUI is not in the Docker image: the
# server downloads the admin console from GitHub, and a fresh image after an
# upgrade means fetching it again. With no outbound HTTPS at that moment the
# fetch fails, /healthz/ready still answers 200 (the mail server genuinely is
# serving) and the admin is handed a "successful" upgrade whose control panel
# 404s with nothing to explain why.
#
# That is reported loudly but is NOT failed on, deliberately. The engine's
# response to a failed verify is to abort and restore, and restoring cannot put
# back a bundle that was never downloaded — it would roll a perfectly good mail
# server back a version to fix a missing web page, then hit the same empty
# GitHub fetch on the next attempt. The upgrade did land; a piece of it needs a
# network path the box did not have. Saying exactly that beats both a silent
# pass and a pointless rollback.
# stalwart_upgrade_admin_ui_code
# HTTP status of the admin console, empty if the container did not answer.
stalwart_upgrade_admin_ui_code() {
runFileOp docker exec stalwart-service curl -fsS -o /dev/null -w '%{http_code}' \
--max-time 5 http://localhost:8080/admin 2>/dev/null | tr -d '\r'
}
# stalwart_upgrade_verify <app> <expected-tag> <deadline-epoch>
# 0 only when /healthz/ready answers 200 and keeps answering it. The admin
# console is reported on afterwards but cannot change the result.
stalwart_upgrade_verify() {
local app="$1" expected="$2" deadline="$3"
local stable=0 stable_needed=3 last=""
while [ "$(date +%s)" -lt "$deadline" ]; do
local code
code="$(runFileOp docker exec stalwart-service curl -fsS -o /dev/null -w '%{http_code}' \
--max-time 5 http://localhost:8080/healthz/ready 2>/dev/null | tr -d '\r')"
last="healthz/ready=${code:-none}"
if [ "$code" = "200" ]; then
stable=$((stable + 1))
# Ready must HOLD: a server that flaps ready/not-ready is mid-restart,
# and one lucky 200 is not evidence the upgrade settled.
if (( stable >= stable_needed )); then
isSuccessful "Stalwart is ready (readiness probe stable) after moving to $expected."
stalwart_upgrade_check_admin_ui "$expected" "$deadline"
return 0
fi
else
stable=0
fi
sleep 5
done
isError "Stalwart did not report ready for $expected before the deadline.${last:+ Last probe: $last}"
return 1
}
# stalwart_upgrade_check_admin_ui <expected-tag> <deadline-epoch>
# Reports whether /admin is being served. ALWAYS returns 0 — named "check" and
# not "verify" so nobody wires it into the gate later by mistake; see the note
# at the top of this file for why a missing console must not trigger a restore.
#
# Polled rather than probed once, because the WebUI fetch runs behind the server
# coming up: readiness can go stable while the bundle is still being pulled, and
# a single early probe would cry 404 at an install that was seconds from fine.
# The window is short and capped by the caller's deadline — the upgrade result
# is already decided by this point, so there is nothing to be gained by holding
# the run open for the full remaining timeout waiting on a web asset.
stalwart_upgrade_check_admin_ui() {
local expected="$1" deadline="$2"
local grace=$(( $(date +%s) + 60 ))
local code=""
(( grace > deadline )) && grace="$deadline"
while [ "$(date +%s)" -lt "$grace" ]; do
code="$(stalwart_upgrade_admin_ui_code)"
if [ -n "$code" ] && [ "$code" != "404" ]; then
isSuccessful "Stalwart's admin console is being served (/admin=$code)."
return 0
fi
sleep 5
done
if [ -z "$code" ]; then
# No status at all is a different fault from a 404: the probe never
# reached the container. Named separately so nobody goes hunting a
# firewall rule for what is actually an exec/curl problem.
isError "The admin console could not be probed — check it by hand before relying on it."
isNotice " No response from 'docker exec stalwart-service curl … /admin'. The upgrade"
isNotice " to $expected itself is fine: the readiness probe above passed."
return 0
fi
# Loud, because a 404 admin panel is exactly the kind of thing that gets
# discovered weeks later — but explicitly NOT a failure, so the wording has
# to carry that or the next person reads it as a broken upgrade.
isError "The admin console is MISSING (/admin=$code) — the upgrade to $expected still succeeded."
isNotice " Stalwart does not bundle the WebUI: it downloads the console from"
isNotice " https://github.com/stalwartlabs/webui/releases/latest when it starts on a new"
isNotice " image. That download did not complete, so /admin and /account will 404 until"
isNotice " it does. Mail delivery is unaffected — only the web interface."
isNotice " Fix: allow outbound HTTPS to github.com from this host, then restart the"
isNotice " container: docker restart stalwart-service"
isNotice " Do NOT roll back — the previous version fetches the same bundle from the"
isNotice " same place and will land in the same state."
return 0
}