stalwart_wait_http hardcoded the /healthz/ prefix and returned a yes/no,
so the admin-console check could not use it and grew its own copy of the
docker exec curl line. Extract stalwart_http_code <path> [max-time] and
build both on it: the wait loop keeps its probe-name signature and its
3s timeout, the console check keeps its 5s and gets the status code back
rather than a verdict, since 404 and no-reply-at-all need saying apart.
Probe commands are byte-identical to before; no behaviour change. The
upgrade verifier keeps its own copy on purpose — verifiers here are
self-contained (see nextcloud's, which inlines the occ idiom rather than
calling the install hook's wrapper) and should not drag a lifecycle file
they have no other use for into an upgrade run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A mail server is two quite different products wearing one name, and until
now LibrePortal only offered the hard one. Installing Stalwart meant being
handed a wall of DNS records, a red error about port 25 and a warning about
reverse DNS — all of it correct, none of it fixable by the installer, and
most of it irrelevant to someone who wanted mailboxes and a shared calendar
on their own network.
CFG_STALWART_MODE now names which one you are running:
private mailboxes, IMAP, CalDAV and CardDAV on your own network. Port 25
is not published at all; the client ports stay bound to the host
but are never opened through the firewall. No MX, no PTR, no
deliverability. Nothing to publish, so nothing is printed.
public the internet mail server, as before.
auto public if Traefik is installed, private if not, resolved at
install and written back so it reads as a real answer afterwards.
DKIM keys are generated in both modes even though private has no use for
them today — that is what makes switching later a setting change rather
than a key ceremony. The WebUI gets a "Mail Exposure" tool that flips the
setting both ways and reconfigures the server, plus a "Show DNS Records"
tool that prints the live zone including current DKIM keys.
Two things this had to get right, both found by testing rather than
reading. Port access lives in the shell as CFG_<APP>_PORT_n, not just in
the config file, and the compose file is built from the parsed shell
values — editing only the file left the config claiming port 25 was
disabled while the container published it anyway. And going public needs
an AcmeProvider to exist before a domain can reference one, so the switch
creates it; note that doing so registers an account with Let's Encrypt.
Verified through real installs: auto resolves to private with no Traefik,
port 25 is genuinely unpublished and absent from the compose file, the
client ports are skipped by the firewall as host-bound, and the tool
round-trips private -> public -> private with the config landing back
exactly where it started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Installing rocketchat failed with
invalid IPv4 address: ParseAddr("IP_DATA_2"): unable to parse IP
ipUpdateComposeTags allocates one IP per SERVICE_TAG_N annotation and fills
IP_TAG_i only where SERVICE_TAG_i exists. The four new apps tagged only their
primary service, so every sidecar — matrix's postgres, mattermost's postgres,
rocketchat's mongo, and fifteen of stoat's sixteen — kept a literal IP_DATA_n
in the deployed compose and docker refused to create the container.
Tag every service that carries an ipv4_address, index-aligned with its IP_TAG.
For stoat that also meant moving caddy from SERVICE_TAG_1 to _6 so the indices
line up with the IPs rather than the reading order.
mastodon had the same latent break (IP_TAG_2 and _3 untagged) and is fixed the
same way — it would have failed on first install for the same reason.
SERVICE_TAG carries the compose *key*, not container_name: 'libreportal app
restart <app> <service>' passes it to 'docker compose restart', which only
understands keys.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Found by running the installer for real rather than testing the hook in
isolation. Stalwart never started: it failed to open its database with
"Permission denied" on /var/lib/stalwart, which meant no mail could be
stored and the setup wizard could not be completed by hand either.
The image runs as its own uid 2000. LibrePortal gives container directories
to the docker install user under rootless and to the manager under rooted,
and 2000 is neither, so the bind mounts were unwritable in both modes. This
was not something the new provisioning introduced — it predates it, and the
app has never been able to hold mail.
Running as container-root maps to whichever host user owns those
directories. Under rootless that is the unprivileged docker install user,
not host root.
Also stop discarding the server's error when setup fails. Both failures
that actually occur — a hostname under a TLD that does not resolve, and the
unwritable data directory above — name themselves precisely, and a bare
"setup failed" turns a one-line fix into guesswork.
Verified end to end through `libreportal app install stalwart` on a clean
install: setup applied, DKIM keys generated, postmaster mailbox created,
and the full record set printed from the server's own zone data.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A new Stalwart drops you into a five-screen wizard — hostname, domain,
storage backend, directory, logging, DNS — before it will do anything.
LibrePortal already knows the two answers that matter and the rest have
sane defaults, so asking is asking a question we can answer ourselves.
v0.16 exposes those wizard fields as a `Bootstrap` singleton, so the whole
thing is one `update` applied through the Stalwart CLI. The CLI is not in
the server image (upstream split it into its own repo), but it publishes a
multi-arch container, so we borrow the server's network namespace and run
it there — nothing installed on the host, nothing to clean up, arm64 works.
Setup now also:
- generates DKIM keys (Ed25519 + RSA) with rotation left switched on, and
requests a TLS certificate. That last one is easy to miss: Traefik only
fronts the admin port, so 25/465/587/993 never see its certificate and
clients would hit a self-signed one on 993.
- creates postmaster@<domain>. The generated zone points DMARC and TLS-RPT
reports there and nothing was creating it, so those reports bounced.
- prints the record set read back from the server rather than composed
here, so it includes the real DKIM public keys, MTA-STS, TLS-RPT and the
SRV records clients autoconfigure from. This hook used to tell the user
to go and fetch DKIM themselves; by that point the keys exist.
Optionally hands DNS to a provider API (Cloudflare/DigitalOcean/DeSEC),
which keeps the whole record set in sync and makes DKIM rotation safe to
leave on. Off by default: the token can write to your zone and lives in
the mail server's database.
Re-running is safe — provisioning is skipped once config.json exists, and
the plans use upsert so they reconcile rather than duplicate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A failed verify makes the engine abort and restore, and a restore cannot
put back a bundle that was never downloaded — it would roll a working
mail server back a version to fix a missing web page, then hit the same
empty GitHub fetch next time. So the console check now warns loudly and
returns 0; readiness stays the only gate.
Renamed to stalwart_upgrade_check_admin_ui so the name cannot be read as
part of the gate, and bounded its poll to a 60s grace window (capped by
the caller's deadline) — the upgrade result is already decided by then,
so there is no reason to hold the run open on a web asset. The unreach-
able-probe branch is advisory for the same reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Stalwart v0.16 does not ship the WebUI in its Docker image — the admin
console is fetched from GitHub on first start. With no outbound HTTPS at
that moment the fetch fails silently: /healthz/ready still answers 200
because the mail server genuinely is serving, so both the installer and
the upgrade verifier reported success while /admin and /account 404'd
with nothing to explain why.
Install hook now probes /admin after the port-25 and PTR checks and, on
404, names the GitHub download as the cause rather than emitting a
generic failure. Upgrade verifier treats stable readiness as necessary
but not sufficient and confirms /admin before returning 0; the console
is polled under the same deadline because the bundle download runs
behind the server coming up, and failing on the first 404 would abort an
upgrade that was seconds from finishing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Stepping 31 -> 32 -> 33 is arithmetic. Knowing 32 FINISHED before
touching 33 is the whole safety story, and it is invisible from outside
the app: Nextcloud runs its migration on boot and sits in maintenance
mode — or fails halfway — while Docker reports the container perfectly
healthy. Advance a rung there and a migration has been skipped on live
data.
Contract: <app>_upgrade_verify <app> <expected-tag> <deadline> -> 0
Returns 0 ONLY on positive confirmation that the app serves at the
expected version with nothing outstanding. Unhealthy, indeterminate and
timed-out all return non-zero — uncertainty is a failure, not a maybe,
because the alternative gambles with data.
nextcloud `occ status`: installed, NOT in maintenance, no pending DB
upgrade, and the running major matches the tag. Maintenance
mid-migration is expected and simply keeps waiting.
mastodon /health serving, ZERO "down" rows in db:migrate:status, and
the version from /api/v1/instance matching. /health alone is
insufficient — Puma answers before migrations finish.
stalwart /healthz/ready (per its documented probes), required to hold
stable rather than flash once. Weaker by design: the probes
confirm serving but report no version, and the file says so
rather than implying more.
updaterVerifyGeneric (running + healthy + no restart during a settle
window) is the fallback for everything else, and is explicitly NOT
sufficient to justify climbing a rung — the engine will refuse to ladder
an app with no declared verifier.
9 tests drive the dangerous states directly: maintenance mode, pending DB
upgrade, and a wrong major all correctly REFUSE to verify; clean states
pass. Those three negatives are the ones that would have corrupted data.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
One container providing SMTP/IMAP/POP3/JMAP plus CalDAV/CardDAV, an admin
UI and spam filtering — chosen over mailcow (owns its own installer, which
is what killed the earlier attempt now sitting in scripts/unused/) and
over Mailu (~7 containers) because a single image with a single data dir
is the only shape that fits the existing conventions cleanly: one anchor
service the updater can version, one path the backup engine can snapshot.
Mail-specific departures from the usual app template, each deliberate:
* Ports are FIXED, not random. Other mail servers connect to :25 by
number and clients expect 465/587/993 — a randomised external port
would silently make the server unreachable. Only the admin UI takes a
random port, since that one really is just a browser behind Traefik.
143/995/4190/443 ship disabled; the port processor comments them out.
* UPDATE_TYPE=manual and the image pinned to v0.16, not :latest.
Stalwart is pre-1.0 and has said the storage schema is still being
finalised, so an unattended minor bump could carry a data migration on
the message store. This is the one app where the auto default is wrong.
* BACKUP_STRATEGY=stop-snapshot-start. The message store is written
continuously; a live copy can land mid-transaction. Seconds of queued
delivery (senders retry) buys a consistent snapshot.
* The install hook checks outbound port 25 and reverse DNS, then prints
the MX/SPF/DMARC records with real values. A mail server whose
container started is not a working mail server, and every remaining
requirement lives at the registrar or the VPS provider.
Admin credentials are seeded via STALWART_RECOVERY_ADMIN from the app
config rather than left to Stalwart's first-run random password, which
would otherwise exist only in the container log.
Icon is a drawn placeholder, not the upstream trademark.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>