477 Commits

Author SHA1 Message Date
librelad
a96a3a69a1 fix(linkding): declare the admin keys its auth adapter writes
linkding_auth.sh persists ADMIN_USER and ADMIN_PASSWORD when the first admin is
created, and keeps the password in step on later resets of that account, but
linkding.config declared neither — so both writes were no-ops and the WebUI
credentials card never had anything to show. Predates the slot work; it only
became visible once authPersistCfg started warning instead of failing silently.

Added empty rather than RANDOMIZED*, because unlike bookstack or nextcloud
nothing seeds a linkding account at install — the first user is created from the
WebUI. A generated password would name an account that does not exist, and the
card would display a password that cannot log in. Unslotted for the same reason:
the slot number marks a value the installer generates, and this one is written at
runtime by the tool.

No AUTH_PROFILE key: nothing reads it (it exists only in a comment in
auth_adapter.sh), and adding an unread key is what was just cleaned up elsewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 21:24:28 +01:00
librelad
cbbb978e50 feat(webui): make Overview board rows clickable end to end
Each action-board row on /apps/overview already carries a navigation
button (Review / View / Open Backups). The whole row now triggers that
same navigation, so the small button is no longer the only hit target.

Rows opt in via a `nav` descriptor, so only plain go-to-that-tab rows
become clickable — rows with no action, and the "Update all" button
inside the updates row, are unaffected (updater actions are matched
first in the delegated handler, so Update all never double-fires).

The row is not given role="button": nesting the real buttons inside a
button role would break them for assistive tech. The inner button stays
the focusable, keyboard-reachable control; the row is a mouse-only
widening, with a hover state that lifts row and button together. A
whole-row click is suppressed when it ends a text selection.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 21:18:45 +01:00
librelad
c9a513abc0 fix(webui): stop inventing a username on password-only app cards
The Login Details card fell back to the literal 'admin' when an app declared no
user or email key, so speedtest — which has a single CFG_SPEEDTEST_PASSWORD_1
and no user concept at all — advertised "User: admin" to anyone reading its
card. There is no such account; the field was fabricated by the fallback.

Show the username row only when a user or email key actually exists, and the
password row only when a password key does. eoCredList already omits any row
whose value is null, so passing undefined for either half renders just the half
that is real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 21:13:30 +01:00
librelad
296c6ddff1 fix(webui): show generated admin passwords again on the app cards
Every app with a generated admin password displayed "(not generated)" in its
Login Details. The credential matcher anchored on
^CFG_<APP>_(ADMIN_)?(PASSWORD)$, which stopped matching when generated keys were
given slot numbers — CFG_MATRIX_ADMIN_PASSWORD_1 and friends no longer hit the
regex, so the lookup fell through to its placeholder.

All ten apps carrying an admin password were affected: adguard, authelia,
bookstack, gitea, invidious, matrix, nextcloud, owncloud, pihole and stalwart.

Allow an optional _<n> tail. The anchor is kept otherwise, so a database or
upstream credential still cannot be mistaken for a login — CFG_<APP>_DB_PASSWORD_1
does not match, which is the case the anchor was added for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 21:09:33 +01:00
librelad
4a95b4c41e fix(webui): make the app detail category tag clickable and its icon legible
The category pill on the app detail header was rendered inert — no click
handler at all — while the identical pill on the app cards navigated to that
category's filter view. Both now come from one AppsManager.renderCategoryTag(),
so the detail pill behaves like the card pill and the two can't drift again.

The pill's glyph was an <img> of a category SVG, and those SVGs hardcode
#1e90ff. That only ever matched the dark-blue theme; on nebula (the default,
accent #00d4ff) and any other theme the icon read as a dark smudge next to its
own label. It's now painted as a CSS mask filled with currentColor, so it
always matches the pill's text on every theme.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 21:00:45 +01:00
librelad
c2fdfaccdb fix(stoat): tell the services the RabbitMQ password we gave the broker
The compose hands RabbitMQ a generated password, but the Stoat services fall
back to the defaults compiled into them — rabbituser/rabbitpass — so api, crond,
pushd and voice-ingress panicked on ACCESS_REFUSED and restarted forever.

The failure was easy to misread: the eleven services that never touch RabbitMQ
came up healthy and the web client answered on port 80, so the stack looked
almost fine while none of the messaging worked.

Write a [rabbit] section into Revolt.toml carrying the same credentials the
broker was given. Verified after the fix: all sixteen containers up, /api
returns the instance descriptor, /autumn answers, and /.well-known/stoat carries
the right URL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 20:56:52 +01:00
librelad
e8a2aa453e fix(auth): resolve slot-numbered credential keys, drop two dead ones
Only two of the four keys flagged as unused actually were. gitea and invidious
ADMIN_PASSWORD are written by their auth adapters through authPersistCfg, which
builds the name as CFG_${app^^}_${key} from a parameter — invisible to a literal
grep, which is why the earlier pass called them dead. They stay.

Worse, the slot rename broke that write path for five apps: adguard, bookstack,
gitea, invidious and nextcloud all persist ADMIN_PASSWORD, and the config now
holds ADMIN_PASSWORD_1. updateConfigOption only rewrites a key that already
exists, so the write became a no-op — the app's password would really change
while the config and the WebUI kept showing the old one.

authPersistCfg now falls back to the numbered slot when the bare key is absent,
so adapters never need to know how a credential is numbered and adding a slot
can't silently disconnect the adapter that writes it. When neither name exists
it warns and returns non-zero instead of failing silently, which surfaces a
pre-existing case: linkding's adapter persists ADMIN_USER and ADMIN_PASSWORD but
its config declares neither, and never did.

Deleted the two that really are dead: CFG_TRAEFIK_ADMIN_PASSWORD_1 (its adapter
uses CFG_TRAEFIK_USER/CFG_TRAEFIK_PASS from the system config) and
CFG_GLUETUN_CONTROL_SERVER_API_KEY_1, plus their WebUI field mappings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 20:55:52 +01:00
librelad
80b94fb21f fix(stoat): write every bind-mounted file before anything that can fail
The install read the generated LiveKit credentials with a plain grep, but
secrets.env is chmod 600 and owned by the docker install user while the hooks
run as the manager — so the read returned nothing, the hook errored out, and
Caddyfile and livekit.yml were never written. Compose then refused to start,
because a bind mount whose source does not exist is not a soft failure.

Read secrets through runFileOp, and reorder so the Caddyfile and the three
URL-bearing files are written first: any step that can fail now comes after
every mount source already exists. The missing LiveKit keys are downgraded from
fatal to a warning for the same reason — losing voice is worth reporting, but it
is no reason to take the other fifteen services down with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 20:54:49 +01:00
librelad
ec98a83b48 fix(gitea,invidious): stop two secrets sharing one placeholder index
Both configs defined a second secret in a later section starting the numbering
over, so CFG_GITEA_METRICS_TOKEN_1 and CFG_GITEA_ADMIN_PASSWORD_1 both read
RANDOMIZEDPASSWORD1 — and the replacer generates one value per distinct
placeholder and seds every occurrence, so the two came out identical. Same for
Invidious's HMAC key and admin password. Predates the slot rename.

Currently latent, since neither admin password is consumed by anything, but it
would silently pair a live secret with whatever gets wired to the other key.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 20:51:31 +01:00
librelad
4eb5fa6434 feat(matrix,stoat): run without a domain, on the LAN or a VPN
Both apps demanded a domain and Traefik. That was over-constrained: LibrePortal
ships WireGuard, Headscale and private ports, so LAN and VPN-only is a
first-class deployment here, and Rocket.Chat and Mattermost already prove chat
apps work fine on http://<lan-ip>:<port>.

The gate on Matrix rested on a mistake of mine: server_name being permanent.
server_name and public_baseurl are independent — the identity can be a domain
you own with no DNS behind it while clients reach the server on a LAN address,
so federation can be switched on later by adding DNS and TLS, with no rebuild
and no lost history. CFG_MATRIX_SERVER_NAME now exposes exactly that, and the
install warns when it falls back to the machine's IP.

What is genuinely lost without a domain is stated where it belongs, at install:
Matrix cannot federate and Element's mobile apps want HTTPS; Stoat cannot do
camera or microphone, because browsers gate getUserMedia on a secure context
and a VPN does not change that, the check being on the URL scheme.

Both now derive their URL from the port that was actually allocated. Since ports
are only assigned during compose-up, each writes a best guess before start and
corrects it afterwards, restarting only when the value really changed.

Three bugs found while proving it works end to end:

- The Synapse image writes /data as its UID/GID env, default 991, which under
  rootless is a host sub-UID owning nothing — so the generated signing key could
  not be moved by the install user. Both the generate container and the service
  now run as the same identity USER_TAG resolves to.

- Element's config.json is bind-mounted as a file, and docker silently creates a
  DIRECTORY when the source is missing. An early return left exactly that
  landmine, which then broke every later run. It is written first now, and a
  stale directory is cleared.

- A successful admin registration was reported as an error: checkSuccess read $?
  after an intervening [[ ]] test rather than the command's own status.

Verified with no domain and no Traefik installed: Synapse answers
/_matrix/client/versions and /health on http://<ip>:<port>, admin login returns
a token, and Element is configured against the corrected base_url.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 20:51:16 +01:00
librelad
c7df07ffc2 containers: trim overlong app card descriptions
Nine LONG_DESCRIPTION values had drifted well past the 90-140 char
range the rest of the catalog uses (stoat was 407). Cut them back
while keeping the caveats that matter — Rocket.Chat's user cap,
Stoat's resource weight, Matrix federation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:57:34 +01:00
librelad
a0f0bc087e fix(matrix,stoat): derive the public host without reading the compose
Both hooks read their host (and matrix its database password) back out of the
deployed docker-compose.yml. That cannot work: install_post_compose runs after
the compose TEMPLATE is copied but before dockerConfigSetupFileWithData fills
the tags, so at that point the file still holds raw placeholders. Matrix aborted
with "Database password was not generated in the compose file" even though the
password had been generated correctly — it just was not in the compose yet.

Derive the host from port_subdomains[0] + domain_full instead, both already in
scope from variables_init_app, applying the same empty/@/root rule as
tagsProcessorPortSubdomains so the computed name and the Traefik rule generated
later cannot drift apart. Matrix takes its database password from
CFG_MATRIX_DB_PASSWORD_1, which is where the secret is generated and remembered
and is the same variable the compose tag is filled from a step later.

The error messages now name the actual missing thing — the domain — rather than
blaming the compose file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:55:32 +01:00
librelad
751a4578d6 webui: trim the field tooltips added for the slot rename
Matches b562059 — the fourteen mapping entries added alongside the slot rename
were written before that landed and ran 88-100 chars against a median of 44.
2026-08-18 19:53:06 +01:00
librelad
65167463f9 fix(chat apps): tag every service so its IP actually substitutes
Installing rocketchat failed with

    invalid IPv4 address: ParseAddr("IP_DATA_2"): unable to parse IP

ipUpdateComposeTags allocates one IP per SERVICE_TAG_N annotation and fills
IP_TAG_i only where SERVICE_TAG_i exists. The four new apps tagged only their
primary service, so every sidecar — matrix's postgres, mattermost's postgres,
rocketchat's mongo, and fifteen of stoat's sixteen — kept a literal IP_DATA_n
in the deployed compose and docker refused to create the container.

Tag every service that carries an ipv4_address, index-aligned with its IP_TAG.
For stoat that also meant moving caddy from SERVICE_TAG_1 to _6 so the indices
line up with the IPs rather than the reading order.

mastodon had the same latent break (IP_TAG_2 and _3 untagged) and is fixed the
same way — it would have failed on first install for the same reason.

SERVICE_TAG carries the compose *key*, not container_name: 'libreportal app
restart <app> <service>' passes it to 'docker compose restart', which only
understands keys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:48:09 +01:00
librelad
4685320353 feat(secrets): real VAPID keypair for mastodon, slot-numbered DB passwords
VAPID: the two values are the halves of one P-256 keypair, not independent
secrets — the browser verifies that a push is signed by the private key matching
the public key it subscribed with. The RANDOMIZED* generators mint each
placeholder on its own, so they produced two unrelated strings and web push could
never have worked. Generate the pair in mastodon_install_post_setup the way stoat
already does, encoded as Mastodon's webpush gem expects: unpadded URL-safe base64
of the 32-byte private scalar and the 65-byte uncompressed public point, sliced
out of the SEC1 DER. Verified by rebuilding the key from the emitted private half
and re-deriving the public point — openssl accepts it and the point matches.

Generated once and never rotated (rotation would invalidate every subscription),
but a pair of the wrong shape is replaced, so an install carrying the old
unrelated strings heals itself on next install — their public half is 42 chars
where a real point is 87.

Slots: CFG_<APP>_DB_PASSWORD -> CFG_<APP>_DB_PASSWORD_1 and likewise for
DB_ROOT_PASSWORD, across mastodon, owncloud, mattermost, matrix, nextcloud and
bookstack, so a database credential is always a numbered slot and a second one is
just _2. Renaming a key means reconciliation drops the old and adds the new
holding its placeholder, so an existing install regenerates unless the value is
carried over first — documented, including that the old file survives as
.<app>.config.bak.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:42:50 +01:00
librelad
5706498565 fix(secrets): move app credentials into <app>.config, fix slot collision
Five apps (mastodon, owncloud, mattermost, matrix, stoat) took their generated
secrets from the compose-side generator tags PASSWORD_TAG_<n>/RANDOM_TAG_<n>/
HEX_TAG_<n>/VAPID_TAG_<n>. Those mint a fresh secret on every templating run, so
a reinstall handed the app a new database password while its data volume kept the
one initdb was given, and the app came back up unable to open its own database.

Move them to <app>.config as RANDOMIZED* placeholders, reaching the compose via
the #LIBREPORTAL|<APP>_<KEY>_TAG| mechanism tags_processor_app_config_values
already provides. No new handler: the tag name is derived from the config key, so
this is a config line plus a tag per secret. Generation is unchanged — still
random on first install; the value is now remembered instead of re-rolled.

Also fixes two things this exposed:

- The RANDOMIZED* replacers matched unanchored. `sort -u` orders slots lexically
  (1, 10, 11, 2), so slot 1's pattern rewrote the prefix inside slot 10's
  placeholder and slots 10+ ended up holding slot 1's secret with a digit glued
  on — derivable, and invisible because the values weren't byte-identical.
  Anchoring with \b makes match order irrelevant. Verified at 20 slots across
  all four placeholder types: 64 keys, 64 distinct values, no prefix collisions.

- generateRandomPassword drew from base64 without constraining the mix; measured
  over 2000 draws, 1 in 40 contained no digit at all. Retry until the result has
  both a digit and a letter, bounded so a pathological length can't spin.

owncloud gains a fix in passing: its compose seeded the admin account from
PASSWORD_TAG_2 while the WebUI displayed CFG_OWNCLOUD_ADMIN_PASSWORD, which was
generated separately and never used. Both now read the same value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:33:40 +01:00
librelad
c11052b753 fix(mastodon): make tag annotations substitutable
The tag manager reads `#LIBREPORTAL|<TAG>|<VALUE>` and takes <VALUE> as the
current literal to search for on that line, so it must equal the string in the
line body. Mastodon used `unconfigured` as the annotation value against bodies
like `PASSWORD_TAG_1_DATA` — nothing matched, nothing was ever substituted, and
the placeholder shipped as the live database password, SECRET_KEY_BASE, OTP
secret and VAPID keypair.

Nothing caught it either: `unconfigured` doesn't match `_DATA`, so
tagsManagerGetTagState reported the tags as configured, and the stale-tag gate
in dockerComposeUp (which tests the annotation value against
`^[A-Z][A-Z0-9_]*_DATA(_[0-9]+)?$`) let the app start.

Adopt the convention every other app already uses — body placeholder identical
to the annotation value, `<KIND>_DATA_<n>`. All 11 tags now substitute, the app
and postgres services agree on the same generated credentials, and an unfilled
tag is visible to the pre-start gate.

Existing 0.1.0 installs keep their literal credentials until re-installed, and
their Postgres was initialised with them, so a plain re-install desynchronises
the compose from the volume. Document both recovery paths in upgrade notes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:36:28 +01:00
librelad
9919eea138 stoat: add the ex-Revolt stack as the closest Discord equivalent
Sixteen containers: MongoDB, Valkey, RabbitMQ, MinIO and eleven Stoat services.
Servers, channels, roles and voice/video through LiveKit — the nearest thing in
the catalogue to Discord itself, at the price of being much the heaviest app in
it. Does not federate.

The compose service keys are deliberately kept identical to upstream's
(database, redis, api, autumn, ...) while container_name is prefixed stoat-.
Compose registers both on the network, so upstream's internal defaults keep
resolving and LibrePortal still gets the prefixed names its port, firewall and
backup layers key on.

Upstream's Caddy is kept as the internal path router and Traefik simply proxies
to it, which is upstream's own supported behind-a-reverse-proxy mode —
reimplementing eight path routes as Traefik labels would be a second copy to
keep in sync for nothing. The install hook is a non-interactive port of
generate_config.sh, and it never rewrites an existing secrets.env:
REVOLT__FILES__ENCRYPTION_KEY decrypts every file ever uploaded, so
regenerating it would orphan the whole media store.

LiveKit's UDP media range is published literally rather than through the port
table, because the firewall rebuild emits /tcp rules only and a range declared
there would produce a wrong rule rather than no rule. Voice falls back to TCP
7881 until the range is opened by hand; the post-install notice says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:28:46 +01:00
librelad
50059ea1b8 rocketchat: add Rocket.Chat with a single-node Mongo replica set
Rocket.Chat tails the Mongo oplog for realtime delivery, and a standalone
mongod has no oplog — so the database has to be a replica set even with one
member.

Uses the official mongo image rather than bitnami/mongodb (which upstream's own
compose uses) because Bitnami moved its catalog behind a paid registry and the
free tags are no longer dependable for a long-lived install. The cost is that
rs.initiate() is not automatic, so the post-start hook runs it once — guarded by
rs.status() so a reinstall over restored data doesn't re-initiate a live set,
and followed by a wait for the member to report itself primary.

Mongo runs without auth: enabling it on a replica set also requires a shared
keyfile for member-to-member auth, which is a lot of moving parts for a database
that is never published outside the docker network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:28:46 +01:00
librelad
c0025d8211 mattermost: add Team Edition as a low-friction chat app
One container against Postgres, with the polished desktop and mobile clients
that make it the least demanding of the four chat options.

Runs as the bind-mount owner via USER_TAG: the image bakes in USER mattermost
(uid 2000) so it never runs as root and cannot chown its own data directory,
which under rootless Docker means it dies on first write.

CFG_MATTERMOST_AUTHELIA stays false — OIDC/SAML is a paid tier here, so
forward-auth would block the native clients from the API without buying single
sign-on in exchange.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:28:31 +01:00
librelad
b9e334dc49 matrix: add Synapse + Element as a federated chat app
Synapse on Postgres plus the Element web client, on two subdomains: the
homeserver on matrix.<domain> (which becomes server_name, so IDs read
@alice:matrix.<domain>) and Element on element.<domain>.

Two hosts rather than one because server_name then matches the host Traefik
already terminates TLS for, so 'serve_server_wellknown: true' is all the
federation delegation needed and nothing has to be published at the apex
domain — which this app has no way to configure.

CFG_MATRIX_AUTHELIA is pinned false and documented: forward-auth in front of
/_matrix locks out every client and every federating peer, since they carry
Matrix access tokens and cannot follow a redirect. Real SSO goes through the
OIDC block in resources/homeserver.yaml instead.

The install hook generates the signing key once via upstream's own 'generate'
command and refuses to regenerate it over an existing install — a new key would
be rejected by every server that had cached the old one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:28:31 +01:00
librelad
63f276523b stalwart: run as container-root so it can write its own data directory
Found by running the installer for real rather than testing the hook in
isolation. Stalwart never started: it failed to open its database with
"Permission denied" on /var/lib/stalwart, which meant no mail could be
stored and the setup wizard could not be completed by hand either.

The image runs as its own uid 2000. LibrePortal gives container directories
to the docker install user under rootless and to the manager under rooted,
and 2000 is neither, so the bind mounts were unwritable in both modes. This
was not something the new provisioning introduced — it predates it, and the
app has never been able to hold mail.

Running as container-root maps to whichever host user owns those
directories. Under rootless that is the unprivileged docker install user,
not host root.

Also stop discarding the server's error when setup fails. Both failures
that actually occur — a hostname under a TLD that does not resolve, and the
unwritable data directory above — name themselves precisely, and a bare
"setup failed" turns a one-line fix into guesswork.

Verified end to end through `libreportal app install stalwart` on a clean
install: setup applied, DKIM keys generated, postmaster mailbox created,
and the full record set printed from the server's own zone data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:14:19 +01:00
librelad
7ed1cddd5c stalwart: answer the setup wizard instead of handing it to the user
A new Stalwart drops you into a five-screen wizard — hostname, domain,
storage backend, directory, logging, DNS — before it will do anything.
LibrePortal already knows the two answers that matter and the rest have
sane defaults, so asking is asking a question we can answer ourselves.

v0.16 exposes those wizard fields as a `Bootstrap` singleton, so the whole
thing is one `update` applied through the Stalwart CLI. The CLI is not in
the server image (upstream split it into its own repo), but it publishes a
multi-arch container, so we borrow the server's network namespace and run
it there — nothing installed on the host, nothing to clean up, arm64 works.

Setup now also:

- generates DKIM keys (Ed25519 + RSA) with rotation left switched on, and
  requests a TLS certificate. That last one is easy to miss: Traefik only
  fronts the admin port, so 25/465/587/993 never see its certificate and
  clients would hit a self-signed one on 993.
- creates postmaster@<domain>. The generated zone points DMARC and TLS-RPT
  reports there and nothing was creating it, so those reports bounced.
- prints the record set read back from the server rather than composed
  here, so it includes the real DKIM public keys, MTA-STS, TLS-RPT and the
  SRV records clients autoconfigure from. This hook used to tell the user
  to go and fetch DKIM themselves; by that point the keys exist.

Optionally hands DNS to a provider API (Cloudflare/DigitalOcean/DeSEC),
which keeps the whole record set in sync and makes DKIM rotation safe to
leave on. Off by default: the token can write to your zone and lives in
the mail server's database.

Re-running is safe — provisioning is skipped once config.json exists, and
the plans use upsert so they reconcile rather than duplicate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:05:03 +01:00
librelad
0eda3104c7 stalwart: report a missing admin console without failing the upgrade
A failed verify makes the engine abort and restore, and a restore cannot
put back a bundle that was never downloaded — it would roll a working
mail server back a version to fix a missing web page, then hit the same
empty GitHub fetch next time. So the console check now warns loudly and
returns 0; readiness stays the only gate.

Renamed to stalwart_upgrade_check_admin_ui so the name cannot be read as
part of the gate, and bounded its poll to a 60s grace window (capped by
the caller's deadline) — the upgrade result is already decided by then,
so there is no reason to hold the run open on a web asset. The unreach-
able-probe branch is advisory for the same reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:59:02 +01:00
librelad
4c80ea018d stalwart: check the admin console, not just readiness
Stalwart v0.16 does not ship the WebUI in its Docker image — the admin
console is fetched from GitHub on first start. With no outbound HTTPS at
that moment the fetch fails silently: /healthz/ready still answers 200
because the mail server genuinely is serving, so both the installer and
the upgrade verifier reported success while /admin and /account 404'd
with nothing to explain why.

Install hook now probes /admin after the port-25 and PTR checks and, on
404, names the GitHub download as the cause rather than emitting a
generic failure. Upgrade verifier treats stable readiness as necessary
but not sufficient and confirms /admin before returning 0; the console
is polled under the same deadline because the bundle download runs
behind the server coming up, and failing on the first 404 would abort an
upgrade that was seconds from finishing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:54:32 +01:00
librelad
4fae22c89e fix(stalwart): use the official logo instead of my placeholder
Replaces the drawn shield-and-envelope stand-in with the real mark from
stalw.art (/favicon.svg), in their #DB2D54.

Padded from the source's 159.95x139.07 to a square 159.95 viewBox with
the art vertically centred, matching every other catalogue icon — all of
which are square, so a non-square box would letterbox in the app tiles.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:23:33 +01:00
librelad
b238a0a6e8 fix(vikunja): real logo mark — the llama, without the wordmark
Second correction: the earlier version was only the outer blue disc.
This keeps all 10 artwork paths (circle, body, ears, muzzle, wool, face)
and drops the <g> holding the 'vikunja' wordmark, since a catalogue tile
wants the mark alone. Cropped from the wide 872.6x256.8 logo viewBox to
the mark's own 0 0 256 256, and the web-app attributes (class,
xml:space) removed.

Extracted programmatically from the supplied SVG rather than
transcribed, so no path data could be mangled by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:18:44 +01:00
librelad
55da05fb4e fix(vikunja): use the real Vikunja logo mark
Replaces my drawn placeholder with the actual logo path supplied by the
maintainer.

Kept as .svg rather than .ico: webuiSyncAppIcon only looks for
<app>.svg then <app>.png, so an .ico would never be copied into the
frontend and the app would silently fall back to the default icon. SVG
is also the right format here — it is a single vector path that scales
to any tile size, where an .ico would bake in fixed raster sizes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:15:55 +01:00
librelad
f596b36a73 chore: remove Focalboard from the catalogue
Retired last commit, deleted now at the maintainer's call. Mattermost
ended support in 2023, the community repo is asking for maintainers, and
the image had not been rebuilt in 1042 days — the staleness signal's
worst case after speedtest. Vikunja covers the same ground and is
actively developed.

Self-contained: every reference lived inside containers/focalboard/ plus
its generated manifest entries, so nothing else needed touching. Git
keeps the history.

Anyone with it already installed keeps a running container and their
data — removing the template only stops NEW installs. Their app will now
report as unknown in the App Center rather than offering an update,
which is the honest state for software with no upstream.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:13:00 +01:00
librelad
b915e0731a fix(vikunja): run as the mount owner, and drop the docker socket
Found by installing it. Two problems, one fatal and one worse.

The container crash-looped: the image runs as uid 1000, which under
rootless Docker maps to host sub-UID 232071 while the bind mounts are
owned by the install user, so Vikunja died on its first write to
/app/vikunja/files and restarted forever. Its own error message
diagnosed it exactly. Fixed with the existing USER_TAG mechanism the
portal container already uses — 0:0 under rootless (container root IS
the install user on the host), the real uid:gid under rooted — rather
than hardcoding either.

More seriously, the compose mounted the docker socket, copied in from a
template that needed it. A task manager has no business talking to the
daemon, and the socket is root-equivalent access on the host. Removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:08:56 +01:00
librelad
cdf76040f1 feat(vikunja): add Vikunja; retire Focalboard rather than swap it
Focalboard is the one audit finding with no successor to follow.
Mattermost ended support in 2023, the community repo is openly asking
for maintainers, and the image has not been rebuilt in 1042 days.

Adds Vikunja as the replacement for NEW installs: lists, kanban, table
and gantt — the same job Focalboard did — from a project rebuilt 14 days
ago. One container on SQLite, no database sidecar, following the
catalogue conventions (tag sentinels throughout, category/title +
backup.db/backup.files labels, traefik block, gluetun markers).

VIKUNJA_SERVICE_PUBLICURL is wired to the existing APP_URL_TAG rather
than a hand-built URL. It is not optional for this app — get it wrong
and creating the first account fails with a bare "unauthorized" — and
APP_URL_TAG already resolves to https://<domain> behind Traefik or
http://<host>:<assigned-port> otherwise, so the port is never guessed.
(First attempt invented a PORT_DATA_1 tag that does not exist; checking
what the processors actually emit found the real mechanism, which
bookstack already uses.)

Focalboard is RETIRED, not deleted. Replacing an app in place would stand
still for anyone already running it — their data does not move to Vikunja
— so it keeps working, and instead:
  * the description says plainly that it is unmaintained, why, and what
    to use instead
  * UPDATE_TYPE drops to manual, because there is nothing to update TO
    and auto-pulling a 2.8-year-old tag is pure churn

Icon is a drawn placeholder, not the upstream trademark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:04:23 +01:00
librelad
298ccf4868 fix(speedtest): follow LibreSpeed to its maintained image
adolfintel/speedtest was last rebuilt 1624 days ago — 4.4 years, the
worst in the catalogue. LibreSpeed itself is fine: the project moved to
the librespeed org, and LinuxServer.io's build was rebuilt 2 days ago.

Chose lscr.io/linuxserver/librespeed over the org's own ghcr.io image
specifically because it is on Docker Hub: our tag enumeration and the
new staleness signal can query Hub but not ghcr, so this keeps the app
visible to the tooling that would catch it going stale again.

Not a tag swap. The LSIO image uses their house conventions, so the env
block is ported rather than copied: PUID/PGID/TZ (matching bookstack,
the catalogue's other LSIO app) and everything under /config instead of
/database. MODE, WEBPORT and ENABLE_ID_OBFUSCATION have no equivalent —
this image is standalone by design and always serves on 80 internally,
which is what our port mapping already assumed.

PASSWORD and DB_TYPE keep their names and meaning, so the existing
config keys still drive the results page and telemetry. DB_NAME is
deliberately left unset: the image picks a path inside /config, and
guessing one risks a database sitting outside the volume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:02:01 +01:00
librelad
b7f3d30170 fix(trilium): follow the project to TriliumNext
zadam/trilium was last rebuilt 810 days ago because the original author
handed the repository to the community project and put his own in
maintenance mode. The project is very much alive — triliumnext/trilium
was rebuilt 24 days ago. Same story as Pi-hole: we were pinned to an
abandoned original, not a dead project.

A genuinely clean swap, verified against upstream's own docs rather than
assumed: identical data path (/home/node/trilium-data), identical port
(8080), and TriliumNext states there are no special migration steps — it
opens an existing zadam database as-is. TRILIUM_DATA_DIR is now set
explicitly, matching upstream's compose.

URL metadata repointed at the maintained repo so the App Center links
somewhere that still gets commits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 03:56:12 +01:00
librelad
614e895b5d fix(unbound): maintained image, and actually recursive this time
mvance/unbound was last rebuilt 668 days ago. Replaced with
madnuttah/unbound (14 days): distroless, runs unprivileged as non-root,
listens on 5335 by default — exactly the "upstream behind a blocker"
shape — and publishes clean semver tags. klutchell/unbound is equally
fresh but defaults to port 53 (fighting Pi-hole/AdGuard for it) and its
tag namespace is CI build soup.

The shipped config was worse than the stale image. It was not a
recursive resolver at all:

    interface: 0.0.0.0@53
    forward-addr: 10.100.0.3@53   # "Local AdGuard" — a hardcoded IP
    forward-addr: 9.9.9.9@853

So it listened on 53 (conflicting with any blocker on the same host),
forwarded to Quad9 — surrendering the "nobody sees my queries" property
that is the only reason to run Unbound in front of a blocker — and
pointed at AdGuard, inverting the dependency: AdGuard should point HERE.

Replaced with a drop-in at conf.d/libreportal.conf. The image's own
unbound.conf ends with `include-toplevel: conf.d/*.conf`, so ours ADDS
to a working recursive config the image author maintains rather than
replacing it — upstream keeps owning the parts that change between
Unbound releases. It contributes access-control (private ranges allow,
everything else REFUSE, so this can never become an open resolver for
amplification attacks), DNSSEC hardening, rebinding protection, and
cache sizing suited to a small VPS. Forwarding is included commented
out, with the trade stated rather than silently chosen.

Ports corrected to 5335:5335 — the old mapping assumed an image
listening on 53 internally. Added the libreportal.category/title labels
the app was missing (no traefik labels: it has no web interface).
Install hook copies the drop-in and repairs a stub directory first, the
same trap that kept Nextcloud's nginx from starting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 03:51:52 +01:00
librelad
ab811440e4 fix(pihole): use the official image, ported to Pi-hole v6
The app shipped cbcrowe/pihole-unbound — a third-party bundle last
rebuilt 841 days ago. Pi-hole itself is fine: the official
pihole/pihole was rebuilt 42 days ago. We packaged an abandoned fork,
not a dead project.

Not a tag swap. The official image is v6, which replaced nearly every
v5 environment variable with an FTLCONF_ equivalent — and the container
ACCEPTS the old names and ignores them, so a v5-style block looks
correct while configuring nothing, including the admin password:

  WEBPASSWORD         -> FTLCONF_webserver_api_password
  WEBTHEME            -> FTLCONF_webserver_interface_theme
  PIHOLE_DNS_         -> FTLCONF_dns_upstreams   (";" separates values)
  DNSSEC              -> FTLCONF_dns_dnssec
  DNSMASQ_LISTENING   -> FTLCONF_dns_listeningMode
  REV_SERVER{,_TARGET,_DOMAIN,_CIDR} -> FTLCONF_dns_revServers, one
                         combined "<enabled>,<cidr>,<target>,<domain>"
  FTLCONF_LOCAL_IPV4  -> gone in v6

listeningMode is ALL rather than the old "single": on a bridge network
queries arrive via the docker gateway, and "single" drops them.

The bundled unbound is gone, so PIHOLE_DNS_=127.0.0.1#5335 pointed at
nothing. New CFG_PIHOLE_UPSTREAM_DNS defaults to Quad9, with the config
documenting how to point it at the unbound app for full recursion. The
freed port slot becomes the (disabled) DHCP port, which v6 supports.

Volumes: v6 keeps config, databases and gravity under /etc/pihole, and
ignores /etc/dnsmasq.d unless explicitly re-enabled — so the old
two-mount layout is replaced by a single ./etc-pihole.

Variable names taken from the official v5->v6 upgrade doc, not memory.
Substitution verified end to end: every placeholder resolves and
revServers renders in the documented format.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 02:30:13 +01:00
librelad
ca266e9391 feat(updater): flag apps whose image upstream has stopped rebuilding
"Up to date" answers one question — has the tag I track moved? — and an
abandoned project answers it reassuringly forever. The tag stays put, the
digest never changes, and the app reports as current while receiving no
security patches at all. Nothing in the UI could tell a healthy stable
app from a dead one.

An audit of all 34 anchor images found five in exactly that state:
speedtest (4.4y since rebuild), focalboard (2.8y — Mattermost dropped
support in 2023), pihole-unbound (2.3y), trilium (2.2y), unbound (1.8y).

The scan now records image_updated_at per app (one cheap Hub call inside
the existing registry window, cached between windows like everything
else) and emits stale_after_days from CFG_UPDATER_STALE_DAYS (365, 0
disables) so the UI and the config agree on one number.

Surfaced as an "unmaintained?" severity chip on the fleet row and a
dated explanation in the app detail. Phrased as an observation rather
than an accusation — plenty of small tools are simply finished — but it
does spell out the security consequence, because that is the part a user
cannot infer from "up to date".

Deliberately NOT a "needs action" row on the Overview board: it is not
fixable by pressing anything, and a permanently amber board teaches
people to ignore the board.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 01:09:03 +01:00
librelad
6b44b7dd59 fix(nextcloud): copy nginx.conf on install, so the web container can start
Found by actually installing it. The compose bind-mounts
./resources/nginx.conf into the web container, but nothing ever copied
that file into the container tree, so Docker created a DIRECTORY in its
place and nginx died with:

  error mounting ".../resources/nginx.conf" to rootfs at
  "/etc/nginx/nginx.conf": not a directory

Worse than a hard failure: the app still recorded as installed. Three of
four containers came up, the DB and the app itself were fine, and only
the web front end was missing — a quiet, partial install.

Apps needing a resource file declare the copy in a hook (authelia does
exactly this); Nextcloud simply never had one. Adds
nextcloud_install_post_compose — after the compose file is written,
before permissions and `up` — which repairs any stub directory left by a
previous attempt and then copies the file.

The stub repair matters: without it the copy lands INSIDE the directory
(resources/nginx.conf/nginx.conf) and the mount fails identically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 00:18:55 +01:00
librelad
046d69bad7 feat(updater): Upgrade button for cross-version moves
Puts the stepped engine behind the "34 available" chip so it is not
CLI-only. Wired end to end: updater_upgrade task type -> task-router ->
updaterUpgrade action -> `libreportal updater upgrade <app> [version]`,
with a label and icon in the tasks list.

The button always confirms, and the dialog states the plan and the
guarantee rather than asking "are you sure?" — which app, from which
version to which, that every step snapshots first and waits for the app
to confirm it is serving with no migration outstanding, that a failure
stops the ladder on the last version that verified, and that it can take
a long time because each release runs its own migration.

Both delegated dispatchers (per-app Updates tab and the fleet Overview
rows) learn the action, so the button works wherever the chip appears.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 00:08:16 +01:00
librelad
598f74c26b feat(updater): per-app upgrade verifiers — the safety half of stepping
Stepping 31 -> 32 -> 33 is arithmetic. Knowing 32 FINISHED before
touching 33 is the whole safety story, and it is invisible from outside
the app: Nextcloud runs its migration on boot and sits in maintenance
mode — or fails halfway — while Docker reports the container perfectly
healthy. Advance a rung there and a migration has been skipped on live
data.

Contract:  <app>_upgrade_verify <app> <expected-tag> <deadline>  -> 0

Returns 0 ONLY on positive confirmation that the app serves at the
expected version with nothing outstanding. Unhealthy, indeterminate and
timed-out all return non-zero — uncertainty is a failure, not a maybe,
because the alternative gambles with data.

  nextcloud  `occ status`: installed, NOT in maintenance, no pending DB
             upgrade, and the running major matches the tag. Maintenance
             mid-migration is expected and simply keeps waiting.
  mastodon   /health serving, ZERO "down" rows in db:migrate:status, and
             the version from /api/v1/instance matching. /health alone is
             insufficient — Puma answers before migrations finish.
  stalwart   /healthz/ready (per its documented probes), required to hold
             stable rather than flash once. Weaker by design: the probes
             confirm serving but report no version, and the file says so
             rather than implying more.

updaterVerifyGeneric (running + healthy + no restart during a settle
window) is the fallback for everything else, and is explicitly NOT
sufficient to justify climbing a rung — the engine will refuse to ladder
an app with no declared verifier.

9 tests drive the dangerous states directly: maintenance mode, pending DB
upgrade, and a wrong major all correctly REFUSE to verify; clean states
pass. Those three negatives are the ones that would have corrupted data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:58:23 +01:00
librelad
8a0c3641d0 fix(nextcloud,mastodon): track a moving stable line like every other app
Both templates were stale in different ways, and the new tag enumeration
surfaced it: nextcloud sat on 31-fpm-alpine with 34 out, mastodon on
v4.2.0 with v4.6 out.

  nextcloud  31-fpm-alpine -> 34-fpm-alpine
  mastodon   v4.2.0        -> v4.6

The mastodon one was the real problem: v4.2.0 is an EXACT patch pin, so
it never moved at all — no security patches, ever. v4.6 is a moving
minor-line tag (the same shape as stalwart's v0.16), so auto-update now
delivers patches within the line.

Deliberately NOT floated to :latest or :stable. Both projects require
stepped upgrades — Nextcloud in particular refuses to skip a major — so
a tag that crosses majors on its own would break the app on a routine
container recreate. A major-pinned, patch-moving tag is the correct
shape here, not a limitation.

Verified against the registry: all three pinned apps now report nothing
newer, and each tag still moves (v0.16 2026-08-10, 34-fpm-alpine
2026-08-03, v4.6 2026-08-06).

Templates only, so this changes NEW installs. Existing installs keep
their tag and will now be told a newer line exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:39:36 +01:00
librelad
7334557706 feat(updater): detect newer release lines, not just newer builds
The digest compare only ever asks about the tag already pinned, so it
answers "has my tag been rebuilt?" and can never answer "does a newer
version exist?". An app on v0.16 reports up to date forever while 0.17
ships. That is the gap between an app that updates and an app that is
current, and it silently affects every pinned app.

Adds tag enumeration for VERSIONED tags only (rolling tags already move
on their own): list the repo's tags, keep those sharing the current tag's
SHAPE, and pick the numerically greatest.

Shape matching is the whole safety story — v0.16 -> v#.# so it can never
"upgrade" you onto v0.16-alpine, 31-fpm-alpine onto 31-apache, or a date
tag onto a semver one. Comparison is component-wise numeric, so 0.10 > 0.9
and 1.0 > 0.99 (a string sort gets both wrong), with 10# forcing base ten
so an upstream "08" cannot be read as octal. 15 unit tests cover it.

Docker Hub only, deliberately: all three pinned apps live there, it needs
no auth, and the generic OCI tags/list wants a per-registry token dance.
Other registries stay quiet rather than guess. Throttled inside the
existing registry window and cached between windows so it cannot flicker.

Surfaced as INFORMATION, never an action: no button applies it, because a
version move can carry a data migration. `update_available` and the "up
to date" badge keep their exact meaning; the new state sits beside them
and points at the Version field.

Against the live registry: stalwart v0.16 is current, nextcloud is on
31-fpm-alpine with 34-fpm-alpine out, mastodon on v4.2.0 with v4.6.5 out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:30:08 +01:00
librelad
fd8ac29621 feat(stalwart): auto-update patches, VERSION config for release moves
Reverses the manual default from 4ee2529, which was over-cautious once
the tag pin is taken into account.

Two things were conflated. Auto-update does not reinstall anything: it
snapshots, `compose pull`, `up -d` — the container is recreated from the
new image and the data volume is untouched. And because the image is
pinned to v0.16, the updater compares the digest of THAT tag, so auto
can only ever apply rebuilds of 0.16 (security/bug patches). It cannot
jump to 0.17. That is the safe half of updating, and there is no good
reason to withhold it.

Adds CFG_STALWART_VERSION=v0.16, which drives the image tag through the
existing #LIBREPORTAL|STALWART_VERSION_TAG| sentinel (verified: setting
it to v0.17 rewrites the image line). Moving between releases is now a
config change a user can make from the app's config page — the roadmap's
config-first version identity, used for real.

Net behaviour: patches land unattended inside the update window; a
version jump stays a deliberate decision, which is what pre-1.0 software
with a settling storage schema warrants.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:13:47 +01:00
librelad
47b614f8b4 fix(apps): stop re-serialising the config section, which killed its listeners
Every themed dropdown on the app config page was dead: it rendered, but
clicking did nothing. Only that page — every other dropdown in the WebUI
worked.

renderAppDetail captured the config section's own innerHTML right after
displayConfigForm() had rendered it...

  const configHTML = document.getElementById('config-section')?.innerHTML;
  ...67 lines later...
  configSection.innerHTML = configHTML;

...and wrote the same string straight back. That is a no-op for the
markup and a catastrophe for behaviour: re-assigning innerHTML re-parses
the subtree, so every listener in it is destroyed.

custom-select.js had already wrapped each <select>, so the captured
string contained the .custom-select wrapper and the custom-select-native
class. The re-inserted copy therefore looked enhanced — which made the
enhancer correctly skip it as already-done — while having no click
handler at all. A dropdown that renders perfectly and does nothing.

The container is never wholesale-rewritten in that function (it updates
header/config/console individually), so the capture-and-restore had no
purpose. Both lines removed, with a comment stating that anything added
there must mutate the section rather than re-assign its innerHTML.

Diagnosed in a real browser rather than by reading: instrumenting
CustomSelect.build() showed the enhancer DID build a widget for the field
while the wrapper in the DOM was not the one it built, and patching the
innerHTML setter named apps-manager.js:785 as the writer. Verified live:
the dropdown opens, and picking an option sets the value (auto) and the
label.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:09:25 +01:00
librelad
07daa7555a fix(webui): load mobile-menu.js; prune orphaned task queue entries
Both found in a user's console log.

1. ReferenceError: setupMobileMenu is not defined (dashboard.js:98)

   core/topbar/js/mobile-menu.js defines that global, and index.html
   never loaded it. dashboard.js called it unguarded as the FIRST line
   of setupEventListeners, so dashboard init threw every page load and
   took loadInstalledApps() with it — and the burger menu was dead on
   mobile. system-loader already guarded its own call with a typeof
   check, which is why this survived unnoticed.

   Loads the script (before dashboard.js) and guards the call, so
   optional nav chrome can never take down the page below it again.

2. Endless 404s on /api/tasks/<id> for tasks that no longer exist

   queue.json is append-only from the enqueue side and nothing ever
   pruned it, so any task file removed afterwards left an id the WebUI
   re-fetched forever, one 404 per poll per orphan. Adds
   cleanupOrphanQueueEntries to the idle housekeeping pass: entries with
   no task file are dropped and logged. Self-heals existing strays.

   (Provoked by my own clean-up of two test tasks earlier in this
   session, but the gap is real and predates it.)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:48:55 +01:00
librelad
5e06e77d7f fix(forms): one bad select no longer kills every dropdown after it
Reported: dropdowns dead on the app config page only, in a private window
(so not cache), with the served frontend confirmed identical to source.

The enhancer ran in a bare forEach with no try/catch anywhere, and
build() inserts its wrapper via `select.parentNode.insertBefore(...)`.
A detached select makes that a null deref, and one throw abandoned the
rest of the pass — every select AFTER it silently stayed native. The
MutationObserver callback had the same exposure for the remaining
mutation records in a batch.

The app config page is the one that can produce a detached select: it
builds its category panels in an async loop, so the observer can see a
node a later render already replaced. That matches "only app config".

Worse than losing the theme: if the throw landed after build() added
.custom-select-native, the select was left opacity:0 / pointer-events:
none behind a button with no listeners — a dropdown that looks right and
does nothing.

Now: detached/unconnected selects are skipped (they get enhanced when
their subtree is attached and the observer fires again), every
enhancement is individually guarded, a failed one is rolled back so it
can never be left invisible, and failures console.warn with the field
name instead of vanishing.

Verified with jsdom against the real file: detached nodes skipped without
throwing, a failure mid-pass leaves later selects working, and the
survivors open and set their value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:40:26 +01:00
librelad
4ee25292d5 feat(stalwart): add Stalwart Mail Server as a catalog app
One container providing SMTP/IMAP/POP3/JMAP plus CalDAV/CardDAV, an admin
UI and spam filtering — chosen over mailcow (owns its own installer, which
is what killed the earlier attempt now sitting in scripts/unused/) and
over Mailu (~7 containers) because a single image with a single data dir
is the only shape that fits the existing conventions cleanly: one anchor
service the updater can version, one path the backup engine can snapshot.

Mail-specific departures from the usual app template, each deliberate:

* Ports are FIXED, not random. Other mail servers connect to :25 by
  number and clients expect 465/587/993 — a randomised external port
  would silently make the server unreachable. Only the admin UI takes a
  random port, since that one really is just a browser behind Traefik.
  143/995/4190/443 ship disabled; the port processor comments them out.

* UPDATE_TYPE=manual and the image pinned to v0.16, not :latest.
  Stalwart is pre-1.0 and has said the storage schema is still being
  finalised, so an unattended minor bump could carry a data migration on
  the message store. This is the one app where the auto default is wrong.

* BACKUP_STRATEGY=stop-snapshot-start. The message store is written
  continuously; a live copy can land mid-transaction. Seconds of queued
  delivery (senders retry) buys a consistent snapshot.

* The install hook checks outbound port 25 and reverse DNS, then prints
  the MX/SPF/DMARC records with real values. A mail server whose
  container started is not a working mail server, and every remaining
  requirement lives at the registrar or the VPS provider.

Admin credentials are seeded via STALWART_RECOVERY_ADMIN from the app
config rather than left to Stalwart's first-run random password, which
would otherwise exist only in the container log.

Icon is a drawn placeholder, not the upstream trademark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:02:26 +01:00
librelad
8153d82282 fix(apps): don't serve a cached config form after a deploy
getFieldMappings/getConfigCategories fetched host-GENERATED files with
default caching, so a browser that had the page open before a release
kept rendering the previous release's config UI — a newly shipped field
(UPDATE_TYPE) simply never appeared, with nothing on screen to hint the
page was stale. Only a hard refresh fixed it.

Adds {cache:'no-store'} to both, plus the one configs.json read in this
file that was missing it while two others already had it. Cost is a
conditional request per config-page open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 16:22:43 +01:00
librelad
66c79f997e feat(updater): install window, honest Check-now, failed-auto surfacing
Four fixes that make the auto-updater a trustworthy background system:

* CFG_UPDATER_WINDOW (default 06:00-08:00 host time, right after the
  05:00 backup cron; HH:MM-HH:MM wraps midnight, 'always' = any time).
  Gates only the enqueue — scans keep running all day, so the Updates
  page stays current and pending updates visibly wait for the window.
  Malformed values fail closed and are rejected by the WebUI validator.

* "Check now" actually checks: an explicit `updater check` sets
  UPDATER_REGISTRY_FORCE=1. The flag existed but nothing ever set it,
  so the button silently reused the 6h digest cache and could not find
  a build the user knew had shipped. Force also overrides interval 0,
  which now means "manual-only" as documented in the roadmap.

* Registry stamp moved from /tmp to <system>/logs: the task processor
  runs under PrivateTmp, so daemon and CLI each kept a separate 6h
  clock and the daemon's reset on every service restart.

* A failed automatic attempt is no longer invisible: the scan emits
  auto_attempted_digest (the one-shot no-retry stamp), and when it
  matches the available build the UI stops promising an install that
  will never come — per-app detail explains, the fleet row gets an
  "auto failed" chip, and the Overview board counts it as needing you.

Also corrects the CFG_TIMEZONE label: it sets the containers' TZ only;
scheduled tasks follow the host clock (timedatectl), and the old
"Timezone for scheduled tasks" wording promised a knob that never
existed. The window + auto_window display state plainly WHEN updates
land, answering "how does the user know when the next update happens".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 21:06:27 +01:00
librelad
7fae6bc308 fix(updater): stop a callee blanking the app name mid-update
First real end-to-end auto-update on a live install failed like this:

  Automatically updating trivy (a recovery snapshot is taken first)
  Snapshotting trivy before update…
  Pulling new image(s) for …
  Update of  failed — rolling back…
  Could not roll  back automatically

The app name went empty after the snapshot. Cause: bash is dynamically
scoped, so a callee assigning an undeclared variable writes the CALLER's
local of that name — and a `while read app` loop leaves it EMPTY at EOF.
webuiBackupAppStatus's dashboard generator runs at the end of every backup
and did exactly that to updaterApplyApp's `app`.

Nothing was damaged: the pull ran against an empty name, failed before
touching the image, and the rollback was a no-op on a nonexistent app.

Fixed both ends. The generator (and three gluetun loops with the same
latent leak) now declare `local app`. updaterApplyApp/updaterRollbackApp
hold the name in `_upd_app` so they no longer depend on every callee's
hygiene, and updaterApplyAll stops leaking its own loop var.

This is exactly the untested path the roadmap flagged: "apply/revert not
yet exercised end-to-end on a live install with a pending update."

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:37:14 +01:00
librelad
cdeb2d1658 feat(updater): per-app UPDATE_TYPE, automatic by default
Adds the decision half of the app updater. Detection (P2) and the
snapshot-first apply/revert (P3) were already real, but nothing ever
pressed the button — every update waited for a click.

  CFG_<APP>_UPDATE_TYPE=auto|manual   per app, default auto (33 templates)
  CFG_UPDATER_AUTO=true|false         master switch, default true

updaterAppPolicy resolves the two the way backupResolveStrategy already
resolves backup strategy: the global switch can only make things more
manual. updaterApplyAuto runs at the end of `updater check` and enqueues
the ordinary updater_apply task for each auto app that has an update —
never applies inline, so an automatic update is the same code path, task
log, History entry and Roll back button as a manual one.

Safety: each attempt stamps its target digest under generated/auto/, so a
build that fails is rolled back and then left alone rather than retried on
every scan; in-flight updater tasks are skipped so scans can't stack.

Tracked end to end: updates.json carries each app's resolved update_type,
History entries carry trigger=manual|auto. The WebUI says whether updates
install themselves, chips only the apps that opted out, labels automatic
history, and — since an auto app's pending update needs no decision — keeps
it off the Overview board's "Needs action" view.

Also fixes artifactApplyAuto enqueueing without --detach: it runs inside
the single-threaded task processor's own poll, so following the new task in
the foreground waits for a task that cannot start until it returns.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:22:04 +01:00