19 Commits

Author SHA1 Message Date
librelad
26e98698d8 feat(storage): per-app placement, app move, and the READONLY marker
Phase 2 and 4 of docs/roadmap/storage-locations.md. Apps can now be
placed on a location and moved between them.

CFG_<APP>_STORAGE lands in all 37 app templates, holding a location NAME
rather than a path: names survive a migrate to a host with different
disks, paths do not. The 11 infrastructure apps that other apps reach by
literal path (traefik, prometheus, grafana, adguard, gluetun, crowdsec,
headscale, dashy, pihole, unbound, wireguard) are pinned. libreportal
itself never gets the key — it is pinned structurally by webuiDir.

Pinning needed no second config key. "Pinned" is not a fact about a value,
it is a statement about whether the field may be edited, so it goes in the
comment beside **ADVANCED** and **DEV** as **READONLY**, and the field
factory renders those disabled. That marker earns its keep beyond this
feature: derived fields already warned in prose that editing them does
nothing (crowdsec.config:72) next to a perfectly editable input.

storage_app_config.sh keeps the comment honest — it carries the resolved
path for hand-recovery and regenerates the dropdown from the registry, but
only writes when something actually changed, since the app .config is
user-editable and lives in the container-owned tree.

app move stops the app (a live copy of a running Postgres is a corrupt
copy), snapshots it, copies, verifies, and only then removes the source.
The copy runs in libreportal-ownership because it must: app data holds
rootless sub-UID files the manager can neither read nor recreate.
Verified against two real ext4 filesystems that a cross-device move
preserves uid 231141 and the payload, and that the source survives every
refusal path — unregistered destination, the WebUI app, a traversal in the
app name, and an occupied destination.

Task titles registered in both tables, with a specific rule so a move
renders as "Nextcloud - Move to bigdisk" rather than the generic fallback
dropping the destination. lp-task-names could not be run to confirm — it
borrows the WebUI container's node and no containers are running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:37:35 +01:00
librelad
8b5e02c760 refactor(storage): resolve every app directory through appDir
The main sweep — ~260 call sites across ~100 files move from string
concatenation on a single root to appDir/storageAppDirs/storageAppConfigs.
On a single-root install the resolved paths are identical, so this is a
no-op until a location is registered.

Enumerators were the interesting half. `for d in "$containers_dir"/*/`
appears in the menus, the registry/artifact scanners and the DNS setup —
and a shell glob cannot list a rootless 751 tree at all, which is the
same bug config_find_file.sh already documents in a comment. Routing them
through storageAppDirs (which enumerates as the owning user) fixes that
alongside the multi-root work.

Three places needed judgement rather than substitution:

db_app_scan.sh deletes database rows and port allocations for apps whose
folder is missing, and reaps "empty" app dirs. With a storage location
unmounted, every app on it looks exactly like that. Each of those
branches now gates on appStorageAvailable first — an app on an unplugged
drive is skipped with a notice, never deleted.

instance_create.sh rewrites cloned hooks so an instance touches its own
directory instead of the base app's. Its sed matched ${containers_dir}<type>,
which this sweep just replaced with $(appDir <type>) — so it would have
silently stopped redirecting, and an instance would have written to the
original's files (the adguard auth adapter case its own comment warns
about). Now matches both appDir forms, verified against bare, quoted,
unrelated-app, legacy and prose cases.

peer_shell/peer_pull streamed and extracted relative to the primary root.
Both now use the app's own root, and peer_shell keeps a single-root
fallback since it runs as a restricted SSH shell with no LibrePortal env.

Also fixes a pre-existing bug found on the way: webui_app_config.sh
tested "$containers_dir/frontend/data/last_update", one level short of the
real tree under the libreportal app dir, so the WebUI refresh trigger
after a config update has never once fired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 04:09:51 +01:00
librelad
934193d901 copy: drop deployment trivia from app descriptions
App descriptions are read by people deciding whether they want the app,
not by people maintaining it. Several were spending their last clause on
facts the reader cannot act on and would not recognise — and in Stoat's
case actively talking the app down: "Heavy (16 containers) and does not
federate" is a maintainer's note, not a description.

Eight rewritten, all the same fault:

  stoat        "Heavy (16 containers) and does not federate", and LiveKit
               named as though the reader would know what it is
  vikunja      "Runs as a single container on SQLite, with no database sidecar"
  stalwart     "in a single container"
  gitea        "written in Go", plus "self-hosted Git service" twice in one line
  vaultwarden  "an alternative implementation of the Bitwarden server API
               written in Rust" — says what it is to a developer, not what it
               does for you
  speedtest    "implemented in Javascript"
  adguard      "resolving blocked domains to a local blackhole address"
  matrix       "Installs Synapse plus the Element web client"

Deliberately kept, because they change whether the app suits you rather
than merely describing how it is built: Rocket.Chat's free-edition user
cap, Mattermost's unlimited users, Navidrome's Subsonic compatibility
(it tells you which phone apps will work), Stalwart's protocol list, and
Gluetun's provider count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 01:07:59 +01:00
librelad
6813621fe9 config: enable multiple instances everywhere it is actually possible
Three apps were instanceable and the rest were silent, so the feature
looked far narrower than it is. Every app has now been checked against
the two rules instance_create.sh enforces, and the answer recorded in
its config rather than left unset.

32 apps are instance-safe and now say so. Six are not, and each says why
in its own words instead of being indistinguishable from an app nobody
had reviewed:

  pihole      a DNS server must own port 53
  unbound     a resolver must own its fixed 5335
  stalwart    a mail server must own 25/465/587/993
  traefik     must own 443, and one Traefik routes every other app
  prometheus  node-exporter and cadvisor carry no "prometheus" prefix
  stoat       pins 7881, and database/redis/rabbit/minio carry no prefix

The first four are genuinely one-per-host: the port is not arbitrary, it
is the protocol. The last two are compose-identity problems and could be
fixed by prefixing those service names, which is a change to make
deliberately rather than in passing.

Recorded as an explicit false with a reason, not left unset, so the next
person reads a decision instead of an absence. The audit was verified not
to pass anything vacuously: every app resolves at least one service name,
so no app reached "eligible" merely because nothing was found to check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 23:38:34 +01:00
librelad
e14e295f3f stalwart: record that the public-side ACME path is not fully verified
The private direction of the mode switch is exercised end to end. The
public one has only ever run against a throwaway .test domain, where Let's
Encrypt rejects the contact address before the provider is created — so
everything past that call is reasoned rather than observed.

The plan shape IS confirmed up to that point: contact is a set, matchOn is
the directory URL, and a domain cannot reference automatic certificate
management without an acmeProviderId. What is unproven is the link holding
once the provider actually exists.

Saying so in the file beats leaving it in a chat log nobody reads.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 00:23:45 +01:00
librelad
213c689cc1 stalwart: one primitive for probing the admin listener
stalwart_wait_http hardcoded the /healthz/ prefix and returned a yes/no,
so the admin-console check could not use it and grew its own copy of the
docker exec curl line. Extract stalwart_http_code <path> [max-time] and
build both on it: the wait loop keeps its probe-name signature and its
3s timeout, the console check keeps its 5s and gets the status code back
rather than a verdict, since 404 and no-reply-at-all need saying apart.

Probe commands are byte-identical to before; no behaviour change. The
upgrade verifier keeps its own copy on purpose — verifiers here are
self-contained (see nextcloud's, which inlines the occ idiom rather than
calling the install hook's wrapper) and should not drag a lifecycle file
they have no other use for into an upgrade run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 00:12:46 +01:00
librelad
f9ec4cc986 refactor(auth): drop the unread AUTH_PROFILE key
Eleven app configs declared CFG_<APP>_AUTH_PROFILE as a "capability tier for the
WebUI auth tools". Nothing read it — not a shell script, not the frontend, and it
was never emitted into apps.json, so the WebUI could not have acted on it even in
principle.

The job it was meant to do is already done, and done better: authAdapterCanDo
tests `declare -F authAdapter_<app>_<method>`, so what an app can do is derived
from the functions it actually implements. A declared tier is a second source of
truth that can only drift — traefik declared single_password while its adapter
implements setPassword only, and linkding declared nothing at all while shipping
a full multi-user adapter, and neither mismatch had any effect.

Removed the key and its comment from all eleven configs, and replaced the stale
contract note in auth_adapter.sh with what the dispatcher really does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 23:13:53 +01:00
librelad
af78ce1681 stalwart: make the mode switch finish the job itself
Switching between private and public wrote the setting, reconfigured the
server and then asked the user to run `libreportal app install stalwart`
to make the ports actually change. That left a window where the WebUI
reported public while port 25 was still closed — or worse, reported
private while 25 was still open and listening. A mode switch that does not
move the ports is not a mode switch.

The tool now runs the install itself. Safe from here: tools are dispatched
inline rather than as their own task, so this is not a nested task and
cannot deadlock on the task lock, and nothing in Stalwart's install hooks
calls back into the tool. Provisioning inside that install is a no-op
because it skips once config.json exists.

Dropped the separate firewall rebuild — the install reallocates the ports
and rebuilds the rules from the result, so doing it beforehand only worked
from the old allocation and was then immediately redone.

Verified both directions on a real install: private -> public publishes 25,
public -> private removes it, the admin port keeps its existing random
allocation across both (no --reset-network, so bookmarked WebUI links do
not move), mailboxes survive with their original creation timestamps, and
re-selecting the current mode is a no-op.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 23:03:23 +01:00
librelad
88e9631b68 stalwart: choose private or public mail, and switch between them later
A mail server is two quite different products wearing one name, and until
now LibrePortal only offered the hard one. Installing Stalwart meant being
handed a wall of DNS records, a red error about port 25 and a warning about
reverse DNS — all of it correct, none of it fixable by the installer, and
most of it irrelevant to someone who wanted mailboxes and a shared calendar
on their own network.

CFG_STALWART_MODE now names which one you are running:

  private  mailboxes, IMAP, CalDAV and CardDAV on your own network. Port 25
           is not published at all; the client ports stay bound to the host
           but are never opened through the firewall. No MX, no PTR, no
           deliverability. Nothing to publish, so nothing is printed.
  public   the internet mail server, as before.
  auto     public if Traefik is installed, private if not, resolved at
           install and written back so it reads as a real answer afterwards.

DKIM keys are generated in both modes even though private has no use for
them today — that is what makes switching later a setting change rather
than a key ceremony. The WebUI gets a "Mail Exposure" tool that flips the
setting both ways and reconfigures the server, plus a "Show DNS Records"
tool that prints the live zone including current DKIM keys.

Two things this had to get right, both found by testing rather than
reading. Port access lives in the shell as CFG_<APP>_PORT_n, not just in
the config file, and the compose file is built from the parsed shell
values — editing only the file left the config claiming port 25 was
disabled while the container published it anyway. And going public needs
an AcmeProvider to exist before a domain can reference one, so the switch
creates it; note that doing so registers an account with Let's Encrypt.

Verified through real installs: auto resolves to private with no Traefik,
port 25 is genuinely unpublished and absent from the compose file, the
client ports are skipped by the firewall as host-bound, and the tool
round-trips private -> public -> private with the config landing back
exactly where it started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 22:52:55 +01:00
librelad
c7df07ffc2 containers: trim overlong app card descriptions
Nine LONG_DESCRIPTION values had drifted well past the 90-140 char
range the rest of the catalog uses (stoat was 407). Cut them back
while keeping the caveats that matter — Rocket.Chat's user cap,
Stoat's resource weight, Matrix federation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:57:34 +01:00
librelad
65167463f9 fix(chat apps): tag every service so its IP actually substitutes
Installing rocketchat failed with

    invalid IPv4 address: ParseAddr("IP_DATA_2"): unable to parse IP

ipUpdateComposeTags allocates one IP per SERVICE_TAG_N annotation and fills
IP_TAG_i only where SERVICE_TAG_i exists. The four new apps tagged only their
primary service, so every sidecar — matrix's postgres, mattermost's postgres,
rocketchat's mongo, and fifteen of stoat's sixteen — kept a literal IP_DATA_n
in the deployed compose and docker refused to create the container.

Tag every service that carries an ipv4_address, index-aligned with its IP_TAG.
For stoat that also meant moving caddy from SERVICE_TAG_1 to _6 so the indices
line up with the IPs rather than the reading order.

mastodon had the same latent break (IP_TAG_2 and _3 untagged) and is fixed the
same way — it would have failed on first install for the same reason.

SERVICE_TAG carries the compose *key*, not container_name: 'libreportal app
restart <app> <service>' passes it to 'docker compose restart', which only
understands keys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:48:09 +01:00
librelad
63f276523b stalwart: run as container-root so it can write its own data directory
Found by running the installer for real rather than testing the hook in
isolation. Stalwart never started: it failed to open its database with
"Permission denied" on /var/lib/stalwart, which meant no mail could be
stored and the setup wizard could not be completed by hand either.

The image runs as its own uid 2000. LibrePortal gives container directories
to the docker install user under rootless and to the manager under rooted,
and 2000 is neither, so the bind mounts were unwritable in both modes. This
was not something the new provisioning introduced — it predates it, and the
app has never been able to hold mail.

Running as container-root maps to whichever host user owns those
directories. Under rootless that is the unprivileged docker install user,
not host root.

Also stop discarding the server's error when setup fails. Both failures
that actually occur — a hostname under a TLD that does not resolve, and the
unwritable data directory above — name themselves precisely, and a bare
"setup failed" turns a one-line fix into guesswork.

Verified end to end through `libreportal app install stalwart` on a clean
install: setup applied, DKIM keys generated, postmaster mailbox created,
and the full record set printed from the server's own zone data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:14:19 +01:00
librelad
7ed1cddd5c stalwart: answer the setup wizard instead of handing it to the user
A new Stalwart drops you into a five-screen wizard — hostname, domain,
storage backend, directory, logging, DNS — before it will do anything.
LibrePortal already knows the two answers that matter and the rest have
sane defaults, so asking is asking a question we can answer ourselves.

v0.16 exposes those wizard fields as a `Bootstrap` singleton, so the whole
thing is one `update` applied through the Stalwart CLI. The CLI is not in
the server image (upstream split it into its own repo), but it publishes a
multi-arch container, so we borrow the server's network namespace and run
it there — nothing installed on the host, nothing to clean up, arm64 works.

Setup now also:

- generates DKIM keys (Ed25519 + RSA) with rotation left switched on, and
  requests a TLS certificate. That last one is easy to miss: Traefik only
  fronts the admin port, so 25/465/587/993 never see its certificate and
  clients would hit a self-signed one on 993.
- creates postmaster@<domain>. The generated zone points DMARC and TLS-RPT
  reports there and nothing was creating it, so those reports bounced.
- prints the record set read back from the server rather than composed
  here, so it includes the real DKIM public keys, MTA-STS, TLS-RPT and the
  SRV records clients autoconfigure from. This hook used to tell the user
  to go and fetch DKIM themselves; by that point the keys exist.

Optionally hands DNS to a provider API (Cloudflare/DigitalOcean/DeSEC),
which keeps the whole record set in sync and makes DKIM rotation safe to
leave on. Off by default: the token can write to your zone and lives in
the mail server's database.

Re-running is safe — provisioning is skipped once config.json exists, and
the plans use upsert so they reconcile rather than duplicate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:05:03 +01:00
librelad
0eda3104c7 stalwart: report a missing admin console without failing the upgrade
A failed verify makes the engine abort and restore, and a restore cannot
put back a bundle that was never downloaded — it would roll a working
mail server back a version to fix a missing web page, then hit the same
empty GitHub fetch next time. So the console check now warns loudly and
returns 0; readiness stays the only gate.

Renamed to stalwart_upgrade_check_admin_ui so the name cannot be read as
part of the gate, and bounded its poll to a 60s grace window (capped by
the caller's deadline) — the upgrade result is already decided by then,
so there is no reason to hold the run open on a web asset. The unreach-
able-probe branch is advisory for the same reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:59:02 +01:00
librelad
4c80ea018d stalwart: check the admin console, not just readiness
Stalwart v0.16 does not ship the WebUI in its Docker image — the admin
console is fetched from GitHub on first start. With no outbound HTTPS at
that moment the fetch fails silently: /healthz/ready still answers 200
because the mail server genuinely is serving, so both the installer and
the upgrade verifier reported success while /admin and /account 404'd
with nothing to explain why.

Install hook now probes /admin after the port-25 and PTR checks and, on
404, names the GitHub download as the cause rather than emitting a
generic failure. Upgrade verifier treats stable readiness as necessary
but not sufficient and confirms /admin before returning 0; the console
is polled under the same deadline because the bundle download runs
behind the server coming up, and failing on the first 404 would abort an
upgrade that was seconds from finishing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:54:32 +01:00
librelad
4fae22c89e fix(stalwart): use the official logo instead of my placeholder
Replaces the drawn shield-and-envelope stand-in with the real mark from
stalw.art (/favicon.svg), in their #DB2D54.

Padded from the source's 159.95x139.07 to a square 159.95 viewBox with
the art vertically centred, matching every other catalogue icon — all of
which are square, so a non-square box would letterbox in the app tiles.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 04:23:33 +01:00
librelad
598f74c26b feat(updater): per-app upgrade verifiers — the safety half of stepping
Stepping 31 -> 32 -> 33 is arithmetic. Knowing 32 FINISHED before
touching 33 is the whole safety story, and it is invisible from outside
the app: Nextcloud runs its migration on boot and sits in maintenance
mode — or fails halfway — while Docker reports the container perfectly
healthy. Advance a rung there and a migration has been skipped on live
data.

Contract:  <app>_upgrade_verify <app> <expected-tag> <deadline>  -> 0

Returns 0 ONLY on positive confirmation that the app serves at the
expected version with nothing outstanding. Unhealthy, indeterminate and
timed-out all return non-zero — uncertainty is a failure, not a maybe,
because the alternative gambles with data.

  nextcloud  `occ status`: installed, NOT in maintenance, no pending DB
             upgrade, and the running major matches the tag. Maintenance
             mid-migration is expected and simply keeps waiting.
  mastodon   /health serving, ZERO "down" rows in db:migrate:status, and
             the version from /api/v1/instance matching. /health alone is
             insufficient — Puma answers before migrations finish.
  stalwart   /healthz/ready (per its documented probes), required to hold
             stable rather than flash once. Weaker by design: the probes
             confirm serving but report no version, and the file says so
             rather than implying more.

updaterVerifyGeneric (running + healthy + no restart during a settle
window) is the fallback for everything else, and is explicitly NOT
sufficient to justify climbing a rung — the engine will refuse to ladder
an app with no declared verifier.

9 tests drive the dangerous states directly: maintenance mode, pending DB
upgrade, and a wrong major all correctly REFUSE to verify; clean states
pass. Those three negatives are the ones that would have corrupted data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:58:23 +01:00
librelad
fd8ac29621 feat(stalwart): auto-update patches, VERSION config for release moves
Reverses the manual default from 4ee2529, which was over-cautious once
the tag pin is taken into account.

Two things were conflated. Auto-update does not reinstall anything: it
snapshots, `compose pull`, `up -d` — the container is recreated from the
new image and the data volume is untouched. And because the image is
pinned to v0.16, the updater compares the digest of THAT tag, so auto
can only ever apply rebuilds of 0.16 (security/bug patches). It cannot
jump to 0.17. That is the safe half of updating, and there is no good
reason to withhold it.

Adds CFG_STALWART_VERSION=v0.16, which drives the image tag through the
existing #LIBREPORTAL|STALWART_VERSION_TAG| sentinel (verified: setting
it to v0.17 rewrites the image line). Moving between releases is now a
config change a user can make from the app's config page — the roadmap's
config-first version identity, used for real.

Net behaviour: patches land unattended inside the update window; a
version jump stays a deliberate decision, which is what pre-1.0 software
with a settling storage schema warrants.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:13:47 +01:00
librelad
4ee25292d5 feat(stalwart): add Stalwart Mail Server as a catalog app
One container providing SMTP/IMAP/POP3/JMAP plus CalDAV/CardDAV, an admin
UI and spam filtering — chosen over mailcow (owns its own installer, which
is what killed the earlier attempt now sitting in scripts/unused/) and
over Mailu (~7 containers) because a single image with a single data dir
is the only shape that fits the existing conventions cleanly: one anchor
service the updater can version, one path the backup engine can snapshot.

Mail-specific departures from the usual app template, each deliberate:

* Ports are FIXED, not random. Other mail servers connect to :25 by
  number and clients expect 465/587/993 — a randomised external port
  would silently make the server unreachable. Only the admin UI takes a
  random port, since that one really is just a browser behind Traefik.
  143/995/4190/443 ship disabled; the port processor comments them out.

* UPDATE_TYPE=manual and the image pinned to v0.16, not :latest.
  Stalwart is pre-1.0 and has said the storage schema is still being
  finalised, so an unattended minor bump could carry a data migration on
  the message store. This is the one app where the auto default is wrong.

* BACKUP_STRATEGY=stop-snapshot-start. The message store is written
  continuously; a live copy can land mid-transaction. Seconds of queued
  delivery (senders retry) buys a consistent snapshot.

* The install hook checks outbound port 25 and reverse DNS, then prints
  the MX/SPF/DMARC records with real values. A mail server whose
  container started is not a working mail server, and every remaining
  requirement lives at the registrar or the VPS provider.

Admin credentials are seeded via STALWART_RECOVERY_ADMIN from the app
config rather than left to Stalwart's first-run random password, which
would otherwise exist only in the container log.

Icon is a drawn placeholder, not the upstream trademark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:02:26 +01:00