backup: stop claiming restored files "are already owned correctly"

resticRestoreSnapshot forgives the un-mappable-uid lchown failures so a restore
is not aborted by them, and reported: "expected, they are already owned
correctly". That is true only for LibrePortal's own files, whose owner is the
backup user restic already runs as. It is false for container-owned data.

_resticUsernsPrefix maps the subuid range and root, but unshare takes one range
per option so the backup user's own GID is never mapped — and app data is
written as <container-uid>:<backup-user>. Every such chown fails with EINVAL and
the file falls back to <backup-user>:<backup-user>. Verified directly:
231543:231543 applies, 231543:1002 does not.

Observed on a 13-app restore: grafana's grafana.db is recorded as 231543:1002
and landed as 1002:1002, so grafana (running as 231543) could not write it at
mode 0640 and died with "attempt to write a readonly database" — under a restore
that reported success. 1626 of that run's 2086 failed chowns were grafana's.

This commit does not fix the mapping — that is the backup engine's ownership
handling rather than the first-run restore path, and the candidate fixes
(newuidmap multi-range maps, or restoring as root via a path-validated helper)
want a decision first. See docs/roadmap/first-run-restore.md §3.5. What it fixes
is the reporting: count the files and say plainly that container-owned data was
not reinstated and the app may fail to write.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
librelad 2026-08-27 13:42:53 +01:00
parent c9779d6581
commit 1fa36abb63
2 changed files with 82 additions and 3 deletions

View File

@ -200,6 +200,64 @@ stoat's livekit publishes a fixed UDP range (5000050100) that it advertises t
clients and so cannot be re-rolled, and a desktop's `kdeconnectd` held 50016.
Worth a fixed-range preflight check of its own; see §9.
### 3.5 — Open: ownership is not actually being reinstated
Found by the same run and **not fixed** — it is a different subsystem (the
backup engine's ownership handling), it predates all of the above, and the
sensible fixes are security-relevant enough to want a decision first.
`grafana` restored and then died with *"attempt to write a readonly database"*.
The snapshot records:
```
-rw-r----- 231543 1002 /libreportal-containers/grafana/grafana_storage/grafana.db
```
and what landed on disk was `1002:1002`. The owner was lost, so grafana —
running as 231543 — cannot write its own database at mode 0640.
The cause is exact. `_resticUsernsPrefix` runs restic under
```
unshare --map-root-user --map-users=231072:231072:65536 --map-groups=231072:231072:65536
```
and `unshare` accepts **one range per option**. So the backup user's own GID
(1002) is never mapped, while LibrePortal writes app data as
`<container-uid>:<backup-user>`. Verified directly:
| target | result |
|---|---|
| `231543:231543` (both in the subuid range) | ✅ applied |
| `231543:1002` (the real case) | ❌ EINVAL → falls back to `1002:1002` |
| `1002:1002` (LibrePortal's own files) | ❌ EINVAL → `1002:1002`, which is correct anyway |
1626 of the 2086 failed chowns in one 13-app restore were grafana's.
`resticRestoreErrorsAreBenign` forgave all of them and reported *"expected, they
are already owned correctly"* — true only for the third row. That message is now
honest about what was not reinstated, which is the only part fixed here.
Ways out, none free:
1. **`newuidmap`/`newgidmap`** can write multi-range maps (`1002 1002 1` plus
`231072 231072 65536`), which is exactly what is needed — but it means
forking, mapping from outside, then continuing, rather than a one-line
prefix. `unshare --map-auto --map-current-user` does **not** work: it maps
the subuid range to low inner ids (container-style, `0 → 231072`), while
restic needs identity.
2. **Restore as real root via a helper**, the way `app-adopt` already moves
trees, with the target path validated against the storage registry. Simplest
and most reliable; widens what root does on the manager's say-so.
3. **Reapply ownership afterwards** from `restic ls -l`, via a root helper.
Contained, but a second pass over every file.
Recommendation: (1) if it can be kept small, else (2) with the same path
validation `app-adopt` uses. Worth noting that apps whose data is captured into
`.lp-backup/files` are rehydrated "via container" as the right uid and so
survive regardless — which is why bookstack, matrix, mattermost and nextcloud
came back healthy and grafana did not.
## 4. The password problem, stated plainly
**An encrypted repository cannot be opened with anything inside itself.** `CFG_BACKUP_LOC_<idx>_PASSWORD` lives in the system config — which is *inside the backup*. So on a fresh machine the user must supply the repository password by hand. There is no way around this and it is not a bug; it is what encryption means.

View File

@ -92,10 +92,31 @@ resticRestoreSnapshot()
rc=$?
printf '%s\n' "$out"
# restic exits non-zero for the un-mappable-uid lchowns even though the files
# themselves landed correctly. Only forgive that exact case.
# restic exits non-zero for un-mappable-uid lchowns even though the file
# CONTENTS landed. Forgive only that case — but do not pretend it is
# nothing, which is what this used to do.
#
# It said "expected, they are already owned correctly". That holds for
# LibrePortal's own files (owner = the backup user, which is who restic runs
# as anyway). It does NOT hold for container-owned data, and the namespace
# this runs in cannot currently map those: unshare takes a single range per
# option, so _resticUsernsPrefix maps the subuid range and root but not the
# backup user's own GID — and LibrePortal writes app data as
# <container-uid>:<backup-user>. Every such chown fails and the file falls
# back to <backup-user>:<backup-user>.
#
# Observed: grafana's grafana.db is recorded in the snapshot as 231543:1002
# and restored as 1002:1002. At mode 0640 the grafana process — running as
# 231543 — then cannot write it, and the app dies with "attempt to write a
# readonly database". The restore reported success.
#
# So: still do not fail the restore (the data is there and some apps are
# rehydrated by other means), but say plainly what was not reinstated.
if [[ $rc -ne 0 && ${#ns_prefix[@]} -gt 0 ]] && resticRestoreErrorsAreBenign "$out"; then
isNotice "Restore reported ownership warnings for LibrePortal's own files — expected, they are already owned correctly."
local _lch
_lch=$(printf '%s\n' "$out" | grep -cE "^ignoring error for .*lchown ")
isNotice "Restore could not reinstate ownership on ${_lch} file(s); they now belong to ${docker_install_user:-the backup user}."
isNotice "LibrePortal's own files are correct that way. Container-owned data is NOT — an app may fail to write (e.g. a read-only database). Check the app after it starts."
rc=0
fi