LibrePortal/scripts/backup/engine/restic_restore.sh
librelad 6018250526 Choose which snapshot to restore, per app and for the settings
A snapshot is one app's data, or the settings tree — never a machine. A
four-snapshot repository is typically two apps plus two versions of the
settings, not four backups to pick between. So the choice belongs on Contents,
after unlocking, where each snapshot has a name and a date rather than being a
hash.

Every row with more than one snapshot gets a picker, defaulting to the newest.
A row with one shows its date as text: a dropdown holding a single entry is a
control that cannot be operated, and it makes a repository with one backup look
like it is hiding something.

The chain already supported this. restorePickSnapshot has always passed any
value that is not the string "latest" straight through as an id; nothing ever
offered the choice. What was missing:

  - restoreInspect returns every snapshot per app and for the settings, not
    just the newest.
  - restoreFirstRunBulk reads an optional RESTORE_SNAPSHOT_CHOICE map instead
    of hardcoding "latest". An associative array rather than an argument,
    because the CLI wrapper pads argv to nine slots and a per-app map cannot
    survive it; the map reaches the host as base64 JSON, validated at the route
    against restic short ids and app names since both hit a command line.
  - backupRestoreSystemConfig takes a snapshot AND a host.

That host was a real bug. It defaulted to this machine's install name, which is
right for "recover my own settings" and wrong for a rebuild — the snapshots
carry the name of the machine being rebuilt FROM. It surfaced the moment a
restore adopted a config with a different install name and the next lookup
found nothing at all.

Verified by restoring both settings snapshots and diffing: 28bedbb0 brings back
a config carrying example.com, cc5b6bcf one with no domains.

Two CSS traps on the picker: appearance stayed `auto`, so the browser painted
its own control and ignored the colours entirely while the computed styles
looked right; and a `background:` shorthand later in the rule silently reset the
background-image, wiping the arrow set three lines above it.

lp-restore-adopt-test asserted configs/* were mode 0755 and started failing on
configs/webui, which libreportal-ownership sets to 0751:container on purpose —
tighter, and perfectly traversable. It asserts "the container user can traverse
it" now. A test that pins an incidental number reports a regression every time
someone improves the thing it is watching.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:44:21 +01:00

153 lines
6.3 KiB
Bash

#!/bin/bash
# Build the `unshare` prefix that lets a NON-ROOT restic recreate the container
# uids a snapshot recorded.
#
# Why this is needed: backups run as the docker install user (runBackupOp — the
# backup engine never gets root). A non-root restic cannot chown a restored file
# to anyone else, so every file came back owned by that user. For LibrePortal's
# own files that is correct; for the ones a CONTAINER owns it is fatal. Under
# rootless, a container process running as uid N appears on the host as
# subuid_start + N - 1 (prometheus' nobody -> 296605, postgres -> 231141), and an
# app whose data dir is no longer owned by its own uid does not start:
# prometheus dies on "open data/queries.active: permission denied", and postgres
# refuses outright unless its data dir is 0700 and its own. Restores therefore
# handed back apps that could not boot.
#
# The fix needs no new privilege. The docker install user already owns a subuid
# range (that is what makes rootless work), so it may enter a user namespace in
# which it is root and those subuids are mappable. Mapping them to THEMSELVES
# means an id recorded in the snapshot is written back as the same host id.
#
# Files recorded as the docker install user's own uid are the one gap: that uid
# is outside the subuid range and is already consumed by the inner-root mapping,
# so restic's lchown for them fails with EINVAL. It is harmless — restic runs as
# inner root, which IS that user on the host, so those files already land with
# exactly the right owner. resticRestoreErrorsAreBenign below is what keeps that
# from being reported as a failed restore.
# True when every error restic reported is the expected "cannot map the backup
# user's own uid" one described above. Anything else — a missing pack, a full
# disk, a permission problem on the target — must still fail the restore.
resticRestoreErrorsAreBenign()
{
local out="$1"
local bad
# Every line restic prints for a failed ownership set, minus the benign form.
bad=$(printf '%s\n' "$out" | grep -E "^ignoring error for " \
| grep -vE "lchown .*: (invalid argument|operation not permitted)$")
[[ -z "$bad" ]]
}
resticRestoreSnapshot()
{
local idx="$1"
local snapshot_id="$2"
local target_dir="$3"
local include_path="$4"
if [[ -z "$snapshot_id" || -z "$target_dir" ]]; then
isError "resticRestoreSnapshot requires snapshot_id and target_dir"
return 1
fi
resticEnvExport "$idx" || return 1
runFileOp mkdir -p "$target_dir"
local args=(restore "$snapshot_id" --target "$target_dir")
[[ -n "$include_path" ]] && args+=(--include "$include_path")
isNotice "Restoring ${snapshot_id:0:8} from $(resticLocationName "$idx")$target_dir"
local ns_prefix=()
mapfile -t ns_prefix < <(backupUsernsPrefix)
# Output is captured (not streamed) so the benign-error check below can read
# it; it is echoed straight back afterwards, so the operator sees the same
# restic report as before.
local out rc
out=$(runBackupOp "${ns_prefix[@]}" restic "${args[@]}" 2>&1)
rc=$?
printf '%s\n' "$out"
# restic exits non-zero for un-mappable-uid lchowns even though the file
# CONTENTS landed. Forgive only that case.
#
# With the namespace mapping fixed (see restic-userns-exec), the only id
# that is still unmappable is the restoring user's own — its slot is spent
# on inner root — and a file stored as <caller>:<caller> lands owned by the
# caller regardless, because that is who inner root is on the outside. So
# these really are correct, which is what the message used to claim before
# the mapping worked and container-owned data was quietly losing its owner.
#
# Still report the count: if this number is large the mapping has stopped
# working again, and the symptom is an app that cannot write its own data.
if [[ $rc -ne 0 && ${#ns_prefix[@]} -gt 0 ]] && resticRestoreErrorsAreBenign "$out"; then
local _lch
_lch=$(printf '%s\n' "$out" | grep -cE "^ignoring error for .*lchown ")
isNotice "Ownership warnings on ${_lch} file(s) owned by ${docker_install_user:-the backup user} — expected; they are restored correctly."
rc=0
fi
resticEnvUnset
return $rc
}
resticRestoreAppLatest()
{
local idx="$1"
local app_name="$2"
local target_dir="$3"
local host="${4:-$CFG_INSTALL_NAME}"
local snapshot_id
snapshot_id=$(resticSnapshotLatestId "$idx" "$app_name" "$host")
if [[ -z "$snapshot_id" ]]; then
isError "No snapshot found in $(resticLocationName "$idx") for app=$app_name host=$host"
return 1
fi
# Prefer the path the SNAPSHOT records over this host's layout: they differ
# whenever the snapshot came from a host with a different --containers-dir,
# or from a different storage location, and an include filter that matches
# nothing restores nothing without saying so.
local include_path=""
if declare -f storageSnapshotSourcePath >/dev/null 2>&1; then
include_path=$(storageSnapshotSourcePath "$idx" "$snapshot_id" "$app_name" 2>/dev/null) || include_path=""
fi
[[ -z "$include_path" ]] && include_path="$(appDir "$app_name")"
resticRestoreSnapshot "$idx" "$snapshot_id" "$target_dir" "$include_path"
}
resticRestoreSystemLatest()
{
local idx="$1"
local target_dir="$2"
local host="${3:-$CFG_INSTALL_NAME}"
# An explicit snapshot, when the caller has one. The wizard offers a choice
# between system-config snapshots, and a picker whose answer is ignored is
# worse than no picker.
local want="${4:-}"
local snapshot_id
if [[ -n "$want" && "$want" != "latest" ]]; then
snapshot_id="$want"
else
resticEnvExport "$idx" || return 1
snapshot_id=$(runBackupOp restic snapshots \
--tag "system=config" --host "$host" \
--latest 1 --json --no-lock 2>/dev/null | \
grep -o '"short_id":"[^"]*"' | head -1 | cut -d'"' -f4)
resticEnvUnset
fi
if [[ -z "$snapshot_id" ]]; then
isError "No system-config snapshot found in $(resticLocationName "$idx") for host=$host"
return 1
fi
# Whole-snapshot restore (the snapshot is just the config tree) into staging.
resticRestoreSnapshot "$idx" "$snapshot_id" "$target_dir"
}