LibrePortal/scripts/app/install/app_install.sh
librelad 770492b7c7 Allocate IPs per missing service, not all-or-nothing
Not a flake. IP allocation lived in the else-branch of "did the database return
any rows for this app", so it ran only when the app held ZERO rows. An app with
even one row skipped the loop entirely, and a service without a row never got an
IP and never would. Its IP_TAG_<n> stayed unfilled, the literal IP_DATA_<n>
reached the compose, and docker refused the app with

  invalid IPv4 address: ParseAddr("IP_DATA_3")

which surfaced as "no container started (image pull failed?)". Nothing repaired
it: reinstalling re-ran the same skip, so the app stayed broken until someone
uninstalled it and wiped the rows.

Partial state is not exotic — an app that GAINS a service in a later version hits
this on its very next install, because the old services still hold rows. That is
the case worth worrying about; Stoat only got there by being installed and
uninstalled repeatedly.

Reproduced deterministically by deleting one row from a healthy 16-service Stoat:
the install reported "No IP allocated for service: stoat-rabbit" as a NOTICE,
then "Success: Updated 15 IP tag system", then failed at compose. After the fix
the same broken state self-heals — "Allocated IP: stoat/stoat-rabbit" — with no
uninstall.

Three more bugs in the same path, all found while tracing it:

- ipFindAvailable tested pool membership with a substring match against the
  newline-joined list of allocated IPs, so .4 read as taken whenever .46 or .147
  existed. Demonstrated: with 3 addresses allocated it excluded 5. Harmless at
  low occupancy, but it silently shrinks the pool as it fills and would report
  exhaustion early. Now an exact whole-line match.

- ipFindAvailable set available_ip="" on an exhausted pool and carried on to
  index the empty array, where RANDOM % 0 is a division-by-zero that would bury
  the real message. ipAllocation did the same and still ran its INSERT, writing
  a row with an empty resource_value — which then satisfied "this service has an
  allocation" forever after, making the service unrepairable. Both now return.

- first_allocated_ip was only assigned inside the allocate branch, so on every
  reinstall (where rows already exist) it came out empty and the trusted-domains
  list shipped with a hole. Now taken from the mapping.

An unfilled tag is also an error rather than a notice now: the compose is
unshippable at that point, and reporting "Success: updated 15 IP tags" is how
this reached the user as a confusing pull failure several steps later. The
install backstop no longer guesses "(image pull failed?)" either — that guess
was written for one cause and misdirects for every other.

Verified: clean install allocates all 16 with no unfilled tags, a deliberately
broken row self-heals, and no duplicate IPs exist across any app.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 01:35:19 +01:00

235 lines
11 KiB
Bash

#!/bin/bash
# Generic per-app install/uninstall/start/stop/restart/edit driver.
#
# The 31 containers/<app>/<app>.sh files used to each define their own
# install<App>() with the SAME 10-step sequence. ~4,000 lines of duplicated
# boilerplate. This is the one place that sequence lives now; per-app
# customisation lands in declarative hooks in containers/<app>/tools/
# <app>_tools.sh (or wherever the app's tools.sh lives — auto-sourced).
#
# Dispatch is driven by the `$<slug>` global variable (set by dockerInstallApp
# in scripts/docker/app/functions/function_install_app.sh — `declare $app=i`).
# Same convention the per-app .sh files used; nothing changes for the caller.
# Actions are letters: c (config edit), u (uninstall), s (stop), r (restart),
# i (install), t (treated like c — legacy alias).
#
# Hook surface — all are `declare -f`-gated, silent no-op when absent:
#
# <slug>_install_pre before any install work. THE ONE HOOK WHOSE
# RETURN VALUE COUNTS — return non-zero and
# the install stops here (see CFG_<APP>_REQUIRES)
# <slug>_install_post_setup after dockerConfigSetupToContainer
# (install folder + .config exist; compose
# file not yet written)
# <slug>_install_post_compose after dockerComposeSetupFile (the compose
# TEMPLATE has been copied into place;
# container not yet up). NOTE: the tags are
# NOT substituted yet — IPs, ports and
# #LIBREPORTAL values are filled later, by
# dockerConfigSetupFileWithData during
# dockerComposeUpdateAndStartApp. A hook that
# needs a settled value must read it from
# CFG_<APP>_* / the port arrays in scope, not
# by grepping the deployed compose.
# <slug>_install_post_start after dockerComposeUpdateAndStartApp
# (container is up; the place for
# wait-for-ready + post-up API calls)
# <slug>_install_message_data echoes extra args for menuShowFinalMessages
# (typically credentials / URLs)
# <slug>_install_post very last thing, after the final message
#
# <slug>_uninstall_pre / _post around dockerUninstallApp
# <slug>_stop_post after dockerComposeDown
# <slug>_restart_post after dockerComposeRestart
#
# Hooks receive $app_name as $1 (and stay un-namespaced — they're already
# slug-prefixed). Return code is ignored unless they isError; the install
# continues regardless. Use that escape hatch for non-fatal app-specific
# refinements (rotate a key, patch a yaml after start, etc.).
# Returns the hook's own exit status when it ran, and 0 when no such hook
# exists. The explicit `return 0` matters: without it an absent hook returns the
# status of the failed `declare -F` test, i.e. non-zero, and any caller that
# gates on the result would treat "app has no hook" as "hook failed".
_appCallHook()
{
local hook_name="$1"; shift
if declare -F "$hook_name" >/dev/null 2>&1; then
"$hook_name" "$@"
return $?
fi
return 0
}
# Standard "post-start integration" steps. Same for every app. Lives in a
# helper so the generic install body stays readable; safe for apps that
# don't tag for monitoring (the helpers no-op gracefully).
_appPostStartIntegrations()
{
local app_name="$1"
appUpdateSpecifics "$app_name"
setupHeadscale "$app_name"
databaseInstallApp "$app_name"
webuiContainerSetup "$app_name" install
# Scrape-target + dashboard re-gather. The compose-level toggle ran
# already (post-compose, so the running container reflects it).
# monitoringRefreshAll is self-correcting and no-ops when Prometheus
# / Grafana aren't installed.
if declare -F monitoringRefreshAll >/dev/null 2>&1; then
monitoringRefreshAll 2>/dev/null || true
fi
}
installApp()
{
local app_slug="$1"
local config_variables="$2"
# APP_NAME comes from the app's CFG_<APP>_APP_NAME (the user's chosen
# subdomain / install name). Fall back to the slug if unset.
local app_name_var="CFG_${app_slug^^}_APP_NAME"
local app_name="${!app_name_var:-$app_slug}"
# Dispatch flags live in the $<slug> global, e.g. linkding=i. Default to
# install if nothing set — installApp called directly without flag = install.
local actions="${!app_slug:-i}"
# Setup phase shared by every action (folder + variables).
if [[ "$actions" == *[cCtTuUsSrRiI]* ]]; then
dockerConfigSetupToContainer silent "$app_slug"
initializeAppVariables "$app_name"
fi
if [[ "$actions" == *[cCtT]* ]]; then
editAppConfig "$app_name"
fi
if [[ "$actions" == *[uU]* ]]; then
_appCallHook "${app_slug}_uninstall_pre" "$app_name"
dockerUninstallApp "$app_name"
_appCallHook "${app_slug}_uninstall_post" "$app_name"
fi
if [[ "$actions" == *[sS]* ]]; then
dockerComposeDown "$app_name"
_appCallHook "${app_slug}_stop_post" "$app_name"
fi
if [[ "$actions" == *[rR]* ]]; then
dockerComposeRestart "$app_name"
_appCallHook "${app_slug}_restart_post" "$app_name"
fi
if [[ "$actions" == *[iI]* ]]; then
isHeader "Install $app_name"
# The ONE hook whose return value is honoured. An app declaring
# CFG_<APP>_REQUIRES uses its _install_pre to refuse when a prerequisite
# is missing; before this gate existed the refusal printed its reasons
# and the install carried straight on, leaving a half-configured app
# whose later steps failed for confusing secondary reasons.
if ! _appCallHook "${app_slug}_install_pre" "$app_name"; then
isError "Install of $app_name stopped — its pre-install checks did not pass."
return 1
fi
((menu_number++))
echo ""
echo "---- $menu_number. Setting up install folder and config for $app_name."
echo ""
dockerConfigSetupToContainer "loud" "$app_name" "install" "$config_variables"
isSuccessful "Install folders and Config files set up for $app_name."
_appCallHook "${app_slug}_install_post_setup" "$app_name"
((menu_number++))
echo ""
echo "---- $menu_number. Setting up the $app_name docker-compose.yml."
echo ""
dockerComposeSetupFile "$app_name"
# Compose-level monitoring toggle MUST run before docker-compose up
# — the compose file is the source of truth for the running
# container, so editing it post-start wouldn't take effect until
# the next restart. Idempotent + no-op for apps without a marker
# block; apps that toggle additional files (authelia config.yml,
# traefik traefik.yml, unbound unbound.conf …) call it again from
# their _install_post_compose hook.
if declare -F monitoringToggleAppConfig >/dev/null 2>&1; then
monitoringToggleAppConfig "$app_name" "docker-compose.yml" 2>/dev/null || true
fi
_appCallHook "${app_slug}_install_post_compose" "$app_name"
# Optional .env handling — apps that ship a .env in their template
# dir get it copied + tag-substituted. No-op for apps without one.
if [[ -f "${install_containers_dir}${app_slug}/.env" ]]; then
local result
result=$(copyResource "$app_name" ".env" "")
checkSuccess "Copying .env for $app_name"
configSetupFileWithData "$app_name" ".env"
fi
((menu_number++))
echo ""
echo "---- $menu_number. Updating file permissions before starting."
echo ""
fixPermissionsBeforeStart "$app_name"
isSuccessful "File permissions updated for $app_name."
((menu_number++))
echo ""
echo "---- $menu_number. Running docker-compose to install + start $app_name."
echo ""
dockerComposeUpdateAndStartApp "$app_name" install
_appCallHook "${app_slug}_install_post_start" "$app_name"
# Reality gate: after `up`, the app MUST have at least one container
# (its compose project == the app dir name). An image-pull failure
# creates none, yet the deep compose call chain swallows that error —
# without this the app is recorded as installed+active with nothing
# running (the silent-fail seen when a low path-MTU black-holed pulls).
# ps -a (not just running) so a created-but-slow-to-start container
# still counts; only a total absence is treated as failure.
if declare -F dockerCommandRun >/dev/null 2>&1 \
&& ! dockerCommandRun "docker ps -a --filter label=com.docker.compose.project=$app_name --format '{{.Names}}' 2>/dev/null" 2>/dev/null | grep -q '[^[:space:]]'; then
# Deliberately does not name a cause. "(image pull failed?)" was a
# guess carried over from the case this backstop was written for, and
# it actively misdirects for every other one — an unsubstituted tag
# that made compose reject the file reads as a registry problem, and
# the real error is further up the log.
isError "$app_name: no container started — not installed."
isNotice "The compose output above says why. Common causes: an image that could not be pulled, or a compose the daemon rejected (e.g. an unsubstituted LibrePortal tag)."
eval "$app_slug=n"
return 1
fi
((menu_number++))
echo ""
echo "---- $menu_number. Running post-install integrations."
echo ""
_appPostStartIntegrations "$app_name"
((menu_number++))
echo ""
echo "---- $menu_number. You can find $app_name files at $containers_dir$app_name"
echo ""
# Final-message data — apps that want extra args (creds, URLs, etc.)
# printed in the menu output echo them from their hook. Word-split is
# intentional: each space-separated token becomes a positional arg.
local msg_data=""
if declare -F "${app_slug}_install_message_data" >/dev/null 2>&1; then
msg_data=$("${app_slug}_install_message_data" "$app_name")
fi
# shellcheck disable=SC2086 # intentional split — hook returns "u p" etc.
menuShowFinalMessages "$app_name" $msg_data
_appCallHook "${app_slug}_install_post" "$app_name"
menu_number=0
fi
# Reset the dispatch flag so a stale value doesn't trip a later call.
eval "$app_slug=n"
}