From c24396d50a2e1066dc2bfc12f4c46360f6d55545 Mon Sep 17 00:00:00 2001 From: Lemon-miaow Date: Sat, 26 Sep 2026 13:40:47 +0800 Subject: [PATCH] =?UTF-8?q?docs:=20PostgreSQL=20=E5=9C=A8=20k3s=20?= =?UTF-8?q?=E5=86=85=E8=BF=90=E8=A1=8C=E5=90=8E=E7=9A=84=E5=8D=87=E7=BA=A7?= =?UTF-8?q?=E3=80=81=E5=9B=9E=E9=80=80=E3=80=81=E5=8D=B8=E8=BD=BD=E4=B8=8E?= =?UTF-8?q?=E6=8E=92=E9=9A=9C?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- README.md | 2 +- README_EN.md | 2 +- docs/operations.md | 127 +++++++++++++++++++++++++++++----------- docs/troubleshooting.md | 109 +++++++++++++++++++++++++--------- 4 files changed, 177 insertions(+), 63 deletions(-) diff --git a/README.md b/README.md index ef17c88..88af886 100644 --- a/README.md +++ b/README.md @@ -33,7 +33,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy, curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash ``` -脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。 +脚本将自动安装 K3s,在 K3s 内部署 PostgreSQL 与控制平面,并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。旧版本装在宿主上的 PostgreSQL 会在重跑时整库迁进 K3s,宿主上的那份停用保留,供回退(见 [运维手册 §4](docs/operations.md#4-upgrading-the-pieces-around-felis))。 动手之前,脚本先检查内存、磁盘、端口、网段冲突、已有的 Kubernetes 和外网连通,把所有问题一次列出并停下,主机上什么都没改(检查项见 [运维手册 §1](docs/operations.md#1-supported-hosts))。 diff --git a/README_EN.md b/README_EN.md index abe9c5f..a51df1c 100644 --- a/README_EN.md +++ b/README_EN.md @@ -36,7 +36,7 @@ every push): curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash ``` -The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain. +The script installs K3s, deploys PostgreSQL and the control plane inside it, and launches a setup wizard. Once done, open your browser at the configured domain. A PostgreSQL an earlier release installed on the host is moved into K3s on the next rerun, and the host copy is stopped and kept for a rollback (see [Operations §4](docs/operations.md#4-upgrading-the-pieces-around-felis)). Before it changes anything, the script checks RAM, disk, ports, network-range clashes, any Kubernetes already there and outbound access, lists every problem at once and stops with the host untouched (the checks are in [operations §1](docs/operations.md#1-supported-hosts)). diff --git a/docs/operations.md b/docs/operations.md index 7fec058..5ee8a4c 100644 --- a/docs/operations.md +++ b/docs/operations.md @@ -12,12 +12,13 @@ Evidence tags follow troubleshooting.md: **[VM-VERIFIED]** was run on a real hos ## 1. Supported hosts `deploy/bootstrap.sh` provisions a single node. It needs systemd, root, and one of the -package managers below; everything else (Docker, k3s, PostgreSQL, the JRE, cloudflared) -it installs. +package managers below; everything else (Docker, k3s, the JRE, cloudflared) it installs. +PostgreSQL runs inside k3s as the `felis-postgres` Deployment, from the official image the +release pins by digest, with its data on the host in `/var/lib/felis/postgres`. | OS family | Package manager | Architectures | Status | |---|---|---|---| -| CentOS Stream 9 (firewalld active, PostgreSQL 13) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall | +| CentOS Stream 9 (firewalld active, SELinux enforcing) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall, with the database on the host PostgreSQL 13 of the releases before felis-postgres; felis-postgres and the move into it [SH-TESTED] | | Ubuntu 24.04 LTS | apt | x86_64 | **[CI]** fresh install, same-commit rerun, and upgrade from the newest release to the pushed commit | | RHEL / Rocky / Alma 9, Fedora | dnf | x86_64, aarch64 | [CODE-ONLY] same code path as CentOS Stream | | Debian 12, other Ubuntu releases | apt | x86_64, aarch64 | [CODE-ONLY] | @@ -34,7 +35,7 @@ cloudflared is left as it is, see §4): | Temurin JRE (Velocity) | 25, patch build pinned | `FELIS_JRE_VERSION`, sha256 per architecture | | Go (nano builds) | 1.26.8 | `GO_PINNED_VERSION`, sha256 per architecture | | Minecraft / Limbo / Paper / Velocity / LuckPerms | `deploy/game-stack.lock` | §15b | -| PostgreSQL | the distribution's package | 13 and 18 are exercised by the `pgint` CI job | +| PostgreSQL | 18.6, the official `postgres` image by digest | `POSTGRES_IMAGE` in `bootstrap.sh`, `defaultPostgresImage` in `internal/platform` | 32-bit hosts are not supported: there is no k3s, JRE or Go build the installer will fetch for them. @@ -62,13 +63,12 @@ host the checks misjudge. Two things the host must keep for as long as the install lives: -- **Its address.** The install is bound to the IPv4 address it was made on (the - database connection string, `pg_hba.conf`, the network policies, the panel - certificate and the default nip.io domain all carry it). Give the host a static - address or a DHCP reservation before installing; the installer warns when the address - is a lease, and the watchdog reports `host-address` when the host loses it - (troubleshooting §13c). The k3s node name is pinned at install time, so a hostname - change is harmless. +- **Its address.** The install is bound to the IPv4 address it was made on (the k3s + node, the network policies, the panel certificate and the default nip.io domain all + carry it). Give the host a static address or a DHCP reservation before installing; + the installer warns when the address is a lease, and the watchdog reports + `host-address` when the host loses it (troubleshooting §13c). The k3s node name is + pinned at install time, so a hostname change is harmless. - **A synchronized clock.** The installer turns NTP on (chrony where nothing else can) and the watchdog reports a clock that stays unsynchronized. Allow outbound UDP 123, or set `FELIS_MANAGE_TIME_SYNC=0` on a host whose clock is kept another way. @@ -134,7 +134,7 @@ server running **[VM-VERIFIED]**: | lobby (Paper, pod limit 1 GiB) | ~0.7–0.85 GB | | login (Limbo, pod limit 512 MiB) | ~0.16 GB | | felis-api, felis-operator, registry gate | ~50 MB each | -| PostgreSQL | ~30 MB plus page cache | +| PostgreSQL (the felis-postgres pod) | ~30 MB plus page cache | | **Total in use** | **~3.4 GB** | Every game server adds the memory its owner gave it: the pod's limit equals its request, @@ -178,6 +178,7 @@ curl -fsSL /deploy/bootstrap.sh | sudo FELIS_VELOCITY_XMX=2G bash | k3s's containerd images | `/var/lib/rancher/k3s/agent/containerd` | 6–9 GB | | Docker's images and build cache | `/var/lib/containerd` (Docker's containerd store) | 5–10 GB after repeated upgrades | | Toolchains and sources | `/opt/felis` | ~2.5 GB | +| Database | `/var/lib/felis/postgres` (felis-postgres's cluster) | tens of MB; the audit log is most of it | | Database bundles | `/var/lib/felis/db-backups` | a few MB each, 14 daily kept | k3s's local-path volumes do not enforce the requested sizes (§9), so every volume shares @@ -201,23 +202,29 @@ With a private repository, fetch it the way the README fetches `bootstrap.sh`. Both modes remove the `felis-*` systemd units and `cloudflared-felis.service`, the Velocity user, `/opt/felis`, `/usr/local/bin/felis`, the installer's cloudflared binary -(unless another unit runs it), the `felis_postgres` and `felis_edge` nftables tables and -the firewalld ports the installer opened. k3s goes with k3s's own `k3s-uninstall.sh` when -the cluster holds nothing but Felis's namespaces; when it runs anything else only -`felis`, `minecraft`, `felis-build` and the MinecraftServer CRD are deleted. +(unless another unit runs it), the `felis_edge` nftables table (and `felis_postgres`, which +releases before the database moved into k3s loaded) and the firewalld ports the installer +opened. k3s goes with k3s's own `k3s-uninstall.sh` when the cluster holds nothing but +Felis's namespaces; when it runs anything else only `felis`, `minecraft`, `felis-build` +and the MinecraftServer CRD are deleted. `--keep-k3s` and `--remove-k3s` override that choice. | | keep data (default) | `--purge` | |---|---|---| | Final database bundle | taken first (`felis db backup -label manual`); a failure stops the uninstall before anything is removed. `--no-backup` skips it | none | -| `felis` database and role | kept | dropped; `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run | +| The database (`/var/lib/felis/postgres`) | kept; felis-postgres is stopped cleanly before k3s goes | deleted with `/var/lib/felis` | +| A host PostgreSQL an earlier release ran the database on | kept as it is: stopped after the move into k3s (below, §4), with its old copy of `felis` | its `felis` database and role are dropped (the server is started for that and stopped again), and `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run | | `/etc/felis` (secrets, `felis.toml`, `offsite.env`, tunnel config) | kept; `bootstrap.done` and the per-run records go | deleted, with the tunnel's credentials file | -| `/var/lib/felis` (database bundles) | kept | deleted | +| `/var/lib/felis` (the database, its bundles) | kept | deleted | | Worlds, archives, registry, uploads | moved to `/var/lib/felis/retained/k3s-storage-/` (with `--keep-k3s`: their volumes switch to `Retain` and stay in place) | deleted | | Felis images, Docker build cache | kept | deleted | -Neither mode removes packages (Docker, PostgreSQL, git, nftables) or the swap file: other -software may use them. On a host that should end up bare: +The two database rows are [SH-TESTED] (`deploy/uninstall_test.sh`); the VM runs above +predate felis-postgres. + +Neither mode removes packages (Docker, git, nftables, and the PostgreSQL server an earlier +release installed) or the swap file: other software may use them. On a host that should +end up bare: ``` sudo swapoff /swapfile && sudo rm /swapfile && sudo sed -i '\|^/swapfile |d' /etc/fstab @@ -232,8 +239,9 @@ DNS records for the panel hostnames, and the Access application. A keep-data uninstall leaves everything a reinstall needs. The installer reuses `/etc/felis/secrets.env`, so the database password and the forwarding and session -secrets are unchanged, and it migrates the kept database instead of creating one -**[VM-VERIFIED]**. +secrets are unchanged, and the installer migrates the kept database instead of creating +one **[VM-VERIFIED]** (with the host database of the releases before felis-postgres). +felis-postgres starts again on the cluster kept in `/var/lib/felis/postgres` [SH-TESTED]. Each step below was run on the reference VM after a keep-data uninstall, and the restored worlds matched their kept `level.dat` checksums **[VM-VERIFIED]**. `kept` names @@ -305,7 +313,7 @@ the version they were installed with unless noted: | Temurin JRE | moves to the pinned patch build | rerun | | k3s | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: moves to the pinned release through that tag's install script, one minor version at a time (a bigger jump stops before anything changes and names the release to go through), never backwards | | cloudflared | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: swaps `/usr/local/bin/cloudflared` for the pinned, sha256-checked release and restarts `cloudflared-felis`; a cloudflared the distribution installed stays with its package manager | -| PostgreSQL | the distribution's package | the package manager for a minor release; a major version needs `pg_upgrade` first (below) | +| PostgreSQL | follows the image the release pins | a minor release comes with a Felis release, and the rerun restarts felis-postgres on it (a few seconds without the API); a major version is a dump and restore (below) | | Docker, git, nftables | distribution packages | the package manager | ```sh @@ -315,8 +323,9 @@ curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap `sudo felis update` reports Felis, Velocity, k3s, cloudflared, the JRE and PostgreSQL against their newest releases; `--k3s`, `--cloudflared`, `--jre` and `--postgres` narrow -it to one. PostgreSQL is compared within its major, since a minor release is a package -update, and a major past its end of life gets a note naming the current one. +it to one. PostgreSQL is read from the felis-postgres container and compared within its +major, since a minor release arrives with a Felis release, and a major past its end of life +gets a note naming the current one. The installer also sets up `felis-update-check.timer`, which runs `felis update --record` once a day around 05:30 (and at boot after a missed run). `--record` stores the result @@ -370,20 +379,72 @@ The env var only carries the password; mail still needs the relay itself, set in ### PostgreSQL major versions [CODE-ONLY] -The installer takes the major the distribution ships (13 on EL9) and never moves it. To -go to a newer one, stop the writers, keep a dump, then use the distribution's upgrade -path: +felis-postgres keeps its cluster in `/var/lib/felis/postgres//docker`. A release that +moves the image to a new major finds the old major's cluster there and stops before it +changes anything: the new server would start an empty cluster beside it. The way across is +a bundle, taken on the release you run now, restored into the new major's empty cluster: ```sh +# On the release you run now: +b="$(sudo felis db backup -label pre-upgrade | sed -n 's/^felis db backup: wrote //p')" +sudo k3s kubectl -n felis scale deploy/felis-postgres --replicas=0 +sudo mv /var/lib/felis/postgres/18 /var/lib/felis/postgres-18.old # the old major's cluster, for a way back + +# Install the new release: it starts an empty cluster on the new major and creates the schema. +curl -fsSL /deploy/bootstrap.sh | sudo bash + +# Put the data back and bring its schema up to the new release. sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=0 -sudo -u postgres pg_dumpall > /root/felis-pg-$(date +%F).sql -# EL9: sudo systemctl stop postgresql; sudo dnf module switch-to postgresql:16 -# sudo dnf install postgresql-upgrade; sudo postgresql-setup --upgrade -# Debian/Ubuntu: sudo pg_upgradecluster main -sudo systemctl start postgresql +sudo felis db restore -yes -no-safety-backup "$b" +sudo felis migrate up -config /etc/felis/felis.host.toml sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=1 ``` +Delete `/var/lib/felis/postgres-18.old` once the new release has run for a while. To go back +instead, scale felis-postgres to 0, move the new major's directory out of +`/var/lib/felis/postgres`, move `postgres-18.old` back as `/var/lib/felis/postgres/18`, and +rerun the older release's installer. + +### The database's move into k3s [SH-TESTED] + +Releases before the move ran the database on a PostgreSQL the installer installed on the +host. The first rerun of a release with felis-postgres moves it, once: + +1. It stops felis-api, felis-operator and the host timers, and heads the host's + `pg_hba.conf` with a block that refuses every connection to `felis` but its own copy + (the original is kept beside it as `pg_hba.conf.pre-pg-move`). +2. It takes a `pre-pg-move` bundle of the host database (`felis db backup`), restores it + into felis-postgres (`felis db restore`) and compares the row count of every table on + both servers. Any failure up to here puts `pg_hba.conf` and the control plane back and + the platform keeps running on the host database, untouched. +3. It stops and disables the host `postgresql` service, which stays installed with its + copy of the data, and writes `/var/lib/felis/postgres-moved`. A host server that also + holds other databases keeps running; its `felis` copy is then reachable over loopback + only. + +From then on the host config points at felis-postgres (`127.0.0.1:15432`, and +`deployment = "felis/felis-postgres"`, through which `felis db` runs `pg_dump`, `psql` +and `pg_restore` inside the pod) and the pods at `felis-postgres.felis.svc:5432`. + +To go back to the host database, for instance to reinstall the release before the move: + +```sh +sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator deploy/felis-postgres --replicas=0 +hba="$(sudo -u postgres psql -XtAc 'SHOW hba_file' 2>/dev/null || echo /var/lib/pgsql/data/pg_hba.conf)" +sudo cp -p "${hba}.pre-pg-move" "$hba" +sudo systemctl enable --now postgresql +sudo rm /var/lib/felis/postgres-moved +curl -fsSL /deploy/bootstrap.sh | sudo bash +``` + +`SHOW hba_file` needs the server running; with it stopped, the fallback path is EL's +(Debian and Ubuntu keep it in `/etc/postgresql//main/`). Whatever the platform wrote +after the move lives only in felis-postgres; take a bundle there first +(`sudo felis db backup`) and restore it onto the host database afterwards if that matters. +Once the move has run for a while, drop the host copy: +`sudo systemctl start postgresql; sudo -u postgres dropdb felis; sudo -u postgres dropuser felis`, +or remove the server package altogether. + ### The MinecraftServer CRD [VM-VERIFIED] Every rerun applies the CRD embedded in the `felis` binary (`felis bootstrap-assets crd`). diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 4caea97..ee4d8ff 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -596,6 +596,25 @@ jsonpath='{.data.platform}' | base64 -d`), and repeat that for the DBs as often as advisories matter to you. The watchdog warning stays until the timer can reach upstream; that is accurate. +**The platform's own images on an air-gapped node.** The registry pod and the database +pod run images from Docker Hub by digest (`REGISTRY_IMAGE` and `POSTGRES_IMAGE` in +`bootstrap.sh`), which the installer pulls into k3s's containerd and pins there so the +kubelet's image GC never collects them. When the pull fails (`could not pull …`), fetch +the same digest on a machine that can, for the node's architecture, keeping the manifest +as it is, and import it on the node: + +```sh +# elsewhere: the reference as bootstrap.sh names it, e.g. docker.io/library/postgres:18.6-trixie@sha256:… +sudo ctr images pull --platform linux/arm64 "$ref" +sudo ctr images export --platform linux/arm64 image.tar "$ref" +# on the node: +sudo k3s ctr images import image.tar +``` + +A `docker save` of the image rewrites its manifest, so its copy never matches the digest +the Deployment names. Rerun the installer afterwards; it finds the image and pins it. +[CODE-ONLY] + **Kaniko is archived upstream** (June 2025); v1.24.0 is its last release and gets no security fixes. To run a maintained fork, copy it under `mirror/` and point the override at it. The other `[registry]` keys in `felis.toml`: @@ -1353,12 +1372,12 @@ space. ## 13c. The host's address, name or clock changed -An install is bound to the address it was made on. `bootstrap.sh` writes that -address into the database connection string, `pg_hba.conf`, the network -policies, the panel certificate and the default `.nip.io` root domain, and -nothing re-addresses a live install. When the host loses the address (a DHCP -lease that came back different, a moved VM), felis-api cannot reach PostgreSQL -and the panel stops answering on its old name. The watchdog reports it as +An install is bound to the address it was made on. `bootstrap.sh` gives that +address to the k3s node and writes it into the network policies, the panel +certificate and the default `.nip.io` root domain, and nothing re-addresses +a live install. When the host loses the address (a DHCP lease that came back +different, a moved VM), the cluster and the proxy's path to the game servers +still name the old one, and the panel stops answering on its old name. The watchdog reports it as `host-address` (critical, after 5 minutes). The installer warns at install time when the address is a DHCP lease. @@ -1665,19 +1684,18 @@ runs changed: |---|---| | `felis-velocity` (the proxy) | its unit, the JRE, `velocity.jar`, `velocity.toml`, the forwarding secret, the felis-link settings or a plugin jar changed, or it was not running. The fingerprint lives in `/etc/felis/velocity.fingerprint`; delete it to force a restart. | | login and lobby pods | the rebuilt limbo or lobby image has a new image ID (`/etc/felis/system-server-images`). The installer then pins that server's `spec.image` to the digest its tag names now (`felis pin-images --system login\|lobby`) and the operator rolls the pod onto it, each on its own; with the registry unreachable it recreates the pod instead. The installer turns off BuildKit's default provenance attestation (`BUILDX_NO_DEFAULT_ATTESTATIONS=1`): it records the build time, which would give every rebuild a new ID. | -| PostgreSQL | first install only (`listen_addresses` needs a restart). A rerun reloads the configuration, which keeps connections open. | +| felis-postgres | the release moved `POSTGRES_IMAGE` or changed the pod; a few seconds without the API. A rerun that changes neither leaves it running. | | felis-api, felis-operator, the registry pod (its gate and GC containers run the felis binary) | the image tag changed (an upgrade), or a same-version rerun rebuilt it. | -**PostgreSQL across reruns.** On hosts without firewalld the installer loads an -nftables table, `inet felis_postgres`, from `felis-postgres-firewall.service`: -port 5432 accepts loopback, the pod network and the node's own address and drops -everything else (`nft list table inet felis_postgres`). firewalld hosts already -keep 5432 closed to the network. The installer also refuses to start a -PostgreSQL whose major version differs from the cluster in the data directory, -and prints the `pg_upgrade` steps; distributions that move the server package to -a new major (Arch, Fedora) would otherwise leave the database unable to start. -On Arch the installer's `pacman -Syu` holds `postgresql` back once a cluster -exists, so the database is upgraded only when you run `pg_upgrade` yourself. +**PostgreSQL across reruns.** The database runs in k3s from the image the release +pins, so a distribution upgrade never moves it. The installer refuses to start a +PostgreSQL whose major version differs from the cluster in `/var/lib/felis/postgres` +and names the dump-and-restore path (docs/operations.md §4). The first rerun of a +release with felis-postgres on a host an earlier release installed moves the +database off the host PostgreSQL (docs/operations.md §4, "The database's move into +k3s"), and removes that release's `felis-postgres-firewall.service`. On Arch the +installer's `pacman -Syu` keeps holding the host `postgresql` package back while its +cluster exists: that cluster is the copy a rollback of the move starts again. `rollout undo` reverts the image only. The upgrade's database migrations stay applied; when they are the problem, restore the `pre-migrate` bundle the upgrade @@ -1766,8 +1784,11 @@ build no server and no whitelist entry names is pruned after 24 hours, and the The PostgreSQL database behind felis-api holds everything that is not a world: accounts, passkeys, Minecraft account links, server ownership, quotas, audit logs, and the `world_backups` index that maps an archive (§10) back to its -owner. Losing it orphans every world archive. It lives on the host (not in -k3s), so it is backed up on the host too. +owner. Losing it orphans every world archive. It runs in k3s as the +`felis-postgres` Deployment, with its cluster on the host in +`/var/lib/felis/postgres`; `felis db` runs `pg_dump`, `psql` and `pg_restore` +inside that pod (the host config's `[database] deployment`) and keeps the +bundles on the host. ### What runs, and where the bundles go @@ -1829,11 +1850,42 @@ sudo journalctl -u felis-db-backup -n 50 --no-pager # why the last run failed sudo felis db backup # take one now (label manual) ``` -Common failures: PostgreSQL down (`pg_dump: ... connection refused`); the -backup directory's disk full (the half-written `.partial` is removed and the -previous bundles stay intact); `pg_dump: server version mismatch` when an -external database is newer than the host's client tools (install the matching -`postgresql` client package). +Common failures: felis-postgres not running (`kubectl exec` reports no running +pod, or `pg_dump: ... connection refused`; next section); `k3s: executable file not +found` from a `felis` that runs with neither `/usr/local/bin` on PATH nor k3s +anywhere else; the backup directory's disk full (the half-written `.partial` is +removed and the previous bundles stay intact); `pg_dump: server version mismatch` +when an external database (no `deployment` in `[database]`) is newer than the +host's client tools (install the matching `postgresql` client package). + +### felis-postgres is not running, or never became ready + +The installer stops with `felis-postgres did not become ready` after 10 minutes +and prints the pod's events; the watchdog reports the Deployment the same way it +reports felis-api. Look at the pod: + +```sh +sudo k3s kubectl -n felis get pods -l app.kubernetes.io/component=postgres -o wide +sudo k3s kubectl -n felis describe deploy/felis-postgres | tail -n 30 +sudo k3s kubectl -n felis logs deploy/felis-postgres --tail=60 +``` + +| What it says | Cause | Fix | +|---|---|---| +| `ImagePullBackOff` / `ErrImagePull` on `postgres` | the node cannot reach Docker Hub and has no copy | §8e, "The platform's own images on an air-gapped node" | +| `CreateContainerConfigError`, `secret "felis-postgres" not found` | the superuser Secret is gone | rerun the installer, which makes a new one. The image reads it only when it creates a cluster; the installer and `felis db` reach the database over the pod's socket | +| the log says `Permission denied` on `/var/lib/postgresql/18/docker` | the hostPath lost its owner (uid 999) or, with SELinux enforcing, its `container_file_t` label (a restore of `/var/lib/felis` by hand, `restorecon` without the installer's rule) | rerun the installer, which sets both; by hand: `sudo chown -R 999:999 /var/lib/felis/postgres` and `sudo restorecon -R /var/lib/felis/postgres` | +| the log says `database files are incompatible with server` | the cluster was made by another major version | docs/operations.md §4, "PostgreSQL major versions" | +| `Pending`, `Insufficient memory` | the node is full | §13b | + +Query the database by hand from inside the pod, as the superuser over its socket: + +```sh +sudo k3s kubectl -n felis exec -it deploy/felis-postgres -c postgres -- psql -U postgres felis +``` + +The copy an earlier release's host PostgreSQL still holds, from before the move +into k3s, stays on the host; docs/operations.md §4 covers going back to it. ### Check a bundle @@ -1897,7 +1949,7 @@ off-site bucket (next sections) plus a fresh install. What the host holds: | Data | On the host | In the bucket | Brought back by | Lost at most | |---|---|---|---|---| -| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | PostgreSQL | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer | +| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | felis-postgres, `/var/lib/felis/postgres` | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer | | Host state (`/etc/felis`: secrets, both `felis.toml` copies, `offsite.env`, panel TLS pair) | `/etc/felis` | inside every bundle | `tar -x` of the bundle's `state/` | as the database | | MinecraftServer objects | k3s | inside every bundle (`k8s/minecraftservers.json`) | `kubectl apply` | as the database | | World archives (reaper, "Back up now", pre-restore snapshots) | `felis-backups` volume | each one within the hour | `fetch-worlds` | archives written in the last hour | @@ -2222,7 +2274,7 @@ sign-in keeps working. Each lock is audited as `auth.otp.locked` and counted in To lift a lock early once you have confirmed the owner locked themselves out: ```sh -sudo -u postgres psql felis -c \ +sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c \ "DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '');" ``` @@ -2238,7 +2290,7 @@ keys on) and `user_agent`. `actor` is display text: a verified email or the username, never an address the caller set without verifying. ```sh -sudo -u postgres psql felis -c " +sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c " SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason FROM audit_logs WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour' @@ -2306,7 +2358,7 @@ something failed, `otp_skip_detail`. Root can edit the row afterwards, so it records attribution without proving it. ```sh -sudo -u postgres psql felis -c " +sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c " SELECT created_at, action, actor, payload->>'verified' AS verified, payload->>'otp_skipped' AS skipped, payload->>'otp_skip_detail' AS detail FROM audit_logs WHERE source = 'break-glass' ORDER BY created_at DESC LIMIT 20;" @@ -2355,6 +2407,7 @@ for 10 seconds (the Free plan's limits). | `pre-migration backup failed, nothing applied` during an upgrade | §16 | | Undo a mistaken change / restore the control-plane database | §16 | | Host lost: rebuild from a database bundle | §16 | +| `felis-postgres did not become ready`; `kubectl exec` finds no database pod | §16 | | `the off-site copy last completed ... ago` / `NO OFF-SITE COPY` / reaper `awaiting_offsite` stays above 0 | §16, §10 | | Sign-in 429 `rate_limited` for everyone at once | §17 | | 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 |