docs: PostgreSQL 在 k3s 内运行后的升级、回退、卸载与排障
This commit is contained in:
4 files changed
+177
-63
No files matched your search
@@ -33,7 +33,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
|
||||
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||
```
|
||||
|
||||
脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。
|
||||
脚本将自动安装 K3s,在 K3s 内部署 PostgreSQL 与控制平面,并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。旧版本装在宿主上的 PostgreSQL 会在重跑时整库迁进 K3s,宿主上的那份停用保留,供回退(见 [运维手册 §4](docs/operations.md#4-upgrading-the-pieces-around-felis))。
|
||||
|
||||
动手之前,脚本先检查内存、磁盘、端口、网段冲突、已有的 Kubernetes 和外网连通,把所有问题一次列出并停下,主机上什么都没改(检查项见 [运维手册 §1](docs/operations.md#1-supported-hosts))。
|
||||
|
||||
|
||||
+1
-1
@@ -36,7 +36,7 @@ every push):
|
||||
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||
```
|
||||
|
||||
The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain.
|
||||
The script installs K3s, deploys PostgreSQL and the control plane inside it, and launches a setup wizard. Once done, open your browser at the configured domain. A PostgreSQL an earlier release installed on the host is moved into K3s on the next rerun, and the host copy is stopped and kept for a rollback (see [Operations §4](docs/operations.md#4-upgrading-the-pieces-around-felis)).
|
||||
|
||||
Before it changes anything, the script checks RAM, disk, ports, network-range clashes, any Kubernetes already there and outbound access, lists every problem at once and stops with the host untouched (the checks are in [operations §1](docs/operations.md#1-supported-hosts)).
|
||||
|
||||
|
||||
+94
-33
@@ -12,12 +12,13 @@ Evidence tags follow troubleshooting.md: **[VM-VERIFIED]** was run on a real hos
|
||||
## 1. Supported hosts
|
||||
|
||||
`deploy/bootstrap.sh` provisions a single node. It needs systemd, root, and one of the
|
||||
package managers below; everything else (Docker, k3s, PostgreSQL, the JRE, cloudflared)
|
||||
it installs.
|
||||
package managers below; everything else (Docker, k3s, the JRE, cloudflared) it installs.
|
||||
PostgreSQL runs inside k3s as the `felis-postgres` Deployment, from the official image the
|
||||
release pins by digest, with its data on the host in `/var/lib/felis/postgres`.
|
||||
|
||||
| OS family | Package manager | Architectures | Status |
|
||||
|---|---|---|---|
|
||||
| CentOS Stream 9 (firewalld active, PostgreSQL 13) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall |
|
||||
| CentOS Stream 9 (firewalld active, SELinux enforcing) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall, with the database on the host PostgreSQL 13 of the releases before felis-postgres; felis-postgres and the move into it [SH-TESTED] |
|
||||
| Ubuntu 24.04 LTS | apt | x86_64 | **[CI]** fresh install, same-commit rerun, and upgrade from the newest release to the pushed commit |
|
||||
| RHEL / Rocky / Alma 9, Fedora | dnf | x86_64, aarch64 | [CODE-ONLY] same code path as CentOS Stream |
|
||||
| Debian 12, other Ubuntu releases | apt | x86_64, aarch64 | [CODE-ONLY] |
|
||||
@@ -34,7 +35,7 @@ cloudflared is left as it is, see §4):
|
||||
| Temurin JRE (Velocity) | 25, patch build pinned | `FELIS_JRE_VERSION`, sha256 per architecture |
|
||||
| Go (nano builds) | 1.26.8 | `GO_PINNED_VERSION`, sha256 per architecture |
|
||||
| Minecraft / Limbo / Paper / Velocity / LuckPerms | `deploy/game-stack.lock` | §15b |
|
||||
| PostgreSQL | the distribution's package | 13 and 18 are exercised by the `pgint` CI job |
|
||||
| PostgreSQL | 18.6, the official `postgres` image by digest | `POSTGRES_IMAGE` in `bootstrap.sh`, `defaultPostgresImage` in `internal/platform` |
|
||||
|
||||
32-bit hosts are not supported: there is no k3s, JRE or Go build the installer will fetch
|
||||
for them.
|
||||
@@ -62,13 +63,12 @@ host the checks misjudge.
|
||||
|
||||
Two things the host must keep for as long as the install lives:
|
||||
|
||||
- **Its address.** The install is bound to the IPv4 address it was made on (the
|
||||
database connection string, `pg_hba.conf`, the network policies, the panel
|
||||
certificate and the default nip.io domain all carry it). Give the host a static
|
||||
address or a DHCP reservation before installing; the installer warns when the address
|
||||
is a lease, and the watchdog reports `host-address` when the host loses it
|
||||
(troubleshooting §13c). The k3s node name is pinned at install time, so a hostname
|
||||
change is harmless.
|
||||
- **Its address.** The install is bound to the IPv4 address it was made on (the k3s
|
||||
node, the network policies, the panel certificate and the default nip.io domain all
|
||||
carry it). Give the host a static address or a DHCP reservation before installing;
|
||||
the installer warns when the address is a lease, and the watchdog reports
|
||||
`host-address` when the host loses it (troubleshooting §13c). The k3s node name is
|
||||
pinned at install time, so a hostname change is harmless.
|
||||
- **A synchronized clock.** The installer turns NTP on (chrony where nothing else can)
|
||||
and the watchdog reports a clock that stays unsynchronized. Allow outbound UDP 123,
|
||||
or set `FELIS_MANAGE_TIME_SYNC=0` on a host whose clock is kept another way.
|
||||
@@ -134,7 +134,7 @@ server running **[VM-VERIFIED]**:
|
||||
| lobby (Paper, pod limit 1 GiB) | ~0.7–0.85 GB |
|
||||
| login (Limbo, pod limit 512 MiB) | ~0.16 GB |
|
||||
| felis-api, felis-operator, registry gate | ~50 MB each |
|
||||
| PostgreSQL | ~30 MB plus page cache |
|
||||
| PostgreSQL (the felis-postgres pod) | ~30 MB plus page cache |
|
||||
| **Total in use** | **~3.4 GB** |
|
||||
|
||||
Every game server adds the memory its owner gave it: the pod's limit equals its request,
|
||||
@@ -178,6 +178,7 @@ curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo FELIS_VELOCITY_XMX=2G bash
|
||||
| k3s's containerd images | `/var/lib/rancher/k3s/agent/containerd` | 6–9 GB |
|
||||
| Docker's images and build cache | `/var/lib/containerd` (Docker's containerd store) | 5–10 GB after repeated upgrades |
|
||||
| Toolchains and sources | `/opt/felis` | ~2.5 GB |
|
||||
| Database | `/var/lib/felis/postgres` (felis-postgres's cluster) | tens of MB; the audit log is most of it |
|
||||
| Database bundles | `/var/lib/felis/db-backups` | a few MB each, 14 daily kept |
|
||||
|
||||
k3s's local-path volumes do not enforce the requested sizes (§9), so every volume shares
|
||||
@@ -201,23 +202,29 @@ With a private repository, fetch it the way the README fetches `bootstrap.sh`.
|
||||
|
||||
Both modes remove the `felis-*` systemd units and `cloudflared-felis.service`, the
|
||||
Velocity user, `/opt/felis`, `/usr/local/bin/felis`, the installer's cloudflared binary
|
||||
(unless another unit runs it), the `felis_postgres` and `felis_edge` nftables tables and
|
||||
the firewalld ports the installer opened. k3s goes with k3s's own `k3s-uninstall.sh` when
|
||||
the cluster holds nothing but Felis's namespaces; when it runs anything else only
|
||||
`felis`, `minecraft`, `felis-build` and the MinecraftServer CRD are deleted.
|
||||
(unless another unit runs it), the `felis_edge` nftables table (and `felis_postgres`, which
|
||||
releases before the database moved into k3s loaded) and the firewalld ports the installer
|
||||
opened. k3s goes with k3s's own `k3s-uninstall.sh` when the cluster holds nothing but
|
||||
Felis's namespaces; when it runs anything else only `felis`, `minecraft`, `felis-build`
|
||||
and the MinecraftServer CRD are deleted.
|
||||
`--keep-k3s` and `--remove-k3s` override that choice.
|
||||
|
||||
| | keep data (default) | `--purge` |
|
||||
|---|---|---|
|
||||
| Final database bundle | taken first (`felis db backup -label manual`); a failure stops the uninstall before anything is removed. `--no-backup` skips it | none |
|
||||
| `felis` database and role | kept | dropped; `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run |
|
||||
| The database (`/var/lib/felis/postgres`) | kept; felis-postgres is stopped cleanly before k3s goes | deleted with `/var/lib/felis` |
|
||||
| A host PostgreSQL an earlier release ran the database on | kept as it is: stopped after the move into k3s (below, §4), with its old copy of `felis` | its `felis` database and role are dropped (the server is started for that and stopped again), and `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run |
|
||||
| `/etc/felis` (secrets, `felis.toml`, `offsite.env`, tunnel config) | kept; `bootstrap.done` and the per-run records go | deleted, with the tunnel's credentials file |
|
||||
| `/var/lib/felis` (database bundles) | kept | deleted |
|
||||
| `/var/lib/felis` (the database, its bundles) | kept | deleted |
|
||||
| Worlds, archives, registry, uploads | moved to `/var/lib/felis/retained/k3s-storage-<stamp>/` (with `--keep-k3s`: their volumes switch to `Retain` and stay in place) | deleted |
|
||||
| Felis images, Docker build cache | kept | deleted |
|
||||
|
||||
Neither mode removes packages (Docker, PostgreSQL, git, nftables) or the swap file: other
|
||||
software may use them. On a host that should end up bare:
|
||||
The two database rows are [SH-TESTED] (`deploy/uninstall_test.sh`); the VM runs above
|
||||
predate felis-postgres.
|
||||
|
||||
Neither mode removes packages (Docker, git, nftables, and the PostgreSQL server an earlier
|
||||
release installed) or the swap file: other software may use them. On a host that should
|
||||
end up bare:
|
||||
|
||||
```
|
||||
sudo swapoff /swapfile && sudo rm /swapfile && sudo sed -i '\|^/swapfile |d' /etc/fstab
|
||||
@@ -232,8 +239,9 @@ DNS records for the panel hostnames, and the Access application.
|
||||
|
||||
A keep-data uninstall leaves everything a reinstall needs. The installer reuses
|
||||
`/etc/felis/secrets.env`, so the database password and the forwarding and session
|
||||
secrets are unchanged, and it migrates the kept database instead of creating one
|
||||
**[VM-VERIFIED]**.
|
||||
secrets are unchanged, and the installer migrates the kept database instead of creating
|
||||
one **[VM-VERIFIED]** (with the host database of the releases before felis-postgres).
|
||||
felis-postgres starts again on the cluster kept in `/var/lib/felis/postgres` [SH-TESTED].
|
||||
|
||||
Each step below was run on the reference VM after a keep-data uninstall, and the
|
||||
restored worlds matched their kept `level.dat` checksums **[VM-VERIFIED]**. `kept` names
|
||||
@@ -305,7 +313,7 @@ the version they were installed with unless noted:
|
||||
| Temurin JRE | moves to the pinned patch build | rerun |
|
||||
| k3s | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: moves to the pinned release through that tag's install script, one minor version at a time (a bigger jump stops before anything changes and names the release to go through), never backwards |
|
||||
| cloudflared | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: swaps `/usr/local/bin/cloudflared` for the pinned, sha256-checked release and restarts `cloudflared-felis`; a cloudflared the distribution installed stays with its package manager |
|
||||
| PostgreSQL | the distribution's package | the package manager for a minor release; a major version needs `pg_upgrade` first (below) |
|
||||
| PostgreSQL | follows the image the release pins | a minor release comes with a Felis release, and the rerun restarts felis-postgres on it (a few seconds without the API); a major version is a dump and restore (below) |
|
||||
| Docker, git, nftables | distribution packages | the package manager |
|
||||
|
||||
```sh
|
||||
@@ -315,8 +323,9 @@ curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap
|
||||
|
||||
`sudo felis update` reports Felis, Velocity, k3s, cloudflared, the JRE and PostgreSQL
|
||||
against their newest releases; `--k3s`, `--cloudflared`, `--jre` and `--postgres` narrow
|
||||
it to one. PostgreSQL is compared within its major, since a minor release is a package
|
||||
update, and a major past its end of life gets a note naming the current one.
|
||||
it to one. PostgreSQL is read from the felis-postgres container and compared within its
|
||||
major, since a minor release arrives with a Felis release, and a major past its end of life
|
||||
gets a note naming the current one.
|
||||
|
||||
The installer also sets up `felis-update-check.timer`, which runs `felis update --record`
|
||||
once a day around 05:30 (and at boot after a missed run). `--record` stores the result
|
||||
@@ -370,20 +379,72 @@ The env var only carries the password; mail still needs the relay itself, set in
|
||||
|
||||
### PostgreSQL major versions [CODE-ONLY]
|
||||
|
||||
The installer takes the major the distribution ships (13 on EL9) and never moves it. To
|
||||
go to a newer one, stop the writers, keep a dump, then use the distribution's upgrade
|
||||
path:
|
||||
felis-postgres keeps its cluster in `/var/lib/felis/postgres/<major>/docker`. A release that
|
||||
moves the image to a new major finds the old major's cluster there and stops before it
|
||||
changes anything: the new server would start an empty cluster beside it. The way across is
|
||||
a bundle, taken on the release you run now, restored into the new major's empty cluster:
|
||||
|
||||
```sh
|
||||
# On the release you run now:
|
||||
b="$(sudo felis db backup -label pre-upgrade | sed -n 's/^felis db backup: wrote //p')"
|
||||
sudo k3s kubectl -n felis scale deploy/felis-postgres --replicas=0
|
||||
sudo mv /var/lib/felis/postgres/18 /var/lib/felis/postgres-18.old # the old major's cluster, for a way back
|
||||
|
||||
# Install the new release: it starts an empty cluster on the new major and creates the schema.
|
||||
curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo bash
|
||||
|
||||
# Put the data back and bring its schema up to the new release.
|
||||
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=0
|
||||
sudo -u postgres pg_dumpall > /root/felis-pg-$(date +%F).sql
|
||||
# EL9: sudo systemctl stop postgresql; sudo dnf module switch-to postgresql:16
|
||||
# sudo dnf install postgresql-upgrade; sudo postgresql-setup --upgrade
|
||||
# Debian/Ubuntu: sudo pg_upgradecluster <old-major> main
|
||||
sudo systemctl start postgresql
|
||||
sudo felis db restore -yes -no-safety-backup "$b"
|
||||
sudo felis migrate up -config /etc/felis/felis.host.toml
|
||||
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=1
|
||||
```
|
||||
|
||||
Delete `/var/lib/felis/postgres-18.old` once the new release has run for a while. To go back
|
||||
instead, scale felis-postgres to 0, move the new major's directory out of
|
||||
`/var/lib/felis/postgres`, move `postgres-18.old` back as `/var/lib/felis/postgres/18`, and
|
||||
rerun the older release's installer.
|
||||
|
||||
### The database's move into k3s [SH-TESTED]
|
||||
|
||||
Releases before the move ran the database on a PostgreSQL the installer installed on the
|
||||
host. The first rerun of a release with felis-postgres moves it, once:
|
||||
|
||||
1. It stops felis-api, felis-operator and the host timers, and heads the host's
|
||||
`pg_hba.conf` with a block that refuses every connection to `felis` but its own copy
|
||||
(the original is kept beside it as `pg_hba.conf.pre-pg-move`).
|
||||
2. It takes a `pre-pg-move` bundle of the host database (`felis db backup`), restores it
|
||||
into felis-postgres (`felis db restore`) and compares the row count of every table on
|
||||
both servers. Any failure up to here puts `pg_hba.conf` and the control plane back and
|
||||
the platform keeps running on the host database, untouched.
|
||||
3. It stops and disables the host `postgresql` service, which stays installed with its
|
||||
copy of the data, and writes `/var/lib/felis/postgres-moved`. A host server that also
|
||||
holds other databases keeps running; its `felis` copy is then reachable over loopback
|
||||
only.
|
||||
|
||||
From then on the host config points at felis-postgres (`127.0.0.1:15432`, and
|
||||
`deployment = "felis/felis-postgres"`, through which `felis db` runs `pg_dump`, `psql`
|
||||
and `pg_restore` inside the pod) and the pods at `felis-postgres.felis.svc:5432`.
|
||||
|
||||
To go back to the host database, for instance to reinstall the release before the move:
|
||||
|
||||
```sh
|
||||
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator deploy/felis-postgres --replicas=0
|
||||
hba="$(sudo -u postgres psql -XtAc 'SHOW hba_file' 2>/dev/null || echo /var/lib/pgsql/data/pg_hba.conf)"
|
||||
sudo cp -p "${hba}.pre-pg-move" "$hba"
|
||||
sudo systemctl enable --now postgresql
|
||||
sudo rm /var/lib/felis/postgres-moved
|
||||
curl -fsSL <raw-url-of-that-release>/deploy/bootstrap.sh | sudo bash
|
||||
```
|
||||
|
||||
`SHOW hba_file` needs the server running; with it stopped, the fallback path is EL's
|
||||
(Debian and Ubuntu keep it in `/etc/postgresql/<major>/main/`). Whatever the platform wrote
|
||||
after the move lives only in felis-postgres; take a bundle there first
|
||||
(`sudo felis db backup`) and restore it onto the host database afterwards if that matters.
|
||||
Once the move has run for a while, drop the host copy:
|
||||
`sudo systemctl start postgresql; sudo -u postgres dropdb felis; sudo -u postgres dropuser felis`,
|
||||
or remove the server package altogether.
|
||||
|
||||
### The MinecraftServer CRD [VM-VERIFIED]
|
||||
|
||||
Every rerun applies the CRD embedded in the `felis` binary (`felis bootstrap-assets crd`).
|
||||
|
||||
+81
-28
@@ -596,6 +596,25 @@ jsonpath='{.data.platform}' | base64 -d`), and repeat that for the DBs as often
|
||||
as advisories matter to you. The watchdog warning stays until the timer can
|
||||
reach upstream; that is accurate.
|
||||
|
||||
**The platform's own images on an air-gapped node.** The registry pod and the database
|
||||
pod run images from Docker Hub by digest (`REGISTRY_IMAGE` and `POSTGRES_IMAGE` in
|
||||
`bootstrap.sh`), which the installer pulls into k3s's containerd and pins there so the
|
||||
kubelet's image GC never collects them. When the pull fails (`could not pull …`), fetch
|
||||
the same digest on a machine that can, for the node's architecture, keeping the manifest
|
||||
as it is, and import it on the node:
|
||||
|
||||
```sh
|
||||
# elsewhere: the reference as bootstrap.sh names it, e.g. docker.io/library/postgres:18.6-trixie@sha256:…
|
||||
sudo ctr images pull --platform linux/arm64 "$ref"
|
||||
sudo ctr images export --platform linux/arm64 image.tar "$ref"
|
||||
# on the node:
|
||||
sudo k3s ctr images import image.tar
|
||||
```
|
||||
|
||||
A `docker save` of the image rewrites its manifest, so its copy never matches the digest
|
||||
the Deployment names. Rerun the installer afterwards; it finds the image and pins it.
|
||||
[CODE-ONLY]
|
||||
|
||||
**Kaniko is archived upstream** (June 2025); v1.24.0 is its last release and
|
||||
gets no security fixes. To run a maintained fork, copy it under `mirror/` and
|
||||
point the override at it. The other `[registry]` keys in `felis.toml`:
|
||||
@@ -1353,12 +1372,12 @@ space.
|
||||
|
||||
## 13c. The host's address, name or clock changed
|
||||
|
||||
An install is bound to the address it was made on. `bootstrap.sh` writes that
|
||||
address into the database connection string, `pg_hba.conf`, the network
|
||||
policies, the panel certificate and the default `<ip>.nip.io` root domain, and
|
||||
nothing re-addresses a live install. When the host loses the address (a DHCP
|
||||
lease that came back different, a moved VM), felis-api cannot reach PostgreSQL
|
||||
and the panel stops answering on its old name. The watchdog reports it as
|
||||
An install is bound to the address it was made on. `bootstrap.sh` gives that
|
||||
address to the k3s node and writes it into the network policies, the panel
|
||||
certificate and the default `<ip>.nip.io` root domain, and nothing re-addresses
|
||||
a live install. When the host loses the address (a DHCP lease that came back
|
||||
different, a moved VM), the cluster and the proxy's path to the game servers
|
||||
still name the old one, and the panel stops answering on its old name. The watchdog reports it as
|
||||
`host-address` (critical, after 5 minutes). The installer warns at install time
|
||||
when the address is a DHCP lease.
|
||||
|
||||
@@ -1665,19 +1684,18 @@ runs changed:
|
||||
|---|---|
|
||||
| `felis-velocity` (the proxy) | its unit, the JRE, `velocity.jar`, `velocity.toml`, the forwarding secret, the felis-link settings or a plugin jar changed, or it was not running. The fingerprint lives in `/etc/felis/velocity.fingerprint`; delete it to force a restart. |
|
||||
| login and lobby pods | the rebuilt limbo or lobby image has a new image ID (`/etc/felis/system-server-images`). The installer then pins that server's `spec.image` to the digest its tag names now (`felis pin-images --system login\|lobby`) and the operator rolls the pod onto it, each on its own; with the registry unreachable it recreates the pod instead. The installer turns off BuildKit's default provenance attestation (`BUILDX_NO_DEFAULT_ATTESTATIONS=1`): it records the build time, which would give every rebuild a new ID. |
|
||||
| PostgreSQL | first install only (`listen_addresses` needs a restart). A rerun reloads the configuration, which keeps connections open. |
|
||||
| felis-postgres | the release moved `POSTGRES_IMAGE` or changed the pod; a few seconds without the API. A rerun that changes neither leaves it running. |
|
||||
| felis-api, felis-operator, the registry pod (its gate and GC containers run the felis binary) | the image tag changed (an upgrade), or a same-version rerun rebuilt it. |
|
||||
|
||||
**PostgreSQL across reruns.** On hosts without firewalld the installer loads an
|
||||
nftables table, `inet felis_postgres`, from `felis-postgres-firewall.service`:
|
||||
port 5432 accepts loopback, the pod network and the node's own address and drops
|
||||
everything else (`nft list table inet felis_postgres`). firewalld hosts already
|
||||
keep 5432 closed to the network. The installer also refuses to start a
|
||||
PostgreSQL whose major version differs from the cluster in the data directory,
|
||||
and prints the `pg_upgrade` steps; distributions that move the server package to
|
||||
a new major (Arch, Fedora) would otherwise leave the database unable to start.
|
||||
On Arch the installer's `pacman -Syu` holds `postgresql` back once a cluster
|
||||
exists, so the database is upgraded only when you run `pg_upgrade` yourself.
|
||||
**PostgreSQL across reruns.** The database runs in k3s from the image the release
|
||||
pins, so a distribution upgrade never moves it. The installer refuses to start a
|
||||
PostgreSQL whose major version differs from the cluster in `/var/lib/felis/postgres`
|
||||
and names the dump-and-restore path (docs/operations.md §4). The first rerun of a
|
||||
release with felis-postgres on a host an earlier release installed moves the
|
||||
database off the host PostgreSQL (docs/operations.md §4, "The database's move into
|
||||
k3s"), and removes that release's `felis-postgres-firewall.service`. On Arch the
|
||||
installer's `pacman -Syu` keeps holding the host `postgresql` package back while its
|
||||
cluster exists: that cluster is the copy a rollback of the move starts again.
|
||||
|
||||
`rollout undo` reverts the image only. The upgrade's database migrations stay
|
||||
applied; when they are the problem, restore the `pre-migrate` bundle the upgrade
|
||||
@@ -1766,8 +1784,11 @@ build no server and no whitelist entry names is pruned after 24 hours, and the
|
||||
The PostgreSQL database behind felis-api holds everything that is not a world:
|
||||
accounts, passkeys, Minecraft account links, server ownership, quotas, audit
|
||||
logs, and the `world_backups` index that maps an archive (§10) back to its
|
||||
owner. Losing it orphans every world archive. It lives on the host (not in
|
||||
k3s), so it is backed up on the host too.
|
||||
owner. Losing it orphans every world archive. It runs in k3s as the
|
||||
`felis-postgres` Deployment, with its cluster on the host in
|
||||
`/var/lib/felis/postgres`; `felis db` runs `pg_dump`, `psql` and `pg_restore`
|
||||
inside that pod (the host config's `[database] deployment`) and keeps the
|
||||
bundles on the host.
|
||||
|
||||
### What runs, and where the bundles go
|
||||
|
||||
@@ -1829,11 +1850,42 @@ sudo journalctl -u felis-db-backup -n 50 --no-pager # why the last run failed
|
||||
sudo felis db backup # take one now (label manual)
|
||||
```
|
||||
|
||||
Common failures: PostgreSQL down (`pg_dump: ... connection refused`); the
|
||||
backup directory's disk full (the half-written `.partial` is removed and the
|
||||
previous bundles stay intact); `pg_dump: server version mismatch` when an
|
||||
external database is newer than the host's client tools (install the matching
|
||||
`postgresql` client package).
|
||||
Common failures: felis-postgres not running (`kubectl exec` reports no running
|
||||
pod, or `pg_dump: ... connection refused`; next section); `k3s: executable file not
|
||||
found` from a `felis` that runs with neither `/usr/local/bin` on PATH nor k3s
|
||||
anywhere else; the backup directory's disk full (the half-written `.partial` is
|
||||
removed and the previous bundles stay intact); `pg_dump: server version mismatch`
|
||||
when an external database (no `deployment` in `[database]`) is newer than the
|
||||
host's client tools (install the matching `postgresql` client package).
|
||||
|
||||
### felis-postgres is not running, or never became ready
|
||||
|
||||
The installer stops with `felis-postgres did not become ready` after 10 minutes
|
||||
and prints the pod's events; the watchdog reports the Deployment the same way it
|
||||
reports felis-api. Look at the pod:
|
||||
|
||||
```sh
|
||||
sudo k3s kubectl -n felis get pods -l app.kubernetes.io/component=postgres -o wide
|
||||
sudo k3s kubectl -n felis describe deploy/felis-postgres | tail -n 30
|
||||
sudo k3s kubectl -n felis logs deploy/felis-postgres --tail=60
|
||||
```
|
||||
|
||||
| What it says | Cause | Fix |
|
||||
|---|---|---|
|
||||
| `ImagePullBackOff` / `ErrImagePull` on `postgres` | the node cannot reach Docker Hub and has no copy | §8e, "The platform's own images on an air-gapped node" |
|
||||
| `CreateContainerConfigError`, `secret "felis-postgres" not found` | the superuser Secret is gone | rerun the installer, which makes a new one. The image reads it only when it creates a cluster; the installer and `felis db` reach the database over the pod's socket |
|
||||
| the log says `Permission denied` on `/var/lib/postgresql/18/docker` | the hostPath lost its owner (uid 999) or, with SELinux enforcing, its `container_file_t` label (a restore of `/var/lib/felis` by hand, `restorecon` without the installer's rule) | rerun the installer, which sets both; by hand: `sudo chown -R 999:999 /var/lib/felis/postgres` and `sudo restorecon -R /var/lib/felis/postgres` |
|
||||
| the log says `database files are incompatible with server` | the cluster was made by another major version | docs/operations.md §4, "PostgreSQL major versions" |
|
||||
| `Pending`, `Insufficient memory` | the node is full | §13b |
|
||||
|
||||
Query the database by hand from inside the pod, as the superuser over its socket:
|
||||
|
||||
```sh
|
||||
sudo k3s kubectl -n felis exec -it deploy/felis-postgres -c postgres -- psql -U postgres felis
|
||||
```
|
||||
|
||||
The copy an earlier release's host PostgreSQL still holds, from before the move
|
||||
into k3s, stays on the host; docs/operations.md §4 covers going back to it.
|
||||
|
||||
### Check a bundle
|
||||
|
||||
@@ -1897,7 +1949,7 @@ off-site bucket (next sections) plus a fresh install. What the host holds:
|
||||
|
||||
| Data | On the host | In the bucket | Brought back by | Lost at most |
|
||||
|---|---|---|---|---|
|
||||
| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | PostgreSQL | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer |
|
||||
| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | felis-postgres, `/var/lib/felis/postgres` | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer |
|
||||
| Host state (`/etc/felis`: secrets, both `felis.toml` copies, `offsite.env`, panel TLS pair) | `/etc/felis` | inside every bundle | `tar -x` of the bundle's `state/` | as the database |
|
||||
| MinecraftServer objects | k3s | inside every bundle (`k8s/minecraftservers.json`) | `kubectl apply` | as the database |
|
||||
| World archives (reaper, "Back up now", pre-restore snapshots) | `felis-backups` volume | each one within the hour | `fetch-worlds` | archives written in the last hour |
|
||||
@@ -2222,7 +2274,7 @@ sign-in keeps working. Each lock is audited as `auth.otp.locked` and counted in
|
||||
To lift a lock early once you have confirmed the owner locked themselves out:
|
||||
|
||||
```sh
|
||||
sudo -u postgres psql felis -c \
|
||||
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c \
|
||||
"DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '<name>');"
|
||||
```
|
||||
|
||||
@@ -2238,7 +2290,7 @@ keys on) and `user_agent`. `actor` is display text: a verified email or the
|
||||
username, never an address the caller set without verifying.
|
||||
|
||||
```sh
|
||||
sudo -u postgres psql felis -c "
|
||||
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c "
|
||||
SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason
|
||||
FROM audit_logs
|
||||
WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour'
|
||||
@@ -2306,7 +2358,7 @@ something failed, `otp_skip_detail`. Root can edit the row afterwards, so it
|
||||
records attribution without proving it.
|
||||
|
||||
```sh
|
||||
sudo -u postgres psql felis -c "
|
||||
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c "
|
||||
SELECT created_at, action, actor, payload->>'verified' AS verified,
|
||||
payload->>'otp_skipped' AS skipped, payload->>'otp_skip_detail' AS detail
|
||||
FROM audit_logs WHERE source = 'break-glass' ORDER BY created_at DESC LIMIT 20;"
|
||||
@@ -2355,6 +2407,7 @@ for 10 seconds (the Free plan's limits).
|
||||
| `pre-migration backup failed, nothing applied` during an upgrade | §16 |
|
||||
| Undo a mistaken change / restore the control-plane database | §16 |
|
||||
| Host lost: rebuild from a database bundle | §16 |
|
||||
| `felis-postgres did not become ready`; `kubectl exec` finds no database pod | §16 |
|
||||
| `the off-site copy last completed ... ago` / `NO OFF-SITE COPY` / reaper `awaiting_offsite` stays above 0 | §16, §10 |
|
||||
| Sign-in 429 `rate_limited` for everyone at once | §17 |
|
||||
| 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 |
|
||||
|
||||
Reference in new issue
Block a user