docs: PostgreSQL 在 k3s 内运行后的升级、回退、卸载与排障
This commit is contained in:
4 files changed
+177
-63
No files matched your search
@@ -33,7 +33,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
|
|||||||
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||||
```
|
```
|
||||||
|
|
||||||
脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。
|
脚本将自动安装 K3s,在 K3s 内部署 PostgreSQL 与控制平面,并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。旧版本装在宿主上的 PostgreSQL 会在重跑时整库迁进 K3s,宿主上的那份停用保留,供回退(见 [运维手册 §4](docs/operations.md#4-upgrading-the-pieces-around-felis))。
|
||||||
|
|
||||||
动手之前,脚本先检查内存、磁盘、端口、网段冲突、已有的 Kubernetes 和外网连通,把所有问题一次列出并停下,主机上什么都没改(检查项见 [运维手册 §1](docs/operations.md#1-supported-hosts))。
|
动手之前,脚本先检查内存、磁盘、端口、网段冲突、已有的 Kubernetes 和外网连通,把所有问题一次列出并停下,主机上什么都没改(检查项见 [运维手册 §1](docs/operations.md#1-supported-hosts))。
|
||||||
|
|
||||||
|
|||||||
+1
-1
@@ -36,7 +36,7 @@ every push):
|
|||||||
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||||
```
|
```
|
||||||
|
|
||||||
The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain.
|
The script installs K3s, deploys PostgreSQL and the control plane inside it, and launches a setup wizard. Once done, open your browser at the configured domain. A PostgreSQL an earlier release installed on the host is moved into K3s on the next rerun, and the host copy is stopped and kept for a rollback (see [Operations §4](docs/operations.md#4-upgrading-the-pieces-around-felis)).
|
||||||
|
|
||||||
Before it changes anything, the script checks RAM, disk, ports, network-range clashes, any Kubernetes already there and outbound access, lists every problem at once and stops with the host untouched (the checks are in [operations §1](docs/operations.md#1-supported-hosts)).
|
Before it changes anything, the script checks RAM, disk, ports, network-range clashes, any Kubernetes already there and outbound access, lists every problem at once and stops with the host untouched (the checks are in [operations §1](docs/operations.md#1-supported-hosts)).
|
||||||
|
|
||||||
|
|||||||
+94
-33
@@ -12,12 +12,13 @@ Evidence tags follow troubleshooting.md: **[VM-VERIFIED]** was run on a real hos
|
|||||||
## 1. Supported hosts
|
## 1. Supported hosts
|
||||||
|
|
||||||
`deploy/bootstrap.sh` provisions a single node. It needs systemd, root, and one of the
|
`deploy/bootstrap.sh` provisions a single node. It needs systemd, root, and one of the
|
||||||
package managers below; everything else (Docker, k3s, PostgreSQL, the JRE, cloudflared)
|
package managers below; everything else (Docker, k3s, the JRE, cloudflared) it installs.
|
||||||
it installs.
|
PostgreSQL runs inside k3s as the `felis-postgres` Deployment, from the official image the
|
||||||
|
release pins by digest, with its data on the host in `/var/lib/felis/postgres`.
|
||||||
|
|
||||||
| OS family | Package manager | Architectures | Status |
|
| OS family | Package manager | Architectures | Status |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| CentOS Stream 9 (firewalld active, PostgreSQL 13) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall |
|
| CentOS Stream 9 (firewalld active, SELinux enforcing) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall, with the database on the host PostgreSQL 13 of the releases before felis-postgres; felis-postgres and the move into it [SH-TESTED] |
|
||||||
| Ubuntu 24.04 LTS | apt | x86_64 | **[CI]** fresh install, same-commit rerun, and upgrade from the newest release to the pushed commit |
|
| Ubuntu 24.04 LTS | apt | x86_64 | **[CI]** fresh install, same-commit rerun, and upgrade from the newest release to the pushed commit |
|
||||||
| RHEL / Rocky / Alma 9, Fedora | dnf | x86_64, aarch64 | [CODE-ONLY] same code path as CentOS Stream |
|
| RHEL / Rocky / Alma 9, Fedora | dnf | x86_64, aarch64 | [CODE-ONLY] same code path as CentOS Stream |
|
||||||
| Debian 12, other Ubuntu releases | apt | x86_64, aarch64 | [CODE-ONLY] |
|
| Debian 12, other Ubuntu releases | apt | x86_64, aarch64 | [CODE-ONLY] |
|
||||||
@@ -34,7 +35,7 @@ cloudflared is left as it is, see §4):
|
|||||||
| Temurin JRE (Velocity) | 25, patch build pinned | `FELIS_JRE_VERSION`, sha256 per architecture |
|
| Temurin JRE (Velocity) | 25, patch build pinned | `FELIS_JRE_VERSION`, sha256 per architecture |
|
||||||
| Go (nano builds) | 1.26.8 | `GO_PINNED_VERSION`, sha256 per architecture |
|
| Go (nano builds) | 1.26.8 | `GO_PINNED_VERSION`, sha256 per architecture |
|
||||||
| Minecraft / Limbo / Paper / Velocity / LuckPerms | `deploy/game-stack.lock` | §15b |
|
| Minecraft / Limbo / Paper / Velocity / LuckPerms | `deploy/game-stack.lock` | §15b |
|
||||||
| PostgreSQL | the distribution's package | 13 and 18 are exercised by the `pgint` CI job |
|
| PostgreSQL | 18.6, the official `postgres` image by digest | `POSTGRES_IMAGE` in `bootstrap.sh`, `defaultPostgresImage` in `internal/platform` |
|
||||||
|
|
||||||
32-bit hosts are not supported: there is no k3s, JRE or Go build the installer will fetch
|
32-bit hosts are not supported: there is no k3s, JRE or Go build the installer will fetch
|
||||||
for them.
|
for them.
|
||||||
@@ -62,13 +63,12 @@ host the checks misjudge.
|
|||||||
|
|
||||||
Two things the host must keep for as long as the install lives:
|
Two things the host must keep for as long as the install lives:
|
||||||
|
|
||||||
- **Its address.** The install is bound to the IPv4 address it was made on (the
|
- **Its address.** The install is bound to the IPv4 address it was made on (the k3s
|
||||||
database connection string, `pg_hba.conf`, the network policies, the panel
|
node, the network policies, the panel certificate and the default nip.io domain all
|
||||||
certificate and the default nip.io domain all carry it). Give the host a static
|
carry it). Give the host a static address or a DHCP reservation before installing;
|
||||||
address or a DHCP reservation before installing; the installer warns when the address
|
the installer warns when the address is a lease, and the watchdog reports
|
||||||
is a lease, and the watchdog reports `host-address` when the host loses it
|
`host-address` when the host loses it (troubleshooting §13c). The k3s node name is
|
||||||
(troubleshooting §13c). The k3s node name is pinned at install time, so a hostname
|
pinned at install time, so a hostname change is harmless.
|
||||||
change is harmless.
|
|
||||||
- **A synchronized clock.** The installer turns NTP on (chrony where nothing else can)
|
- **A synchronized clock.** The installer turns NTP on (chrony where nothing else can)
|
||||||
and the watchdog reports a clock that stays unsynchronized. Allow outbound UDP 123,
|
and the watchdog reports a clock that stays unsynchronized. Allow outbound UDP 123,
|
||||||
or set `FELIS_MANAGE_TIME_SYNC=0` on a host whose clock is kept another way.
|
or set `FELIS_MANAGE_TIME_SYNC=0` on a host whose clock is kept another way.
|
||||||
@@ -134,7 +134,7 @@ server running **[VM-VERIFIED]**:
|
|||||||
| lobby (Paper, pod limit 1 GiB) | ~0.7–0.85 GB |
|
| lobby (Paper, pod limit 1 GiB) | ~0.7–0.85 GB |
|
||||||
| login (Limbo, pod limit 512 MiB) | ~0.16 GB |
|
| login (Limbo, pod limit 512 MiB) | ~0.16 GB |
|
||||||
| felis-api, felis-operator, registry gate | ~50 MB each |
|
| felis-api, felis-operator, registry gate | ~50 MB each |
|
||||||
| PostgreSQL | ~30 MB plus page cache |
|
| PostgreSQL (the felis-postgres pod) | ~30 MB plus page cache |
|
||||||
| **Total in use** | **~3.4 GB** |
|
| **Total in use** | **~3.4 GB** |
|
||||||
|
|
||||||
Every game server adds the memory its owner gave it: the pod's limit equals its request,
|
Every game server adds the memory its owner gave it: the pod's limit equals its request,
|
||||||
@@ -178,6 +178,7 @@ curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo FELIS_VELOCITY_XMX=2G bash
|
|||||||
| k3s's containerd images | `/var/lib/rancher/k3s/agent/containerd` | 6–9 GB |
|
| k3s's containerd images | `/var/lib/rancher/k3s/agent/containerd` | 6–9 GB |
|
||||||
| Docker's images and build cache | `/var/lib/containerd` (Docker's containerd store) | 5–10 GB after repeated upgrades |
|
| Docker's images and build cache | `/var/lib/containerd` (Docker's containerd store) | 5–10 GB after repeated upgrades |
|
||||||
| Toolchains and sources | `/opt/felis` | ~2.5 GB |
|
| Toolchains and sources | `/opt/felis` | ~2.5 GB |
|
||||||
|
| Database | `/var/lib/felis/postgres` (felis-postgres's cluster) | tens of MB; the audit log is most of it |
|
||||||
| Database bundles | `/var/lib/felis/db-backups` | a few MB each, 14 daily kept |
|
| Database bundles | `/var/lib/felis/db-backups` | a few MB each, 14 daily kept |
|
||||||
|
|
||||||
k3s's local-path volumes do not enforce the requested sizes (§9), so every volume shares
|
k3s's local-path volumes do not enforce the requested sizes (§9), so every volume shares
|
||||||
@@ -201,23 +202,29 @@ With a private repository, fetch it the way the README fetches `bootstrap.sh`.
|
|||||||
|
|
||||||
Both modes remove the `felis-*` systemd units and `cloudflared-felis.service`, the
|
Both modes remove the `felis-*` systemd units and `cloudflared-felis.service`, the
|
||||||
Velocity user, `/opt/felis`, `/usr/local/bin/felis`, the installer's cloudflared binary
|
Velocity user, `/opt/felis`, `/usr/local/bin/felis`, the installer's cloudflared binary
|
||||||
(unless another unit runs it), the `felis_postgres` and `felis_edge` nftables tables and
|
(unless another unit runs it), the `felis_edge` nftables table (and `felis_postgres`, which
|
||||||
the firewalld ports the installer opened. k3s goes with k3s's own `k3s-uninstall.sh` when
|
releases before the database moved into k3s loaded) and the firewalld ports the installer
|
||||||
the cluster holds nothing but Felis's namespaces; when it runs anything else only
|
opened. k3s goes with k3s's own `k3s-uninstall.sh` when the cluster holds nothing but
|
||||||
`felis`, `minecraft`, `felis-build` and the MinecraftServer CRD are deleted.
|
Felis's namespaces; when it runs anything else only `felis`, `minecraft`, `felis-build`
|
||||||
|
and the MinecraftServer CRD are deleted.
|
||||||
`--keep-k3s` and `--remove-k3s` override that choice.
|
`--keep-k3s` and `--remove-k3s` override that choice.
|
||||||
|
|
||||||
| | keep data (default) | `--purge` |
|
| | keep data (default) | `--purge` |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| Final database bundle | taken first (`felis db backup -label manual`); a failure stops the uninstall before anything is removed. `--no-backup` skips it | none |
|
| Final database bundle | taken first (`felis db backup -label manual`); a failure stops the uninstall before anything is removed. `--no-backup` skips it | none |
|
||||||
| `felis` database and role | kept | dropped; `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run |
|
| The database (`/var/lib/felis/postgres`) | kept; felis-postgres is stopped cleanly before k3s goes | deleted with `/var/lib/felis` |
|
||||||
|
| A host PostgreSQL an earlier release ran the database on | kept as it is: stopped after the move into k3s (below, §4), with its old copy of `felis` | its `felis` database and role are dropped (the server is started for that and stopped again), and `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run |
|
||||||
| `/etc/felis` (secrets, `felis.toml`, `offsite.env`, tunnel config) | kept; `bootstrap.done` and the per-run records go | deleted, with the tunnel's credentials file |
|
| `/etc/felis` (secrets, `felis.toml`, `offsite.env`, tunnel config) | kept; `bootstrap.done` and the per-run records go | deleted, with the tunnel's credentials file |
|
||||||
| `/var/lib/felis` (database bundles) | kept | deleted |
|
| `/var/lib/felis` (the database, its bundles) | kept | deleted |
|
||||||
| Worlds, archives, registry, uploads | moved to `/var/lib/felis/retained/k3s-storage-<stamp>/` (with `--keep-k3s`: their volumes switch to `Retain` and stay in place) | deleted |
|
| Worlds, archives, registry, uploads | moved to `/var/lib/felis/retained/k3s-storage-<stamp>/` (with `--keep-k3s`: their volumes switch to `Retain` and stay in place) | deleted |
|
||||||
| Felis images, Docker build cache | kept | deleted |
|
| Felis images, Docker build cache | kept | deleted |
|
||||||
|
|
||||||
Neither mode removes packages (Docker, PostgreSQL, git, nftables) or the swap file: other
|
The two database rows are [SH-TESTED] (`deploy/uninstall_test.sh`); the VM runs above
|
||||||
software may use them. On a host that should end up bare:
|
predate felis-postgres.
|
||||||
|
|
||||||
|
Neither mode removes packages (Docker, git, nftables, and the PostgreSQL server an earlier
|
||||||
|
release installed) or the swap file: other software may use them. On a host that should
|
||||||
|
end up bare:
|
||||||
|
|
||||||
```
|
```
|
||||||
sudo swapoff /swapfile && sudo rm /swapfile && sudo sed -i '\|^/swapfile |d' /etc/fstab
|
sudo swapoff /swapfile && sudo rm /swapfile && sudo sed -i '\|^/swapfile |d' /etc/fstab
|
||||||
@@ -232,8 +239,9 @@ DNS records for the panel hostnames, and the Access application.
|
|||||||
|
|
||||||
A keep-data uninstall leaves everything a reinstall needs. The installer reuses
|
A keep-data uninstall leaves everything a reinstall needs. The installer reuses
|
||||||
`/etc/felis/secrets.env`, so the database password and the forwarding and session
|
`/etc/felis/secrets.env`, so the database password and the forwarding and session
|
||||||
secrets are unchanged, and it migrates the kept database instead of creating one
|
secrets are unchanged, and the installer migrates the kept database instead of creating
|
||||||
**[VM-VERIFIED]**.
|
one **[VM-VERIFIED]** (with the host database of the releases before felis-postgres).
|
||||||
|
felis-postgres starts again on the cluster kept in `/var/lib/felis/postgres` [SH-TESTED].
|
||||||
|
|
||||||
Each step below was run on the reference VM after a keep-data uninstall, and the
|
Each step below was run on the reference VM after a keep-data uninstall, and the
|
||||||
restored worlds matched their kept `level.dat` checksums **[VM-VERIFIED]**. `kept` names
|
restored worlds matched their kept `level.dat` checksums **[VM-VERIFIED]**. `kept` names
|
||||||
@@ -305,7 +313,7 @@ the version they were installed with unless noted:
|
|||||||
| Temurin JRE | moves to the pinned patch build | rerun |
|
| Temurin JRE | moves to the pinned patch build | rerun |
|
||||||
| k3s | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: moves to the pinned release through that tag's install script, one minor version at a time (a bigger jump stops before anything changes and names the release to go through), never backwards |
|
| k3s | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: moves to the pinned release through that tag's install script, one minor version at a time (a bigger jump stops before anything changes and names the release to go through), never backwards |
|
||||||
| cloudflared | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: swaps `/usr/local/bin/cloudflared` for the pinned, sha256-checked release and restarts `cloudflared-felis`; a cloudflared the distribution installed stays with its package manager |
|
| cloudflared | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: swaps `/usr/local/bin/cloudflared` for the pinned, sha256-checked release and restarts `cloudflared-felis`; a cloudflared the distribution installed stays with its package manager |
|
||||||
| PostgreSQL | the distribution's package | the package manager for a minor release; a major version needs `pg_upgrade` first (below) |
|
| PostgreSQL | follows the image the release pins | a minor release comes with a Felis release, and the rerun restarts felis-postgres on it (a few seconds without the API); a major version is a dump and restore (below) |
|
||||||
| Docker, git, nftables | distribution packages | the package manager |
|
| Docker, git, nftables | distribution packages | the package manager |
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
@@ -315,8 +323,9 @@ curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap
|
|||||||
|
|
||||||
`sudo felis update` reports Felis, Velocity, k3s, cloudflared, the JRE and PostgreSQL
|
`sudo felis update` reports Felis, Velocity, k3s, cloudflared, the JRE and PostgreSQL
|
||||||
against their newest releases; `--k3s`, `--cloudflared`, `--jre` and `--postgres` narrow
|
against their newest releases; `--k3s`, `--cloudflared`, `--jre` and `--postgres` narrow
|
||||||
it to one. PostgreSQL is compared within its major, since a minor release is a package
|
it to one. PostgreSQL is read from the felis-postgres container and compared within its
|
||||||
update, and a major past its end of life gets a note naming the current one.
|
major, since a minor release arrives with a Felis release, and a major past its end of life
|
||||||
|
gets a note naming the current one.
|
||||||
|
|
||||||
The installer also sets up `felis-update-check.timer`, which runs `felis update --record`
|
The installer also sets up `felis-update-check.timer`, which runs `felis update --record`
|
||||||
once a day around 05:30 (and at boot after a missed run). `--record` stores the result
|
once a day around 05:30 (and at boot after a missed run). `--record` stores the result
|
||||||
@@ -370,20 +379,72 @@ The env var only carries the password; mail still needs the relay itself, set in
|
|||||||
|
|
||||||
### PostgreSQL major versions [CODE-ONLY]
|
### PostgreSQL major versions [CODE-ONLY]
|
||||||
|
|
||||||
The installer takes the major the distribution ships (13 on EL9) and never moves it. To
|
felis-postgres keeps its cluster in `/var/lib/felis/postgres/<major>/docker`. A release that
|
||||||
go to a newer one, stop the writers, keep a dump, then use the distribution's upgrade
|
moves the image to a new major finds the old major's cluster there and stops before it
|
||||||
path:
|
changes anything: the new server would start an empty cluster beside it. The way across is
|
||||||
|
a bundle, taken on the release you run now, restored into the new major's empty cluster:
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
|
# On the release you run now:
|
||||||
|
b="$(sudo felis db backup -label pre-upgrade | sed -n 's/^felis db backup: wrote //p')"
|
||||||
|
sudo k3s kubectl -n felis scale deploy/felis-postgres --replicas=0
|
||||||
|
sudo mv /var/lib/felis/postgres/18 /var/lib/felis/postgres-18.old # the old major's cluster, for a way back
|
||||||
|
|
||||||
|
# Install the new release: it starts an empty cluster on the new major and creates the schema.
|
||||||
|
curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo bash
|
||||||
|
|
||||||
|
# Put the data back and bring its schema up to the new release.
|
||||||
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=0
|
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=0
|
||||||
sudo -u postgres pg_dumpall > /root/felis-pg-$(date +%F).sql
|
sudo felis db restore -yes -no-safety-backup "$b"
|
||||||
# EL9: sudo systemctl stop postgresql; sudo dnf module switch-to postgresql:16
|
sudo felis migrate up -config /etc/felis/felis.host.toml
|
||||||
# sudo dnf install postgresql-upgrade; sudo postgresql-setup --upgrade
|
|
||||||
# Debian/Ubuntu: sudo pg_upgradecluster <old-major> main
|
|
||||||
sudo systemctl start postgresql
|
|
||||||
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=1
|
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=1
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Delete `/var/lib/felis/postgres-18.old` once the new release has run for a while. To go back
|
||||||
|
instead, scale felis-postgres to 0, move the new major's directory out of
|
||||||
|
`/var/lib/felis/postgres`, move `postgres-18.old` back as `/var/lib/felis/postgres/18`, and
|
||||||
|
rerun the older release's installer.
|
||||||
|
|
||||||
|
### The database's move into k3s [SH-TESTED]
|
||||||
|
|
||||||
|
Releases before the move ran the database on a PostgreSQL the installer installed on the
|
||||||
|
host. The first rerun of a release with felis-postgres moves it, once:
|
||||||
|
|
||||||
|
1. It stops felis-api, felis-operator and the host timers, and heads the host's
|
||||||
|
`pg_hba.conf` with a block that refuses every connection to `felis` but its own copy
|
||||||
|
(the original is kept beside it as `pg_hba.conf.pre-pg-move`).
|
||||||
|
2. It takes a `pre-pg-move` bundle of the host database (`felis db backup`), restores it
|
||||||
|
into felis-postgres (`felis db restore`) and compares the row count of every table on
|
||||||
|
both servers. Any failure up to here puts `pg_hba.conf` and the control plane back and
|
||||||
|
the platform keeps running on the host database, untouched.
|
||||||
|
3. It stops and disables the host `postgresql` service, which stays installed with its
|
||||||
|
copy of the data, and writes `/var/lib/felis/postgres-moved`. A host server that also
|
||||||
|
holds other databases keeps running; its `felis` copy is then reachable over loopback
|
||||||
|
only.
|
||||||
|
|
||||||
|
From then on the host config points at felis-postgres (`127.0.0.1:15432`, and
|
||||||
|
`deployment = "felis/felis-postgres"`, through which `felis db` runs `pg_dump`, `psql`
|
||||||
|
and `pg_restore` inside the pod) and the pods at `felis-postgres.felis.svc:5432`.
|
||||||
|
|
||||||
|
To go back to the host database, for instance to reinstall the release before the move:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator deploy/felis-postgres --replicas=0
|
||||||
|
hba="$(sudo -u postgres psql -XtAc 'SHOW hba_file' 2>/dev/null || echo /var/lib/pgsql/data/pg_hba.conf)"
|
||||||
|
sudo cp -p "${hba}.pre-pg-move" "$hba"
|
||||||
|
sudo systemctl enable --now postgresql
|
||||||
|
sudo rm /var/lib/felis/postgres-moved
|
||||||
|
curl -fsSL <raw-url-of-that-release>/deploy/bootstrap.sh | sudo bash
|
||||||
|
```
|
||||||
|
|
||||||
|
`SHOW hba_file` needs the server running; with it stopped, the fallback path is EL's
|
||||||
|
(Debian and Ubuntu keep it in `/etc/postgresql/<major>/main/`). Whatever the platform wrote
|
||||||
|
after the move lives only in felis-postgres; take a bundle there first
|
||||||
|
(`sudo felis db backup`) and restore it onto the host database afterwards if that matters.
|
||||||
|
Once the move has run for a while, drop the host copy:
|
||||||
|
`sudo systemctl start postgresql; sudo -u postgres dropdb felis; sudo -u postgres dropuser felis`,
|
||||||
|
or remove the server package altogether.
|
||||||
|
|
||||||
### The MinecraftServer CRD [VM-VERIFIED]
|
### The MinecraftServer CRD [VM-VERIFIED]
|
||||||
|
|
||||||
Every rerun applies the CRD embedded in the `felis` binary (`felis bootstrap-assets crd`).
|
Every rerun applies the CRD embedded in the `felis` binary (`felis bootstrap-assets crd`).
|
||||||
|
|||||||
+81
-28
@@ -596,6 +596,25 @@ jsonpath='{.data.platform}' | base64 -d`), and repeat that for the DBs as often
|
|||||||
as advisories matter to you. The watchdog warning stays until the timer can
|
as advisories matter to you. The watchdog warning stays until the timer can
|
||||||
reach upstream; that is accurate.
|
reach upstream; that is accurate.
|
||||||
|
|
||||||
|
**The platform's own images on an air-gapped node.** The registry pod and the database
|
||||||
|
pod run images from Docker Hub by digest (`REGISTRY_IMAGE` and `POSTGRES_IMAGE` in
|
||||||
|
`bootstrap.sh`), which the installer pulls into k3s's containerd and pins there so the
|
||||||
|
kubelet's image GC never collects them. When the pull fails (`could not pull …`), fetch
|
||||||
|
the same digest on a machine that can, for the node's architecture, keeping the manifest
|
||||||
|
as it is, and import it on the node:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
# elsewhere: the reference as bootstrap.sh names it, e.g. docker.io/library/postgres:18.6-trixie@sha256:…
|
||||||
|
sudo ctr images pull --platform linux/arm64 "$ref"
|
||||||
|
sudo ctr images export --platform linux/arm64 image.tar "$ref"
|
||||||
|
# on the node:
|
||||||
|
sudo k3s ctr images import image.tar
|
||||||
|
```
|
||||||
|
|
||||||
|
A `docker save` of the image rewrites its manifest, so its copy never matches the digest
|
||||||
|
the Deployment names. Rerun the installer afterwards; it finds the image and pins it.
|
||||||
|
[CODE-ONLY]
|
||||||
|
|
||||||
**Kaniko is archived upstream** (June 2025); v1.24.0 is its last release and
|
**Kaniko is archived upstream** (June 2025); v1.24.0 is its last release and
|
||||||
gets no security fixes. To run a maintained fork, copy it under `mirror/` and
|
gets no security fixes. To run a maintained fork, copy it under `mirror/` and
|
||||||
point the override at it. The other `[registry]` keys in `felis.toml`:
|
point the override at it. The other `[registry]` keys in `felis.toml`:
|
||||||
@@ -1353,12 +1372,12 @@ space.
|
|||||||
|
|
||||||
## 13c. The host's address, name or clock changed
|
## 13c. The host's address, name or clock changed
|
||||||
|
|
||||||
An install is bound to the address it was made on. `bootstrap.sh` writes that
|
An install is bound to the address it was made on. `bootstrap.sh` gives that
|
||||||
address into the database connection string, `pg_hba.conf`, the network
|
address to the k3s node and writes it into the network policies, the panel
|
||||||
policies, the panel certificate and the default `<ip>.nip.io` root domain, and
|
certificate and the default `<ip>.nip.io` root domain, and nothing re-addresses
|
||||||
nothing re-addresses a live install. When the host loses the address (a DHCP
|
a live install. When the host loses the address (a DHCP lease that came back
|
||||||
lease that came back different, a moved VM), felis-api cannot reach PostgreSQL
|
different, a moved VM), the cluster and the proxy's path to the game servers
|
||||||
and the panel stops answering on its old name. The watchdog reports it as
|
still name the old one, and the panel stops answering on its old name. The watchdog reports it as
|
||||||
`host-address` (critical, after 5 minutes). The installer warns at install time
|
`host-address` (critical, after 5 minutes). The installer warns at install time
|
||||||
when the address is a DHCP lease.
|
when the address is a DHCP lease.
|
||||||
|
|
||||||
@@ -1665,19 +1684,18 @@ runs changed:
|
|||||||
|---|---|
|
|---|---|
|
||||||
| `felis-velocity` (the proxy) | its unit, the JRE, `velocity.jar`, `velocity.toml`, the forwarding secret, the felis-link settings or a plugin jar changed, or it was not running. The fingerprint lives in `/etc/felis/velocity.fingerprint`; delete it to force a restart. |
|
| `felis-velocity` (the proxy) | its unit, the JRE, `velocity.jar`, `velocity.toml`, the forwarding secret, the felis-link settings or a plugin jar changed, or it was not running. The fingerprint lives in `/etc/felis/velocity.fingerprint`; delete it to force a restart. |
|
||||||
| login and lobby pods | the rebuilt limbo or lobby image has a new image ID (`/etc/felis/system-server-images`). The installer then pins that server's `spec.image` to the digest its tag names now (`felis pin-images --system login\|lobby`) and the operator rolls the pod onto it, each on its own; with the registry unreachable it recreates the pod instead. The installer turns off BuildKit's default provenance attestation (`BUILDX_NO_DEFAULT_ATTESTATIONS=1`): it records the build time, which would give every rebuild a new ID. |
|
| login and lobby pods | the rebuilt limbo or lobby image has a new image ID (`/etc/felis/system-server-images`). The installer then pins that server's `spec.image` to the digest its tag names now (`felis pin-images --system login\|lobby`) and the operator rolls the pod onto it, each on its own; with the registry unreachable it recreates the pod instead. The installer turns off BuildKit's default provenance attestation (`BUILDX_NO_DEFAULT_ATTESTATIONS=1`): it records the build time, which would give every rebuild a new ID. |
|
||||||
| PostgreSQL | first install only (`listen_addresses` needs a restart). A rerun reloads the configuration, which keeps connections open. |
|
| felis-postgres | the release moved `POSTGRES_IMAGE` or changed the pod; a few seconds without the API. A rerun that changes neither leaves it running. |
|
||||||
| felis-api, felis-operator, the registry pod (its gate and GC containers run the felis binary) | the image tag changed (an upgrade), or a same-version rerun rebuilt it. |
|
| felis-api, felis-operator, the registry pod (its gate and GC containers run the felis binary) | the image tag changed (an upgrade), or a same-version rerun rebuilt it. |
|
||||||
|
|
||||||
**PostgreSQL across reruns.** On hosts without firewalld the installer loads an
|
**PostgreSQL across reruns.** The database runs in k3s from the image the release
|
||||||
nftables table, `inet felis_postgres`, from `felis-postgres-firewall.service`:
|
pins, so a distribution upgrade never moves it. The installer refuses to start a
|
||||||
port 5432 accepts loopback, the pod network and the node's own address and drops
|
PostgreSQL whose major version differs from the cluster in `/var/lib/felis/postgres`
|
||||||
everything else (`nft list table inet felis_postgres`). firewalld hosts already
|
and names the dump-and-restore path (docs/operations.md §4). The first rerun of a
|
||||||
keep 5432 closed to the network. The installer also refuses to start a
|
release with felis-postgres on a host an earlier release installed moves the
|
||||||
PostgreSQL whose major version differs from the cluster in the data directory,
|
database off the host PostgreSQL (docs/operations.md §4, "The database's move into
|
||||||
and prints the `pg_upgrade` steps; distributions that move the server package to
|
k3s"), and removes that release's `felis-postgres-firewall.service`. On Arch the
|
||||||
a new major (Arch, Fedora) would otherwise leave the database unable to start.
|
installer's `pacman -Syu` keeps holding the host `postgresql` package back while its
|
||||||
On Arch the installer's `pacman -Syu` holds `postgresql` back once a cluster
|
cluster exists: that cluster is the copy a rollback of the move starts again.
|
||||||
exists, so the database is upgraded only when you run `pg_upgrade` yourself.
|
|
||||||
|
|
||||||
`rollout undo` reverts the image only. The upgrade's database migrations stay
|
`rollout undo` reverts the image only. The upgrade's database migrations stay
|
||||||
applied; when they are the problem, restore the `pre-migrate` bundle the upgrade
|
applied; when they are the problem, restore the `pre-migrate` bundle the upgrade
|
||||||
@@ -1766,8 +1784,11 @@ build no server and no whitelist entry names is pruned after 24 hours, and the
|
|||||||
The PostgreSQL database behind felis-api holds everything that is not a world:
|
The PostgreSQL database behind felis-api holds everything that is not a world:
|
||||||
accounts, passkeys, Minecraft account links, server ownership, quotas, audit
|
accounts, passkeys, Minecraft account links, server ownership, quotas, audit
|
||||||
logs, and the `world_backups` index that maps an archive (§10) back to its
|
logs, and the `world_backups` index that maps an archive (§10) back to its
|
||||||
owner. Losing it orphans every world archive. It lives on the host (not in
|
owner. Losing it orphans every world archive. It runs in k3s as the
|
||||||
k3s), so it is backed up on the host too.
|
`felis-postgres` Deployment, with its cluster on the host in
|
||||||
|
`/var/lib/felis/postgres`; `felis db` runs `pg_dump`, `psql` and `pg_restore`
|
||||||
|
inside that pod (the host config's `[database] deployment`) and keeps the
|
||||||
|
bundles on the host.
|
||||||
|
|
||||||
### What runs, and where the bundles go
|
### What runs, and where the bundles go
|
||||||
|
|
||||||
@@ -1829,11 +1850,42 @@ sudo journalctl -u felis-db-backup -n 50 --no-pager # why the last run failed
|
|||||||
sudo felis db backup # take one now (label manual)
|
sudo felis db backup # take one now (label manual)
|
||||||
```
|
```
|
||||||
|
|
||||||
Common failures: PostgreSQL down (`pg_dump: ... connection refused`); the
|
Common failures: felis-postgres not running (`kubectl exec` reports no running
|
||||||
backup directory's disk full (the half-written `.partial` is removed and the
|
pod, or `pg_dump: ... connection refused`; next section); `k3s: executable file not
|
||||||
previous bundles stay intact); `pg_dump: server version mismatch` when an
|
found` from a `felis` that runs with neither `/usr/local/bin` on PATH nor k3s
|
||||||
external database is newer than the host's client tools (install the matching
|
anywhere else; the backup directory's disk full (the half-written `.partial` is
|
||||||
`postgresql` client package).
|
removed and the previous bundles stay intact); `pg_dump: server version mismatch`
|
||||||
|
when an external database (no `deployment` in `[database]`) is newer than the
|
||||||
|
host's client tools (install the matching `postgresql` client package).
|
||||||
|
|
||||||
|
### felis-postgres is not running, or never became ready
|
||||||
|
|
||||||
|
The installer stops with `felis-postgres did not become ready` after 10 minutes
|
||||||
|
and prints the pod's events; the watchdog reports the Deployment the same way it
|
||||||
|
reports felis-api. Look at the pod:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sudo k3s kubectl -n felis get pods -l app.kubernetes.io/component=postgres -o wide
|
||||||
|
sudo k3s kubectl -n felis describe deploy/felis-postgres | tail -n 30
|
||||||
|
sudo k3s kubectl -n felis logs deploy/felis-postgres --tail=60
|
||||||
|
```
|
||||||
|
|
||||||
|
| What it says | Cause | Fix |
|
||||||
|
|---|---|---|
|
||||||
|
| `ImagePullBackOff` / `ErrImagePull` on `postgres` | the node cannot reach Docker Hub and has no copy | §8e, "The platform's own images on an air-gapped node" |
|
||||||
|
| `CreateContainerConfigError`, `secret "felis-postgres" not found` | the superuser Secret is gone | rerun the installer, which makes a new one. The image reads it only when it creates a cluster; the installer and `felis db` reach the database over the pod's socket |
|
||||||
|
| the log says `Permission denied` on `/var/lib/postgresql/18/docker` | the hostPath lost its owner (uid 999) or, with SELinux enforcing, its `container_file_t` label (a restore of `/var/lib/felis` by hand, `restorecon` without the installer's rule) | rerun the installer, which sets both; by hand: `sudo chown -R 999:999 /var/lib/felis/postgres` and `sudo restorecon -R /var/lib/felis/postgres` |
|
||||||
|
| the log says `database files are incompatible with server` | the cluster was made by another major version | docs/operations.md §4, "PostgreSQL major versions" |
|
||||||
|
| `Pending`, `Insufficient memory` | the node is full | §13b |
|
||||||
|
|
||||||
|
Query the database by hand from inside the pod, as the superuser over its socket:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sudo k3s kubectl -n felis exec -it deploy/felis-postgres -c postgres -- psql -U postgres felis
|
||||||
|
```
|
||||||
|
|
||||||
|
The copy an earlier release's host PostgreSQL still holds, from before the move
|
||||||
|
into k3s, stays on the host; docs/operations.md §4 covers going back to it.
|
||||||
|
|
||||||
### Check a bundle
|
### Check a bundle
|
||||||
|
|
||||||
@@ -1897,7 +1949,7 @@ off-site bucket (next sections) plus a fresh install. What the host holds:
|
|||||||
|
|
||||||
| Data | On the host | In the bucket | Brought back by | Lost at most |
|
| Data | On the host | In the bucket | Brought back by | Lost at most |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | PostgreSQL | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer |
|
| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | felis-postgres, `/var/lib/felis/postgres` | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer |
|
||||||
| Host state (`/etc/felis`: secrets, both `felis.toml` copies, `offsite.env`, panel TLS pair) | `/etc/felis` | inside every bundle | `tar -x` of the bundle's `state/` | as the database |
|
| Host state (`/etc/felis`: secrets, both `felis.toml` copies, `offsite.env`, panel TLS pair) | `/etc/felis` | inside every bundle | `tar -x` of the bundle's `state/` | as the database |
|
||||||
| MinecraftServer objects | k3s | inside every bundle (`k8s/minecraftservers.json`) | `kubectl apply` | as the database |
|
| MinecraftServer objects | k3s | inside every bundle (`k8s/minecraftservers.json`) | `kubectl apply` | as the database |
|
||||||
| World archives (reaper, "Back up now", pre-restore snapshots) | `felis-backups` volume | each one within the hour | `fetch-worlds` | archives written in the last hour |
|
| World archives (reaper, "Back up now", pre-restore snapshots) | `felis-backups` volume | each one within the hour | `fetch-worlds` | archives written in the last hour |
|
||||||
@@ -2222,7 +2274,7 @@ sign-in keeps working. Each lock is audited as `auth.otp.locked` and counted in
|
|||||||
To lift a lock early once you have confirmed the owner locked themselves out:
|
To lift a lock early once you have confirmed the owner locked themselves out:
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
sudo -u postgres psql felis -c \
|
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c \
|
||||||
"DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '<name>');"
|
"DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '<name>');"
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -2238,7 +2290,7 @@ keys on) and `user_agent`. `actor` is display text: a verified email or the
|
|||||||
username, never an address the caller set without verifying.
|
username, never an address the caller set without verifying.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
sudo -u postgres psql felis -c "
|
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c "
|
||||||
SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason
|
SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason
|
||||||
FROM audit_logs
|
FROM audit_logs
|
||||||
WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour'
|
WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour'
|
||||||
@@ -2306,7 +2358,7 @@ something failed, `otp_skip_detail`. Root can edit the row afterwards, so it
|
|||||||
records attribution without proving it.
|
records attribution without proving it.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
sudo -u postgres psql felis -c "
|
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c "
|
||||||
SELECT created_at, action, actor, payload->>'verified' AS verified,
|
SELECT created_at, action, actor, payload->>'verified' AS verified,
|
||||||
payload->>'otp_skipped' AS skipped, payload->>'otp_skip_detail' AS detail
|
payload->>'otp_skipped' AS skipped, payload->>'otp_skip_detail' AS detail
|
||||||
FROM audit_logs WHERE source = 'break-glass' ORDER BY created_at DESC LIMIT 20;"
|
FROM audit_logs WHERE source = 'break-glass' ORDER BY created_at DESC LIMIT 20;"
|
||||||
@@ -2355,6 +2407,7 @@ for 10 seconds (the Free plan's limits).
|
|||||||
| `pre-migration backup failed, nothing applied` during an upgrade | §16 |
|
| `pre-migration backup failed, nothing applied` during an upgrade | §16 |
|
||||||
| Undo a mistaken change / restore the control-plane database | §16 |
|
| Undo a mistaken change / restore the control-plane database | §16 |
|
||||||
| Host lost: rebuild from a database bundle | §16 |
|
| Host lost: rebuild from a database bundle | §16 |
|
||||||
|
| `felis-postgres did not become ready`; `kubectl exec` finds no database pod | §16 |
|
||||||
| `the off-site copy last completed ... ago` / `NO OFF-SITE COPY` / reaper `awaiting_offsite` stays above 0 | §16, §10 |
|
| `the off-site copy last completed ... ago` / `NO OFF-SITE COPY` / reaper `awaiting_offsite` stays above 0 | §16, §10 |
|
||||||
| Sign-in 429 `rate_limited` for everyone at once | §17 |
|
| Sign-in 429 `rate_limited` for everyone at once | §17 |
|
||||||
| 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 |
|
| 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 |
|
||||||
|
|||||||
Reference in new issue
Block a user