docs: PostgreSQL 在 k3s 内运行后的升级、回退、卸载与排障

This commit is contained in:
Lemon-miaow committed 2026-09-26 13:40:47 +08:00
1 parent 36b8a8b444
commit c24396d50a
4 files changed
+177 -63

No files matched your search

+1 -1
View File
@@ -33,7 +33,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
```
脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。
脚本将自动安装 K3s,在 K3s 内部署 PostgreSQL 与控制平面,并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。旧版本装在宿主上的 PostgreSQL 会在重跑时整库迁进 K3s,宿主上的那份停用保留,供回退(见 [运维手册 §4](docs/operations.md#4-upgrading-the-pieces-around-felis))。
动手之前,脚本先检查内存、磁盘、端口、网段冲突、已有的 Kubernetes 和外网连通,把所有问题一次列出并停下,主机上什么都没改(检查项见 [运维手册 §1](docs/operations.md#1-supported-hosts))。
+1 -1
View File
@@ -36,7 +36,7 @@ every push):
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
```
The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain.
The script installs K3s, deploys PostgreSQL and the control plane inside it, and launches a setup wizard. Once done, open your browser at the configured domain. A PostgreSQL an earlier release installed on the host is moved into K3s on the next rerun, and the host copy is stopped and kept for a rollback (see [Operations §4](docs/operations.md#4-upgrading-the-pieces-around-felis)).
Before it changes anything, the script checks RAM, disk, ports, network-range clashes, any Kubernetes already there and outbound access, lists every problem at once and stops with the host untouched (the checks are in [operations §1](docs/operations.md#1-supported-hosts)).
+94 -33
View File
@@ -12,12 +12,13 @@ Evidence tags follow troubleshooting.md: **[VM-VERIFIED]** was run on a real hos
## 1. Supported hosts
`deploy/bootstrap.sh` provisions a single node. It needs systemd, root, and one of the
package managers below; everything else (Docker, k3s, PostgreSQL, the JRE, cloudflared)
it installs.
package managers below; everything else (Docker, k3s, the JRE, cloudflared) it installs.
PostgreSQL runs inside k3s as the `felis-postgres` Deployment, from the official image the
release pins by digest, with its data on the host in `/var/lib/felis/postgres`.
| OS family | Package manager | Architectures | Status |
|---|---|---|---|
| CentOS Stream 9 (firewalld active, PostgreSQL 13) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall |
| CentOS Stream 9 (firewalld active, SELinux enforcing) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall, with the database on the host PostgreSQL 13 of the releases before felis-postgres; felis-postgres and the move into it [SH-TESTED] |
| Ubuntu 24.04 LTS | apt | x86_64 | **[CI]** fresh install, same-commit rerun, and upgrade from the newest release to the pushed commit |
| RHEL / Rocky / Alma 9, Fedora | dnf | x86_64, aarch64 | [CODE-ONLY] same code path as CentOS Stream |
| Debian 12, other Ubuntu releases | apt | x86_64, aarch64 | [CODE-ONLY] |
@@ -34,7 +35,7 @@ cloudflared is left as it is, see §4):
| Temurin JRE (Velocity) | 25, patch build pinned | `FELIS_JRE_VERSION`, sha256 per architecture |
| Go (nano builds) | 1.26.8 | `GO_PINNED_VERSION`, sha256 per architecture |
| Minecraft / Limbo / Paper / Velocity / LuckPerms | `deploy/game-stack.lock` | §15b |
| PostgreSQL | the distribution's package | 13 and 18 are exercised by the `pgint` CI job |
| PostgreSQL | 18.6, the official `postgres` image by digest | `POSTGRES_IMAGE` in `bootstrap.sh`, `defaultPostgresImage` in `internal/platform` |
32-bit hosts are not supported: there is no k3s, JRE or Go build the installer will fetch
for them.
@@ -62,13 +63,12 @@ host the checks misjudge.
Two things the host must keep for as long as the install lives:
- **Its address.** The install is bound to the IPv4 address it was made on (the
database connection string, `pg_hba.conf`, the network policies, the panel
certificate and the default nip.io domain all carry it). Give the host a static
address or a DHCP reservation before installing; the installer warns when the address
is a lease, and the watchdog reports `host-address` when the host loses it
(troubleshooting §13c). The k3s node name is pinned at install time, so a hostname
change is harmless.
- **Its address.** The install is bound to the IPv4 address it was made on (the k3s
node, the network policies, the panel certificate and the default nip.io domain all
carry it). Give the host a static address or a DHCP reservation before installing;
the installer warns when the address is a lease, and the watchdog reports
`host-address` when the host loses it (troubleshooting §13c). The k3s node name is
pinned at install time, so a hostname change is harmless.
- **A synchronized clock.** The installer turns NTP on (chrony where nothing else can)
and the watchdog reports a clock that stays unsynchronized. Allow outbound UDP 123,
or set `FELIS_MANAGE_TIME_SYNC=0` on a host whose clock is kept another way.
@@ -134,7 +134,7 @@ server running **[VM-VERIFIED]**:
| lobby (Paper, pod limit 1 GiB) | ~0.7–0.85 GB |
| login (Limbo, pod limit 512 MiB) | ~0.16 GB |
| felis-api, felis-operator, registry gate | ~50 MB each |
| PostgreSQL | ~30 MB plus page cache |
| PostgreSQL (the felis-postgres pod) | ~30 MB plus page cache |
| **Total in use** | **~3.4 GB** |
Every game server adds the memory its owner gave it: the pod's limit equals its request,
@@ -178,6 +178,7 @@ curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo FELIS_VELOCITY_XMX=2G bash
| k3s's containerd images | `/var/lib/rancher/k3s/agent/containerd` | 6–9 GB |
| Docker's images and build cache | `/var/lib/containerd` (Docker's containerd store) | 5–10 GB after repeated upgrades |
| Toolchains and sources | `/opt/felis` | ~2.5 GB |
| Database | `/var/lib/felis/postgres` (felis-postgres's cluster) | tens of MB; the audit log is most of it |
| Database bundles | `/var/lib/felis/db-backups` | a few MB each, 14 daily kept |
k3s's local-path volumes do not enforce the requested sizes (§9), so every volume shares
@@ -201,23 +202,29 @@ With a private repository, fetch it the way the README fetches `bootstrap.sh`.
Both modes remove the `felis-*` systemd units and `cloudflared-felis.service`, the
Velocity user, `/opt/felis`, `/usr/local/bin/felis`, the installer's cloudflared binary
(unless another unit runs it), the `felis_postgres` and `felis_edge` nftables tables and
the firewalld ports the installer opened. k3s goes with k3s's own `k3s-uninstall.sh` when
the cluster holds nothing but Felis's namespaces; when it runs anything else only
`felis`, `minecraft`, `felis-build` and the MinecraftServer CRD are deleted.
(unless another unit runs it), the `felis_edge` nftables table (and `felis_postgres`, which
releases before the database moved into k3s loaded) and the firewalld ports the installer
opened. k3s goes with k3s's own `k3s-uninstall.sh` when the cluster holds nothing but
Felis's namespaces; when it runs anything else only `felis`, `minecraft`, `felis-build`
and the MinecraftServer CRD are deleted.
`--keep-k3s` and `--remove-k3s` override that choice.
| | keep data (default) | `--purge` |
|---|---|---|
| Final database bundle | taken first (`felis db backup -label manual`); a failure stops the uninstall before anything is removed. `--no-backup` skips it | none |
| `felis` database and role | kept | dropped; `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run |
| The database (`/var/lib/felis/postgres`) | kept; felis-postgres is stopped cleanly before k3s goes | deleted with `/var/lib/felis` |
| A host PostgreSQL an earlier release ran the database on | kept as it is: stopped after the move into k3s (below, §4), with its old copy of `felis` | its `felis` database and role are dropped (the server is started for that and stopped again), and `listen_addresses` and `pg_hba.conf` go back to how they were. Checked before anything is removed: a role that still owns another database (the `felis_pgint` the PG contract tests use, CONTRIBUTING.md) or holds grants elsewhere stops the purge up front with the list and the `ALTER DATABASE … OWNER TO postgres` to run |
| `/etc/felis` (secrets, `felis.toml`, `offsite.env`, tunnel config) | kept; `bootstrap.done` and the per-run records go | deleted, with the tunnel's credentials file |
| `/var/lib/felis` (database bundles) | kept | deleted |
| `/var/lib/felis` (the database, its bundles) | kept | deleted |
| Worlds, archives, registry, uploads | moved to `/var/lib/felis/retained/k3s-storage-<stamp>/` (with `--keep-k3s`: their volumes switch to `Retain` and stay in place) | deleted |
| Felis images, Docker build cache | kept | deleted |
Neither mode removes packages (Docker, PostgreSQL, git, nftables) or the swap file: other
software may use them. On a host that should end up bare:
The two database rows are [SH-TESTED] (`deploy/uninstall_test.sh`); the VM runs above
predate felis-postgres.
Neither mode removes packages (Docker, git, nftables, and the PostgreSQL server an earlier
release installed) or the swap file: other software may use them. On a host that should
end up bare:
```
sudo swapoff /swapfile && sudo rm /swapfile && sudo sed -i '\|^/swapfile |d' /etc/fstab
@@ -232,8 +239,9 @@ DNS records for the panel hostnames, and the Access application.
A keep-data uninstall leaves everything a reinstall needs. The installer reuses
`/etc/felis/secrets.env`, so the database password and the forwarding and session
secrets are unchanged, and it migrates the kept database instead of creating one
**[VM-VERIFIED]**.
secrets are unchanged, and the installer migrates the kept database instead of creating
one **[VM-VERIFIED]** (with the host database of the releases before felis-postgres).
felis-postgres starts again on the cluster kept in `/var/lib/felis/postgres` [SH-TESTED].
Each step below was run on the reference VM after a keep-data uninstall, and the
restored worlds matched their kept `level.dat` checksums **[VM-VERIFIED]**. `kept` names
@@ -305,7 +313,7 @@ the version they were installed with unless noted:
| Temurin JRE | moves to the pinned patch build | rerun |
| k3s | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: moves to the pinned release through that tag's install script, one minor version at a time (a bigger jump stops before anything changes and names the release to go through), never backwards |
| cloudflared | left alone | rerun with `FELIS_UPGRADE_DEPS=1`: swaps `/usr/local/bin/cloudflared` for the pinned, sha256-checked release and restarts `cloudflared-felis`; a cloudflared the distribution installed stays with its package manager |
| PostgreSQL | the distribution's package | the package manager for a minor release; a major version needs `pg_upgrade` first (below) |
| PostgreSQL | follows the image the release pins | a minor release comes with a Felis release, and the rerun restarts felis-postgres on it (a few seconds without the API); a major version is a dump and restore (below) |
| Docker, git, nftables | distribution packages | the package manager |
```sh
@@ -315,8 +323,9 @@ curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap
`sudo felis update` reports Felis, Velocity, k3s, cloudflared, the JRE and PostgreSQL
against their newest releases; `--k3s`, `--cloudflared`, `--jre` and `--postgres` narrow
it to one. PostgreSQL is compared within its major, since a minor release is a package
update, and a major past its end of life gets a note naming the current one.
it to one. PostgreSQL is read from the felis-postgres container and compared within its
major, since a minor release arrives with a Felis release, and a major past its end of life
gets a note naming the current one.
The installer also sets up `felis-update-check.timer`, which runs `felis update --record`
once a day around 05:30 (and at boot after a missed run). `--record` stores the result
@@ -370,20 +379,72 @@ The env var only carries the password; mail still needs the relay itself, set in
### PostgreSQL major versions [CODE-ONLY]
The installer takes the major the distribution ships (13 on EL9) and never moves it. To
go to a newer one, stop the writers, keep a dump, then use the distribution's upgrade
path:
felis-postgres keeps its cluster in `/var/lib/felis/postgres/<major>/docker`. A release that
moves the image to a new major finds the old major's cluster there and stops before it
changes anything: the new server would start an empty cluster beside it. The way across is
a bundle, taken on the release you run now, restored into the new major's empty cluster:
```sh
# On the release you run now:
b="$(sudo felis db backup -label pre-upgrade | sed -n 's/^felis db backup: wrote //p')"
sudo k3s kubectl -n felis scale deploy/felis-postgres --replicas=0
sudo mv /var/lib/felis/postgres/18 /var/lib/felis/postgres-18.old # the old major's cluster, for a way back
# Install the new release: it starts an empty cluster on the new major and creates the schema.
curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo bash
# Put the data back and bring its schema up to the new release.
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=0
sudo -u postgres pg_dumpall > /root/felis-pg-$(date +%F).sql
# EL9: sudo systemctl stop postgresql; sudo dnf module switch-to postgresql:16
# sudo dnf install postgresql-upgrade; sudo postgresql-setup --upgrade
# Debian/Ubuntu: sudo pg_upgradecluster <old-major> main
sudo systemctl start postgresql
sudo felis db restore -yes -no-safety-backup "$b"
sudo felis migrate up -config /etc/felis/felis.host.toml
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator --replicas=1
```
Delete `/var/lib/felis/postgres-18.old` once the new release has run for a while. To go back
instead, scale felis-postgres to 0, move the new major's directory out of
`/var/lib/felis/postgres`, move `postgres-18.old` back as `/var/lib/felis/postgres/18`, and
rerun the older release's installer.
### The database's move into k3s [SH-TESTED]
Releases before the move ran the database on a PostgreSQL the installer installed on the
host. The first rerun of a release with felis-postgres moves it, once:
1. It stops felis-api, felis-operator and the host timers, and heads the host's
`pg_hba.conf` with a block that refuses every connection to `felis` but its own copy
(the original is kept beside it as `pg_hba.conf.pre-pg-move`).
2. It takes a `pre-pg-move` bundle of the host database (`felis db backup`), restores it
into felis-postgres (`felis db restore`) and compares the row count of every table on
both servers. Any failure up to here puts `pg_hba.conf` and the control plane back and
the platform keeps running on the host database, untouched.
3. It stops and disables the host `postgresql` service, which stays installed with its
copy of the data, and writes `/var/lib/felis/postgres-moved`. A host server that also
holds other databases keeps running; its `felis` copy is then reachable over loopback
only.
From then on the host config points at felis-postgres (`127.0.0.1:15432`, and
`deployment = "felis/felis-postgres"`, through which `felis db` runs `pg_dump`, `psql`
and `pg_restore` inside the pod) and the pods at `felis-postgres.felis.svc:5432`.
To go back to the host database, for instance to reinstall the release before the move:
```sh
sudo k3s kubectl -n felis scale deploy/felis-api deploy/felis-operator deploy/felis-postgres --replicas=0
hba="$(sudo -u postgres psql -XtAc 'SHOW hba_file' 2>/dev/null || echo /var/lib/pgsql/data/pg_hba.conf)"
sudo cp -p "${hba}.pre-pg-move" "$hba"
sudo systemctl enable --now postgresql
sudo rm /var/lib/felis/postgres-moved
curl -fsSL <raw-url-of-that-release>/deploy/bootstrap.sh | sudo bash
```
`SHOW hba_file` needs the server running; with it stopped, the fallback path is EL's
(Debian and Ubuntu keep it in `/etc/postgresql/<major>/main/`). Whatever the platform wrote
after the move lives only in felis-postgres; take a bundle there first
(`sudo felis db backup`) and restore it onto the host database afterwards if that matters.
Once the move has run for a while, drop the host copy:
`sudo systemctl start postgresql; sudo -u postgres dropdb felis; sudo -u postgres dropuser felis`,
or remove the server package altogether.
### The MinecraftServer CRD [VM-VERIFIED]
Every rerun applies the CRD embedded in the `felis` binary (`felis bootstrap-assets crd`).
+81 -28
View File
@@ -596,6 +596,25 @@ jsonpath='{.data.platform}' | base64 -d`), and repeat that for the DBs as often
as advisories matter to you. The watchdog warning stays until the timer can
reach upstream; that is accurate.
**The platform's own images on an air-gapped node.** The registry pod and the database
pod run images from Docker Hub by digest (`REGISTRY_IMAGE` and `POSTGRES_IMAGE` in
`bootstrap.sh`), which the installer pulls into k3s's containerd and pins there so the
kubelet's image GC never collects them. When the pull fails (`could not pull …`), fetch
the same digest on a machine that can, for the node's architecture, keeping the manifest
as it is, and import it on the node:
```sh
# elsewhere: the reference as bootstrap.sh names it, e.g. docker.io/library/postgres:18.6-trixie@sha256:…
sudo ctr images pull --platform linux/arm64 "$ref"
sudo ctr images export --platform linux/arm64 image.tar "$ref"
# on the node:
sudo k3s ctr images import image.tar
```
A `docker save` of the image rewrites its manifest, so its copy never matches the digest
the Deployment names. Rerun the installer afterwards; it finds the image and pins it.
[CODE-ONLY]
**Kaniko is archived upstream** (June 2025); v1.24.0 is its last release and
gets no security fixes. To run a maintained fork, copy it under `mirror/` and
point the override at it. The other `[registry]` keys in `felis.toml`:
@@ -1353,12 +1372,12 @@ space.
## 13c. The host's address, name or clock changed
An install is bound to the address it was made on. `bootstrap.sh` writes that
address into the database connection string, `pg_hba.conf`, the network
policies, the panel certificate and the default `<ip>.nip.io` root domain, and
nothing re-addresses a live install. When the host loses the address (a DHCP
lease that came back different, a moved VM), felis-api cannot reach PostgreSQL
and the panel stops answering on its old name. The watchdog reports it as
An install is bound to the address it was made on. `bootstrap.sh` gives that
address to the k3s node and writes it into the network policies, the panel
certificate and the default `<ip>.nip.io` root domain, and nothing re-addresses
a live install. When the host loses the address (a DHCP lease that came back
different, a moved VM), the cluster and the proxy's path to the game servers
still name the old one, and the panel stops answering on its old name. The watchdog reports it as
`host-address` (critical, after 5 minutes). The installer warns at install time
when the address is a DHCP lease.
@@ -1665,19 +1684,18 @@ runs changed:
|---|---|
| `felis-velocity` (the proxy) | its unit, the JRE, `velocity.jar`, `velocity.toml`, the forwarding secret, the felis-link settings or a plugin jar changed, or it was not running. The fingerprint lives in `/etc/felis/velocity.fingerprint`; delete it to force a restart. |
| login and lobby pods | the rebuilt limbo or lobby image has a new image ID (`/etc/felis/system-server-images`). The installer then pins that server's `spec.image` to the digest its tag names now (`felis pin-images --system login\|lobby`) and the operator rolls the pod onto it, each on its own; with the registry unreachable it recreates the pod instead. The installer turns off BuildKit's default provenance attestation (`BUILDX_NO_DEFAULT_ATTESTATIONS=1`): it records the build time, which would give every rebuild a new ID. |
| PostgreSQL | first install only (`listen_addresses` needs a restart). A rerun reloads the configuration, which keeps connections open. |
| felis-postgres | the release moved `POSTGRES_IMAGE` or changed the pod; a few seconds without the API. A rerun that changes neither leaves it running. |
| felis-api, felis-operator, the registry pod (its gate and GC containers run the felis binary) | the image tag changed (an upgrade), or a same-version rerun rebuilt it. |
**PostgreSQL across reruns.** On hosts without firewalld the installer loads an
nftables table, `inet felis_postgres`, from `felis-postgres-firewall.service`:
port 5432 accepts loopback, the pod network and the node's own address and drops
everything else (`nft list table inet felis_postgres`). firewalld hosts already
keep 5432 closed to the network. The installer also refuses to start a
PostgreSQL whose major version differs from the cluster in the data directory,
and prints the `pg_upgrade` steps; distributions that move the server package to
a new major (Arch, Fedora) would otherwise leave the database unable to start.
On Arch the installer's `pacman -Syu` holds `postgresql` back once a cluster
exists, so the database is upgraded only when you run `pg_upgrade` yourself.
**PostgreSQL across reruns.** The database runs in k3s from the image the release
pins, so a distribution upgrade never moves it. The installer refuses to start a
PostgreSQL whose major version differs from the cluster in `/var/lib/felis/postgres`
and names the dump-and-restore path (docs/operations.md §4). The first rerun of a
release with felis-postgres on a host an earlier release installed moves the
database off the host PostgreSQL (docs/operations.md §4, "The database's move into
k3s"), and removes that release's `felis-postgres-firewall.service`. On Arch the
installer's `pacman -Syu` keeps holding the host `postgresql` package back while its
cluster exists: that cluster is the copy a rollback of the move starts again.
`rollout undo` reverts the image only. The upgrade's database migrations stay
applied; when they are the problem, restore the `pre-migrate` bundle the upgrade
@@ -1766,8 +1784,11 @@ build no server and no whitelist entry names is pruned after 24 hours, and the
The PostgreSQL database behind felis-api holds everything that is not a world:
accounts, passkeys, Minecraft account links, server ownership, quotas, audit
logs, and the `world_backups` index that maps an archive (§10) back to its
owner. Losing it orphans every world archive. It lives on the host (not in
k3s), so it is backed up on the host too.
owner. Losing it orphans every world archive. It runs in k3s as the
`felis-postgres` Deployment, with its cluster on the host in
`/var/lib/felis/postgres`; `felis db` runs `pg_dump`, `psql` and `pg_restore`
inside that pod (the host config's `[database] deployment`) and keeps the
bundles on the host.
### What runs, and where the bundles go
@@ -1829,11 +1850,42 @@ sudo journalctl -u felis-db-backup -n 50 --no-pager # why the last run failed
sudo felis db backup # take one now (label manual)
```
Common failures: PostgreSQL down (`pg_dump: ... connection refused`); the
backup directory's disk full (the half-written `.partial` is removed and the
previous bundles stay intact); `pg_dump: server version mismatch` when an
external database is newer than the host's client tools (install the matching
`postgresql` client package).
Common failures: felis-postgres not running (`kubectl exec` reports no running
pod, or `pg_dump: ... connection refused`; next section); `k3s: executable file not
found` from a `felis` that runs with neither `/usr/local/bin` on PATH nor k3s
anywhere else; the backup directory's disk full (the half-written `.partial` is
removed and the previous bundles stay intact); `pg_dump: server version mismatch`
when an external database (no `deployment` in `[database]`) is newer than the
host's client tools (install the matching `postgresql` client package).
### felis-postgres is not running, or never became ready
The installer stops with `felis-postgres did not become ready` after 10 minutes
and prints the pod's events; the watchdog reports the Deployment the same way it
reports felis-api. Look at the pod:
```sh
sudo k3s kubectl -n felis get pods -l app.kubernetes.io/component=postgres -o wide
sudo k3s kubectl -n felis describe deploy/felis-postgres | tail -n 30
sudo k3s kubectl -n felis logs deploy/felis-postgres --tail=60
```
| What it says | Cause | Fix |
|---|---|---|
| `ImagePullBackOff` / `ErrImagePull` on `postgres` | the node cannot reach Docker Hub and has no copy | §8e, "The platform's own images on an air-gapped node" |
| `CreateContainerConfigError`, `secret "felis-postgres" not found` | the superuser Secret is gone | rerun the installer, which makes a new one. The image reads it only when it creates a cluster; the installer and `felis db` reach the database over the pod's socket |
| the log says `Permission denied` on `/var/lib/postgresql/18/docker` | the hostPath lost its owner (uid 999) or, with SELinux enforcing, its `container_file_t` label (a restore of `/var/lib/felis` by hand, `restorecon` without the installer's rule) | rerun the installer, which sets both; by hand: `sudo chown -R 999:999 /var/lib/felis/postgres` and `sudo restorecon -R /var/lib/felis/postgres` |
| the log says `database files are incompatible with server` | the cluster was made by another major version | docs/operations.md §4, "PostgreSQL major versions" |
| `Pending`, `Insufficient memory` | the node is full | §13b |
Query the database by hand from inside the pod, as the superuser over its socket:
```sh
sudo k3s kubectl -n felis exec -it deploy/felis-postgres -c postgres -- psql -U postgres felis
```
The copy an earlier release's host PostgreSQL still holds, from before the move
into k3s, stays on the host; docs/operations.md §4 covers going back to it.
### Check a bundle
@@ -1897,7 +1949,7 @@ off-site bucket (next sections) plus a fresh install. What the host holds:
| Data | On the host | In the bucket | Brought back by | Lost at most |
|---|---|---|---|---|
| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | PostgreSQL | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer |
| Control-plane database (accounts, passkeys, ownership, quotas, audit, submissions, the `world_backups` index) | felis-postgres, `/var/lib/felis/postgres` | every bundle, copied within the hour of being written | `fetch-db`, `db restore` | changes since the newest bundle: up to a day plus an hour with the daily timer |
| Host state (`/etc/felis`: secrets, both `felis.toml` copies, `offsite.env`, panel TLS pair) | `/etc/felis` | inside every bundle | `tar -x` of the bundle's `state/` | as the database |
| MinecraftServer objects | k3s | inside every bundle (`k8s/minecraftservers.json`) | `kubectl apply` | as the database |
| World archives (reaper, "Back up now", pre-restore snapshots) | `felis-backups` volume | each one within the hour | `fetch-worlds` | archives written in the last hour |
@@ -2222,7 +2274,7 @@ sign-in keeps working. Each lock is audited as `auth.otp.locked` and counted in
To lift a lock early once you have confirmed the owner locked themselves out:
```sh
sudo -u postgres psql felis -c \
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c \
"DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '<name>');"
```
@@ -2238,7 +2290,7 @@ keys on) and `user_agent`. `actor` is display text: a verified email or the
username, never an address the caller set without verifying.
```sh
sudo -u postgres psql felis -c "
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c "
SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason
FROM audit_logs
WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour'
@@ -2306,7 +2358,7 @@ something failed, `otp_skip_detail`. Root can edit the row afterwards, so it
records attribution without proving it.
```sh
sudo -u postgres psql felis -c "
sudo k3s kubectl -n felis exec deploy/felis-postgres -c postgres -- psql -U postgres felis -c "
SELECT created_at, action, actor, payload->>'verified' AS verified,
payload->>'otp_skipped' AS skipped, payload->>'otp_skip_detail' AS detail
FROM audit_logs WHERE source = 'break-glass' ORDER BY created_at DESC LIMIT 20;"
@@ -2355,6 +2407,7 @@ for 10 seconds (the Free plan's limits).
| `pre-migration backup failed, nothing applied` during an upgrade | §16 |
| Undo a mistaken change / restore the control-plane database | §16 |
| Host lost: rebuild from a database bundle | §16 |
| `felis-postgres did not become ready`; `kubectl exec` finds no database pod | §16 |
| `the off-site copy last completed ... ago` / `NO OFF-SITE COPY` / reaper `awaiting_offsite` stays above 0 | §16, §10 |
| Sign-in 429 `rate_limited` for everyone at once | §17 |
| 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 |