docs: 发布产物安装、FELIS_ARTIFACT_DIR 与回退提示

This commit is contained in:
Lemon-miaow committed 2026-09-26 15:23:54 +08:00
1 parent d34a118723
commit 2ca2f5ca86
3 files changed
+106 -21

No files matched your search

+3 -1
View File
@@ -33,7 +33,9 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
```
脚本将自动安装 K3s,在 K3s 内部署 PostgreSQL 与控制平面,并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。旧版本装在宿主上的 PostgreSQL 会在重跑时整库迁进 K3s,宿主上的那份停用保留,供回退(见 [运维手册 §4](docs/operations.md#4-upgrading-the-pieces-around-felis))。
脚本将自动安装 K3s,在 K3s 内部署 PostgreSQL 与控制平面,并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。
安装发布版时,二进制、全部镜像与 Velocity 插件都取自该版本在 CI 里预构建好的 release 附件,逐个核对 `SHA256SUMS` 后导入,主机上无需 Docker、Gradle 或 Go,也不从 Docker Hub 拉取;某个附件缺失或校验不符时,只有那一个镜像退回到本机构建,并给出提示(见 [故障排查 §15c](docs/troubleshooting.md))。内网或受限网络的主机可以先把 release 附件拷到本机,再用 `FELIS_ARTIFACT_DIR=<绝对路径>` 安装(见 [运维手册 §1](docs/operations.md#1-supported-hosts))。旧版本装在宿主上的 PostgreSQL 会在重跑时整库迁进 K3s,宿主上的那份停用保留,供回退(见 [运维手册 §4](docs/operations.md#4-upgrading-the-pieces-around-felis))。
动手之前,脚本先检查内存、磁盘、端口、网段冲突、已有的 Kubernetes 和外网连通,把所有问题一次列出并停下,主机上什么都没改(检查项见 [运维手册 §1](docs/operations.md#1-supported-hosts))。
+70 -15
View File
@@ -12,14 +12,16 @@ Evidence tags follow troubleshooting.md: **[VM-VERIFIED]** was run on a real hos
## 1. Supported hosts
`deploy/bootstrap.sh` provisions a single node. It needs systemd, root, and one of the
package managers below; everything else (Docker, k3s, the JRE, cloudflared) it installs.
package managers below; everything else (k3s, the JRE, cloudflared, and Docker when an image
has to be built on the host; see "Where the binary and the images come from" below) it
installs.
PostgreSQL runs inside k3s as the `felis-postgres` Deployment, from the official image the
release pins by digest, with its data on the host in `/var/lib/felis/postgres`.
| OS family | Package manager | Architectures | Status |
|---|---|---|---|
| CentOS Stream 9 (firewalld active, SELinux enforcing) | dnf | aarch64 | **[VM-VERIFIED]** install, same-version rerun, upgrade, uninstall and reinstall, with the database on the host PostgreSQL 13 of the releases before felis-postgres; felis-postgres and the move into it [SH-TESTED] |
| Ubuntu 24.04 LTS | apt | x86_64 | **[CI]** fresh install, same-commit rerun, and upgrade from the newest release to the pushed commit |
| Ubuntu 24.04 LTS | apt | x86_64 | **[CI]** fresh install and same-commit rerun from the pushed commit's release assets, and upgrade from the newest release onto them; the on-host build weekly |
| RHEL / Rocky / Alma 9, Fedora | dnf | x86_64, aarch64 | [CODE-ONLY] same code path as CentOS Stream |
| Debian 12, other Ubuntu releases | apt | x86_64, aarch64 | [CODE-ONLY] |
| openSUSE Leap / Tumbleweed | zypper | x86_64, aarch64 | [CODE-ONLY] |
@@ -45,10 +47,11 @@ once, then stops with nothing touched **[SH-TESTED]**:
- the architecture, systemd as init, and the memory cgroup controller k3s needs;
- RAM: under 1.75 GiB is refused (a "2 GB" VPS passes), under 3.5 GiB is a warning;
- free disk on each filesystem it writes to, summed when they share one: about 23 GiB
on a bare host, 7 GiB for a rerun, a directory that already holds data (Docker's cache,
a reused k3s) counting at the rerun size; a filesystem that would end over 85%, where
k3s starts deleting cached images, is a warning;
- free disk on each filesystem it writes to, summed when they share one: on a bare host
about 17 GiB installing a release, 15 GiB from `FELIS_ARTIFACT_DIR` and 23 GiB when it
builds the images itself; 7 GiB for a rerun; a directory that already holds data
(Docker's cache, a reused k3s) counting at the rerun size; a filesystem that would end
over 85%, where k3s starts deleting cached images, is a warning;
- the ports it will listen on: the game port, the panel NodePort, k3s's 6443/6444 and
10248–10259 and the registry's loopback 5000. A port held by the installer's own
proxy or k3s is a rerun and passes;
@@ -56,7 +59,10 @@ once, then stops with nothing touched **[SH-TESTED]**:
- the node address or a routed network inside k3s's `10.42.0.0/16` and `10.43.0.0/16`
(a Docker network there is the usual case); a wider route such as a `10.0.0.0/8` VPN
is a warning;
- HTTPS to the hosts it downloads from (GitHub, PaperMC's download API, Docker Hub).
- HTTPS to the hosts it downloads from: GitHub and PaperMC's download API always, Docker
Hub when it builds images on the host. Installing a release, an unreachable Docker Hub
is a warning (it is needed only if an asset turns out unusable); from
`FELIS_ARTIFACT_DIR` it is not checked.
`FELIS_PREFLIGHT=warn` reports the same problems as warnings and installs anyway, for a
host the checks misjudge.
@@ -86,6 +92,51 @@ and gains no failover. A multi-node shape would need, at least, storage that can
pod to another node and leader election in felis-operator (controller-runtime's
`LeaderElection`) so a second replica can stand by.
### Where the binary and the images come from
A release install (the default channel, and the setup console) takes everything Felis
builds from that release's assets, each checked against the release's `SHA256SUMS` before
it is used: the `felis` binary, the control-plane image, the limbo, lobby and paper images,
the registry and PostgreSQL images (at the digests `bootstrap.sh` pins), and
`felis-velocity.jar`. The images go into k3s's containerd with `k3s ctr images import` and
from there into the in-cluster registry, so the host needs no Docker, Gradle, Go or Docker
Hub for them. k3s's own images come from k3s's GitHub release
(`k3s-airgap-images-<arch>.tar.zst`, checked against k3s's sha256 list) before k3s first
starts. An upgrade downloads only the image tars holding an image the host lacks; they wait
in `/var/lib/felis/artifacts` until the registry has the images, and are deleted then.
`deploy/build-release-artifacts.sh` documents every asset. The decisions are **[SH-TESTED]**;
the import into a real k3s is **[CODE-ONLY]** until the e2e job runs.
The installer builds on the host instead, installing Docker for it and stopping Docker once
the images are in the registry, when:
- the source is not a release: `FELIS_VERSION_BOOTSTRAP=dev`, a pinned `FELIS_REF`, or
`FELIS_SKIP_FETCH`;
- `FELIS_GAME_STACK=latest`, for the login, lobby and paper images (the rest still come
from the release);
- the release publishes no `SHA256SUMS` (one cut before release assets existed, or still
uploading), or an asset is missing, fails its checksum or is malformed. Only that image is
built (the registry and PostgreSQL images are pulled from Docker Hub instead), and a
warning names it; troubleshooting §15c lists the messages.
`FELIS_ARTIFACT_DIR=<absolute path>` installs from a directory instead of the release: a
release's assets downloaded there (every `felis-*` file and `SHA256SUMS`), or the directory
`deploy/build-release-artifacts.sh <version> <dir>` wrote. Nothing of Felis's own is
downloaded or built (except the game images under `FELIS_GAME_STACK=latest`, which no release
ships), so an asset the directory lacks, or one failing its checksum, stops the install; k3s,
the JRE, cloudflared and Velocity still come from GitHub and PaperMC. It
cannot be combined with `FELIS_REF` or `FELIS_SKIP_FETCH`, which name a source too.
```
# on a machine with access: the release's assets for the host's architecture
gh release download v1.4.0 --repo FelisMC/Felis --dir felis-v1.4.0 \
--pattern 'felis-*linux-amd64*' --pattern felis-velocity.jar --pattern SHA256SUMS
# on the host, after copying the directory over
sudo FELIS_ARTIFACT_DIR=/root/felis-v1.4.0 bash bootstrap.sh
```
`SHA256SUMS` lists both architectures; the files of the other one may be left out.
### While felis-api restarts
An installer rerun that changes felis-api, a node restart or a crashed pod takes the API
@@ -140,9 +191,10 @@ server running **[VM-VERIFIED]**:
Every game server adds the memory its owner gave it: the pod's limit equals its request,
and the JVM heap is derived from it (§1a). Quotas cap it per user (panel → 管理 → 配额).
The installer's own peak is the image builds (Docker plus a Gradle container); it stops
Docker afterwards so that memory goes back to the servers. On a host under 2 GB of RAM
without swap it adds a 2 GiB `/swapfile`.
A release install builds nothing (§1). When the installer builds on the host its peak is
the image builds (Docker plus a Gradle container); it stops Docker afterwards so that memory
goes back to the servers. On a host under 2 GB of RAM without swap it adds a 2 GiB
`/swapfile`.
### Recommendations
@@ -176,7 +228,8 @@ curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo FELIS_VELOCITY_XMX=2G bash
| World archives | the `felis-backups` volume (`FELIS_BACKUP_STORAGE`, default 10Gi requested) | about one compressed world per backup kept |
| In-cluster registry | the `registry` volume (default 10Gi requested) | 2–3 GB for the stock images; grows with custom builds, pruned daily (§9) |
| k3s's containerd images | `/var/lib/rancher/k3s/agent/containerd` | 6–9 GB |
| Docker's images and build cache | `/var/lib/containerd` (Docker's containerd store) | 5–10 GB after repeated upgrades |
| Docker's images and build cache | `/var/lib/containerd` (Docker's containerd store), on a host that built its images (§1) | 5–10 GB after repeated upgrades |
| Release assets during an install | `/var/lib/felis/artifacts` | up to ~2 GB, deleted once the images are in the registry |
| Toolchains and sources | `/opt/felis` | ~2.5 GB |
| Database | `/var/lib/felis/postgres` (felis-postgres's cluster) | tens of MB; the audit log is most of it |
| Database bundles | `/var/lib/felis/db-backups` | a few MB each, 14 daily kept |
@@ -184,8 +237,9 @@ curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo FELIS_VELOCITY_XMX=2G bash
k3s's local-path volumes do not enforce the requested sizes (§9), so every volume shares
the root filesystem. Give the host at least **40 GB**, and 60 GB or more once worlds and
custom images accumulate. The watchdog mails the owners when a watched filesystem passes
its threshold, and §13b covers a full disk. `docker builder prune -af` (with Docker
started) reclaims the build cache when space is short; the next upgrade rebuilds it.
its threshold, and §13b covers a full disk. On a host that built its images,
`docker builder prune -af` (with Docker started) reclaims the build cache when space is
short; the next upgrade rebuilds it.
## 3. Uninstall
@@ -201,8 +255,9 @@ curl -fsSL <raw-url>/deploy/uninstall.sh | sudo bash -s -- --purge # remove th
With a private repository, fetch it the way the README fetches `bootstrap.sh`.
Both modes remove the `felis-*` systemd units and `cloudflared-felis.service`, the
Velocity user, `/opt/felis`, `/usr/local/bin/felis`, the installer's cloudflared binary
(unless another unit runs it), the `felis_edge` nftables table (and `felis_postgres`, which
Velocity user, `/opt/felis`, `/usr/local/bin/felis`, the release assets an interrupted
install left in `/var/lib/felis/artifacts`, the installer's cloudflared binary (unless
another unit runs it), the `felis_edge` nftables table (and `felis_postgres`, which
releases before the database moved into k3s loaded) and the firewalld ports the installer
opened. k3s goes with k3s's own `k3s-uninstall.sh` when the cluster holds nothing but
Felis's namespaces; when it runs anything else only `felis`, `minecraft`, `felis-build`
+33 -5
View File
@@ -598,8 +598,11 @@ reach upstream; that is accurate.
**The platform's own images on an air-gapped node.** The registry pod and the database
pod run images from Docker Hub by digest (`REGISTRY_IMAGE` and `POSTGRES_IMAGE` in
`bootstrap.sh`), which the installer pulls into k3s's containerd and pins there so the
kubelet's image GC never collects them. When the pull fails (`could not pull …`), fetch
`bootstrap.sh`), which the installer imports into k3s's containerd from the release's
`felis-image-base-linux-<arch>.tar` (or, without one, pulls) and pins there so the
kubelet's image GC never collects them. On an air-gapped node the simplest way is to
install from the release's assets copied to the node (`FELIS_ARTIFACT_DIR`, operations
§1). When the pull fails (`could not pull …`) and no release assets are at hand, fetch
the same digest on a machine that can, for the node's architecture, keeping the manifest
as it is, and import it on the node:
@@ -1616,8 +1619,9 @@ annotated Service endpoints picks them up as is.
## 15. Control-plane upgrades, and rolling back a bad one
There is no in-place updater: an upgrade is re-running the installer
(`curl -fsSL <installer URL> | sudo bash`), which rebuilds/re-imports the image
and re-applies the bundle. `felis update --panel` prints that command with the
(`curl -fsSL <installer URL> | sudo bash`), which imports the release's images
(or rebuilds them on the host, operations §1 "Where the binary and the images come
from") and re-applies the bundle. `felis update --panel` prints that command with the
script read at the newest release's tag, so the installer and the binary it
downloads come from the same release. (`sudo felis setup` is not this path; on a completed
install it only opens the config console.) The channel is not persisted across
@@ -1637,10 +1641,12 @@ Two properties of the control plane matter when you do:
| Download | Check |
|---|---|
| `felis-linux-<arch>` (release channel) | its sha256 must match the release's `SHA256SUMS`; a release without one, or a mismatch, is compiled from the same tag instead |
| the image tars `felis-image-*-linux-<arch>.tar`, their listing `felis-images-linux-<arch>.txt`, `felis-velocity.jar` (release channel, `FELIS_ARTIFACT_DIR`) | each sha256 must match the release's `SHA256SUMS`; the listing must give one known role per line with sha256 digests, and each image must be in containerd under the digest the listing names once its tar is imported. An asset that fails is built on the host instead (§15c); from `FELIS_ARTIFACT_DIR` the install stops |
| k3s's own images (`k3s-airgap-images-<arch>.tar.zst`) | against the k3s release's `sha256sum-<arch>.txt`; a mismatch leaves k3s pulling them from Docker Hub |
| k3s (fresh install, or `FELIS_UPGRADE_DEPS=1`) | the install script is read at `FELIS_K3S_VERSION`'s tag (default `v1.36.4+k3s1`), and it checks the binary against that release's sha256 list |
| cloudflared (when absent, or `FELIS_UPGRADE_DEPS=1`) | release `FELIS_CLOUDFLARED_VERSION` (default `2026.9.1`) against a pinned sha256; another version needs `FELIS_CLOUDFLARED_SHA256` |
| Go toolchain (nano, source builds) | pinned sha256 per architecture; another version needs `FELIS_GO_SHA256` |
| the registry image | pinned by digest (`registry:2.8.3@sha256:a3d8…`) |
| the registry and PostgreSQL images | pinned by digest (`registry:2.8.3@sha256:a3d8…`, `postgres:18.6-trixie@sha256:5a5a…`); a release's copy must carry that name and digest |
| Limbo, its spawn schematic, Paper, LuckPerms, Velocity | the builds and sha256s in `deploy/game-stack.lock`; each image build and the proxy install refuse a download that hashes differently (§15b) |
| the Temurin JRE the proxy runs on | release `25.0.4.1+1` against a pinned sha256 per architecture |
| base images of the felis, limbo, lobby and paper images | pinned by digest in each `Dockerfile` |
@@ -1779,6 +1785,28 @@ build no server and no whitelist entry names is pruned after 24 hours, and the
| A running server restarted during an installer re-run | it was pinned in place: the operator rolled it onto the pinned ref, the build it already ran | nothing; it happens once per server |
| Create/edit refused with `the registry no longer holds build …` | the image names a digest the pruner deleted: nothing referenced it for 24 hours (§9) | pick a current tag; whitelist the versioned tag of a build you want kept |
## 15c. The installer builds on the host although it installs a release
A release install takes its images and the Velocity plugin from the release's assets
(operations §1, "Where the binary and the images come from"). When one cannot be used the
installer names it and the reason, and builds that image with Docker instead (the registry
and PostgreSQL images are pulled from Docker Hub). Nothing unchecked is used either way
**[SH-TESTED]**.
| Message | Meaning | What to do |
|---|---|---|
| `release vX publishes no SHA256SUMS … building them on this host instead` | the release predates release assets, or release.yml is still uploading them | nothing for an old release; for a new one, rerun once the release page lists `SHA256SUMS` |
| `SHA256SUMS lists no <file>` or `could not download <file> from release vX` | the release lacks that asset (a partial upload) | rerun later; the host build is correct meanwhile |
| `downloaded <file> hashes to …, but release vX's SHA256SUMS says …` | the download was corrupted, or the asset was replaced after `SHA256SUMS` was written | rerun: the bad copy is gone and is fetched again; the same mismatch every time means the asset itself is bad, so report it |
| `the release's image listing … is malformed` | `felis-images-linux-<arch>.txt` does not parse | report it; every image is built on the host |
| `the release's registry image is …, but this installer runs …` | the release's base tar carries another digest than this `bootstrap.sh` pins: the installer and the release are from different versions | run the installer read at the release's tag (`felis update --panel` prints that command) |
| `k3s containerd holds no <image> from <tar>` | the tar was imported but did not hold the image under the listed digest | `sudo k3s ctr images ls \| grep felis` shows what it holds; report it |
| `FELIS_ARTIFACT_DIR: …`, and the install stops | an asset is missing from the directory or fails its checksum; nothing is built from a directory | copy the named file from the release again and rerun |
An install that stops part way leaves the downloaded tars in `/var/lib/felis/artifacts`; the
rerun reuses those that still match `SHA256SUMS` and deletes the directory once the images
are in the registry.
## 16. Control-plane database backups and disaster recovery
The PostgreSQL database behind felis-api holds everything that is not a world: