feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)

- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
This commit is contained in:
Lemon-miaow committed 2026-09-22 22:03:48 +08:00
1 parent 0a2d654e68
commit 87a9f4eb25
10 files changed
+165 -37

No files matched your search

+2 -2
View File
@@ -19,14 +19,14 @@ IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:f
| 19 | **reaper CronJob 渲染位置错误,永远无法调度**:`felis manifests` 把 CronJob 渲染到控制 ns,却引用 minecraft ns 的备份 PVC(Pod 不能跨 ns 挂 PVC:真机 `FailedScheduling: persistentvolumeclaim "felis-backups" not found`);修正 ns 后又发现 ServerAccount 也不能跨 ns 使用(`serviceaccount "felis-reaper" not found`)。而 reaper 的 Role/RoleBinding 本就在 minecraft ns | 真机三层取证(PVC/SA/调度)| CronJob 与 SA、RoleBinding subject 全部移到 `MinecraftNamespace`(提交 `e4f2cff`+`c839454`)。**修复后完整演练通过**(见下) |
| 6 | **默认安装无备份能力 + 回收链路不可达/不可用**:(a) 没有任何环节渲染归档 PVC,`FELIS_BACKUP_PVC` 永远为空 → backup/restore 恒 503;(b) bootstrap 从不下传 reaper 三旗标 → 官方安装路径根本无法启用回收;(c) 即便启用,stock k3s 的 `<root>/<pvc>` 布局不存在,且 k3s storage root 是 `0700 root:root`、reaper pod 以 uid 1000 运行 → 真机 `lstat /worlds/…: permission denied`(fail-closed 跳过,但纯空转);(d) README 口径(“定时备份”)不实 | 真机全链路:默认渲染 PVC+env → 建服→marker→备份 202→jobs 端点 running→succeeded→改 marker→restore→读回原值 ✅;再建 resolvecheck 世界(20d idle,marker)→ CronJob 手动 Job:`evaluated=2 reaped=1`,归档含 marker、PVC+宿主目录回收、`world_backups` 得 `inactive_15d` 行、servers 行/CR 保留 ✅ | `platform/workloads.go` 渲染归档 PVC(minecraft ns、RWO 10Gi、默认 SC),`felis manifests --backup-pvc` 默认 `felis-backups`(`=` 空为显式关闭),bootstrap 统一下传 env/旗标并写 [archive] local_path;`resolveWorldDir` 新增精确 local-path 目录解析(读 PVC `spec.volumeName`,非 glob,杜绝陈旧 PV 目录顶替);ReaperRole 增 `pvc:get`;bootstrap 设 `FELIS_WORLDS_HOST_PATH` 时给 uid 1000 授 traverse(setfacl/o+x);README/故障手册改写。提交 `fd33fd0`+`2b87a5a` |
| 8 | **磁盘打满灾难链**:DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后数分钟才恢复调度。恢复依赖人工 `docker save | k3s ctr images import -` | 三轮真机 drill(原缺陷复现 + 两轮修复验证):①自定义 1e6 类:游戏 pod 先走、控制面随后仍被驱逐(kubelet 日志逐条列出 ranked/evicted),随后镜像被 GC → ErrImagePull;②内建 `system-cluster-critical`(2e9):填盘至 1.7G free,kubelet 对 felis-api/operator/registry 全部报 **“cannot evict a critical pod”**,三者在整个 DiskPressure 期间保持 Running;游戏 pod(0)被驱逐;③释放磁盘后:`DiskPressure` 经 ~5 分钟(`eviction-pressure-transition-period`)转 False——即“恢复慢”的主因;被 GC 的游戏镜像按 runbook 重导入后 25s 恢复 ✅ | 控制面四类 pod 挂内建 `system-cluster-critical`(用户自定义类值上限 1e9,达不到 2e9 临界阈值;preemption 保留=管理面可调度,已记录权衡);troubleshooting §13b 固化事件链、5 分钟条件过渡与镜像恢复路径。提交见下 |
| 9 | **升级策略 Recreate**:单副本 + Recreate:任何控制面升级=停机;坏升级(错 tag)中断约 95s 且需人工 `rollout undo`(无自动回滚) | 复核部署模板(Recreate + 单副本 + 无 leader election 的注释理由成立) | 不改策略(Recreate 是对无 leader election 的正确取舍),改为固化 runbook:troubleshooting §15 = 升级即重跑安装器;停机窗口=rollout 时长;坏镜像在安装器 180s 等待内以 `kubectl describe` 诊断呈现;回滚 `kubectl rollout undo`(镜像被 GC 时先按 §13b 重导入)。提交见下 |
| 10 | **备份语义**:归档包含整个 /data(jar、libraries、cache),167MB;是否符合 world backup 定位待评估 | 真机 tar 清单(cache/、config/、eula.txt、server.properties、jar)+ 恢复语义(overlay+prune) | 评估结论:**整卷备份是正确语义**(回滚=整服状态回滚,含配置/插件),保留行为;文档化:troubleshooting §10「backup contains」+ OpenAPI/README 口径改为“整服数据卷”。提交见下 |
## 待决策台账(未修)
| # | 主题 | 说明 |
|---|------|------|
| 7 | 异步失败不可感知 | backup/restore 失败后无状态出口:restore 行不变、backup 无行;只有集群侧 Job/日志可查。建议状态字段或 `?failed` 查询 |
| 9 | 升级策略 Recreate | 单副本 + Recreate:任何控制面升级=停机;坏升级(实测错 tag)服务中断约 95s 且需人工 `rollout undo`(无自动回滚)。建议 runbook/文档化 |
| 10 | 备份语义 | 归档包含整个 /data(jar、libraries、cache),167MB;是否符合"world backup"定位待评估 |
| 11 | PG 断连表现 | 会话查询失败报 401 而非 503(fail-closed 但误导;用户以为没登录) |
| 12 | ready 门滞后 | 容器 Ready 后 6~10s 内 API 仍 409 not_running |
| 13 | setup token 截断 | 43 字符 token + 长域名,80 列终端下 TUI 截断显示(复现:tmux 80 列) |
+1 -1
View File
@@ -17,7 +17,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
- **即开即玩**:玩家尝试连接时自动唤醒服务器,空闲后自动休眠,像游戏主机一样省资源。
- **Web 控制面板**:浏览器中查看服务器状态、在线玩家与资源用量,管理备份与恢复。
- **备份与恢复**:一键把世界打包进集群内的归档库,支持从任意备份点回滚;默认安装就已启用(归档 PVC 与路径由安装器一并生成)。
- **备份与恢复**:一键把整服数据(世界、配置、插件/模组,即整个 /data 卷)打包进集群内的归档库,支持从任意备份点回滚;默认安装就已启用(归档 PVC 与路径由安装器一并生成)。
- **智慧回收(可选开启)**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间;安装时设置 `FELIS_WORLDS_HOST_PATH`(k3s 默认 `/var/lib/rancher/k3s/storage`)即启用每日回收,不设置则不删任何世界。
- **多核心支持**:兼容 Paper、Fabric、Forge、NeoForge,经由 Velocity 代理统一入口。
- **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建并部署。
+1 -1
View File
@@ -17,7 +17,7 @@ Table of Contents
- **Wake on Join**: Servers start automatically when a player connects, and stop when idle — like hibernate for your server.
- **Web Dashboard**: Monitor server status, online players, and resource usage from your browser, with backup and restore management.
- **Backup & Restore**: One-click world snapshots into the cluster's archive store, with rollback from any backup point — enabled by default (the installer renders the archive PVC and its path).
- **Backup & Restore**: One-click snapshots of a server's whole data volume (worlds, config, plugins/mods — the entire /data volume) into the cluster's archive store, with rollback from any backup point — enabled by default (the installer renders the archive PVC and its path).
- **World Reaper** (opt in): Worlds idle for more than 15 days are automatically backed up and removed to free disk space. Enable it by setting `FELIS_WORLDS_HOST_PATH` at install time (on k3s: `/var/lib/rancher/k3s/storage`); without it, no world is ever deleted.
- **Multi-core Support**: Compatible with Paper, Fabric, Forge, and NeoForge, federated behind a Velocity proxy.
- **Modpack Submission**: Players submit custom modpacks; admin approval triggers automatic build and deployment.
+8
View File
@@ -402,6 +402,14 @@ func buildConfig(cfg *config.Config) build.Config {
return build.Config{
Namespace: cfg.Registry.BuildNamespace,
RegistryURL: cfg.Registry.URL,
// Empty overrides fall back to the build package's defaults, so an
// install that has not imported kaniko/trivy keeps the compiled-in refs
// (and fails loudly on pull rather than silently building with the wrong
// image).
KanikoImage: cfg.Registry.KanikoImage,
TrivyImage: cfg.Registry.TrivyImage,
CPULimit: cfg.Registry.BuildCPULimit,
MemLimit: cfg.Registry.BuildMemLimit,
}
}
+26
View File
@@ -44,6 +44,32 @@ func TestAuthSourcesFromConfig(t *testing.T) {
}
}
// TestBuildConfig_ProjectsOverrides pins the [registry] overrides reaching the
// build subsystem: unset fields must stay EMPTY (the build package's compiled-in
// defaults apply there, not here), and set fields must pass through verbatim —
// an air-gapped install points these at its imported mirrors.
func TestBuildConfig_ProjectsOverrides(t *testing.T) {
empty := buildConfig(&config.Config{})
if empty.KanikoImage != "" || empty.TrivyImage != "" || empty.CPULimit != "" || empty.MemLimit != "" {
t.Errorf("empty registry config must project empty overrides (defaults live in internal/build), got %+v", empty)
}
full := buildConfig(&config.Config{Registry: config.RegistryConfig{
URL: "registry.felis.svc:5000",
BuildNamespace: "felis-build",
KanikoImage: "reg/kaniko:v1",
TrivyImage: "reg/trivy:v1",
BuildCPULimit: "1",
BuildMemLimit: "2Gi",
}})
if full.KanikoImage != "reg/kaniko:v1" || full.TrivyImage != "reg/trivy:v1" ||
full.CPULimit != "1" || full.MemLimit != "2Gi" {
t.Errorf("registry overrides did not reach build.Config: %+v", full)
}
if full.Namespace != "felis-build" || full.RegistryURL != "registry.felis.svc:5000" {
t.Errorf("namespace/registry url must keep projecting: %+v", full)
}
}
// TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every
// felis-api listener: the shared factory must bound the header and idle phases
// (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because
+13 -1
View File
@@ -44,7 +44,19 @@ A grep across `*.md` and `*.go` returns both sets; only the Go ones are seams.
apply is under way.
- `internal/submit/blobstore.go:40` — the uploads PVC is mounted into felis-api but
not into the Kaniko build Pod, so a submitted context is durable at the derived
location without yet being readable by the build that consumes it.
location without yet being readable by the build that consumes it. Audited
2026-09-22: this is not a missing volume line — a PVC cannot cross namespaces
(uploads live in the control namespace; build Pods run in `felis-build`), so the
fix is a transport, not a mount. The `s3://` lane does not close it either: the
build Job carries no AWS credentials (no env, and the weak SA's token is
deliberately unmounted, so no IAM either). Options on the table: (a) object
storage with credentials plumbed into the build Pod as a per-build Secret plus an
egress allowance; (b) a context-handoff PVC/Job pair in `felis-build` fed from
the API side; (c) a node-local path both sides mount (single-node only, and it
hands an arbitrary Dockerfile a filesystem view — needs its own security review).
Kaniko/Trivy images are external-only by default; `[registry] kaniko_image /
trivy_image / build_cpu_limit / build_mem_limit` now override them for mirrored
or air-gapped installs.
## Built; only its I/O is unverifiable from this repo
+5 -4
View File
@@ -2771,11 +2771,12 @@ paths:
post:
tags: [backups]
operationId: backupNow
summary: Back up a server's world on demand (owner-or-admin; server must be stopped).
summary: Back up a server's data volume on demand (owner-or-admin; server must be stopped).
description: >-
Snapshots the server's world into the archive store as a first-class
world_backups row (reason "manual"), restorable later like an inactivity
backup. The world PVC is RWO and held by a running server, so the server must
Snapshots the server's whole data volume (worlds, config, plugins/mods,
jars, libraries — not just world folders) into the archive store as a
first-class world_backups row (reason "manual"), restorable later like an
inactivity backup. A restore replaces the volume with the archive. The world PVC is RWO and held by a running server, so the server must
be fully stopped first (409 not_stopped otherwise). The backup runs
asynchronously as a Job, so success is 202 (backing_up).
x-felis-face: [external]
+85 -28
View File
@@ -391,6 +391,32 @@ Inspect:
kubectl logs -n felis-build job/<build-job>
```
### 8e. Build Pods never start: executor images and air-gapped installs
The build Job runs Kaniko and Trivy from external registries by default
(`gcr.io/kaniko-project/executor:latest`, `aquasec/trivy:latest`). On a box whose
build namespace cannot reach those registries (the egress policy allows only
DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box
has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the
build stays `building` until its deadline. Point the overrides at images the box
CAN pull — typically imports into the node's containerd, pushed through the
internal registry — in `felis.toml`:
```toml
[registry]
url = "registry.felis.svc:5000"
build_namespace = "felis-build"
kaniko_image = "registry.felis.svc:5000/mirror/kaniko:v1.23.2"
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.58.1"
build_cpu_limit = "2"
build_mem_limit = "4Gi"
```
then restart `felis-api` (it renders the Job from this config). Unset fields keep
the defaults. Note the user-modpack context topologies are a separate,
still-open seam (see `docs/deferred-seams.md`); this section only makes the
executors reachable.
---
## 9. Registry push/pull failures (spec §15)
@@ -426,6 +452,16 @@ reaps a world only when `now - last_active_at > 15d` (`inactive_15d`); the 15-da
deadline is **hard-fixed in code** (only `warn_before` / `retention` /
`max_local_bytes` are configurable from `felis.toml [archive]`).
### What a "backup" contains
A backup tars the server's ENTIRE data volume — the same volume the server mounts
at `/data`: world folders, `server.properties`, plugins/mods, configs, jars,
libraries, logs and cache, not just the `world/` directory. A restore replaces the
volume's contents with the archive (files added since the backup are pruned), so a
restore also rolls config/plugin changes back. Sizes are dominated by
libraries/cache on stock Paper servers (~170MB for a fresh instance before any
world growth) — do not size the archive PVC as if only world data were stored.
### The backup-before-delete invariant
The reap sequence (all [GO-TESTED] hermetically) preserves the world unless a
@@ -583,8 +619,8 @@ A full disk is the one failure this platform cannot ride out by itself, because
the images exist only in the node's containerd (air-gapped by design), so a
GC'd image has no pull source.
Eviction. Every control-plane pod (api, operator, reaper, registry) runs under
the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
**Eviction.** Every control-plane pod (api, operator, reaper, registry) runs
under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
node-pressure eviction refuses to touch those pods — the log shows
*"Eviction manager: cannot evict a critical pod"* for each of them — while
game-server pods at the default priority 0 are evicted first. A drill that filled
@@ -596,39 +632,32 @@ critical threshold. The built-in class allows preemption (its policy is fixed),
so a control-plane pod that cannot fit may preempt a game pod — deliberate: the
management plane must be placeable.
The pressure condition clears slowly. After you free space, the node can stay
`DiskPressure:True` for up to ~5 minutes (`--eviction-pressure-transition-period`
defaults to 5m, to stop the condition flapping); pods that need scheduling wait
for it. This is the bulk of the "recovery takes minutes" observation, not a
stuck node.
**The pressure condition clears slowly.** After you free space, the node can stay
`DiskPressure:True` for up to ~5 minutes
(`--eviction-pressure-transition-period` defaults to 5m, to stop the condition
flapping); pods that need scheduling wait for it. This is the bulk of the
"recovery takes minutes" observation, not a stuck node.
But the *game* images can still be GC'd. If game pods were evicted, the kubelet
may garbage-collect their images (unused > 2 minutes under imagefs pressure), and
those pods then sit in `ImagePullBackOff` after recovery — re-import as above
(`docker save felis-limbo:demo felis-lobby:demo | k3s ctr images import -`, then
delete the stuck pods). Verified: both system servers returned to Running in
~25s after the import.
**The images may be gone.** If pods were evicted, the kubelet can garbage-collect
their images (unused > 2 minutes under imagefs pressure). Those pods then sit in
`ImagePullBackOff`/`ErrImagePull` for a tag that plainly exists —
`k3s ctr images ls` shows it missing. Recovery:
Symptoms of the image-GC stage: pods stuck `ImagePullBackOff`/`ErrImagePull`
with `kubectl describe pod` showing a pull attempt for a tag that plainly
exists (`k3s ctr images ls` will show it missing — the kubelet GC removed it
under imagefs pressure).
Recovery:
1. Free disk on the node (`df -h /var/lib/rancher`, the biggest consumers are
1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are
`k3s ctr images ls -q` and the world/backup PVCs under
`/var/lib/rancher/k3s/storage`).
2. Re-import the images by re-running the installer (it rebuilds imports from
2. Re-import the images by re-running the installer (it rebuilds/re-imports from
the local Docker store, which the kubelet GC does not touch):
`curl -fsSL <installer URL> | sudo bash` (or `sudo felis setup`), then
`kubectl -n felis rollout status deploy/felis-api`.
3. Delete now-unschedulable stuck pods so they retry with the re-imported image.
3. Delete the stuck pods so they retry against the re-imported image.
If the API itself is down and you only need the images back without a full
installer run: `docker save felis:<tag> | k3s ctr images import -` restores one
image from the Docker store (that store is deliberately a second copy; treat it
as the recovery path, not as free space).
For a single image without a full installer run:
`docker save felis:<tag> | k3s ctr images import -` — the Docker store is
deliberately a second copy; treat it as the recovery path, not as free space.
Verified end to end in the drill: `docker save felis-limbo:demo
felis-lobby:demo | k3s ctr images import -` plus pod deletion had both system
servers Running ~25s later.
---
@@ -648,6 +677,32 @@ All four mandated metrics have real producers; scrape them when triaging:
---
## 15. Control-plane upgrades, and rolling back a bad one
There is no in-place updater: an upgrade is re-running the installer
(`curl -fsSL <installer URL> | sudo bash`, or `sudo felis setup`), which
rebuilds/re-imports the image and re-applies the bundle. Two properties of the
control plane matter when you do:
- Both Deployments use strategy **Recreate** (single replica, no leader election:
two overlapping instances would fight over the same cluster). An upgrade takes
the panel/API down for the rollout window — seconds normally, longer if the new
image still has to be imported.
- If the new pod cannot start (bad tag, missing image), the installer's rollout
wait fails after 180s and prints `kubectl describe` diagnostics: you see
`ErrImagePull`/`ImagePullBackOff` there instead of a silent hang.
Roll back with:
```
kubectl -n felis rollout undo deploy/felis-api
kubectl -n felis rollout status deploy/felis-api
```
(the same for `felis-operator` and `registry`). `rollout undo` returns to the
previous ReplicaSet, whose image is normally still on the node; if it was GC'd
(§13b), re-import it first.
## Quick reference: symptom → section
| Symptom | Section |
@@ -663,10 +718,12 @@ All four mandated metrics have real producers; scrape them when triaging:
| Local password login rejected | §5c |
| Internal callers 401 (service token) | §6 |
| Link/claim 400/409/412/403/404 | §7 |
| Build push 400 / SA denied / egress hang / Failed | §8 |
| Build push 400 / SA denied / egress hang / Failed / executor ImagePullBackOff | §8, §8e |
| Registry push/pull unreachable | §9 |
| World deleted unexpectedly / backup skipped | §10 |
| Idle auto-stop not firing; player count 0 | §11 |
| A config field seems ignored | §12 |
| PVC left behind after delete | §13 |
| Node out of disk; pods evicted / ImagePullBackOff | §13b |
| Which metric to scrape | §14 |
| Upgrade / roll back a bad control-plane image | §15 |
+13
View File
@@ -119,6 +119,19 @@ type K8sConfig struct {
type RegistryConfig struct {
URL string `toml:"url"`
BuildNamespace string `toml:"build_namespace"`
// KanikoImage / TrivyImage / BuildCPULimit / BuildMemLimit override the
// build subsystem's compiled-in defaults (gcr.io/kaniko-project/executor and
// aquasec/trivy, 2 CPU / 4Gi per build container). The defaults assume the
// build namespace can reach those registries; on an air-gapped or mirrored
// install there IS no such reach (the build egress policy allows only DNS,
// the internal registry and explicit package mirrors), so the operator must
// point these at whatever their box can actually pull — typically images
// imported into the node's containerd alongside the felis image. Empty keeps
// the default.
KanikoImage string `toml:"kaniko_image"`
TrivyImage string `toml:"trivy_image"`
BuildCPULimit string `toml:"build_cpu_limit"`
BuildMemLimit string `toml:"build_mem_limit"`
// UserUploadsContext is the object-store base under which a user-submitted
// modpack's Kaniko build context is pinned. It belongs to the §16 build
// subsystem's input domain (the build-context store), introduced by the
+11
View File
@@ -47,6 +47,10 @@ metallb_pool = "192.0.2.200-250"
[registry]
url = "registry.felis.svc:5000"
build_namespace = "felis-build"
kaniko_image = "registry.felis.svc:5000/mirror/kaniko:v1.23.2"
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.58.1"
build_cpu_limit = "1"
build_mem_limit = "2Gi"
[archive]
store = "tarLocal"
@@ -84,6 +88,13 @@ func TestLoadValid(t *testing.T) {
if cfg.Archive.S3.Bucket != "felis-backups" {
t.Errorf("s3 bucket = %q", cfg.Archive.S3.Bucket)
}
if cfg.Registry.KanikoImage != "registry.felis.svc:5000/mirror/kaniko:v1.23.2" ||
cfg.Registry.TrivyImage != "registry.felis.svc:5000/mirror/trivy:0.58.1" {
t.Errorf("build image overrides = %q / %q", cfg.Registry.KanikoImage, cfg.Registry.TrivyImage)
}
if cfg.Registry.BuildCPULimit != "1" || cfg.Registry.BuildMemLimit != "2Gi" {
t.Errorf("build resource overrides = %q / %q", cfg.Registry.BuildCPULimit, cfg.Registry.BuildMemLimit)
}
}
func TestLoadAppliesDefaults(t *testing.T) {