docs(audit): idle auto-stop postmortem — ledger #25 and self-checks in §11

Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
This commit is contained in:
Lemon-miaow committed 2026-09-23 04:14:11 +08:00
1 parent f650bf892a
commit 089d4f3a80
2 files changed
+36 -1

No files matched your search

+10 -1
View File
@@ -26,6 +26,7 @@ IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:f
| 22 | **`role='owner'` 是死信**:迁移 0011 加了 owner 角色并把全部用户管理路由压在 `IsOwner()` 上,但**没有任何代码写过 `owner`**——breakGlass(`UpsertOwner`)、setup MC 绑定(`CompleteOwnerSetup`)、重置路径一律写 `admin` → 全新安装的整个 owner 层(用户列表/创建/编辑/禁用/删除/配额/会话)不可达;且新角色没进各处 staff 判定(op-login 只认 admin、玩家邮箱门只拒 admin、reclaim 保护与 AdminExists 只认 admin);面板一旦有 owner 行还可被降级/删除/禁用 | 真机:owner 会话加载 `/api/v1/users` 403;promote 后 200。op-login:修复前 owner start 中立不发码,修复后铸请求+finish 得 `role=owner` 会话 | `UpsertOwner`/`CompleteOwnerSetup` 改写 `role='owner'`(冲突臂重断言,即 0011 文档的 promote 路径);op-login 双端改 `staffRole`;玩家邮箱门改 `role != 'user'`;`IsProtectedAdminLink`/`AdminExists` 计入 owner;面板新增守卫:owner 行不可降级/删除/禁用(用户名/邮箱编辑仍可)。提交 `e0d2378` | | 22 | **`role='owner'` 是死信**:迁移 0011 加了 owner 角色并把全部用户管理路由压在 `IsOwner()` 上,但**没有任何代码写过 `owner`**——breakGlass(`UpsertOwner`)、setup MC 绑定(`CompleteOwnerSetup`)、重置路径一律写 `admin` → 全新安装的整个 owner 层(用户列表/创建/编辑/禁用/删除/配额/会话)不可达;且新角色没进各处 staff 判定(op-login 只认 admin、玩家邮箱门只拒 admin、reclaim 保护与 AdminExists 只认 admin);面板一旦有 owner 行还可被降级/删除/禁用 | 真机:owner 会话加载 `/api/v1/users` 403;promote 后 200。op-login:修复前 owner start 中立不发码,修复后铸请求+finish 得 `role=owner` 会话 | `UpsertOwner`/`CompleteOwnerSetup` 改写 `role='owner'`(冲突臂重断言,即 0011 文档的 promote 路径);op-login 双端改 `staffRole`;玩家邮箱门改 `role != 'user'`;`IsProtectedAdminLink`/`AdminExists` 计入 owner;面板新增守卫:owner 行不可降级/删除/禁用(用户名/邮箱编辑仍可)。提交 `e0d2378` |
| 23 | **配额门非原子 + storage 缓存被清零**:(a) audit #4:`QuotaCheck` 与 `ClaimServer` 两条语句,同一用户并发认领两台无主服可双双通过 `max_servers`(deferred-seams 曾把这条挂为“只能在真 PG 上关闭”);(b) 更隐蔽:PATCH 资源时 `UpdateServerResources(..., 0)` 把本不能改的 storage 缓存写 0,而缓存列是配额聚合的**唯一**输入 → 此后该服的 storage 维度在配额里凭空消失 | (a) 新增 pgint 并发测试:修复前两台全赢;修复后恰 1 赢 + 1 `ErrQuotaExceeded`,DB 只 1 行 owned;(b) hermetic 测试 `TestPatchServerPreservesStorageCache`(修复前 `resourceUpdates` 里 storage=0) | (a) 门槛进 `ClaimServer`:同一事务内 `pg_advisory_xact_lock(hashtext(user_id))` + 四维重查(与 `QuotaCheck` 共用 `quotaAllows` 防漂移)+ 行 `FOR UPDATE`,两个 claim handler 把 `ErrQuotaExceeded` 映射为与串行一致的 403;(b) resize 前读取现值并透传 storage。提交 `bb68fef`;deferred-seams 对应条目核销 | | 23 | **配额门非原子 + storage 缓存被清零**:(a) audit #4:`QuotaCheck` 与 `ClaimServer` 两条语句,同一用户并发认领两台无主服可双双通过 `max_servers`(deferred-seams 曾把这条挂为“只能在真 PG 上关闭”);(b) 更隐蔽:PATCH 资源时 `UpdateServerResources(..., 0)` 把本不能改的 storage 缓存写 0,而缓存列是配额聚合的**唯一**输入 → 此后该服的 storage 维度在配额里凭空消失 | (a) 新增 pgint 并发测试:修复前两台全赢;修复后恰 1 赢 + 1 `ErrQuotaExceeded`,DB 只 1 行 owned;(b) hermetic 测试 `TestPatchServerPreservesStorageCache`(修复前 `resourceUpdates` 里 storage=0) | (a) 门槛进 `ClaimServer`:同一事务内 `pg_advisory_xact_lock(hashtext(user_id))` + 四维重查(与 `QuotaCheck` 共用 `quotaAllows` 防漂移)+ 行 `FOR UPDATE`,两个 claim handler 把 `ErrQuotaExceeded` 映射为与串行一致的 403;(b) resize 前读取现值并透传 storage。提交 `bb68fef`;deferred-seams 对应条目核销 |
| 24 | **reaper 警告信从不真正投递**:`felis reaper` 从不装配任何 Warner(`r.Warner` 恒 nil),而 `maybeWarn` 对 nil warner / 投递失败一律照样 `MarkWarned` + `warned++` → 每位有主的服都在**无人收到提醒**的情况下 15 天后被静默回收,跑批日志还谎报"已警告 N 台"。红线⑤的"best-effort 不阻塞回收"被误读成了"失败也要记成已通知" | hermetic 单测 3 例(坏 notifier / nil warner → 不盖章、成功才盖章)。真机 drill(auditfix20):13d idle 实收 1 封(收件人=owner 的验证邮箱、subject 含 3d);重跑不重发;14.5d 补发 1d(elif 档位);SMTP 端口打坏 → 日志 `warn delivery failed; will retry next run` 且**不盖章**,恢复后补发;邮箱未验证 → `has no verified email` 不盖章;owner NULL → 静默跳过;边界 11d23h 不发 / 12d1m 发;全程 `world_backups` 保持 4 条不动(纯警告零备份副作用) | `mail.SendNotice` + `mailWarner`:查 owner 的 verified email → `SendNotice`;nil/失败不盖章、下次重试;reaper pod 模板加 optional `FELIS_SMTP_PASSWORD` env;`felis setup` 复制 felis-smtp 镜像并在「configure email」刷新 minecraft ns 的 felis-smtp+felis-config 镜像;docs/troubleshooting §10 更新。提交 `8e7c7bb` | | 24 | **reaper 警告信从不真正投递**:`felis reaper` 从不装配任何 Warner(`r.Warner` 恒 nil),而 `maybeWarn` 对 nil warner / 投递失败一律照样 `MarkWarned` + `warned++` → 每位有主的服都在**无人收到提醒**的情况下 15 天后被静默回收,跑批日志还谎报"已警告 N 台"。红线⑤的"best-effort 不阻塞回收"被误读成了"失败也要记成已通知" | hermetic 单测 3 例(坏 notifier / nil warner → 不盖章、成功才盖章)。真机 drill(auditfix20):13d idle 实收 1 封(收件人=owner 的验证邮箱、subject 含 3d);重跑不重发;14.5d 补发 1d(elif 档位);SMTP 端口打坏 → 日志 `warn delivery failed; will retry next run` 且**不盖章**,恢复后补发;邮箱未验证 → `has no verified email` 不盖章;owner NULL → 静默跳过;边界 11d23h 不发 / 12d1m 发;全程 `world_backups` 保持 4 条不动(纯警告零备份副作用) | `mail.SendNotice` + `mailWarner`:查 owner 的 verified email → `SendNotice`;nil/失败不盖章、下次重试;reaper pod 模板加 optional `FELIS_SMTP_PASSWORD` env;`felis setup` 复制 felis-smtp 镜像并在「configure email」刷新 minecraft ns 的 felis-smtp+felis-config 镜像;docs/troubleshooting §10 更新。提交 `8e7c7bb` |
| 25 | **idle auto-stop 从未触发(三层复合缺陷)**:条件齐备的空载服永远不停。① CRD status schema 未声明 `emptySince`,apiserver **pruning** 掉计时戳(`unknown field "status.emptySince"`),计时器每次读回都是 nil;② 即使戳幸存,Running 空载稳态**没有任何 watch 事件**(玩家进出不碰 CRD、RCON 只在 Reconcile 内探),盖章一次后 reconcile 链停摆,无人叫醒;③ auto-stop 用整对象 `Update` 写 spec 且 OperatorRole 从未有 `minecraftservers:patch/update` → 403 `cannot update resource`。三层任一都让功能永久失效,而单测(fake client 不剪 schema、手动驱动、无 RBAC)全部覆盖不到 | 真机逐层实锤:schema 修复后 `empty=2026-09-22T20:02:13Z` 首次成功持久化;静置 88s+ 无动作、operator 日志 90s 零行(②实锤);修复前日志 5 条 403(③实锤);全修后场景 1:超时戳触发即 Stopped;场景 2:起服 → 20:11:43 盖章 → **全程无干预** → 20:12:17 自动 Stopped;翻转 Running→Stopped→Running 收敛 Running;`idle=null` 后不再自停 | ① schema 补 `emptySince`(`c04a3f0`);② `reconcileRunning` 尾部返回 RequeueAfter——空载=到点精确唤醒、有人=30s 探针周期,+3 个单测(`1c89a5e`);③ auto-stop 改 merge-patch(防 status clobber,与 reaper 的 Stop 同型)+ OperatorRole 补 `patch` + rbac 测试锚点(`f650bf8`)。真机 auditfix22 全通;troubleshooting §11 增自检三连 |
## 待决策台账(未修) ## 待决策台账(未修)
@@ -88,6 +89,14 @@ IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:f
- **边界**:idle=11d23h(< 12d 阈值)不发不盖章;idle=12d1m 发。`http_code` 冒烟:升级 auditfix20 后 panel 200 / op-login 200。 - **边界**:idle=11d23h(< 12d 阈值)不发不盖章;idle=12d1m 发。`http_code` 冒烟:升级 auditfix20 后 panel 200 / op-login 200。
- **环境还原**:warntest CRD+行删除、drill Secret/Pod 删除、operator 恢复;`servers` 表回到 resolvecheck+test-one,CRD 回到 lobby/login/resolvecheck/test-one。 - **环境还原**:warntest CRD+行删除、drill Secret/Pod 删除、operator 恢复;`servers` 表回到 resolvecheck+test-one,CRD 回到 lobby/login/resolvecheck/test-one。
### 本轮新增真机证据(第七批:idle auto-stop 三层修复,auditfix21/22)
- **发现路径**:operator 深挖时先怀疑"计时器无驱动",真机实验立刻抓到 pruning(操作日志 `unknown field "status.emptySince"`)+ 触发后 88s 无动作 + RBAC 403,三层各自独立、各自足以致死。
- **场景 1(超时点唤醒 + 写权限)**:EmptySince 已超时 11 分钟 → annotate 触发一次 → 秒级内 desiredState=Stopped、sts 0/replicas、EmptySince 清、phase=Stopped。
- **场景 2(完整自驱,决定性)**:desired=Running → pod 起 → 20:11:43 盖章 → 静置无干预 → **20:12:17 自动 Stopped**(30s 到点后 ~4s 完成 stop+scaledown+markStopped)。
- **翻转混沌**:3 秒内 Running→Stopped→Running,最终收敛 Running ready(期间 409 竞争为控制器正常噪音,controller-runtime 重试自愈)。
- **收尾**:`spec.idle` 删除后再无自停;test-one 回 Stopped、空 sts、无 EmptySince;VM 镜像 auditfix22 = `f650bf8`。
## 结论:离"生产可用"还差什么(按优先级) ## 结论:离"生产可用"还差什么(按优先级)
1. ~~构建链路的上下文通道~~ ✅ **已修**(`f79e5eb`/`02fd2de`,真机全链路含拉回校验;Trivy DB 需按 §8e 镜像一次)。 1. ~~构建链路的上下文通道~~ ✅ **已修**(`f79e5eb`/`02fd2de`,真机全链路含拉回校验;Trivy DB 需按 §8e 镜像一次)。
@@ -108,6 +117,6 @@ IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:f
- 面板会话 cookie:`/tmp/felis-cookies.json`;API 助手:`/tmp/fcurl.sh` - 面板会话 cookie:`/tmp/felis-cookies.json`;API 助手:`/tmp/fcurl.sh`
- 测试服:`test-one`(minecraft ns,stopped);合法备份 `bk-47ee2e7e96e5a4ca9d0e51b805518bac` - 测试服:`test-one`(minecraft ns,stopped);合法备份 `bk-47ee2e7e96e5a4ca9d0e51b805518bac`
- CDP 调试口:Mac `127.0.0.1:9333`(独立 Chrome,profile `/tmp/felis-chrome2`);WebAuthn 虚拟认证器需在**同一 CDP 会话**内完成仪式,且先 `Page.bringToFront`(否则 NotAllowedError: page does not have focus) - CDP 调试口:Mac `127.0.0.1:9333`(独立 Chrome,profile `/tmp/felis-chrome2`);WebAuthn 虚拟认证器需在**同一 CDP 会话**内完成仪式,且先 `Page.bringToFront`(否则 NotAllowedError: page does not have focus)
- 已部署到 VM:`felis-api`/`felis-operator`/`felis-reaper`(CronJob) 镜像 = `felis:auditfix20`(含 #20–#24 全部修复;迁移 0020 已应用,`schema_migrations` max=20;CronJob 仍 `suspend=true`,SMTP 密码 env 待 `felis setup`「configure email」刷新时落地) - 已部署到 VM:`felis-api`/`felis-operator`/`felis-reaper`(CronJob) 镜像 = `felis:auditfix22`(含 #20–#25 全部修复;迁移 0020 已应用,`schema_migrations` max=20;CronJob 仍 `suspend=true`,SMTP 密码 env 待 `felis setup`「configure email」刷新时落地;operator Role 已手动补 `minecraftservers:patch`,`felis install/setup` 重渲染时收敛)
- 构建链路 drill 现成条件:`felis-build` 里有 `felis-service-token`(Secret 复制);`felis-config` 里 `kaniko_image/trivy_image` 指向 k3s containerd 已导入的 pin tag、`trivy_db_repository = registry.felis.svc:5000/mirror/trivy-db:2`(镜像配方 §8e;VM 上 docker daemon 的 `insecure-registries` 已含 registry ClusterIP) - 构建链路 drill 现成条件:`felis-build` 里有 `felis-service-token`(Secret 复制);`felis-config` 里 `kaniko_image/trivy_image` 指向 k3s containerd 已导入的 pin tag、`trivy_db_repository = registry.felis.svc:5000/mirror/trivy-db:2`(镜像配方 §8e;VM 上 docker daemon 的 `insecure-registries` 已含 registry ClusterIP)
- VM 内部面:`k3s kubectl -n felis port-forward svc/felis-api-internal 18081:8081`(Pod 重建后转发会悬死,需重启);reaper CronJob(minecraft ns)已应用但 `suspend=true` - VM 内部面:`k3s kubectl -n felis port-forward svc/felis-api-internal 18081:8081`(Pod 重建后转发会悬死,需重启);reaper CronJob(minecraft ns)已应用但 `suspend=true`
+26
View File
@@ -600,6 +600,32 @@ Both fields set and still nothing happens? Then the probe is failing rather than
disabled: the server would be stuck in `Starting` with `RconNotReachable` disabled: the server would be stuck in `Starting` with `RconNotReachable`
(`reconciler.go:156`), which is §1's symptom, not this one. (`reconciler.go:156`), which is §1's symptom, not this one.
This path used to fail even with everything configured correctly, through three
stacked defects proven and fixed on a live cluster (auditfix21/22): the
`emptySince` stamp was pruned by a missing CRD status field, a quiescent empty
server produced no watch events to re-check the timer, and the Role lacked the
`minecraftservers:patch` grant the stop write needs. If auto-stop ever looks
dead again, check these three in order (each is now pinned by a test):
```sh
# ① The stamp must persist — should print a timestamp, not an empty string,
# a few seconds after a server goes Ready with zero players.
kubectl get minecraftserver <name> -o jsonpath='{.status.emptySince}'
# ② The operator must be able to write spec.desiredState (403 in the operator
# log = missing patch grant on Role felis-operator).
kubectl auth can-i patch minecraftservers -n <ns> --as=system:serviceaccount:<ctl-ns>:felis-operator
# ③ A wake-up must be scheduled: while empty, expect whatever you set
# as emptySecondsBeforeStop to elapse and the box to flip to Stopped without
# any external action.
```
While players are online the operator re-probes on a 30s cadence so it notices
the moment the last one leaves; while empty it schedules a wake-up exactly at
the deadline. Quiet operator logs on an idle server are normal — the action is
the scheduled wake-up, not a stream of reconciles.
Note the reaper's `last_active_at` (§10) is a *different* subsystem (Postgres Note the reaper's `last_active_at` (§10) is a *different* subsystem (Postgres
business layer, bumped by join events) — it keeps worlds alive against the business layer, bumped by join events) — it keeps worlds alive against the
reaper, but it does **not** auto-stop empty running servers. reaper, but it does **not** auto-stop empty running servers.