diff --git a/AUDIT-2026-09-22.md b/AUDIT-2026-09-22.md index fdfba26..e9e08b3 100644 --- a/AUDIT-2026-09-22.md +++ b/AUDIT-2026-09-22.md @@ -18,7 +18,7 @@ IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:f | 18 | **绑定码并发兑换 500**:`RedeemPlayerBindCode` 无 `FOR UPDATE`(同文件 `VerifyLinkCode` 有),且裸 INSERT。6 路并发同码兑换 → 3×HTTP 500(`users_username_key` 唯一冲突)+ 1×400 + 2×200;跨码并发同 UUID 同样会撞 | 真机四组并发测试 + API 日志 `unmapped error ... duplicate key` | 同码:码行 `FOR UPDATE`(输家干净地 400 invalid_code);跨码:`INSERT users ... ON CONFLICT (username) DO NOTHING`+重读、`account_links ON CONFLICT (mc_uuid) DO NOTHING`(两路汇聚同一 user)。复测:同码×6=1×200+5×400、双码×2=2×200 同 ID、DB 干净、日志零 unmapped | | 19 | **reaper CronJob 渲染位置错误,永远无法调度**:`felis manifests` 把 CronJob 渲染到控制 ns,却引用 minecraft ns 的备份 PVC(Pod 不能跨 ns 挂 PVC:真机 `FailedScheduling: persistentvolumeclaim "felis-backups" not found`);修正 ns 后又发现 ServerAccount 也不能跨 ns 使用(`serviceaccount "felis-reaper" not found`)。而 reaper 的 Role/RoleBinding 本就在 minecraft ns | 真机三层取证(PVC/SA/调度)| CronJob 与 SA、RoleBinding subject 全部移到 `MinecraftNamespace`(提交 `e4f2cff`+`c839454`)。**修复后完整演练通过**(见下) | | 6 | **默认安装无备份能力 + 回收链路不可达/不可用**:(a) 没有任何环节渲染归档 PVC,`FELIS_BACKUP_PVC` 永远为空 → backup/restore 恒 503;(b) bootstrap 从不下传 reaper 三旗标 → 官方安装路径根本无法启用回收;(c) 即便启用,stock k3s 的 `/` 布局不存在,且 k3s storage root 是 `0700 root:root`、reaper pod 以 uid 1000 运行 → 真机 `lstat /worlds/…: permission denied`(fail-closed 跳过,但纯空转);(d) README 口径(“定时备份”)不实 | 真机全链路:默认渲染 PVC+env → 建服→marker→备份 202→jobs 端点 running→succeeded→改 marker→restore→读回原值 ✅;再建 resolvecheck 世界(20d idle,marker)→ CronJob 手动 Job:`evaluated=2 reaped=1`,归档含 marker、PVC+宿主目录回收、`world_backups` 得 `inactive_15d` 行、servers 行/CR 保留 ✅ | `platform/workloads.go` 渲染归档 PVC(minecraft ns、RWO 10Gi、默认 SC),`felis manifests --backup-pvc` 默认 `felis-backups`(`=` 空为显式关闭),bootstrap 统一下传 env/旗标并写 [archive] local_path;`resolveWorldDir` 新增精确 local-path 目录解析(读 PVC `spec.volumeName`,非 glob,杜绝陈旧 PV 目录顶替);ReaperRole 增 `pvc:get`;bootstrap 设 `FELIS_WORLDS_HOST_PATH` 时给 uid 1000 授 traverse(setfacl/o+x);README/故障手册改写。提交 `fd33fd0`+`2b87a5a` | -| 8 | **磁盘打满灾难链**:DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后约 8 分钟才恢复调度。恢复依赖人工 `docker save felis:* | k3s ctr images import -`(docker 存储是唯一副本) | 早前真机 drill;本次复核 k3s 默认驱逐顺序 | 控制面(api/operator/reaper/registry)全部挂 `felis-control-plane` PriorityClass(value 1,000,000,`preemptionPolicy: Never`)——节点压力按优先级升序驱逐,游戏 pod(0)先走,控制面留下;镜像 GC 无法用代码根治(air-gap 无回源),新增 troubleshooting §13b 固化恢复路径(重跑安装器重建导入;单镜像可 `docker save | k3s ctr images import -`)。提交见下 | +| 8 | **磁盘打满灾难链**:DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后数分钟才恢复调度。恢复依赖人工 `docker save | k3s ctr images import -` | 三轮真机 drill(原缺陷复现 + 两轮修复验证):①自定义 1e6 类:游戏 pod 先走、控制面随后仍被驱逐(kubelet 日志逐条列出 ranked/evicted),随后镜像被 GC → ErrImagePull;②内建 `system-cluster-critical`(2e9):填盘至 1.7G free,kubelet 对 felis-api/operator/registry 全部报 **“cannot evict a critical pod”**,三者在整个 DiskPressure 期间保持 Running;游戏 pod(0)被驱逐;③释放磁盘后:`DiskPressure` 经 ~5 分钟(`eviction-pressure-transition-period`)转 False——即“恢复慢”的主因;被 GC 的游戏镜像按 runbook 重导入后 25s 恢复 ✅ | 控制面四类 pod 挂内建 `system-cluster-critical`(用户自定义类值上限 1e9,达不到 2e9 临界阈值;preemption 保留=管理面可调度,已记录权衡);troubleshooting §13b 固化事件链、5 分钟条件过渡与镜像恢复路径。提交见下 | ## 待决策台账(未修) diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 84972e3..e910be8 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -583,13 +583,31 @@ A full disk is the one failure this platform cannot ride out by itself, because the images exist only in the node's containerd (air-gapped by design), so a GC'd image has no pull source. -Eviction ordering. Kubelet's node-pressure eviction removes pods in ascending -priority. Every control-plane pod (api, operator, reaper, registry) carries the -bundle's `felis-control-plane` PriorityClass (value 1,000,000, -`preemptionPolicy: Never`), while game-server pods run at the default 0 — so a -burst of running servers is evicted first and the control plane keeps serving -status/console until pressure is genuinely extreme. The class never *preempts*: -a scheduling decision will not kill a running game server to restart the api. +Eviction. Every control-plane pod (api, operator, reaper, registry) runs under +the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's +node-pressure eviction refuses to touch those pods — the log shows +*"Eviction manager: cannot evict a critical pod"* for each of them — while +game-server pods at the default priority 0 are evicted first. A drill that filled +the disk to 1.7G free saw exactly this: login/lobby evicted, the whole control +plane still Running (before the fix the same drill evicted the api, operator and +registry too, and the image-GC stage below followed). User-defined +PriorityClasses cannot substitute: the API caps them at 1e9, below kubelet's +critical threshold. The built-in class allows preemption (its policy is fixed), +so a control-plane pod that cannot fit may preempt a game pod — deliberate: the +management plane must be placeable. + +The pressure condition clears slowly. After you free space, the node can stay +`DiskPressure:True` for up to ~5 minutes (`--eviction-pressure-transition-period` +defaults to 5m, to stop the condition flapping); pods that need scheduling wait +for it. This is the bulk of the "recovery takes minutes" observation, not a +stuck node. + +But the *game* images can still be GC'd. If game pods were evicted, the kubelet +may garbage-collect their images (unused > 2 minutes under imagefs pressure), and +those pods then sit in `ImagePullBackOff` after recovery — re-import as above +(`docker save felis-limbo:demo felis-lobby:demo | k3s ctr images import -`, then +delete the stuck pods). Verified: both system servers returned to Running in +~25s after the import. Symptoms of the image-GC stage: pods stuck `ImagePullBackOff`/`ErrImagePull` with `kubectl describe pod` showing a pull attempt for a tag that plainly diff --git a/internal/platform/bundle.go b/internal/platform/bundle.go index 419cd8a..9a5713b 100644 --- a/internal/platform/bundle.go +++ b/internal/platform/bundle.go @@ -34,9 +34,10 @@ type Object interface { // Deployments and the in-cluster registry (Deployment + Service + PVC), which // make the SAs and NetworkPolicy peers refer to something real, plus the // world-archive PVC that backs backup/restore when a backup PVC is named (see -// workloads.go). Workloads also carries the one cluster-scoped object, the -// felis-control-plane PriorityClass (node-pressure eviction shield — a -// PriorityClass is not namespaced by design); everything else is namespaced. +// workloads.go). Every object here is namespaced; the control plane's +// node-pressure eviction shield is the BUILT-IN system-cluster-critical +// PriorityClass the pod templates reference (workloads.go controlPlanePriorityName), +// not an object this bundle renders. // The reaper CronJob is also part of Workloads, rendered only when the retention // storage topology is supplied (WorldsHostPath + BackupPVC + ArchiveLocalPath — // workloads.go documents the gate and the shape-asserted hostPath caveat). The diff --git a/internal/platform/workloads.go b/internal/platform/workloads.go index 55d1caf..4d15799 100644 --- a/internal/platform/workloads.go +++ b/internal/platform/workloads.go @@ -7,7 +7,6 @@ import ( appsv1 "k8s.io/api/apps/v1" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" - schedulingv1 "k8s.io/api/scheduling/v1" "k8s.io/apimachinery/pkg/api/resource" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/util/intstr" @@ -177,14 +176,13 @@ func InternalAPIBaseURL(controlNamespace string) string { // felis-operator Deployment, and the in-cluster registry (Deployment + Service + // PVC), the world-archive PVC when p.BackupPVC names it (it backs the // backup/restore Jobs and the reaper), plus the reaper CronJob when -// reaperEnabled(p). Every pod template carries the felis-control-plane -// PriorityClass (controlPlanePriorityClass below), the node-pressure eviction -// shield. FelisImage is required — `felis manifests` enforces it (fail-loud), so -// a rendered bundle always names a concrete image. +// reaperEnabled(p). Every pod template carries the built-in +// system-cluster-critical PriorityClass (controlPlanePriorityName), the +// node-pressure eviction shield. FelisImage is required — `felis manifests` +// enforces it (fail-loud), so a rendered bundle always names a concrete image. func Workloads(p Params) []Object { p = p.withDefaults() objs := []Object{ - controlPlanePriorityClass(), APIDeployment(p), apiService(p), apiInternalService(p), @@ -204,35 +202,30 @@ func Workloads(p Params) []Object { } // controlPlanePriorityName is the PriorityClass every control-plane workload runs -// under (api, operator, reaper, registry). Node-pressure eviction (a full disk -// being the realistic case on a game box) removes pods in ASCENDING priority, and -// a classless pod is priority 0 — the same as the game servers, which are exactly -// the pods a burst of joins has just filled the node with. A full disk then takes -// the api/operator down too, and recovery needs a human: with no reachable -// registry on an air-gapped box the images are gone, so the fix is a re-run of the -// installer to re-import them. The class is an eviction shield, not a preemption -// lever: PreemptionPolicy=Never, so a busy node never loses a running game server -// merely to schedule the api. Value 1,000,000 sits above every game pod (0) and -// far below the kubelet's system-reserved classes (2,000,000,000). -const ( - controlPlanePriorityName = "felis-control-plane" - controlPlanePriorityValue int32 = 1_000_000 - controlPlanePriorityReason = "Felis control plane (api/operator/reaper/registry) survives node-pressure eviction before game servers" -) - -// controlPlanePriorityClass renders the cluster-scoped PriorityClass the -// control-plane pod templates reference (the ONE cluster-scoped object in the -// bundle; a PriorityClass is not namespaced by design). -func controlPlanePriorityClass() *schedulingv1.PriorityClass { - never := corev1.PreemptNever - return &schedulingv1.PriorityClass{ - TypeMeta: metav1.TypeMeta{APIVersion: "scheduling.k8s.io/v1", Kind: "PriorityClass"}, - ObjectMeta: metav1.ObjectMeta{Name: controlPlanePriorityName}, - Value: controlPlanePriorityValue, - PreemptionPolicy: &never, - Description: controlPlanePriorityReason, - } -} +// under (api, operator, reaper, registry): the BUILT-IN system-cluster-critical, +// not a class we render. Two live findings force it: +// +// - ordering alone is not enough. With the disk at k3s's eviction threshold +// (nodefs.available<5%) a drill watched the kubelet evict the game pods first +// — and then the api, operator and registry too, because evicting them was +// never what would reclaim the disk. The images live only in the node's +// containerd (air-gapped), so once their pods were gone the kubelet GC'd them +// and recovery degenerated into ImagePullBackOff plus a manual re-import. +// - a class we render cannot reach the critical threshold: user-defined +// PriorityClasses are capped at 1e9, while kubelet's +// `IsCriticalPodBasedOnPriority` — and with it the eviction manager's +// "cannot evict a critical pod" refusal — needs ≥ 2e9. Only the built-in +// classes carry those values. +// +// system-cluster-critical keeps the panel/console alive through a full-disk +// incident (which is what lets an operator see it and act) and makes the control +// plane schedule before game pods after a reboot. Its preemptionPolicy is +// PreemptLowerPriority (built-in classes fix it; the API rejects a per-pod +// override), so a control-plane pod that cannot fit MAY preempt a game pod: +// accepted deliberately — the management plane must be placeable, and the +// alternative is an unreachable box. Game pods stay at the default 0 and are +// still the first casualties of node pressure. +const controlPlanePriorityName = "system-cluster-critical" // reaperEnabled reports whether the retention CronJob should render. It needs all // three storage coordinates: WorldsHostPath (where worlds live, mounted to read diff --git a/internal/platform/workloads_test.go b/internal/platform/workloads_test.go index 1e60414..f498109 100644 --- a/internal/platform/workloads_test.go +++ b/internal/platform/workloads_test.go @@ -6,7 +6,6 @@ import ( appsv1 "k8s.io/api/apps/v1" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" - schedulingv1 "k8s.io/api/scheduling/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/labels" ) @@ -460,8 +459,8 @@ func TestRegistry_DeploymentServicePVC(t *testing.T) { // internal Service must be present or the login pod's felis-api:8081 path is dead. func TestWorkloads_BundleContents(t *testing.T) { objs := Workloads(testParams()) - if len(objs) != 9 { - t.Fatalf("Workloads returned %d objects, want 9", len(objs)) + if len(objs) != 8 { + t.Fatalf("Workloads returned %d objects, want 8", len(objs)) } var haveInternalSvc bool for _, o := range objs { @@ -479,24 +478,20 @@ func TestWorkloads_BundleContents(t *testing.T) { } // TestControlPlanePriorityClass pins the node-pressure eviction shield: every -// control-plane pod template (api/operator/reaper/registry) references the one -// PriorityClass, which outranks the default-0 game pods by eviction order while -// never preempting them, and is the bundle's only cluster-scoped object. Without -// this, a full disk evicts the api alongside the game pods and (no reachable -// registry on an air-gapped box) recovery needs a human re-importing images. +// control-plane pod template (api/operator/reaper/registry) runs under the +// BUILT-IN system-cluster-critical class (value 2e9), at which kubelet's +// eviction manager refuses to evict the pod. A live drill showed the whole +// cascade with plain ordering: disk pressure evicted the game pods and then the +// control plane, whose images (air-gapped, containerd-only) were GC'd → +// ImagePullBackOff plus a manual re-import. User-defined classes are capped at +// 1e9 (API-enforced), so the built-in class is the only way to reach the +// critical threshold. func TestControlPlanePriorityClass(t *testing.T) { objs := Workloads(reaperParams()) // include the reaper CronJob's pod template - var class *schedulingv1.PriorityClass - clusterScoped := 0 deployments, cronJobs := 0, 0 for _, o := range objs { - if o.GetNamespace() == "" { - clusterScoped++ - } switch v := o.(type) { - case *schedulingv1.PriorityClass: - class = v case *appsv1.Deployment: deployments++ if v.Spec.Template.Spec.PriorityClassName != controlPlanePriorityName { @@ -509,25 +504,13 @@ func TestControlPlanePriorityClass(t *testing.T) { } } } - if class == nil { - t.Fatal("Workloads must render the control-plane PriorityClass") - } - if class.Name != controlPlanePriorityName || class.Value != controlPlanePriorityValue { - t.Errorf("class = %s/%d, want %s/%d", class.Name, class.Value, controlPlanePriorityName, controlPlanePriorityValue) - } - if class.PreemptionPolicy == nil || *class.PreemptionPolicy != corev1.PreemptNever { - t.Errorf("class preemptionPolicy = %v, want Never (an eviction shield, never a lever against running game servers)", class.PreemptionPolicy) - } - if clusterScoped != 1 { - t.Errorf("bundle has %d cluster-scoped objects, want exactly the PriorityClass", clusterScoped) - } if deployments != 3 || cronJobs != 1 { t.Errorf("scanned %d deployments / %d cronjobs, want 3 / 1 — a pod template escaped the class check", deployments, cronJobs) } - // The class must outrank the default game-pod priority (0); equality would make - // the eviction order nondeterministic between the api and a full game server. - if controlPlanePriorityValue <= 0 { - t.Errorf("control-plane priority %d must exceed the default 0 game-pod priority", controlPlanePriorityValue) + // The name must be the built-in critical class: any custom class is capped at + // 1e9 by the API server and would be evictable. + if controlPlanePriorityName != "system-cluster-critical" { + t.Errorf("control-plane priority class = %q; only the built-in critical classes reach the 2e9 eviction-refusal threshold", controlPlanePriorityName) } }