From fe310743a29a2871d37307928fdf143c115883e9 Mon Sep 17 00:00:00 2001 From: Lemon-miaow Date: Tue, 22 Sep 2026 21:08:56 +0800 Subject: [PATCH] fix(platform): give the control plane a PriorityClass eviction shield (#8) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A full disk made kubelet's node-pressure eviction pick control-plane pods alongside game pods (both priority 0), and with the images existing only in the node's containerd (air-gapped), losing the api meant a manual image re-import. Every control-plane pod template (api/operator/reaper/ registry) now names the bundle's cluster-scoped felis-control-plane PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order only, never preempting a running game server. The image-GC half is not code-fixable on an air-gapped box; troubleshooting gains 13b with the recovery path (re-run the installer to rebuild imports, or docker save | k3s ctr images import - for one image). --- AUDIT-2026-09-22.md | 4 +- docs/troubleshooting.md | 37 ++++++++++++++++++ internal/platform/bundle.go | 4 +- internal/platform/workloads.go | 42 ++++++++++++++++++++- internal/platform/workloads_test.go | 58 ++++++++++++++++++++++++++++- 5 files changed, 138 insertions(+), 7 deletions(-) diff --git a/AUDIT-2026-09-22.md b/AUDIT-2026-09-22.md index 8c9cfe3..fdfba26 100644 --- a/AUDIT-2026-09-22.md +++ b/AUDIT-2026-09-22.md @@ -17,14 +17,14 @@ IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:f | 17 | **`ListPendingOpLogins` PG 少列**:接口注释承诺「joined to its staff username」,handler 输出 `username`/`created_at`,fake 正确填充;PG SQL 未 JOIN 也未取 `created_at` → 真机 pending 列表 username 为空、created_at 为 `0001-01-01` | 真机 internal `/op-login/pending` 响应 | `pgrepo.go` SQL 改为 JOIN users + 取 created_at。提交见分支 | | 18 | **绑定码并发兑换 500**:`RedeemPlayerBindCode` 无 `FOR UPDATE`(同文件 `VerifyLinkCode` 有),且裸 INSERT。6 路并发同码兑换 → 3×HTTP 500(`users_username_key` 唯一冲突)+ 1×400 + 2×200;跨码并发同 UUID 同样会撞 | 真机四组并发测试 + API 日志 `unmapped error ... duplicate key` | 同码:码行 `FOR UPDATE`(输家干净地 400 invalid_code);跨码:`INSERT users ... ON CONFLICT (username) DO NOTHING`+重读、`account_links ON CONFLICT (mc_uuid) DO NOTHING`(两路汇聚同一 user)。复测:同码×6=1×200+5×400、双码×2=2×200 同 ID、DB 干净、日志零 unmapped | | 19 | **reaper CronJob 渲染位置错误,永远无法调度**:`felis manifests` 把 CronJob 渲染到控制 ns,却引用 minecraft ns 的备份 PVC(Pod 不能跨 ns 挂 PVC:真机 `FailedScheduling: persistentvolumeclaim "felis-backups" not found`);修正 ns 后又发现 ServerAccount 也不能跨 ns 使用(`serviceaccount "felis-reaper" not found`)。而 reaper 的 Role/RoleBinding 本就在 minecraft ns | 真机三层取证(PVC/SA/调度)| CronJob 与 SA、RoleBinding subject 全部移到 `MinecraftNamespace`(提交 `e4f2cff`+`c839454`)。**修复后完整演练通过**(见下) | +| 6 | **默认安装无备份能力 + 回收链路不可达/不可用**:(a) 没有任何环节渲染归档 PVC,`FELIS_BACKUP_PVC` 永远为空 → backup/restore 恒 503;(b) bootstrap 从不下传 reaper 三旗标 → 官方安装路径根本无法启用回收;(c) 即便启用,stock k3s 的 `/` 布局不存在,且 k3s storage root 是 `0700 root:root`、reaper pod 以 uid 1000 运行 → 真机 `lstat /worlds/…: permission denied`(fail-closed 跳过,但纯空转);(d) README 口径(“定时备份”)不实 | 真机全链路:默认渲染 PVC+env → 建服→marker→备份 202→jobs 端点 running→succeeded→改 marker→restore→读回原值 ✅;再建 resolvecheck 世界(20d idle,marker)→ CronJob 手动 Job:`evaluated=2 reaped=1`,归档含 marker、PVC+宿主目录回收、`world_backups` 得 `inactive_15d` 行、servers 行/CR 保留 ✅ | `platform/workloads.go` 渲染归档 PVC(minecraft ns、RWO 10Gi、默认 SC),`felis manifests --backup-pvc` 默认 `felis-backups`(`=` 空为显式关闭),bootstrap 统一下传 env/旗标并写 [archive] local_path;`resolveWorldDir` 新增精确 local-path 目录解析(读 PVC `spec.volumeName`,非 glob,杜绝陈旧 PV 目录顶替);ReaperRole 增 `pvc:get`;bootstrap 设 `FELIS_WORLDS_HOST_PATH` 时给 uid 1000 授 traverse(setfacl/o+x);README/故障手册改写。提交 `fd33fd0`+`2b87a5a` | +| 8 | **磁盘打满灾难链**:DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后约 8 分钟才恢复调度。恢复依赖人工 `docker save felis:* | k3s ctr images import -`(docker 存储是唯一副本) | 早前真机 drill;本次复核 k3s 默认驱逐顺序 | 控制面(api/operator/reaper/registry)全部挂 `felis-control-plane` PriorityClass(value 1,000,000,`preemptionPolicy: Never`)——节点压力按优先级升序驱逐,游戏 pod(0)先走,控制面留下;镜像 GC 无法用代码根治(air-gap 无回源),新增 troubleshooting §13b 固化恢复路径(重跑安装器重建导入;单镜像可 `docker save | k3s ctr images import -`)。提交见下 | ## 待决策台账(未修) | # | 主题 | 说明 | |---|------|------| -| 6 | 默认安装无备份能力 | backup/restore 端点默认 503(需 `FELIS_BACKUP_PVC`+PVC),reaper CronJob 需 `--backup-pvc/--archive-local-path/--worlds-host-path` 三旗标渲染,bootstrap 一个都不传;文档无说明;README 与"自动备份"口径不符。另:归档 3 个月过期依赖 reaper 清理,未启用则磁盘只增不减 | | 7 | 异步失败不可感知 | backup/restore 失败后无状态出口:restore 行不变、backup 无行;只有集群侧 Job/日志可查。建议状态字段或 `?failed` 查询 | -| 8 | 磁盘打满灾难链 | DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后约 8 分钟才恢复调度。恢复靠 `docker save felis:* | k3s ctr images import -`(docker 守护进程存储是唯一副本,需固化回源路径)。建议:PriorityClass、镜像入内置 registry、磁盘告警 | | 9 | 升级策略 Recreate | 单副本 + Recreate:任何控制面升级=停机;坏升级(实测错 tag)服务中断约 95s 且需人工 `rollout undo`(无自动回滚)。建议 runbook/文档化 | | 10 | 备份语义 | 归档包含整个 /data(jar、libraries、cache),167MB;是否符合"world backup"定位待评估 | | 11 | PG 断连表现 | 会话查询失败报 401 而非 503(fail-closed 但误导;用户以为没登录) | diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 35375a5..84972e3 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -577,6 +577,43 @@ changes, because retention was never conditional in the first place. --- +## 13b. Node runs out of disk: what survives, and how to recover + +A full disk is the one failure this platform cannot ride out by itself, because +the images exist only in the node's containerd (air-gapped by design), so a +GC'd image has no pull source. + +Eviction ordering. Kubelet's node-pressure eviction removes pods in ascending +priority. Every control-plane pod (api, operator, reaper, registry) carries the +bundle's `felis-control-plane` PriorityClass (value 1,000,000, +`preemptionPolicy: Never`), while game-server pods run at the default 0 — so a +burst of running servers is evicted first and the control plane keeps serving +status/console until pressure is genuinely extreme. The class never *preempts*: +a scheduling decision will not kill a running game server to restart the api. + +Symptoms of the image-GC stage: pods stuck `ImagePullBackOff`/`ErrImagePull` +with `kubectl describe pod` showing a pull attempt for a tag that plainly +exists (`k3s ctr images ls` will show it missing — the kubelet GC removed it +under imagefs pressure). + +Recovery: + +1. Free disk on the node (`df -h /var/lib/rancher`, the biggest consumers are + `k3s ctr images ls -q` and the world/backup PVCs under + `/var/lib/rancher/k3s/storage`). +2. Re-import the images by re-running the installer (it rebuilds imports from + the local Docker store, which the kubelet GC does not touch): + `curl -fsSL | sudo bash` (or `sudo felis setup`), then + `kubectl -n felis rollout status deploy/felis-api`. +3. Delete now-unschedulable stuck pods so they retry with the re-imported image. + +If the API itself is down and you only need the images back without a full +installer run: `docker save felis: | k3s ctr images import -` restores one +image from the Docker store (that store is deliberately a second copy; treat it +as the recovery path, not as free space). + +--- + ## 14. Metrics for diagnosis (spec §23) All four mandated metrics have real producers; scrape them when triaging: diff --git a/internal/platform/bundle.go b/internal/platform/bundle.go index bc27a97..419cd8a 100644 --- a/internal/platform/bundle.go +++ b/internal/platform/bundle.go @@ -34,7 +34,9 @@ type Object interface { // Deployments and the in-cluster registry (Deployment + Service + PVC), which // make the SAs and NetworkPolicy peers refer to something real, plus the // world-archive PVC that backs backup/restore when a backup PVC is named (see -// workloads.go). +// workloads.go). Workloads also carries the one cluster-scoped object, the +// felis-control-plane PriorityClass (node-pressure eviction shield — a +// PriorityClass is not namespaced by design); everything else is namespaced. // The reaper CronJob is also part of Workloads, rendered only when the retention // storage topology is supplied (WorldsHostPath + BackupPVC + ArchiveLocalPath — // workloads.go documents the gate and the shape-asserted hostPath caveat). The diff --git a/internal/platform/workloads.go b/internal/platform/workloads.go index b4819ce..55d1caf 100644 --- a/internal/platform/workloads.go +++ b/internal/platform/workloads.go @@ -7,6 +7,7 @@ import ( appsv1 "k8s.io/api/apps/v1" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" + schedulingv1 "k8s.io/api/scheduling/v1" "k8s.io/apimachinery/pkg/api/resource" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/util/intstr" @@ -176,11 +177,14 @@ func InternalAPIBaseURL(controlNamespace string) string { // felis-operator Deployment, and the in-cluster registry (Deployment + Service + // PVC), the world-archive PVC when p.BackupPVC names it (it backs the // backup/restore Jobs and the reaper), plus the reaper CronJob when -// reaperEnabled(p). FelisImage is required — `felis manifests` enforces it -// (fail-loud), so a rendered bundle always names a concrete image. +// reaperEnabled(p). Every pod template carries the felis-control-plane +// PriorityClass (controlPlanePriorityClass below), the node-pressure eviction +// shield. FelisImage is required — `felis manifests` enforces it (fail-loud), so +// a rendered bundle always names a concrete image. func Workloads(p Params) []Object { p = p.withDefaults() objs := []Object{ + controlPlanePriorityClass(), APIDeployment(p), apiService(p), apiInternalService(p), @@ -199,6 +203,37 @@ func Workloads(p Params) []Object { return objs } +// controlPlanePriorityName is the PriorityClass every control-plane workload runs +// under (api, operator, reaper, registry). Node-pressure eviction (a full disk +// being the realistic case on a game box) removes pods in ASCENDING priority, and +// a classless pod is priority 0 — the same as the game servers, which are exactly +// the pods a burst of joins has just filled the node with. A full disk then takes +// the api/operator down too, and recovery needs a human: with no reachable +// registry on an air-gapped box the images are gone, so the fix is a re-run of the +// installer to re-import them. The class is an eviction shield, not a preemption +// lever: PreemptionPolicy=Never, so a busy node never loses a running game server +// merely to schedule the api. Value 1,000,000 sits above every game pod (0) and +// far below the kubelet's system-reserved classes (2,000,000,000). +const ( + controlPlanePriorityName = "felis-control-plane" + controlPlanePriorityValue int32 = 1_000_000 + controlPlanePriorityReason = "Felis control plane (api/operator/reaper/registry) survives node-pressure eviction before game servers" +) + +// controlPlanePriorityClass renders the cluster-scoped PriorityClass the +// control-plane pod templates reference (the ONE cluster-scoped object in the +// bundle; a PriorityClass is not namespaced by design). +func controlPlanePriorityClass() *schedulingv1.PriorityClass { + never := corev1.PreemptNever + return &schedulingv1.PriorityClass{ + TypeMeta: metav1.TypeMeta{APIVersion: "scheduling.k8s.io/v1", Kind: "PriorityClass"}, + ObjectMeta: metav1.ObjectMeta{Name: controlPlanePriorityName}, + Value: controlPlanePriorityValue, + PreemptionPolicy: &never, + Description: controlPlanePriorityReason, + } +} + // reaperEnabled reports whether the retention CronJob should render. It needs all // three storage coordinates: WorldsHostPath (where worlds live, mounted to read // them), BackupPVC (where archives are written) and ArchiveLocalPath (the mount @@ -509,6 +544,7 @@ func reaperCronJob(p Params) *batchv1.CronJob { ObjectMeta: metav1.ObjectMeta{Labels: labels}, Spec: corev1.PodSpec{ ServiceAccountName: SAReaper, + PriorityClassName: controlPlanePriorityName, RestartPolicy: corev1.RestartPolicyNever, SecurityContext: hardenedPodSecurityContext(), Containers: []corev1.Container{container}, @@ -544,6 +580,7 @@ func controlPlaneDeployment(p Params, sa string, container corev1.Container, vol ObjectMeta: metav1.ObjectMeta{Labels: labels}, Spec: corev1.PodSpec{ ServiceAccountName: sa, + PriorityClassName: controlPlanePriorityName, SecurityContext: hardenedPodSecurityContext(), Containers: []corev1.Container{container}, Volumes: volumes, @@ -591,6 +628,7 @@ func registryDeployment(p Params) *appsv1.Deployment { ObjectMeta: metav1.ObjectMeta{Labels: labels}, Spec: corev1.PodSpec{ AutomountServiceAccountToken: boolPtr(false), + PriorityClassName: controlPlanePriorityName, SecurityContext: hardenedPodSecurityContext(), Containers: []corev1.Container{container}, Volumes: []corev1.Volume{ diff --git a/internal/platform/workloads_test.go b/internal/platform/workloads_test.go index ad0c8e2..1e60414 100644 --- a/internal/platform/workloads_test.go +++ b/internal/platform/workloads_test.go @@ -6,6 +6,7 @@ import ( appsv1 "k8s.io/api/apps/v1" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" + schedulingv1 "k8s.io/api/scheduling/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/labels" ) @@ -459,8 +460,8 @@ func TestRegistry_DeploymentServicePVC(t *testing.T) { // internal Service must be present or the login pod's felis-api:8081 path is dead. func TestWorkloads_BundleContents(t *testing.T) { objs := Workloads(testParams()) - if len(objs) != 8 { - t.Fatalf("Workloads returned %d objects, want 8", len(objs)) + if len(objs) != 9 { + t.Fatalf("Workloads returned %d objects, want 9", len(objs)) } var haveInternalSvc bool for _, o := range objs { @@ -477,6 +478,59 @@ func TestWorkloads_BundleContents(t *testing.T) { } } +// TestControlPlanePriorityClass pins the node-pressure eviction shield: every +// control-plane pod template (api/operator/reaper/registry) references the one +// PriorityClass, which outranks the default-0 game pods by eviction order while +// never preempting them, and is the bundle's only cluster-scoped object. Without +// this, a full disk evicts the api alongside the game pods and (no reachable +// registry on an air-gapped box) recovery needs a human re-importing images. +func TestControlPlanePriorityClass(t *testing.T) { + objs := Workloads(reaperParams()) // include the reaper CronJob's pod template + + var class *schedulingv1.PriorityClass + clusterScoped := 0 + deployments, cronJobs := 0, 0 + for _, o := range objs { + if o.GetNamespace() == "" { + clusterScoped++ + } + switch v := o.(type) { + case *schedulingv1.PriorityClass: + class = v + case *appsv1.Deployment: + deployments++ + if v.Spec.Template.Spec.PriorityClassName != controlPlanePriorityName { + t.Errorf("deployment %s pod priority = %q, want %q", v.Name, v.Spec.Template.Spec.PriorityClassName, controlPlanePriorityName) + } + case *batchv1.CronJob: + cronJobs++ + if v.Spec.JobTemplate.Spec.Template.Spec.PriorityClassName != controlPlanePriorityName { + t.Errorf("CronJob %s pod priority = %q, want %q", v.Name, v.Spec.JobTemplate.Spec.Template.Spec.PriorityClassName, controlPlanePriorityName) + } + } + } + if class == nil { + t.Fatal("Workloads must render the control-plane PriorityClass") + } + if class.Name != controlPlanePriorityName || class.Value != controlPlanePriorityValue { + t.Errorf("class = %s/%d, want %s/%d", class.Name, class.Value, controlPlanePriorityName, controlPlanePriorityValue) + } + if class.PreemptionPolicy == nil || *class.PreemptionPolicy != corev1.PreemptNever { + t.Errorf("class preemptionPolicy = %v, want Never (an eviction shield, never a lever against running game servers)", class.PreemptionPolicy) + } + if clusterScoped != 1 { + t.Errorf("bundle has %d cluster-scoped objects, want exactly the PriorityClass", clusterScoped) + } + if deployments != 3 || cronJobs != 1 { + t.Errorf("scanned %d deployments / %d cronjobs, want 3 / 1 — a pod template escaped the class check", deployments, cronJobs) + } + // The class must outrank the default game-pod priority (0); equality would make + // the eviction order nondeterministic between the api and a full game server. + if controlPlanePriorityValue <= 0 { + t.Errorf("control-plane priority %d must exceed the default 0 game-pod priority", controlPlanePriorityValue) + } +} + // cronPodSpec returns the single container and pod template of a CronJob's Job // template, failing if the shape is not a single-container pod (the reaper's shape). func cronPodSpec(t *testing.T, cj *batchv1.CronJob) (corev1.PodSpec, corev1.Container) {