diff --git a/AUDIT-2026-09-22.md b/AUDIT-2026-09-22.md new file mode 100644 index 0000000..ec72bc7 --- /dev/null +++ b/AUDIT-2026-09-22.md @@ -0,0 +1,49 @@ +# Felis 生产就绪审计 — 2026-09-22(真机 E2E + 混沌) + +分支:`audit-fixes-20260922`(已推送)。环境:CentOS Stream 9 / aarch64 / k3s v1.36.4, +IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:fede:69ec`), +面板经 `socat TCP6:443 → 127.0.0.1:30443` 中继(手动启动,重启 VM 后需重开)。 + +## 已修复并验证(分支内) + +| # | 缺陷 | 证据 | 修复 | +|---|------|------|------| +| 1 | **失败恢复后重试被静默吞掉**:restore Job 固定名 + `ErrAlreadyExists` 一律当"幂等成功";失败 Job 占名 10 分钟(TTL),期间重试返回 202 `restoring` 但什么都不跑 | 真机:坏 ref 制造失败 → 立刻合法重试 → Job 原地不动、无新 Pod | `k8sjobs.go`:撞名时检查已完成(成功/失败)→ 删除+**等 finalizer 释放**(有界 10s)+ 重建;进行中仍幂等吸收。单测 3 个。**真机复验:重试 4s 完成恢复** | +| 2 | **备份 Job 全部 FailedMount**:Job 在 minecraft 命名空间挂 `felis-config`,而 bootstrap 只在 felis 命名空间创建该 Secret | 事件:`MountVolume.SetUp failed: secret "felis-config" not found` | `felis setup` 用既有 `ensureSecretReplica` 把 felis-config(`felis.toml`) 复制到 minecraft;VM 上手工复制后备份成功(167MB 归档) | +| 3 | RBAC 缺 `jobs:get/delete`(修复 #1 需要) | Role 检查 | `APIMinecraftRole` jobs → create/get/delete,测试同步更新 | +| 4 | gofmt 9 文件漂移 + CI 无 gofmt 门禁;staticcheck 9 处 | 基线扫描 | 全部修复;CI 加 gofmt job | +| 5 | 依赖漏洞:pgx v5.7.1(GO-2026-5004,可达 pgrepo.go)、x/net、x/text | govulncheck | 升 pgx v5.9.2 / x/net v0.55.0 / x/text v0.39.0 | +| 16 | **`ConsumeLoginEmailOTP` PG 实现与接口契约漂移**:契约/`fakeRepo` 说「同 VerifyEmailOTP 生命周期(扣尝试/锁定/ErrOTPInvalid/ErrOTPLocked)」,PG 却是单条 UPDATE+`ErrNotFound` → 错码/重放/过期在 login 门、op-login finish、migration confirm 三个入口全部 500;且尝试次数永不累计、`otpMaxAttempts` 锁定失效 | 真机:login 门重放正确码 → **HTTP 500**;修复前错码不扣次。修复后复测:5 次错码 400 且 attempts=5(正确码因锁定也 400、码未消费)、新码可用、重放 400 | `pgrepo.go` 改置为 VerifyEmailOTP 同构事务(FOR UPDATE、先锁后比、mismatch 扣次、match 消费),无 users 写副作用。提交 `52549f7` | +| 17 | **`ListPendingOpLogins` PG 少列**:接口注释承诺「joined to its staff username」,handler 输出 `username`/`created_at`,fake 正确填充;PG SQL 未 JOIN 也未取 `created_at` → 真机 pending 列表 username 为空、created_at 为 `0001-01-01` | 真机 internal `/op-login/pending` 响应 | `pgrepo.go` SQL 改为 JOIN users + 取 created_at。提交见分支 | + +## 待决策台账(未修) + +| # | 主题 | 说明 | +|---|------|------| +| 6 | 默认安装无备份能力 | backup/restore 端点默认 503(需 `FELIS_BACKUP_PVC`+PVC),reaper CronJob 需 `--backup-pvc/--archive-local-path/--worlds-host-path` 三旗标渲染,bootstrap 一个都不传;文档无说明;README 与"自动备份"口径不符。另:归档 3 个月过期依赖 reaper 清理,未启用则磁盘只增不减 | +| 7 | 异步失败不可感知 | backup/restore 失败后无状态出口:restore 行不变、backup 无行;只有集群侧 Job/日志可查。建议状态字段或 `?failed` 查询 | +| 8 | 磁盘打满灾难链 | DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后约 8 分钟才恢复调度。恢复靠 `docker save felis:* | k3s ctr images import -`(docker 守护进程存储是唯一副本,需固化回源路径)。建议:PriorityClass、镜像入内置 registry、磁盘告警 | +| 9 | 升级策略 Recreate | 单副本 + Recreate:任何控制面升级=停机;坏升级(实测错 tag)服务中断约 95s 且需人工 `rollout undo`(无自动回滚)。建议 runbook/文档化 | +| 10 | 备份语义 | 归档包含整个 /data(jar、libraries、cache),167MB;是否符合"world backup"定位待评估 | +| 11 | PG 断连表现 | 会话查询失败报 401 而非 503(fail-closed 但误导;用户以为没登录) | +| 12 | ready 门滞后 | 容器 Ready 后 6~10s 内 API 仍 409 not_running | +| 13 | setup token 截断 | 43 字符 token + 长域名,80 列终端下 TUI 截断显示(复现:tmux 80 列) | +| 14 | 日志噪音 | controller-runtime 未 SetLogger,首用打印整段堆栈;TLS handshake EOF 噪音(kubelet 探针) | +| 15 | reaper 启用未演练 | worlds-host-path/PVC 与命名空间的耦合(control ns CronJob 挂 minecraft ns PVC)需一并与 #6 验收 | + +## 已验证事实(正向清单) + +- 安装→hook 发码→Owner 绑定→passkey(虚拟认证器)→面板管理员全链路 ✅ +- 建服(POST /servers)→ 唤醒(operator 拉 StatefulSet pod)→ RCON `list` → SSE 控制台 → 停止 ✅ +- 文件编辑:列目录/读/写(wire 为 base64)/256KiB 413/路径穿越 5 变体全拦截/运行中 409 ✅ +- **备份→篡改→恢复数据演练**:v1→备份→v2→恢复→读回 v1 ✅(G2 数据可恢复) +- **毒档案 fail-closed**:穿越/绝对路径/符号链接条目 → `archive entry escapes target` 退出码 1,零写入 ✅ +- 混沌:PG 掉线(healthz 仍 200、恢复后连接池自愈)、API pod 击杀(~2s 中断)、整机重启(32s 回归、会话/CRD/停止态全保留)✅ +- `felis update` 报告(k3s/velocity 有更新、私有仓库 404 优雅处理)✅ + +## 复现入口速查 + +- 面板会话 cookie:`/tmp/felis-cookies.json`;API 助手:`/tmp/fcurl.sh` +- 测试服:`test-one`(minecraft ns,stopped);合法备份 `bk-47ee2e7e96e5a4ca9d0e51b805518bac` +- CDP 调试口:Mac `127.0.0.1:9333`(独立 Chrome,profile `/tmp/felis-chrome2`) +- 分支已部署到 VM:`felis-api` 镜像 = `felis:auditfix2`(含全部修复) diff --git a/internal/api/pgrepo.go b/internal/api/pgrepo.go index 25960ed..becf772 100644 --- a/internal/api/pgrepo.go +++ b/internal/api/pgrepo.go @@ -1981,12 +1981,14 @@ func (p *PGRepo) OpLoginRequestByID(ctx context.Context, id string) (*OpLoginReq // ListPendingOpLogins returns the live (pending, unconsumed, unexpired at now) // requests oldest-first, for the in-game admin's approval prompt. A resolved or // expired request drops out of the list, so an admin only ever sees actionable -// attempts. +// attempts. The username is joined because the approval prompt names the staff +// account; created_at orders the list and lets the prompt show how long a request +// has been waiting. func (p *PGRepo) ListPendingOpLogins(ctx context.Context, now time.Time) ([]OpLoginRequest, error) { - const q = `SELECT id, user_id, email, expires_at - FROM op_login_requests - WHERE consumed_at IS NULL AND approved_at IS NULL AND expires_at > $1 - ORDER BY created_at` + const q = `SELECT r.id, r.user_id, u.username, r.email, r.expires_at, r.created_at + FROM op_login_requests r JOIN users u ON u.id = r.user_id + WHERE r.consumed_at IS NULL AND r.approved_at IS NULL AND r.expires_at > $1 + ORDER BY r.created_at` rows, err := p.db.QueryContext(ctx, q, now) if err != nil { return nil, err @@ -1995,7 +1997,7 @@ func (p *PGRepo) ListPendingOpLogins(ctx context.Context, now time.Time) ([]OpLo var out []OpLoginRequest for rows.Next() { var r OpLoginRequest - if err := rows.Scan(&r.ID, &r.UserID, &r.Email, &r.ExpiresAt); err != nil { + if err := rows.Scan(&r.ID, &r.UserID, &r.Username, &r.Email, &r.ExpiresAt, &r.CreatedAt); err != nil { return nil, err } r.Status = "pending"