Author SHA1 Message Date
dependabot[bot] a252f04b0f chore(deps): bump actions/setup-go from 5.6.0 to 7.0.0
Bumps [actions/setup-go](https://github.com/actions/setup-go) from 5.6.0 to 7.0.0.
- [Release notes](https://github.com/actions/setup-go/releases)
- [Commits](https://github.com/actions/setup-go/compare/40f1582b2485089dde7abd97c1529aa768e1baff...b7ad1dad31e06c5925ef5d2fc7ad053ef454303e)

---
updated-dependencies:
- dependency-name: actions/setup-go
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <[email protected]>
2026-09-24 16:45:30 +00:00
Lemon-miaow 3e61fde06d fix(bootstrap): 关闭 BuildKit 默认 attestation,无改动重跑不再重启 login/lobby;同版本重跑同时重启 registry gate 2026-09-25 00:43:02 +08:00
Lemon-miaow f0ac5b48ce fix(cfsetup): 隧道凭据文件校验 JSON 与 TunnelID,损坏或不匹配时移开并重新获取,失败页提示重试可收敛 2026-09-25 00:37:07 +08:00
Lemon-miaow 1141ebcd49 ci: 加 -race、staticcheck、govulncheck、shellcheck 与真 PostgreSQL 集成测试门禁,release 复用 ci 门禁 2026-09-25 00:35:09 +08:00
Lemon-miaow 82545548e7 feat(bootstrap): game stack 按 lock 文件固定构建并校验 sha256,基础镜像按 digest 固定,JRE 固定补丁版本 2026-09-25 00:24:02 +08:00
Lemon-miaow aace66e8d3 fix(bootstrap): 重跑只重启有变化的 Velocity、系统服与 PostgreSQL,无 firewalld 时用 nftables 收口 5432,PG 大版本不符拒绝启动 2026-09-25 00:06:47 +08:00
Lemon-miaow b1627e9ba0 style(pgint): gofmt 2026-09-24 23:52:30 +08:00
Lemon-miaow a4fa16ae69 feat(bootstrap): 控制面镜像按版本打 tag,升级后 rollout undo 可回到上一版本 2026-09-24 23:52:30 +08:00
Lemon-miaow 80fbc779d9 feat(operator): 平台升级不再重启运行中的游戏服,registry 清理保留游戏 pod 在用的镜像 2026-09-24 23:52:30 +08:00
Lemon-miaow 07eb682358 docs(upgrade): 列出安装器对下载物的校验与离线导入 registry 镜像的做法 2026-09-24 23:39:39 +08:00
Lemon-miaow 3ad817eeff fix(update): 重跑命令从最新 release tag 读取安装脚本 2026-09-24 23:39:39 +08:00
Lemon-miaow 26dc1722f4 feat(bootstrap): 先校验 SHA256SUMS 再运行下载的 felis,固定 k3s、cloudflared 与 registry 镜像版本 2026-09-24 23:39:38 +08:00
Lemon-miaow c4a4f1f25c ci(release): 发布 SHA256SUMS、SBOM 与构建来源证明,action 固定到 commit,不再覆盖已发布资产 2026-09-24 23:39:38 +08:00
Lemon-miaow 8296153a38 docs(build): 说明构建工具镜像、定时刷新与离线环境做法 2026-09-24 23:25:21 +08:00
Lemon-miaow bf18712a6c feat(build): kaniko/trivy 与扫描库改用 registry 内的 mirror 副本,定时刷新并在过期时告警 2026-09-24 23:25:21 +08:00
Lemon-miaow c0643af118 feat(imagepush): 从公共 registry 拷贝单平台镜像与 OCI 制品 2026-09-24 23:25:04 +08:00
Lemon-miaow 80c4ea8966 docs(registry): 说明 manifest 清理、GC 窗口、PVC 容量与上传上限 2026-09-24 23:10:40 +08:00
Lemon-miaow e75448a118 feat(platform): registry/uploads/backup PVC 容量可配置 2026-09-24 23:10:40 +08:00
Lemon-miaow 720ab81bf8 feat(submit): 上传存储全局上限、磁盘余量检查,被拒上下文 7 天后回收 2026-09-24 23:10:40 +08:00
Lemon-miaow 9baa0a10af fix(registrygate): 启动后先等写入静默再开放 GC 窗口 2026-09-24 23:10:28 +08:00
Lemon-miaow 151c9d2e30 feat(registry): api 定期删除无引用 manifest,gate 提供 manifest 索引 2026-09-24 22:58:07 +08:00
Lemon-miaow 05e8c64e47 fix(imagepin): 带 digest 的镜像引用也向 registry 确认仍存在 2026-09-24 22:58:07 +08:00
Lemon-miaow 7f772bccbb feat(registry): 开启 manifest 删除,GC sidecar 在 gate 只读窗口内回收 blob 2026-09-24 22:46:37 +08:00
Lemon-miaow c49336ba21 fix(bootstrap): 元数据下载在 TLS 握手中断时整体重试 2026-09-24 22:26:30 +08:00
Lemon-miaow 5cadbd40a9 fix(submit): 待审上限在事务内加咨询锁原子检查,跨副本不超额 2026-09-24 22:26:30 +08:00
Lemon-miaow a27d76ae1d docs(build): 修正上下文存储与构建镜像拉取的过时描述 2026-09-24 22:26:29 +08:00
Lemon-miaow e836a73a8a fix(build): 构建并发上限与排队,构建命名空间加 ResourceQuota,SyncAll 逐个容错 2026-09-24 22:26:21 +08:00
Lemon-miaow e521cf5976 fix(submit): 审阅绑定上下文 sha256,批准须带摘要,构建只从内部 API 取上下文并校验字节 2026-09-24 22:24:50 +08:00
Lemon-miaow 13b65e19ec fix(build): 构建 pod 等出口策略生效再运行,加 seccomp、可选 user namespace 与磁盘上限,上下文解包限总字节与条目数 2026-09-24 20:42:14 +08:00
Lemon-miaow 9bda3a52fa fix(backup): 手动备份按服冷却、每服保留上限与独立保留期,容量驱逐不再删除回收世界的唯一副本 2026-09-24 19:58:59 +08:00
Lemon-miaow f69cdec9d4 fix(reaper): 有服务器处理失败或过期备份删不掉时以非零码退出,让 Job 失败告警触达运维 2026-09-24 19:34:09 +08:00
Lemon-miaow 22becd6859 fix(reaper): 认领时重置活跃时钟与预警,回收只复用本次认领后的 reaper 归档 2026-09-24 19:31:23 +08:00
Lemon-miaow 47890ca913 feat(offsite): 世界归档与数据库备份加密同步到异地 S3,reaper 确认异地副本后才删除世界 2026-09-24 19:25:16 +08:00
Lemon-miaow fa8db4c7e8 docs(audit): 告警条目补记 watchdog 邮件告警与真机演练 2026-09-24 18:36:07 +08:00
Lemon-miaow d17524cd67 feat(watchdog): 主机侧巡检定时器按异常邮件通知平台所有者,operator 增加 phase 与 build_info 指标、卡死存活探针与告警规则 2026-09-24 18:06:19 +08:00
Lemon-miaow d50492b86f chore(panel): mock 服务器带固定镜像与内存存储以演示编辑对话框 2026-09-24 17:08:53 +08:00
Lemon-miaow 0c29ab3a96 fix(api): 未改动任何字段的 PATCH 返回空 patched 数组而非 null 2026-09-24 17:04:02 +08:00
Lemon-miaow 215bfd78d7 fix(images): 服务器镜像在创建时固定到仓库 digest,更换镜像需确认备份,安装器重建前先固定旧服并推送不可变版本标签 2026-09-24 16:57:38 +08:00
Lemon-miaow 857b6da886 fix(operator): 删除依赖 rcon-cli 的 preStop 死代码,缩容前由 operator 经 RCON 执行 save-all flush 2026-09-24 16:39:22 +08:00
Lemon-miaow 5521e498a9 fix(platform): minecraft 命名空间强制 PodSecurity baseline,reaper 世界根目录改走静态 hostPath PV 2026-09-24 16:28:12 +08:00
Lemon-miaow 346a93921e fix(operator): 游戏 Pod 改以 UID 1000 运行并丢弃全部能力,prepare-data 初始化容器修正旧存档属主 2026-09-24 16:23:49 +08:00
Lemon-miaow c1796bea17 feat(servers): 新建服务器默认空闲 10 分钟自动停服,面板可调,converge 回填旧服,系统服不休眠 2026-09-24 16:12:53 +08:00
Lemon-miaow bd909d5858 fix(operator): 读不出玩家数时暂停空闲自动停服,支持 1.12/Bukkit/EssentialsX 与颜色码 2026-09-24 16:06:55 +08:00
Lemon-miaow 857c73a66a fix(audit): 审计按账号 id 归属并记录来源 IP/UA,登录失败与限速入审计和指标,写入失败计数告警 2026-09-24 16:02:44 +08:00
Lemon-miaow c4e4953f3d fix(auth): 公开登录门按来源限速并设全站发信上限,冷却表定期清理 2026-09-24 15:51:42 +08:00
Lemon-miaow 15f729ffea fix(auth): 邮箱验证码按账号累计错误次数封顶并通知 2026-09-24 15:37:01 +08:00
Lemon-miaow 18e2397272 fix(panel): 加错误边界与新版本 chunk 自动刷新,无 WebGL 时仪表盘降级为平面视图 2026-09-24 15:27:29 +08:00
Lemon-miaow 248fa5d100 fix(panel): 服主在控制台看得到 LuckPerms 入口,归属判定统一走 /me/servers 2026-09-24 15:23:20 +08:00
Lemon-miaow c7db7d4126 feat(db): 控制面 PG 定时备份、迁移前快照与原子恢复 2026-09-24 15:19:42 +08:00
Lemon-miaow abfe60d62d fix(api): wake 与回档/备份/改文件按服务器互斥 2026-09-24 14:39:02 +08:00
Lemon-miaow 7819e5de50 feat(netpol): 锁定游戏服出站并为 registry 加入站围栏 2026-09-24 14:25:17 +08:00
Lemon-miaow 3424852a39 feat(registry): 写入改走鉴权网关,构建先扫描再推送 2026-09-24 14:25:17 +08:00
Lemon-miaow 8f684cedc0 feat(paper): 大厅菜单从代理动态拉取服务器列表并分页 2026-09-24 13:39:38 +08:00
Lemon-miaow 955433ba81 fix(limbo): 经 felis:control 放行玩家并在 felis-api 故障时退避重试 2026-09-24 13:39:38 +08:00
Lemon-miaow 4648175773 fix(velocity): felis:control 按来源服务器限权并关闭 bungeecord 通道 2026-09-24 13:39:38 +08:00
Lemon-miaow 3247b9e61a docs(audit): batch-47 — v0.1.0 stable release fired end-to-end (release download → no-checkout full bootstrap → felis update) 2026-09-24 11:36:22 +08:00
Lemon-miaow b9c97ffac9 docs(audit): batch-46 — upload-lane quotas/lifecycle, tracker #8/#1, release-pipeline smoke
Real-machine closure for everything batched since batch 45: #75/#76 verified
end to end on auditfix77 (cooldowns 429, pending cap 403, 2GiB budget 403 via
the real accounting path, failure-does-not-burn-window, withdraw/admin-delete
row+blob double clean, panel two-step confirms in CDP), tracker #8 and #1
landed and verified (#77 panel redirect, #78 felis converge), plus #79 docs.
Also: pgint's new assertions first run on real PG (17/17) and the release
pipeline's first-ever run — tag v0.1.0-rc1, multiarch assets published and
executed on the target arch, /releases/latest deliberately untouched.
2026-09-24 11:11:07 +08:00
Lemon-miaow b9ebc872ad docs(troubleshooting): the [INERT] legend points at a field that no longer exists
The legend said "§12 lists the one field this still applies to", but the last
inert field (spec.storage.retainOnDelete) was removed rather than implemented
(§13), and §12 has said "every field below is read by a controller" since.
Reworded so a reader whose change looks ignored follows the condition question
instead of hunting for a dead field.
2026-09-24 10:35:35 +08:00
Lemon-miaow c57daaf861 feat(cli): add felis converge for fields a newer desired spec never delivered (tracker #1)
Provisioning is create-if-absent, so a field the desired spec gained after an
install (spec.rcon, spec.startup.healthHTTPPort, a derived env key) never
reaches the existing login/lobby CR while every re-run of setup reports
success — the reported 'configuration updates never reach an installed
deployment' symptom. converge is the explicit pass: it fills exactly the
zero-valued whitelist fields and the derived env (including a missing key,
which refreshDerivedEnv deliberately never adds), and never overwrites a
non-zero value. The timing stays with the operator because enabling RCON or
the HTTP readiness gate on a pre-listener image would wedge that server in
Starting until it was marked Failed.

Tests: fills predated fields while operator edits survive / non-zero values
left alone / absent + foreign + unset-image guards. usage table updated so the
router-parity test passes; troubleshooting gains §12b.
2026-09-24 10:35:14 +08:00
Lemon-miaow 43699b46db fix(panel): route a setup-locked session to the wizard, not a permission error (tracker #8)
A session that still owes forced onboarding gets 403 setup_required from every
protected route, but the panel rendered it as the generic forbidden line — the
one step that unlocks the app read as missing authorization. api.ts now
announces the code on a window event (client module has no router) and the App
shell, inside the Router, navigates to /setup; the wizard resumes from the
surviving session with or without a token. Other 403s are untouched.

Panel tests: +2 (fires on setup_required, silent on any other 403).
2026-09-24 10:32:38 +08:00
Lemon-miaow 854320ac3f feat(submit): give the upload lane a lifecycle — withdraw + admin delete (#76)
Nothing ever removed a submission: users could not retract a pending row, no
route deleted blobs or rows, and the reaper never touches uploads — so every
upload accumulated on the 5 GiB PVC forever and the only cleanup was SQL or
kubectl against the store.

- Blobs.Delete on both transports (local: RemoveAll of the id-namespaced dir,
  id re-validated at the boundary; S3: idempotent object DELETE).
- Store: DeleteSubmission (admin, any status) and DeletePendingSubmission
  (owner+pending CAS — a reviewed row can never be withdrawn out from under
  its build).
- Manager.Delete / Manager.Withdraw delete the ROW first (under the CAS for
  withdraw) and the blob after, so a live row can never point at a reaped
  blob; a cleanup failure names the orphan explicitly instead of failing mute.
- API: DELETE /me/submissions/{id} (withdraw, app tier) and
  DELETE /api/v1/submissions/{id} (admin) both return the row as it was;
  audit events submission.withdraw / submission.delete; openapi documents both
  paths; admin route pinned in the admin-only table.
- Panel: two-step withdraw on a pending row (frees the pending slot and the
  storage budget); two-step delete on every admin row; zh/en copy; wire tests.

Unit: submit (withdraw happy path / wrong owner / reviewed row / no transport /
blob-cleanup failure), local+S3 delete idempotence, api handlers (200/404/409/
503 + route tier); pgint: withdraw CAS + admin delete exactly-once.
go vet/go test/gofmt clean; panel vitest 120 + typecheck green.
2026-09-24 10:24:23 +08:00
Lemon-miaow ad4d256d8f fix(submit): bound the untrusted upload lane — per-user caps + throttles (#75)
A logged-in user could file submissions without bound and stream a 1 GiB
context per submission. The only limits were the single-blob size cap and the
5 GiB uploads PVC (platform/workloads.go); nothing counted a user's rows or
bytes, so one account could fill the volume and every other user's upload
would start failing.

- Create: per-user pending_review cap (default 5) — the review queue cannot
  be parked full of one account's rows. Check-then-insert, documented soft.
- UploadContext: per-user stored-context budget (default 2 GiB) charged
  against the blob store's REAL sizes (new Blobs.Size on local/S3 stores), so
  the sum cannot drift from the volume; the write is capped at the remaining
  budget, so the excess is refused before it is persisted, and a re-upload is
  charged only for its new bytes.
- API: per-user create/upload throttles (30s/15s, cmd/felis-wired) on a
  dedicated cooldown keyspace, reserve→release so a failed attempt never
  burns the window and a burst collapses to one winner; ErrQuotaExceeded →
  403 submission_quota_exceeded (distinct from the 400 an oversize blob
  gets), 429 submission_cooldown for the throttles.
- Panel: zh/en copy for both codes; openapi documents 403/429 on the two
  user routes; pgint covers the pending-queue count.

Unit tests: submit package (cap, budget boundary/exact-fit/replacement,
oversize-vs-quota split) and api handlers (quota 403 both paths, throttle
429 + recovery + failure-release). go vet/go test/gofmt clean; panel
vitest 118 + typecheck green.
2026-09-24 10:14:42 +08:00
Lemon-miaow 3ed8bd7be9 docs(audit): batch-45 — modpack base-image chain (#70/#71/#72), submission-to-server capstone, tooling fixes (#69/#73/#74) 2026-09-24 03:19:21 +08:00
Lemon-miaow 96aa8176cd test(api): make the console-disconnect teardown case wait for the write, not just the read (#74) 2026-09-24 03:16:40 +08:00
Lemon-miaow 5cbfa893a5 docs: correct the modpack feature claim — approval auto-builds; deployment is an explicit server-image choice (#69) 2026-09-24 03:11:33 +08:00
Lemon-miaow 2961beb8fe test(bootstrap): keep the worlds-root warning case hermetic on hosts that already run k3s (#73) 2026-09-24 03:11:10 +08:00
Lemon-miaow b14bacbfc6 fix(build): mirror the Trivy Java DB — jar-bearing builds failed closed at the scan gate (#72) 2026-09-24 03:11:01 +08:00
Lemon-miaow 6e47730501 fix(build): allow kaniko to unpack base-image layers — drop-ALL removed the caps the tar apply needs (#71) 2026-09-24 03:05:07 +08:00
Lemon-miaow ac403b9cd3 fix(build): let kaniko pull base images from the plain-HTTP registry — push-only insecure flags broke every FROM registry.felis.svc build (#70) 2026-09-24 02:56:56 +08:00
Lemon-miaow 7791ed74f2 docs(audit): batch-44 — modpack-scale (200 MiB) context drill + concurrent builds, zero defects 2026-09-24 02:50:31 +08:00
Lemon-miaow 6abce99b49 docs(audit): batch-43 addendum — #68 context-fetch retry, live-verified; long-poll closure 2026-09-24 02:44:46 +08:00
Lemon-miaow b76d0acec7 fix(build): retry the context fetch through a control-plane restart (#68) 2026-09-24 02:27:45 +08:00
Lemon-miaow 580032056e docs(audit): forty-third batch — build-job reaping (#67) + lifecycle/fault-injection soak
Six CR-level run/stop cycles with three chaos injections (operator pod
kill, postgres restart, api pod kill) all converged (Running 23-29s,
Stopped 3s, pods gone per cycle); PG outage keeps the 503-not-401
session semantics; 30-request burst all 200; ownerless wake correctly
403s.  The build lane's missing Job TTL (#67, fixed in 2755e41) was
verified live on auditfix62: a real build Job carries ttl=604800 and an
old Job patched to ttl=30s was reaped, pod and all, within 40s.
2026-09-24 02:23:18 +08:00
Lemon-miaow 2755e41ff3 fix(build): reap finished build Jobs — they accumulated forever
Every other Job family carries a TTLSecondsAfterFinished (fileedit 2m,
backup/restore 10m) but the build lane never set one: each build left a
completed Job + Pod in felis-build indefinitely (5 already on the drill
cluster, oldest 26h), growing etcd and — since completed pods count
against the node's pod budget (110 on stock k3s) — eventually blocking
new builds.  The code even anticipated GC it never got ("JobUnknown
means the Job was not found (e.g. GC'd)").

Set a deliberately long TTL (7 days): the kaniko log is the admin
failure-triage surface (GET /images/build/{id}/logs), so the window
keeps a week of logs while bounding steady-state pods; terminal builds
are idempotent under Sync, so a late log-404 is the only cost.
2026-09-24 02:03:20 +08:00
Lemon-miaow d5e623c7ba docs(audit): batch-42 addendum — static review of the never-run release pipeline
The release workflow has never executed (no tags/releases exist yet).
Its three load-bearing contracts were verified against their other
halves: the asset names vs bootstrap's download_release_binary (and its
'felis ${REF}' convergence check), the local-exporter path vs the
Dockerfile's final COPY, and the version assertion vs cmdVersion's first
line.  The one un-replicable check (file(1) machine type) was probed
with a cross-compiled linux/arm64 binary: 'ARM aarch64' matches.  No
pre-tag fixes needed.
2026-09-24 01:33:56 +08:00
Lemon-miaow aa82321907 docs(audit): forty-second batch — coverage sweep, README_EN (#66), §28 diagram alignment
Module-by-module coverage matrix against the ledger: repo hygiene clean,
no TODO/FIXME debt, CLI dispatch test-guarded, CI syntax-checks every
tracked shell script, all 22 internal packages accounted for (CRD types
and game-image dirs covered indirectly), panel's 17 page routes drilled
in earlier batches.  Fixes this batch: #66 README_EN was missing the
private-repo install workaround and the rerun/upgrade notes (an English
reader 404s on step one); the §28 claim/link sequence diagrams were
realigned with the audited implementations.  A local-only scripts/
sync.sh conflict (plugins/ exclusion vs go:embed) was reproduced and
fixed on the workstation — private per .git/info/exclude, not committed.
2026-09-24 01:27:42 +08:00
Lemon-miaow 62e8c87d0e docs(diagrams): align §28 claim/link sequences with the audited implementations
The claim-transaction note still described the pre-audit-#4 shape (SELECT
EXISTS plus a quota pre-check outside the transaction); the implemented
ClaimServer serializes on pg_advisory_xact_lock(user_id), takes the row
FOR UPDATE and re-runs the four-dimension gate inside the same
transaction — the diagram's QuotaAvailable step is only a fast path.
The link flow now selects the code FOR UPDATE, treats a same-user
re-verify as idempotent, and lets a live caller take over a retired
(soft-deleted) owner's link — the 409 is only for a different, live user.
Both re-read from internal/api/pgrepo.go and the handler mappings.
2026-09-24 01:26:32 +08:00
Lemon-miaow 62c5a8a408 docs(readme): port the private-repo install path and upgrade notes to the English README (#66)
README.md grew a required workaround — the repository is private, so the
plain raw.githubusercontent one-liner returns 404.  The credentialed form
(token handed to curl via --config - so it never touches argv, sudo -E so
the installer inherits it) and the rerun/upgrade notes (the rerun upgrades
felis-api; the channel is not inherited, so main-followers need
FELIS_VERSION_BOOTSTRAP=dev) never made it into README_EN.md, which is
linked as the English entry point.  An English reader following that page
could not install at all.  Bring it back in sync with the Chinese one.
2026-09-24 01:24:27 +08:00
Lemon-miaow 5f9380dfca docs(audit): forty-first batch — loader mod layer closed
The last module family nobody ever built (#64 gradlew exec bit, #65
license metadata), closed with three real dedicated-server boots
(fabric/forge/neoforge: mod loads, /link registers, console refused)
and a permanent JDK-17 gate in CI + release.  Reachability table now
#1–#65; conclusions item 20; remaining-queue note updated.  CI run
35892544563 = 5/5 with the new mods job green on its first run.
2026-09-24 01:14:49 +08:00
Lemon-miaow aa5abfa911 ci(plugins): gate the three loader mods (JDK 17 wrapper builds) in CI and release
Nothing ever compiled these modules — no CI job, no install path — which is
why the gradlew exec-bit bug (previous commit) shipped unnoticed.  Add
plugins/test-mods.sh: runs each module's vendored wrapper under JDK 17
(they target the Java-17 Minecraft lines; paper/limbo stay on the JDK 21
gate).  ci.yml gains a `mods` job (temurin 17 + setup-gradle, wrapper
pinned per module), release.yml gates both plugin gates before shipping.
README: status now records the server-boot verification, the Building
section pins limbo's `-PlimboVersion=<release>` (the `+` default is
unresolvable from the LOOHP repo) and documents both gates.
2026-09-24 00:59:24 +08:00
Lemon-miaow c2fe6a6bf4 fix(plugins): loader mod metadata — AGPL-3.0-only license, real issue tracker
The three loader mods declared `license = "MIT"` (and Forge/NeoForge a
placeholder `issueTrackerURL = https://example.invalid/felis`) while the
repository is AGPL-3.0-only (README.md:73, LICENSE).  The mods were added
2026-06-26, the LICENSE landed 2026-07-12 — stale leftovers that nothing
ever read back: no CI job, no install path.  Fabric loader prints the
license from fabric.mod.json at boot and both mods.toml files are parsed
by their loaders, so the wrong claim is user-visible.  Align both with
reality — boot-verified on real fabric/forge/neoforge dedicated servers
(AUDIT-2026-09-22.md, batch 41).
2026-09-24 00:59:19 +08:00
Lemon-miaow f6048f268b fix(plugins): commit the gradlew exec bit for the loader mods
All three vendored wrappers were tracked as 100644, so the README's documented
'plugins/<loader>/gradlew -p plugins/<loader> build' failed on every fresh
clone with 'Permission denied' (exit 126) — and since no install path or CI job
ever ran them, nothing caught it.

git update-index --chmod=+x for the three files; the VM then built all three
modules for the first time (results in the gate commit and the audit).
2026-09-24 00:47:32 +08:00
Lemon-miaow 2d04c448c2 docs(audit): fortieth batch — java plugin layer into CI, demo-up single origin, #62 screened out
#63 fixed in f5a76cf (demo-up delegates to bootstrap; T1/T2/T3 real-machine);
plugins/test.sh + the new CI plugins job verified on the VM and in CI
(c59b387, run 35888009965, 4/4 jobs); #62 re-checked against all three
install paths and screened out as unreachable/nonexistent; plugin-layer
real-machine E2E via the MC status/login probes recorded.
2026-09-24 00:25:43 +08:00
Lemon-miaow c59b38775b ci(plugins): run the plugin self-tests and build the shipped plugin jars
The three framework-free test mains under plugins/*/test were never run by
anything — not CI, not the plugin builds — and the velocity/paper/limbo jars
were only ever compiled by deploy/bootstrap.sh on a live host. CI gains a
'plugins' job (JDK 21 plus the Gradle 8.14 the plugin Dockerfiles pin) running
plugins/test.sh: the three mains (InviteCardTest's jars fetched from Maven
Central, pinned and digest-checked) and the three production builds.

The first real run surfaced and fixed two untested assumptions: InviteCardTest's
documented javac line omitted examination-api (adventure-api's Component
signatures reference Examinable, so javac needs it too), and limbo's '+' version
default cannot resolve — LOOHP's repository serves no maven-metadata — so the
script resolves the current release off the Limbo CI artifact name (the same
source bootstrap reads) and plugins/README.md stops advertising a bare
'gradle -p plugins/limbo build' that can never work.

Verified in gradle:8.14-jdk21 on the VM: mains OK (32/36/48 checks);
velocity/paper/limbo BUILD SUCCESSFUL.
2026-09-24 00:19:56 +08:00
Lemon-miaow f5a76cf88a fix(deploy): make demo-up a thin wrapper — bootstrap is the one origin for the game stack (#63)
demo-up.sh grew its own image lane in July, before bootstrap could build the
game stack. The copy has since drifted from the installer it duplicates: it
pinned Paper 1.21.8 while bootstrap derives one version from the Limbo login
gate (both hops of a login must speak one protocol), it never built
felis-velocity.jar (a proxy without it silently routes nothing), it imported
images under local tags the [velocity] wiring no longer points at, and it left
the docker daemon running.

It now runs deploy/bootstrap.sh (unless SKIP_BOOTSTRAP=1), checks that the
earlier run really left the full stack behind, and hands over to 'felis setup'
— so a half-installed base is a clear error instead of a proxy that accepts
logins and routes nowhere.

Verified on the VM: healthy base passes to the handoff (exit 0); a hidden
felis-velocity.jar dies naming the jar and the remedy (exit 1, jar restored).
2026-09-24 00:19:49 +08:00
Lemon-miaow 1184d6b787 docs(audit): thirty-ninth batch — reaper multi-node drill on a cloned second node 2026-09-23 23:58:58 +08:00
Lemon-miaow 811f2b8898 docs(audit): thirty-eighth batch — #56–#59 evidence archived, server detail and ops sweep clean 2026-09-23 23:50:47 +08:00
Lemon-miaow 4c842f5d70 docs(audit): thirty-seventh batch — admin interaction sweep clean, #61 LuckPerms write promise scoped 2026-09-23 23:29:13 +08:00
Lemon-miaow b4ef42d947 fix(panel): scope the LuckPerms write promise to names the server can resolve (#61) 2026-09-23 23:28:28 +08:00
Lemon-miaow ecd42ed484 docs(audit): thirty-sixth batch — panel error localization (#60), Run4c admin write-path review clean 2026-09-23 23:12:14 +08:00
Lemon-miaow c9af5481e7 fix(panel): localize every user-reachable API error code (#60) 2026-09-23 23:09:28 +08:00
Lemon-miaow a2ff2a1102 fix(panel): make the LuckPerms page honest about what the server said (#59)
Live with real LuckPerms 5.5.85: every lp command returns an empty RCON body
(list/plugins answer normally; a standalone RCON client sees the same, and
creategroup/permission-set still persist), so the read projection can never
populate and the page asserted "no parent groups / no explicit nodes" for a
state it could not actually read. The raw reply now rides the same disclosure
the rosters carry, a silent entry-less reply shows an explicit notice instead
of the false empty claims, and the write history's placeholder no longer
dresses up a fabricated "[RCON] ..." line as output.
2026-09-23 22:43:29 +08:00
Lemon-miaow 4d4cdd6ea7 fix(api,panel): refuse file operations on a server without a world volume (#58)
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
2026-09-23 22:17:08 +08:00
Lemon-miaow 70c988e702 fix(panel): render the console command's reply (#57)
sendCommand returns the RCON reply and the component's own contract says it
"displays its plain-text reply", but the reply was discarded and the pod log
does not echo command output, so a sent command produced no visible result at
all (verified live with 'list'). The last command echo + reply now render
above the prompt, terminal-style.
2026-09-23 22:08:19 +08:00
Lemon-miaow 0790f8dfd3 fix(panel): surface the RCON reply for player-access mutations (#56)
The whitelist/ban/kick mutations reported a canned success message and threw
away the server's reply, so a refused command still read as done: live, the
vanilla server answers "That player does not exist" for a name it has never
seen (any player who has not joined yet), while the panel said the player had
been whitelisted/banned. api.ts documents these replies as "surfaced verbatim
as confirmation"; now they are. The localized string stays as the fallback
for a silent server.
2026-09-23 22:08:19 +08:00
Lemon-miaow 9c3a1c5be1 docs(audit): thirty-fifth batch — NetworkPolicy live enforcement matrix green (whitelist closed loop), velocity refresh loop verified 2026-09-23 21:43:23 +08:00
Lemon-miaow cbb0f11288 docs(audit): thirty-fourth batch — felis nano sweep (red 13 -> green 60/60), #55 bad-source ladder stall 2026-09-23 21:30:53 +08:00
Lemon-miaow 9dad61f508 fix(nano): screen unusable profiles per source instead of stopping the ladder (#55)
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.

Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.

Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
2026-09-23 21:21:21 +08:00
Lemon-miaow 971ae01caf docs(audit): thirty-third batch — /updates window API+UI sweep green, #53 repo move, #54 update guidance red/green
Records: (1) the /updates maintenance-window API sweep (unset null, 400s/415 for the invalid set, write->read-back->survives API pod restart, 401/403 auth) and the CDP panel sweep (status transitions, validation copy, clear, audit x3, zero console errors); (2) #53 red/green with the live fix53 binary ('MliroLirrorsIngenuity/Felis' -> 'FelisMC/Felis' in the 404 line) plus the token'd check (old path 301 / new path 200) and the fact the new home has no stable release yet; (3) #54 red/green with the live fix54 binary (installer one-liner + single trailer replace 'run: sudo felis setup'), the setup-vs-installer evidence, and the doc/test synchronization. Reachability table gains #53 (2) and #54 (2, docs); stats 54 total; repro-entry notes updated to the new remote, host binary v0.0.0+fix54 and the batch's artifacts.
2026-09-23 21:04:02 +08:00
Lemon-miaow 397a400d57 fix(update): point the apply guidance at the installer, not felis setup (#54)
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.

Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
2026-09-23 21:01:32 +08:00
Lemon-miaow 2c6739ad76 fix(updater,install,docs): follow the move to FelisMC/Felis (#53)
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.

Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
2026-09-23 20:58:34 +08:00
Lemon-miaow cec9a98305 docs(audit): thirty-second batch — S3 storage wizard sweep, #52 mirror lag red/green, cleanup 2026-09-23 20:39:09 +08:00
Lemon-miaow de7fb2c936 fix(setup): converge the workload felis-config mirror on every apply path (#52)
felis setup's in-TUI applies (storage / connection / edge) refreshed only the
control-namespace felis-config Secret; the workload-namespace mirror kept the
render from the previous run's startup pass until the next setup or installer
run. Found live: after 's -> Local' the minecraft copy still carried
user_uploads_context = s3://felis-wizard-uploads while the control copy and
both tomls were local. The 'configure email' path already overwrote both
mirrors, so storage/connection were the odd ones out.

Move the mirror refresh into applyFelisConfigSecret — the single choke point
every apply path calls — best-effort with a warning, since a control-plane
default install may not have the workload namespace at all. The smtp helper
drops its now-duplicate felis-config block.
2026-09-23 20:31:48 +08:00
Lemon-miaow 01988305a8 docs(audit): thirty-first batch — first-install walkthrough, #51 replica refresh red/green 2026-09-23 20:15:46 +08:00
Lemon-miaow 328e570309 fix(setup,install): refresh the workload namespace's felis-config mirror (#51) 2026-09-23 20:07:20 +08:00
Lemon-miaow cf5a790ea8 docs(audit): thirtieth batch — #50 smtp carry hoard, healed by two live re-runs 2026-09-23 19:57:00 +08:00
Lemon-miaow 4d3c85fd06 fix(bootstrap): [smtp] carry stops hoarding the auth_source comment block (#50) 2026-09-23 19:45:26 +08:00
Lemon-miaow 28fe7c43a9 docs(audit): twenty-ninth batch — image durability live drills; #46–#49
- registry hosting + loopback pull path landed (a9b275a/13d64e0/fa0e8d7) and
  drilled live: three installer re-runs, then GC simulations on the control
  plane (rolled felis-api pulled back in 25ms) and a game image (lobby-0,
  182MB in 10ms).
- #46 registry OOM (475MB-layer push killed the 256Mi template; dmesg evidence)
  fixed and re-verified: oom-kill count unchanged across a full rebuild+push.
- #47 AppleDouble ._*.sql embedding broke felis migrate on a Mac-staged tree;
  .dockerignore fix probed live with a planted junk file.
- #48 per-image docker start/stop tripped systemd start-limit-hit mid-batch;
  one wrap per batch, re-run mirrors all four.
- #49 installer re-runs silently reverted operator [registry]/[archive] config;
  carry-forward landed + live-verified into host toml, pod toml and the Secret,
  and the carried pins drove a successful POST /images/build.
- reachability table extended to #49 (① 20 | ② 19 | ③ 3+ | ④ 4 | 决策 3).
2026-09-23 19:35:13 +08:00
Lemon-miaow 72553cb414 docs(troubleshooting): the installer leaves docker stopped — start it before manual pushes
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
2026-09-23 19:35:13 +08:00
Lemon-miaow 20a95da487 docs(update): the plugins note is a rebuild + registry re-mirror now, not a node re-import
The felis-paper/felis-limbo jars are baked into the lobby/limbo images; with
the images hosted in the in-cluster registry, the extra step is pushing the
rebuilt image there (which is also what survives an image GC), not a bare
containerd import. The installer re-run does both.
2026-09-23 19:33:04 +08:00
Lemon-miaow c7e585e21d fix(bootstrap): mirror the image batch under ONE docker start/stop
Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.

push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
2026-09-23 19:23:44 +08:00
Lemon-miaow 5fa8b7412e fix(build): keep macOS ._*/.DS_Store junk out of the image
A Mac-staged tree (BSD tar materializes extended attributes as ._<name>
sidecars) went through the docker build and one landed in
internal/store/migrations/ — //go:embed-ed into the binary, where every
`felis migrate` then died with 'migration "._0004..." has a non-numeric
version'. Observed live wiring up the auditfix42 image: the installer's own
run_migrations failed on it. Exclude the sidecars and .DS_Store from the
build context; deploy/*.yaml and plugins/ have the same exposure.
2026-09-23 19:19:04 +08:00
Lemon-miaow b8e554dac7 fix(bootstrap): carry the operator's [archive] keys across re-runs too
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.

persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
2026-09-23 19:09:57 +08:00
Lemon-miaow 765a8923a4 fix(bootstrap): re-runs keep the operator's [registry] overrides
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.

Found while landing the registry-hosting work, which depends on those same
keys surviving.

- persisted_registry_block carries the operator-owned [registry] keys and the
  [registry.s3] subtable forward, same first-readable-file rule as
  persisted_smtp_block; url/build_namespace stay installer-owned (they must
  match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
  as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
  idempotence (a second re-run writes a byte-identical file).
2026-09-23 19:05:54 +08:00
Lemon-miaow fa0e8d7d97 docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
2026-09-23 19:03:11 +08:00
Lemon-miaow 13d64e0000 feat(bootstrap): host every built image in the internal registry — GC-durable pulls
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:

- Every image the installer builds is now a registry ref
  (registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
  containerd under that exact name (first boot needs no registry round-trip)
  and mirrored into the registry after deploy_bundle (push_image_to_registry:
  push endpoint 127.0.0.1:5000, and only the path after the host matters to the
  registry — a push there lands where kubelet's mirrored pull looks). A ref
  outside the registry is warned about, not silently unmirrored.

- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
  registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
  registry Deployment binds (node containerd cannot dial the Service VIP — live
  drill: "Empty reply"). k3s regenerates containerd config only at agent start,
  so a CONTENT change restarts k3s and an identical file (every re-run)
  restarts nothing.

- import_registry_image caches registry:2 into containerd so the registry
  Deployment can start on a box that cannot reach Docker Hub.

- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
  'felis-paper:demo') at the registry refs — a user server created from those
  rows must not strand when GC collects the bare tag. Only recommended rows
  still holding the old seed are touched; enabled is preserved; a pre-existing
  target row wins over a duplicate.

bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
2026-09-23 19:02:58 +08:00
Lemon-miaow a9b275abbb fix(platform): registry OOM (audit #46) + loopback hostPort — the node-side pull path
Two changes to the registry Deployment, both prerequisite to GC-durable images:

- Dedicated resource template: the control plane's 256Mi memory limit was a
  live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
  mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
  same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.

- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
  reach the Service VIP (live stack: "Empty reply"), so the node-side pull
  path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
  http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
  the plain-HTTP registry off every other interface.

Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
2026-09-23 18:55:23 +08:00
Lemon-miaow 92c06ac8dc docs(audit): twenty-eighth batch — #45 blind-review fix live-verified (context download, byte-exact); reachability 1-45 2026-09-23 17:01:19 +08:00
Lemon-miaow 168a37542b feat(api,panel): reviewer context download for submissions; dockerfile field documented as audit-only (audit #45) 2026-09-23 16:41:55 +08:00
Lemon-miaow edd9d63f5e docs(audit): twenty-seventh batch — alert module live drill (real build failure -> pending -> firing), cleanup, #44 reachability 2026-09-23 16:35:23 +08:00
Lemon-miaow 43df08b52a feat(alerts): ship Felis alert rules with promtool unit tests; document scraping & rules (troubleshooting §14) 2026-09-23 16:17:51 +08:00
Lemon-miaow 94f71eea19 feat(api): serve felis_* metrics on the internal face (build-failure counter's only scrape path) 2026-09-23 16:17:50 +08:00
Lemon-miaow 17ede3c6aa docs(audit): twenty-sixth batch ledger — setup wizard re-run screens, #44, build-pin drift incident & re-verify 2026-09-23 16:04:20 +08:00
Lemon-miaow ae6e9256c6 docs(troubleshooting): 8e — apply build-image overrides through the config Secret (restart alone does not) 2026-09-23 15:56:44 +08:00
Lemon-miaow abb5910d2f fix(cli): setup re-run keeps its already-set-up framing after connect/storage reconfigure 2026-09-23 15:56:44 +08:00
Lemon-miaow 5450ec786f docs(audit): reachability grading for findings #1-#43 (who actually hits each one) 2026-09-23 15:45:19 +08:00
Lemon-miaow 24a6ab3d1e docs(audit): twenty-fifth batch ledger — breakGlass console screens & backup/restore gates (#39–#43, live-verified) 2026-09-23 07:50:27 +08:00
Lemon-miaow ac3a557566 fix(cli): Sync picker hides system servers; keep the two 409 refusals apart
Two defects from the live Sync drill:

- The picker listed the system servers (login/lobby), which the backup API can
  never accept (reserved names, no servers row): the pick died in name
  validation with a raw "server name is reserved" error. backupPickable now
  filters them out; the halt picker keeps them on purpose (break-glass retains
  full power over system servers).
- backupErrorFromResponse mapped every 409 to the stopped gate, so the new
  world-volume refusal would have displayed the wrong reason. The 409 arm now
  keys on the body's error code; a code-less body still reads as the stopped
  gate.

Live (auditfix38): the picker shows only user servers; a world-less pick shows
the API's own "no world volume yet — start it once" text; the not_stopped text
is unchanged.
2026-09-23 07:49:07 +08:00
Lemon-miaow 508a1c02da fix(api): refuse backup/restore before a missing world volume
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
2026-09-23 07:49:00 +08:00
Lemon-miaow 55d515d41f fix(provisioning): keep the Owner seat single; clash on the operator name stays retryable
Two defects live-drilled in the break-glass staff provisioning:

- An Owner reset that typed any username other than the occupied seat took
  UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
  existing seat — possibly the compromised account the reset was meant to
  replace — live; every owner row is undeletable through the panel, so the tier
  could never converge back to one. provisionOwner now refuses with
  ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
  form); bootstrap still mints, and the seat's own username still resets in
  place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
  console keys its rename prompt off api.ErrConflict — the "choose another
  name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
  the contract; PGRepo had drifted). Map the unique violation to ErrConflict
  and pin it in pgint.

Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
2026-09-23 07:48:54 +08:00
Lemon-miaow f6dbfd3625 docs(audit): twenty-fourth batch ledger — live S3 upload-channel drill (0 defects, reverted clean) 2026-09-23 07:05:42 +08:00
Lemon-miaow f378953982 docs(audit): twenty-third batch ledger — reaper node pin (#38 + multi-node gap) 2026-09-23 07:00:33 +08:00
Lemon-miaow daf760220b fix(cli): pin the reaper to its storage node; drop the stale uid-1000 note
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
2026-09-23 06:58:29 +08:00
Lemon-miaow a31eca65c3 docs(audit): twentieth–twenty-second batch ledger — build outcome visibility, files page, fleet system services 2026-09-23 06:53:40 +08:00
Lemon-miaow 2f90851c03 fix(panel): mark platform system services read-only in the fleet table
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
2026-09-23 06:51:45 +08:00
Lemon-miaow 0a36b3fda9 feat(panel): add the server files page for the world-volume repair lever
The backend could list/read/write a stopped server's world volume since the
file-editor slice, but the panel had no entry, so the one repair path for a
server that will not boot (a wrong line in server.properties) was API-only.
New /servers/:name/files page: breadcrumb browser, editor dialog with the
base64 []byte codec, binary files open read-only, the stopped gate is owned
up front (with a stop action) instead of letting every call 409, and a
doorway card on the console. i18n files namespace + wire-shape tests.
2026-09-23 06:42:45 +08:00
Lemon-miaow 72c4aa3895 fix(submissions): surface each linked build's outcome to the submitter
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
2026-09-23 06:31:13 +08:00
Lemon-miaow 4933c075b0 docs(audit): nineteenth-batch ledger — passkey unbind panel entry 2026-09-23 06:25:01 +08:00
Lemon-miaow 11ac4f50e6 feat(panel): expose owner passkey unbind in the user danger zone
DELETE /users/{id}/passkeys shipped as the owner-tier remediation for a
lost or compromised authenticator, but nothing in the panel reached it.
Add the danger-zone action with a confirm dialog; the account keeps its
other doors (email OTP, in-game op-login re-enrollment), so this severs
a credential without locking anyone out. Wire-shape test pins the call.
2026-09-23 06:24:45 +08:00
Lemon-miaow 35d93d7612 docs(audit): eighteenth-batch ledger — #35 world-executor identity defect and the backups-page completion 2026-09-23 06:20:44 +08:00
Lemon-miaow 97a64c8a33 feat(panel): add back up now and recent operations to the backups page
The backups page could list and restore archives but not create one,
and nothing surfaced backup/restore Job outcomes — a failed 202 was
visible only through kubectl. Add a Back up now action (enabled only on
a stopped server, the backend's own gate; a raced 409 is surfaced in
its words) and a Recent operations card fed by GET /servers/{name}/jobs
that shows running/succeeded/failed with the Job's failure message,
re-reads on an interval while a Job is running, and persists across
reloads. Wire-shape tests pin both endpoints.
2026-09-23 06:20:16 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00
Lemon-miaow f21aef3cfa docs(audit): seventeenth-batch ledger — hasJoined multiplexer drill (fake Yggdrasil) 2026-09-23 05:58:26 +08:00
Lemon-miaow 4298cd5de1 docs(audit): sixteenth-batch ledger — internal-face residual endpoints swept clean 2026-09-23 05:54:32 +08:00
Lemon-miaow 2bd25be712 docs(audit): fifteenth-batch ledger — #34 live closure and executor image refresh 2026-09-23 05:52:12 +08:00
Lemon-miaow bb9798e32c fix(api): in-game identity resolution and link takeover ignore dead accounts
UserByMCUUID now resolves only live accounts: claim, menu, wake
authorization, op-login vouch and the QR link-status poll treat a
disabled or soft-deleted link holder exactly like an unlinked UUID
instead of a retired identity. VerifyLinkCode lets a soft-deleted
link be taken over by a fresh in-game code (the deleted account is
gone, e.g. a migrated source), while a disabled holder still 409s so
the lockout is not bypassable; failed attempts still do not consume
the code. Fake repo and pgint coverage pin both branches.
2026-09-23 05:40:29 +08:00
Lemon-miaow 3ffa3f5318 docs(audit): fourteenth-batch ledger — dead-account resurrection (#33) and its live closure 2026-09-23 05:33:54 +08:00
Lemon-miaow 58535890c4 fix(api): dead accounts cannot log in, hold sessions, or keep identity assets 2026-09-23 05:28:32 +08:00
Lemon-miaow ba9d98f7ce docs(audit): thirteenth-batch ledger — CLI, direct build, panel CDP sweep, #31/#32 2026-09-23 05:15:51 +08:00
Lemon-miaow 6907961ce0 fix(panel): the build page trusts the server-side owner tier and drops its mock build seeds 2026-09-23 05:11:49 +08:00
Lemon-miaow d0b1f9694e docs(audit): twelfth-batch ledger — users admin matrix, defect #30 fix on live 2026-09-23 04:57:36 +08:00
Lemon-miaow 1918da29be fix(api): the quota/link admin sub-resources require a live user (404, not FK 500) 2026-09-23 04:56:06 +08:00
Lemon-miaow fbb6b0c180 docs(audit): eleventh-batch ledger — submission negative matrix, internal context fetch, op-login remint 2026-09-23 04:50:11 +08:00
Lemon-miaow ffe5dc14a8 docs(seams): close the deferred entries that are now live-verified 2026-09-23 04:46:02 +08:00
Lemon-miaow d4bb8d344b docs(audit): tenth-batch ledger — configure-email mirror fix (#29) and auditfix25 deployment 2026-09-23 04:44:12 +08:00
Lemon-miaow ed722d55f8 fix(setup): the workload-ns SMTP mirror must carry the target namespace
'smtpSecretManifest' hardcoded namespace=felis, so the 'configure email'
refresh of the minecraft-namespace copies failed before it began: kubectl
refuses a manifest whose namespace conflicts with -n (found live: 'the
namespace from the provided object "felis" does not match the namespace
"minecraft"'), and the felis-config mirror never ran at all because the
smtp apply returned early. A later SMTP change could therefore never reach
the reaper's pre-reap warnings — the exact failure the refresh was added to
close.

Render the Secret with the caller's namespace (felis for the control-plane
apply, the workload namespace for the mirror). Regression test pins both.
2026-09-23 04:42:59 +08:00
Lemon-miaow 311b1a7ec4 docs(audit): ninth-batch ledger — cfsetup #28 and the unused-endpoint sweep
Records the Access-policy upsert fix and the first end-to-end runs of
account/migrate (all four steps + negative matrix + retire assertions),
passkey credential management, the access player-management group, the
updates window, and fleet — plus the environment restore notes.
2026-09-23 04:41:01 +08:00
Lemon-miaow 36b954d347 docs(backupjob): the backup Job name is not deterministic anymore
The comment described a deterministic-name collision that the unique random
suffix made near-impossible; align it with Backuper.Backup and jobspec's
contract (ErrAlreadyExists survives only as the defensive no-op).
2026-09-23 04:41:01 +08:00
Lemon-miaow 30857df5b2 fix(cfsetup): upsert the Access policy — never swallow already-exists over a broader rule set
CreateAccessPolicy treated a Cloudflare "policy_already_exists" as idempotent
success and kept whatever policy was there. On a re-run with a changed
identity — or against a hand-made broader policy — op.console would stay
guarded by something weaker than the fail-closed body this package builds and
guards, while Setup reported success. The fail-closed validation only ever ran
on the policy we built, never on the one that stayed live.

Now it upserts by name: lookup, PUT the guarded body over the existing policy,
POST only when absent (a racing POST re-looks up and PUTs). apiPost/apiPut
share one apiWrite; three httptest cases pin update-over-existing, create-when-
absent, and the race fallback.
2026-09-23 04:30:53 +08:00
Lemon-miaow faa508e87a docs(audit): operator self-healing postmortem — ledger #26/#27 with live drills
RCON-secret deletion lockup and the stale start anchor (with its permanent
Provisioned=False) get their full live evidence trail, plus the sts
accidental-deletion drill. Deployed image note bumped to auditfix24.
2026-09-23 04:29:02 +08:00
Lemon-miaow 82b5a606f7 fix(operator): end the start attempt on success — stale anchor caused false StartupTimeout
Found live while validating the RCON-secret heal: a server that had already
recovered to Ready was marked Failed(StartupTimeout) minutes later, the moment
an unrelated pod rollout briefly dropped readyReplicas. The anchor
(status.startRequestedAt) was never cleared on success, so its 300s budget
kept ticking under a healthy server and any later blip spent it.

markRunningReady now clears the anchor: every start-or-recovery attempt gets
its own budget. It also flips ConditionProvisioned back to True — markFailed
sets it False and nothing ever reset it, leaving a permanent failure flag on
recovered servers that every conditions consumer would read.

Unit tests pin both: anchor cleared on Ready, Provisioned recovers from
Failed to Running.
2026-09-23 04:23:59 +08:00
Lemon-miaow 56f3abdb36 fix(operator): heal a deleted RCON Secret instead of locking the server out
Deleting the per-server RCON Secret used to leave a running pod authenticating
with the lost password while the operator re-minted a fresh one and probed
with it: the RCON gate failed forever (live: 96s+ of RconNotReachable, headed
for ReadinessTimeout) and nothing re-triggered a pod restart — the server only
came back when the pod was deleted by hand.

Two changes pair up:
- Owns(&corev1.Secret{}) so the deletion is noticed at all (a quiet Running
  server emits no other events; the Secret is controller-owned, so the watch
  maps it back to the CR).
- The pod template now carries a fingerprint of the current password
  (RconSecretAnnotation). Re-creation changes the fingerprint, the
  StatefulSet rolls, and the new pod picks the new password up; while the
  Secret is untouched the value is stable so no spurious rolls.

Unit tests pin stability across reconciles and the change-on-recreation roll.
2026-09-23 04:20:49 +08:00
Lemon-miaow 089d4f3a80 docs(audit): idle auto-stop postmortem — ledger #25 and self-checks in §11
Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
2026-09-23 04:14:11 +08:00
Lemon-miaow f650bf892a fix(operator): idle auto-stop couldn't write — patch the spec, and grant the patch
Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:

- The stop used a whole-object Update while the same reconcile loop writes
  status; that risks clobbering a concurrent status write. Switch to the
  reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
  markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
  closed with 403 (visible in the operator log as 'cannot update resource
  "minecraftservers"'). Grant patch and pin it in the RBAC scope test.

With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
2026-09-23 04:08:58 +08:00
Lemon-miaow 1c89a5eeeb fix(operator): wake up for idle auto-stop — the timer had no driver
EmptySince was stamped and then never revisited: player joins/leaves do not
touch the CRD, RCON is only probed inside Reconcile, and a steady Running
server produces no watch events (its status update goes out unchanged and is a
no-op). Live, an empty server with a 30s grace sat Running for minutes with
zero reconciles in the log — the auto-stop existed only on paper.

reconcileRunning now returns a RequeueAfter for idle-enabled servers: exactly
at the deadline while empty, or a 30s probe cadence while occupied so the
moment the last player leaves is noticed. New unit tests pin all three:
deadline requeue, occupied cadence, and no requeue when disabled.
2026-09-23 04:05:31 +08:00
Lemon-miaow c04a3f083e fix(crd): persist status.emptySince — the field was pruned away by the schema
The operator stamps EmptySince to time the idle auto-stop window, but the CRD's
status schema never declared it. Kubernetes pruned the field on every write
(apiserver warning: unknown field "status.emptySince"), so the timer reset to
nil on every read and idle auto-stop could never fire — a defect invisible to
the fake-client unit tests, which do not enforce the CRD schema. Found live:
the stamp was silently dropped the moment it was set.

Schema now declares emptySince (date-time) like its sibling timestamps.
2026-09-23 04:05:22 +08:00
Lemon-miaow 72f0b4a258 docs(audit): sixth-batch ledger — reaper warning path drilled end-to-end on the live cluster
The auditfix20 image (8e7c7bb) was exercised against a real SMTP sink with a
dedicated warntest server: delivery content, tier precedence, dedupe, retry
semantics (bad relay / unverified email / unowned), threshold boundaries, and
zero backup side effects. Ledger #24 records the defect and the evidence; the
deployed image note is bumped to auditfix20.
2026-09-23 03:56:07 +08:00
Lemon-miaow 8e7c7bbf24 fix(reaper): deliver pre-reap warnings for real — and never fake a delivery
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:

- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
  `warning suppressed — no warner wired` and does NOT stamp; a delivery error
  logs and retries on the next daily run (bounded by the warning window). The
  stamps are no longer burned by notices nobody received.

- A real channel: mail.SendNotice (the second and last message shape the mail
  package sends) plus a mailWarner that resolves the owner's VERIFIED email
  and mails the notice through the configured [smtp] relay. `felis reaper`
  wires it when [smtp] is set (same password_ref convention as felis-api) and
  prints exactly what happens when it is not.

Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.

Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
2026-09-23 03:47:19 +08:00
Lemon-miaow 1d0ec61c9d docs(audit): fifth-batch ledger — quota atomic gate closed (audit #4)
The deferred-seams entry that waited for a real-Postgres harness is struck:
ClaimServer owns the gate now, proven red-then-green by the pgint concurrency
test (two claims, one win, one 403), and the storage-cache zeroing found in the
same pass is recorded with its hermetic test. auditfix19 is live on the VM.
2026-09-23 03:39:25 +08:00
Lemon-miaow bb68fefe04 fix(quota): make the claim gate atomic, and stop zeroing the storage cache
Two defects in the §9.3 quota path, both invisible to the hermetic suite:

- Audit #4's TOCTOU was real and documented: QuotaCheck and ClaimServer were
  separate statements, so two concurrent claims by one user for two different
  ownerless servers both read count < max_servers and both won. The gate now
  lives inside ClaimServer, in the SAME transaction as the ownership write,
  under pg_advisory_xact_lock(hashtext(user_id)) — the aggregate read, the
  four-dimension re-check (shared with QuotaCheck via one helper so the two
  cannot drift), and the UPDATE are one serialized decision. The loser gets
  ErrQuotaExceeded, which both claim handlers map to the same 403 the
  sequential path gives; the server row is additionally taken FOR UPDATE so
  same-server races still resolve to exactly one winner.

- The server PATCH path called UpdateServerResources(..., 0) for storage even
  though a resources patch cannot change storage. The cached columns are the
  ONLY input to the quota aggregate, so every resource patch silently dropped
  that server's storage contribution from its owner's cap. The handler now
  reads the current spec and passes storage through.

Red-then-green: the new pgint test drives two real concurrent claims against
max_servers=1 (before: both win; now: exactly one win + one gated 403, and the
DB shows one owned row); the hermetic suite pins the 403 mapping and the
storage-preserving cache write.
2026-09-23 03:37:59 +08:00
Lemon-miaow d829267f1c docs(audit): fourth-batch ledger — PG contract tests land, verified-email and owner-role defects reconciled
#20–#22 recorded with live evidence: the pgint harness caught the
never-shipped verified-email index on its first run; migration 0020 applied
and the 409 drill replayed live; the admin email edit now drops the stale
proof; the owner role is written by both provisioning paths and its staff
doors, guards, and reclaim protections were drilled end to end on auditfix18.
The remaining-work item "PG-level contract tests" is checked off.
2026-09-23 03:32:48 +08:00
Lemon-miaow e0d23780d8 fix(auth): make the owner role real — provisioning, staff doors, panel guards
Found live while verifying the admin email-edit fix: the Owner account could
not load /api/v1/users at all. Root cause: migration 0011 adds the 'owner'
role and gates every user-administration route on it, but NOTHING ever wrote
it. break-glass (UpsertOwner), the setup MC-bind (CompleteOwnerSetup), and the
re-provision path all forced 'admin', so in a fresh install the entire
owner tier — list/create/edit/disable/delete users, quotas, sessions — was
unreachable. The role was a dead letter in the other direction too: staff
predicates that predate the role did not know it.

- UpsertOwner and CompleteOwnerSetup now write role='owner'; the username-
  conflict arm re-asserts it, which is also the documented pre-0011 promotion
  path ("re-provision via break-glass"). InsertOperator stays plain 'admin'.
- Staff doors learn the role: op-login start/finish admit the Owner; the
  player email door refuses it like any staff account; the in-game approver
  check already used staffRole.
- Reclaim protection: IsProtectedAdminLink (and the break-glass bootstrap
  switch AdminExists) count admin OR owner — the Owner must never be displaced
  by a Mojang-priority reclaim.
- Panel guards make migration 0011's claim true now that owner rows exist: an
  owner can never be demoted, deleted, or disabled through the API (only the
  local break-glass console resets the identity); username/email edits still
  work.

Tests: pgint pins both provisioning paths, the protected-link predicate and
the reset/promote semantics; hermetic suites cover the owner-admitting staff
door, the owner-refusing player door, the three panel guards, and break-glass
attribution.
2026-09-23 03:30:29 +08:00
Lemon-miaow d1ec40f738 fix(users): an admin email edit must clear the stale verification
UpdateUser wrote a new address but kept email_verified, so patching a verified
account asserted a proof of an address nobody had proven — and the
pre-session login mails and resolves on exactly that flag, so a typo'd edit
could hand the account's sign-in codes to the wrong mailbox.

Changing the address now clears the flag in the same write; a no-op patch that
passes the same value keeps it. The fake mirrors the semantics, and the pgint
suite pins both halves (same value keeps proof, new value drops it).
2026-09-23 03:19:55 +08:00
Lemon-miaow b6ef27cd2d fix(auth): enforce the verified-email uniqueness that email login assumes
The design has claimed since migration 0010 that at most one account can hold
a PROVEN email address, with ErrEmailTaken as the 409 a second verifier sees.
Neither half ever shipped: no migration created users_verified_email_unique,
and VerifyEmailOTP had no guard at all — the sentinel was defined but never
returned, so two accounts could both verify one address. The damage is not
cosmetic: the pre-session login resolves accounts BY verified email, so the
duplicate decided which identity a mailed sign-in code belonged to.

- Migration 0020 creates the partial unique index (lower(email) WHERE
  email_verified) the comments have been citing — the database-level backstop.
- VerifyEmailOTP now refuses the take-over with ErrEmailTaken BEFORE consuming
  the code (the address, not the code, is the problem), charges no attempt,
  and maps a lost cross-user race (unique violation) to the same answer.
- The verify handler answers 409 email_taken instead of a generic 500.

Covered by the pgint suite (sequential double-verify refused with the code
still live, a direct duplicate write still loses to the index, the refused
account can still prove its own address) and a hermetic 409 case.
2026-09-23 03:19:35 +08:00
Lemon-miaow 2a55a0d265 test(pgint): verify the business stores against a real Postgres
The hermetic suites encode the store contracts against fakes; PGRepo drifted
behind them three times (attempt accounting, a missing JOIN, a missing FOR
UPDATE) while every unit test stayed green. This harness replays the real
embedded migrations onto a throwaway database — its name must contain "pgint"
or the harness refuses to run — and exercises the SQL directly: sessions, the
onboarding email-OTP lifecycle (supersede/expiry/lockout), the pre-session
login consume, the op-login state machine, link and bind-code redemption,
submissions, and builds with the image admission round trip.

Run it after touching SQL under internal/api/pgrepo.go, internal/submit, or
internal/build; CONTRIBUTING.md carries the one-liner.
2026-09-23 03:18:56 +08:00
Lemon-miaow d2c656533e docs(audit): third-batch ledger — probes verified, build lane proven end to end, #7/#11-#15 reconciled
- The control-plane probes (0c8e29b) and the user-modpack build lane
  (f79e5eb + 02fd2de) get their live evidence recorded, including the three
  drill-only defects the lane fixed (Job scheduling, Kaniko Dockerfile
  ownership, Trivy DB egress).
- The stale 'unfixed' rows #7/#11-#15 are reconciled with their commits.
- Remaining-work list re-stated: image durability, PG contract tests, the
  panel's /jobs block, multi-node reaper placement, alerting, and the newly
  found player-visible build status gap.
2026-09-22 23:05:13 +08:00
Lemon-miaow 02fd2de502 fix(build): three drill-driven fixes so the lane actually completes on a starter node
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:

- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
  Pod never scheduled on the platform's own starter node: FailedScheduling /
  Insufficient memory, Pending forever. Requests are now a small floor
  (250m / 512Mi, never above a configured cap) while the limits stay the
  safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
  the source owner; a 65532-owned context (the distroless felis image uid)
  fails that under the pod's dropped capabilities ('copying dockerfile:
  chown /kaniko/Dockerfile: operation not permitted'). The fetch container
  now extracts as root — the uid Kaniko already runs as — so the copy
  succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
  step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
  renders --db-repository, and docs/troubleshooting.md §8e now carries the
  verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
  internal registry; --insecure already covers its plain HTTP).

Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
2026-09-22 23:01:25 +08:00
Lemon-miaow f79e5ebb5e feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:

- submit: derived context refs become the internal-face URL
  /api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
  gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
  route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
  image's new fetch-context entrypoint) that streams the blob with the
  namespace-local service-token Secret — never mounted into Kaniko — and
  extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
  reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
  build namespace gets the token Secret through the existing replica mechanism
  (bootstrap.sh + felis setup); the build egress lock opens exactly the control
  namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
  escapes/symlinks/non-gzip).

Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
2026-09-22 22:45:09 +08:00
Lemon-miaow 0c8e29b05a fix(platform): give every control-plane Deployment real probes (#8 follow-up)
The api, operator and registry Deployments shipped with no liveness/readiness
probes at all: a wedged process stayed 'Running' forever, and the operator had
no health listener to probe in the first place. Kaniko build evidence on a
fresh install showed the only cluster-wide red after a disk-pressure pass was
Deployment status that never reflected health.

- felis-api: readiness /readyz (DB + K8s API round-trip) and liveness /healthz
  on the internal face (:8081), the only listener that serves both endpoints;
  liveness deliberately avoids /readyz so a DB blip cannot restart the api.
- felis-operator: new --health-probe-bind-address (:8081) with controller-
  runtime's /healthz + /readyz (registered ping checks; an unregistered handler
  map would 404), plus the matching container port and probes.
- registry: /v2/ probes on the pinned port, so a broken storage backend stops
  reading as 'Running'.

Tests pin paths, ports, and that each probe targets a declared container port.
2026-09-22 22:22:37 +08:00
Lemon-miaow edefc34a5b docs(audit): night-2 ledger — default-install backup/reaper verified live, eviction shield re-drilled, OOM + k3s SIGKILL chaos, and the prioritized remaining-work list 2026-09-22 22:15:38 +08:00
Lemon-miaow 87a9f4eb25 feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
2026-09-22 22:03:48 +08:00
Lemon-miaow 0a2d654e68 fix(platform): control plane runs system-cluster-critical, so eviction refuses it (#8)
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.

Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
2026-09-22 21:55:54 +08:00
Lemon-miaow fe310743a2 fix(platform): give the control plane a PriorityClass eviction shield (#8)
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
2026-09-22 21:08:56 +08:00
Lemon-miaow 2b87a5a13b fix(install): grant the reaper traverse on the worlds root (#6)
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.

Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
2026-09-22 21:03:17 +08:00
Lemon-miaow fd33fd05e1 fix(install): backups exist on a default install; retention resolves real world dirs (#6)
Three faces of one gap, all on the supported install path:

- Backup/restore answered 503 out of the box: nothing ever rendered the
  archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
  PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
  'felis manifests' names it by default (--backup-pvc= is the explicit
  no-store shape); bootstrap passes it through so the generated felis.toml
  [archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
  now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
  env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
  nothing deleted).
- Even when enabled it could not find a world on a stock install:
  resolveWorldDir now also resolves the exact local-path directory
  <pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
  glob, so a stale deleted PV's bytes can't be archived in place of the
  current world). Reaper Role gains persistentvolumeclaims:get (weaker
  than the delete it already held).

README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
2026-09-22 20:48:07 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow a2df2f242b fix(operator): re-probe RCON every 2s while Starting
The readiness gate is status-driven; at a 5s re-probe cadence the observed
'container Ready but API still 409 not_running' window was 6~10s. Halving the
cadence halves the worst case; probes still only run while unreachable.
2026-09-22 20:31:53 +08:00
Lemon-miaow 2a8f897e61 fix(api): a session-store outage answers 503, not 401
Resolving a session cookie failed identically whether the credential was
missing or Postgres was unreachable: local_auth_enabled read errors fell into
the fail-closed 'disabled' branch and SessionUser errors into 'invalid
session', both surfacing as 401 'authentication required' — a lie that reads
as 'log in again' during an outage. Split the enabled-read into
(enabled, error), tag non-ErrNotFound store failures with errAuthBackend, and
map that to a new 503 auth_unavailable in requireExternal. Fail-closed is
unchanged: missing setting / bad value / missing session stay 401.
2026-09-22 20:30:43 +08:00
Lemon-miaow abce381faa fix(tui): wrap the one-time setup URL so narrow terminals can't truncate it
The setup URL carries a 43-char token and overruns 80 columns; the TUI
renderer clipped it. Break it at the query '=' boundary (token on its own
line) with a shared wrapDisplayURL helper used by both the Owner wizard and
the mc-bind wizard; unit test pins the no-loss concatenation.
2026-09-22 20:28:01 +08:00
Lemon-miaow a415246adc fix(operator): give controller-runtime a logger instead of a goroutine stack
Without SetLogger, the first reconcile prints
'[controller-runtime] log.SetLogger(...) was never called; logs will not be
displayed' followed by a full stack trace (live-observed in felis-operator).
Route it through logr.FromSlogHandler(slog.Default()) so its messages are
ordinary stderr lines; go-logr/logr promoted to a direct dependency.
2026-09-22 20:25:39 +08:00
347 changed files with 41823 additions and 2113 deletions

No files matched your search

+11
View File
@@ -31,3 +31,14 @@ Dockerfile
*.key *.key
felis felis
felis.exe felis.exe
# macOS materializes extended attributes as ._<name> sidecars (BSD tar uploads,
# Finder copies, network volumes) and leaves .DS_Store behind. Neither is
# source, and one is actively harmful: a ._*.sql beside the migrations is
# //go:embed-ed into the binary and makes every `felis migrate` fail
# ("non-numeric version") — observed live on a Mac-staged tree. Same exposure
# for any tree the other //go:embed patterns walk (deploy/, plugins/).
._*
**/._*
.DS_Store
**/.DS_Store
+23
View File
@@ -0,0 +1,23 @@
# The workflows pin every action to a commit SHA, and every Dockerfile pins its base images
# by digest. This keeps those pins moving: Dependabot reads the "# vX.Y.Z" comment next to
# each action SHA, and the tag in front of each image digest, and opens a PR that bumps both
# together.
version: 2
updates:
- package-ecosystem: github-actions
directory: /
schedule:
interval: weekly
- package-ecosystem: docker
directories:
- /
- /deploy/limbo
- /deploy/lobby
- /deploy/paper
schedule:
interval: weekly
# A new major is a runtime change (Paper 26.x needs Java 25, Limbo's jar is Java 21
# bytecode), so only digests and minors are proposed; majors move by hand.
ignore:
- dependency-name: "*"
update-types: ["version-update:semver-major"]
+113 -6
View File
@@ -14,10 +14,14 @@
# PR is what asks for the answer. # PR is what asks for the answer.
name: ci name: ci
#
# release.yml calls this workflow (workflow_call) before it builds anything, so a tag passes
# exactly these gates and there is one list of them.
on: on:
push: push:
branches: [main] branches: [main]
pull_request: pull_request:
workflow_call:
permissions: permissions:
contents: read contents: read
@@ -30,9 +34,9 @@ jobs:
go: go:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
- uses: actions/setup-go@v5 - uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with: with:
go-version-file: go.mod go-version-file: go.mod
@@ -43,12 +47,66 @@ jobs:
echo "gofmt needed on:"; echo "$unformatted"; exit 1 echo "gofmt needed on:"; echo "$unformatted"; exit 1
fi fi
- run: go vet ./... - run: go vet ./...
- run: go test ./... # -race: felis-api and the operator are mostly goroutines (watchers, the
# registry pruner, the backup scheduler, the rate limiters).
- run: go test -race ./...
# The version is pinned here and bumped by hand; Dependabot does not read `go run`.
- name: staticcheck
run: go run honnef.co/go/tools/cmd/[email protected] ./...
# Separate from the go job so a newly published advisory reads as what it is. govulncheck
# exits non-zero only for vulnerable code this module can actually reach, standard
# library included: setup-go installs the newest patch of go.mod's Go line, so a finding
# there means the Dockerfile's golang digest (which ships the release) needs a bump too.
vuln:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
- uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
go-version-file: go.mod
- run: go run golang.org/x/vuln/cmd/[email protected] ./...
# The business stores' SQL against a real PostgreSQL (internal/pgint): the unit suites run
# on fakes, and PGRepo drifted from them three times while those stayed green. 13 is the
# oldest server a supported distribution installs (EL9), 18 the newest (Arch).
pgint:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
postgres: ['13', '18']
services:
postgres:
image: postgres:${{ matrix.postgres }}
env:
POSTGRES_USER: felis
POSTGRES_PASSWORD: pgint
POSTGRES_DB: felis_pgint
ports:
- 5432:5432
options: >-
--health-cmd "pg_isready -U felis -d felis_pgint"
--health-interval 2s
--health-timeout 5s
--health-retries 30
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
- uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
go-version-file: go.mod
- run: go test -race -tags pgint -count=1 ./internal/pgint/
env:
FELIS_TEST_PG_URL: postgres://felis:pgint@localhost:5432/felis_pgint?sslmode=disable
shell: shell:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
# bootstrap.sh is the only thing that ever runs on a fresh host, and nothing here can # bootstrap.sh is the only thing that ever runs on a fresh host, and nothing here can
# run it — it wants root, a package manager and k3s. Syntax plus the extracted-block # run it — it wants root, a package manager and k3s. Syntax plus the extracted-block
@@ -66,12 +124,22 @@ jobs:
esac esac
done done
# A pinned release rather than the runner image's copy, so a runner update cannot
# change what fails. Warnings and errors fail the job; style notes (info) do not.
- name: shellcheck
run: |
curl -fsSL -o shellcheck.tar.xz \
https://github.com/koalaman/shellcheck/releases/download/v0.11.0/shellcheck-v0.11.0.linux.x86_64.tar.xz
echo "8c3be12b05d5c177a04c29e3c78ce89ac86f1595681cab149b65b97c4e227198 shellcheck.tar.xz" | sha256sum -c
tar -xJf shellcheck.tar.xz
./shellcheck-v0.11.0/shellcheck -S warning $(git ls-files '*.sh')
- run: sh deploy/bootstrap_test.sh - run: sh deploy/bootstrap_test.sh
panel: panel:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
# The Dockerfile's `FROM node:<major>` is the only place the panel's Node version is # The Dockerfile's `FROM node:<major>` is the only place the panel's Node version is
# declared — there is no .nvmrc and no engines field. Reading it here rather than # declared — there is no .nvmrc and no engines field. Reading it here rather than
@@ -84,7 +152,7 @@ jobs:
[ -n "$version" ] || { echo "Dockerfile has no 'FROM ... node:<major>' line"; exit 1; } [ -n "$version" ] || { echo "Dockerfile has no 'FROM ... node:<major>' line"; exit 1; }
echo "version=${version}" >> "$GITHUB_OUTPUT" echo "version=${version}" >> "$GITHUB_OUTPUT"
- uses: actions/setup-node@v4 - uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0
with: with:
node-version: ${{ steps.node.outputs.version }} node-version: ${{ steps.node.outputs.version }}
cache: npm cache: npm
@@ -98,3 +166,42 @@ jobs:
- run: npm run typecheck - run: npm run typecheck
working-directory: panel working-directory: panel
plugins:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
# The other jobs never touch the Java layer: the plugin jars were only ever
# compiled by bootstrap on a live host, and the three test mains under
# plugins/*/test were run by hand. JDK 21 plus the Gradle major the plugin
# Dockerfiles pin (8.14) is that same toolchain, in CI.
- uses: actions/setup-java@cf277c60eb25467037889841efdb72551f06f6c3 # v4.9.1
with:
distribution: temurin
java-version: '21'
- uses: gradle/actions/setup-gradle@ed408507eac070d1f99cc633dbcf757c94c7933a # v4.4.3
with:
gradle-version: '8.14'
- run: bash plugins/test.sh
mods:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
# The three loader mods (Minecraft 1.20.1 / 1.20.4, Java-17 lines) compile
# through their vendored Gradle wrappers, which fetch their own Gradle. Until
# this job nothing ever built them: no install path touches them, and their
# gradlew scripts were committed without the exec bit, so the README's
# one-liners failed on a fresh clone.
- uses: actions/setup-java@cf277c60eb25467037889841efdb72551f06f6c3 # v4.9.1
with:
distribution: temurin
java-version: '17'
- uses: gradle/actions/setup-gradle@ed408507eac070d1f99cc633dbcf757c94c7933a # v4.4.3
- run: bash plugins/test-mods.sh
+108 -18
View File
@@ -17,6 +17,16 @@
# quietly ships a release whose panel is that placeholder. The Dockerfile runs the npm # quietly ships a release whose panel is that placeholder. The Dockerfile runs the npm
# build first, and is the same recipe bootstrap uses, so there is one way to build felis # build first, and is the same recipe bootstrap uses, so there is one way to build felis
# rather than two that can drift. # rather than two that can drift.
#
# SHA256SUMS is a contract with bootstrap too: download_release_binary refuses a binary whose
# hash is not listed there, BEFORE it runs it. A release without the file installs by source
# build instead.
#
# The write token never meets the test suite: `gates` (ci.yml) and `build` run the tests,
# Gradle and the Docker build (each of which executes third-party code) with a read-only
# token, and `build` hands the binaries over as a workflow artifact; `publish` holds contents:write and runs only
# pinned actions and gh. Every action is pinned to a commit SHA (the tag in the trailing
# comment is for humans); .github/dependabot.yml proposes the bumps.
name: release name: release
on: on:
@@ -24,28 +34,29 @@ on:
tags: ['v*'] tags: ['v*']
permissions: permissions:
contents: write # gh release create/upload contents: read
jobs: jobs:
release: # A tag that ships red is worse than a tag that fails to ship. These are ci.yml's gates,
# called rather than copied: Go (race, vet, staticcheck), govulncheck, the PostgreSQL
# contract suite, shellcheck and the bootstrap tests, the panel, and the Java layer the
# binary EMBEDS (bootstrap_asset.go ships the plugin sources, so a tag whose plugins do
# not compile turns every install of that release into a failed bootstrap).
gates:
uses: ./.github/workflows/ci.yml
build:
needs: gates
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
# A tag that ships red is worse than a tag that fails to ship.
- run: go vet ./...
- run: go test ./...
# Both architectures, because bootstrap's default release channel DOWNLOADS these # Both architectures, because bootstrap's default release channel DOWNLOADS these
# rather than compiling on the target host — an arm64 host with no asset silently # rather than compiling on the target host — an arm64 host with no asset silently
# falls back to a slow source build. Neither stage is emulated: the Dockerfile pins # falls back to a slow source build. Neither stage is emulated: the Dockerfile pins
# both build stages to $BUILDPLATFORM and the Go stage cross-compiles via TARGETARCH, # both build stages to $BUILDPLATFORM and the Go stage cross-compiles via TARGETARCH,
# so the second architecture costs about a minute. # so the second architecture costs about a minute.
- uses: docker/setup-buildx-action@v3 - uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0
- name: Build the stamped binaries - name: Build the stamped binaries
run: | run: |
@@ -75,8 +86,67 @@ jobs:
file ./felis-linux-arm64 | grep -q 'ARM aarch64' \ file ./felis-linux-arm64 | grep -q 'ARM aarch64' \
|| { echo "felis-linux-arm64 is not an arm64 ELF — TARGETARCH did not reach the go build"; exit 1; } || { echo "felis-linux-arm64 is not an arm64 ELF — TARGETARCH did not reach the go build"; exit 1; }
# --verify-tag refuses to invent a release for a tag that is not pushed. The upload - name: Checksum the binaries
# fallback makes a re-run converge rather than failing on an existing release. run: sha256sum felis-linux-amd64 felis-linux-arm64 | tee SHA256SUMS
# A CycloneDX SBOM per binary: the Go modules (and versions) linked into it, read
# from the build info the linker embeds.
- uses: anchore/sbom-action@e22c389904149dbc22b58101806040fa8d37a610 # v0.24.0
with:
file: felis-linux-amd64
format: cyclonedx-json
output-file: felis-linux-amd64.cdx.json
upload-artifact: false
upload-release-assets: false
- uses: anchore/sbom-action@e22c389904149dbc22b58101806040fa8d37a610 # v0.24.0
with:
file: felis-linux-arm64
format: cyclonedx-json
output-file: felis-linux-arm64.cdx.json
upload-artifact: false
upload-release-assets: false
- uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: release-assets
path: |
felis-linux-amd64
felis-linux-arm64
felis-linux-amd64.cdx.json
felis-linux-arm64.cdx.json
SHA256SUMS
if-no-files-found: error
retention-days: 7
publish:
needs: build
runs-on: ubuntu-latest
permissions:
contents: write # gh release create/upload
id-token: write # the Sigstore certificate behind the provenance attestation
attestations: write
steps:
- uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
name: release-assets
# The artifact store sits between the two jobs, so check the handover too.
- run: sha256sum -c SHA256SUMS
# Signed SLSA provenance: which workflow run, commit and repository produced each
# binary. Check one with `gh attestation verify felis-linux-amd64 --repo FelisMC/Felis`.
# GitHub only stores attestations for private repositories on Enterprise Cloud, and a
# failure here would block the release, so a private repository skips the step and
# relies on SHA256SUMS alone.
- name: Attest build provenance
if: ${{ !github.event.repository.private }}
uses: actions/attest-build-provenance@977bb373ede98d70efdf65b84cb5f73e068dcc2a # v3.0.0
with:
subject-path: |
felis-linux-amd64
felis-linux-arm64
# --verify-tag refuses to invent a release for a tag that is not pushed.
# #
# The prerelease flag has to be passed explicitly: the trigger glob is v*, so v1.2.3-rc1 # The prerelease flag has to be passed explicitly: the trigger glob is v*, so v1.2.3-rc1
# lands here too, and gh does not read semver out of the tag name. Published as a full # lands here too, and gh does not read semver out of the tag name. Published as a full
@@ -84,13 +154,33 @@ jobs:
# channel installs from and `felis update` polls — so every fresh install would get the # channel installs from and `felis update` polls — so every fresh install would get the
# RC binary and every deployed felis-api would error on the felis component until a # RC binary and every deployed felis-api would error on the felis component until a
# stable tag was cut. Flagged, GitHub keeps latest pointing at the last stable release. # stable tag was cut. Flagged, GitHub keeps latest pointing at the last stable release.
#
# A re-run (the release already exists) uploads only what is missing and never
# replaces a published asset: hosts may already have installed it, and their
# SHA256SUMS check would start failing against a swapped file. An asset that is
# there with different bytes stops the job; cut a new tag instead.
- name: Publish the release - name: Publish the release
env: env:
GH_TOKEN: ${{ github.token }} GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
run: | run: |
assets="felis-linux-amd64 felis-linux-arm64 felis-linux-amd64.cdx.json felis-linux-arm64.cdx.json SHA256SUMS"
flags="" flags=""
case "$GITHUB_REF_NAME" in *-*) flags="--prerelease" ;; esac case "$GITHUB_REF_NAME" in *-*) flags="--prerelease" ;; esac
gh release create "$GITHUB_REF_NAME" --verify-tag --generate-notes $flags \ if ! gh release view "$GITHUB_REF_NAME" >/dev/null 2>&1; then
./felis-linux-amd64 ./felis-linux-arm64 \ # shellcheck disable=SC2086 # word-splitting the list is the point
|| gh release upload "$GITHUB_REF_NAME" \ gh release create "$GITHUB_REF_NAME" --verify-tag --generate-notes $flags $assets
./felis-linux-amd64 ./felis-linux-arm64 --clobber exit 0
fi
# The REST payload's per-asset "digest" is GitHub's own sha256 of the stored file.
published="$(gh api "repos/${GH_REPO}/releases/tags/${GITHUB_REF_NAME}" --jq '.assets[] | "\(.name) \(.digest)"')"
for a in $assets; do
have="$(printf '%s\n' "$published" | awk -v n="$a" '$1 == n { print $2 }')"
want="sha256:$(sha256sum < "$a" | cut -d' ' -f1)"
if [ -z "$have" ]; then
gh release upload "$GITHUB_REF_NAME" "$a"
elif [ "$have" != "$want" ]; then
echo "::error::$a is already published with $have; this run built $want. Published assets are never replaced."
exit 1
fi
done
+967 -15
View File
File diff suppressed because it is too large. Load diff
+14
View File
@@ -67,6 +67,20 @@ go test ./internal/api
go test ./cmd/felis go test ./cmd/felis
``` ```
The hermetic suites run against in-memory fakes; the business stores' SQL is
verified separately against a real Postgres, on a throwaway database whose name
must contain `pgint` (the harness drops and recreates its schema and replays the
embedded migrations):
```bash
FELIS_TEST_PG_URL='postgres://felis:***@127.0.0.1:5432/felis_pgint?sslmode=disable' \
go test -tags pgint ./internal/pgint/ -v
```
Run it after touching anything under `internal/api/pgrepo.go`, `internal/submit`,
or `internal/build` that speaks SQL: the fakes encode the contract, and this
suite exists to catch the drift between the fakes and the real queries.
Build the CLI: Build the CLI:
```bash ```bash
+7 -3
View File
@@ -21,14 +21,18 @@
# minutes. The FINAL stage is deliberately NOT pinned — it must stay on the target platform # minutes. The FINAL stage is deliberately NOT pinned — it must stay on the target platform
# or the published arm64 image would carry amd64 layers. It contains only COPY, which # or the published arm64 image would carry amd64 layers. It contains only COPY, which
# BuildKit performs itself, so it needs no QEMU either; adding a RUN there would. # BuildKit performs itself, so it needs no QEMU either; adding a RUN there would.
FROM --platform=$BUILDPLATFORM node:22-bookworm AS panel #
# Every base image here and in deploy/{limbo,lobby,paper} is pinned by digest, so a rebuild
# of one release uses the same bytes; .github/dependabot.yml proposes the bumps (tag and
# digest together).
FROM --platform=$BUILDPLATFORM node:22-bookworm@sha256:363e1587494626837fa7f9a23bdb453d13b0ff3c67c705c2805cfc69c2d2fad7 AS panel
WORKDIR /panel WORKDIR /panel
COPY panel/package*.json ./ COPY panel/package*.json ./
RUN npm ci RUN npm ci
COPY panel/ ./ COPY panel/ ./
RUN npm run build RUN npm run build
FROM --platform=$BUILDPLATFORM golang:1.26 AS build FROM --platform=$BUILDPLATFORM golang:1.26@sha256:6c2a5538f964f1c82f97ad14988bf05de100d922d159d0e398b54c7b0ca0c6c9 AS build
WORKDIR /src WORKDIR /src
ARG TARGETOS=linux ARG TARGETOS=linux
ARG TARGETARCH ARG TARGETARCH
@@ -55,7 +59,7 @@ ARG FELIS_VERSION=dev
RUN CGO_ENABLED=0 GOOS="$TARGETOS" GOARCH="${TARGETARCH:-$(go env GOARCH)}" \ RUN CGO_ENABLED=0 GOOS="$TARGETOS" GOARCH="${TARGETARCH:-$(go env GOARCH)}" \
go build -trimpath -ldflags="-s -w -X main.version=${FELIS_VERSION}" -o /out/felis ./cmd/felis go build -trimpath -ldflags="-s -w -X main.version=${FELIS_VERSION}" -o /out/felis ./cmd/felis
FROM gcr.io/distroless/static-debian12:nonroot FROM gcr.io/distroless/static-debian12:nonroot@sha256:afa5c872c891853ca7fcf1f12c3edb23f7eeef36189728842dd51042ff57f7ab
ENV PATH=/usr/local/bin:/usr/bin:/bin ENV PATH=/usr/local/bin:/usr/bin:/bin
COPY --chmod=0755 --from=build /out/felis /usr/local/bin/felis COPY --chmod=0755 --from=build /out/felis /usr/local/bin/felis
# distroless "nonroot" is uid 65532; the rendered PodSecurityContext pins # distroless "nonroot" is uid 65532; the rendered PodSecurityContext pins
+6 -5
View File
@@ -17,10 +17,11 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
- **即开即玩**:玩家尝试连接时自动唤醒服务器,空闲后自动休眠,像游戏主机一样省资源。 - **即开即玩**:玩家尝试连接时自动唤醒服务器,空闲后自动休眠,像游戏主机一样省资源。
- **Web 控制面板**:浏览器中查看服务器状态、在线玩家与资源用量,管理备份与恢复。 - **Web 控制面板**:浏览器中查看服务器状态、在线玩家与资源用量,管理备份与恢复。
- **自动备份与恢复**:定时将世界打包存档,支持从任意备份点一键回滚。 - **备份与恢复**:一键把整服数据(世界、配置、插件/模组,即整个 /data 卷)打包进集群内的归档库,支持从任意备份点回滚;默认安装就已启用(归档 PVC 与路径由安装器一并生成)。
- **智慧回收**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间。 - **控制面数据库备份**:账号、服务器归属、配额与存档索引所在的数据库每天自动备份,每次升级迁移前先快照,出错可用 `felis db restore` 整库原子回滚;面板「维护与备份」页显示备份是否新鲜(见 [故障排查 §16](docs/troubleshooting.md))。
- **智慧回收(可选开启)**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间;安装时设置 `FELIS_WORLDS_HOST_PATH`(k3s 默认 `/var/lib/rancher/k3s/storage`)即启用每日回收,不设置则不删任何世界。
- **多核心支持**:兼容 Paper、Fabric、Forge、NeoForge,经由 Velocity 代理统一入口。 - **多核心支持**:兼容 Paper、Fabric、Forge、NeoForge,经由 Velocity 代理统一入口。
- **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建并部署。 - **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建;构建产物进入镜像白名单,可直接选用为服务器镜像完成部署。
- **Passkey 登录**:支持指纹、面容、硬件密钥等无密码认证方式。 - **Passkey 登录**:支持指纹、面容、硬件密钥等无密码认证方式。
- **零信任安全**:面板流量由 Cloudflare Access 保护,集群内 API 不暴露到公网。 - **零信任安全**:面板流量由 Cloudflare Access 保护,集群内 API 不暴露到公网。
@@ -29,7 +30,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
在准备好的 Linux 主机上执行: 在准备好的 Linux 主机上执行:
```bash ```bash
curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/deploy/bootstrap.sh | sudo bash curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
``` ```
脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。 脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。
@@ -41,7 +42,7 @@ curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/de
> export FELIS_GITHUB_TOKEN=<对本仓库有读权限的 token> > export FELIS_GITHUB_TOKEN=<对本仓库有读权限的 token>
> printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \ > printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \
> | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \ > | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \
> https://api.github.com/repos/MliroLirrorsIngenuity/Felis/contents/deploy/bootstrap.sh \ > https://api.github.com/repos/FelisMC/Felis/contents/deploy/bootstrap.sh \
> | sudo -E bash > | sudo -E bash
> ``` > ```
> >
+26 -4
View File
@@ -17,10 +17,11 @@ Table of Contents
- **Wake on Join**: Servers start automatically when a player connects, and stop when idle — like hibernate for your server. - **Wake on Join**: Servers start automatically when a player connects, and stop when idle — like hibernate for your server.
- **Web Dashboard**: Monitor server status, online players, and resource usage from your browser, with backup and restore management. - **Web Dashboard**: Monitor server status, online players, and resource usage from your browser, with backup and restore management.
- **Auto Backup & Restore**: Scheduled world backups with one-click rollback from any backup point. - **Backup & Restore**: One-click snapshots of a server's whole data volume (worlds, config, plugins/mods — the entire /data volume) into the cluster's archive store, with rollback from any backup point — enabled by default (the installer renders the archive PVC and its path).
- **World Reaper**: Worlds idle for more than 15 days are automatically backed up and removed to free disk space. - **Control-plane database backups**: The database holding accounts, server ownership, quotas and the archive index is backed up daily and snapshotted before every upgrade migrates it; `felis db restore` rolls it back atomically, and the panel's Maintenance & Backups page shows whether the newest backup is fresh (see [troubleshooting §16](docs/troubleshooting.md)).
- **World Reaper** (opt in): Worlds idle for more than 15 days are automatically backed up and removed to free disk space. Enable it by setting `FELIS_WORLDS_HOST_PATH` at install time (on k3s: `/var/lib/rancher/k3s/storage`); without it, no world is ever deleted.
- **Multi-core Support**: Compatible with Paper, Fabric, Forge, and NeoForge, federated behind a Velocity proxy. - **Multi-core Support**: Compatible with Paper, Fabric, Forge, and NeoForge, federated behind a Velocity proxy.
- **Modpack Submission**: Players submit custom modpacks; admin approval triggers automatic build and deployment. - **Modpack Submission**: Players submit custom modpacks; admin approval triggers an automatic build, and the result is whitelisted as a server image you can select to deploy.
- **Passkey Login**: Passwordless authentication via fingerprint, face recognition, or hardware security keys. - **Passkey Login**: Passwordless authentication via fingerprint, face recognition, or hardware security keys.
- **Zero Trust Security**: Panel traffic protected by Cloudflare Access; the internal API is never exposed to the internet. - **Zero Trust Security**: Panel traffic protected by Cloudflare Access; the internal API is never exposed to the internet.
@@ -29,11 +30,32 @@ Table of Contents
On a prepared Linux host, run: On a prepared Linux host, run:
```bash ```bash
curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/deploy/bootstrap.sh | sudo bash curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
``` ```
The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain. The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain.
> **This repository is currently private**, so the command above returns 404. Use the
> credentialed form instead; the installer itself needs the same token to resolve and
> download the release, so pass it through with `sudo -E`:
>
> ```bash
> export FELIS_GITHUB_TOKEN=<a token with read access to this repository>
> printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \
> | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \
> https://api.github.com/repos/FelisMC/Felis/contents/deploy/bootstrap.sh \
> | sudo -E bash
> ```
>
> The token reaches `curl --config -` over stdin instead of the command line: argv is
> readable by any local user via `/proc`, and that is exactly why the installer's
> internal `github_api` uses the same form.
Rerunning this command is also how you upgrade felis-api to a newer version (`felis setup`
cannot — it uses the binary already installed on the host). The rerun keeps the installed
root domain but **not** the channel: if this host follows main, also
`export FELIS_VERSION_BOOTSTRAP=dev`.
## Build from Source ## Build from Source
Felis is built with Go and Node.js: Felis is built with Go and Node.js:
+1
View File
@@ -24,6 +24,7 @@ var bootstrapAssets embed.FS
// otherwise be baked into every felis binary. Keep them explicit — add a source // otherwise be baked into every felis binary. Keep them explicit — add a source
// directory here, never a parent. // directory here, never a parent.
// //
//go:embed deploy/game-stack.lock
//go:embed deploy/limbo/Dockerfile deploy/limbo/entrypoint.sh //go:embed deploy/limbo/Dockerfile deploy/limbo/entrypoint.sh
//go:embed deploy/lobby/Dockerfile deploy/lobby/entrypoint.sh //go:embed deploy/lobby/Dockerfile deploy/lobby/entrypoint.sh
//go:embed deploy/paper/Dockerfile deploy/paper/entrypoint.sh //go:embed deploy/paper/Dockerfile deploy/paper/entrypoint.sh
+108
View File
@@ -2,6 +2,7 @@ package felis
import ( import (
"io/fs" "io/fs"
"os"
"regexp" "regexp"
"strings" "strings"
"testing" "testing"
@@ -205,6 +206,113 @@ func requireEmbedded(t *testing.T, path string) {
} }
} }
// The lock file is the install's only source of upstream builds, and bootstrap.sh reads it
// with a strict KEY=value parser that dies on anything unexpected, so a malformed lock is a
// failed install on every host. Check the shipped copy the same way here.
func TestGameStackLockIsComplete(t *testing.T) {
lock := map[string]string{}
for line := range strings.SplitSeq(readGameStackFile(t, "deploy/game-stack.lock"), "\n") {
if line == "" || strings.HasPrefix(line, "#") {
continue
}
k, v, ok := strings.Cut(line, "=")
if !ok {
t.Fatalf("not a KEY=value line: %q", line)
}
lock[k] = v
}
m := regexp.MustCompile(`GAME_STACK_LOCK_KEYS="([^"]*)"`).FindStringSubmatch(BootstrapScript())
if m == nil {
t.Fatal("bootstrap.sh no longer declares GAME_STACK_LOCK_KEYS")
}
keys := strings.Fields(m[1])
sha := regexp.MustCompile(`^[0-9a-f]{64}$`)
for _, k := range keys {
v, ok := lock[k]
if !ok || v == "" {
t.Errorf("game-stack.lock does not set %s", k)
continue
}
if strings.HasSuffix(k, "_SHA256") && !sha.MatchString(v) {
t.Errorf("%s=%q is not a lowercase sha256", k, v)
}
// A moving URL pins nothing: the digest check would start failing the day
// upstream publishes the next build.
if strings.HasSuffix(k, "_URL") && strings.Contains(v, "lastSuccessfulBuild") {
t.Errorf("%s names a moving build: %s", k, v)
}
}
for k := range lock {
if !strings.Contains(" "+m[1]+" ", " "+k+" ") {
t.Errorf("game-stack.lock sets %s, which bootstrap.sh refuses as an unknown key", k)
}
}
// Fill's URLs are content-addressed; a lock whose digest disagrees with its own URL
// was edited by hand and half-way.
for _, name := range []string{"PAPER", "VELOCITY"} {
if !strings.Contains(lock[name+"_JAR_URL"], "/objects/"+lock[name+"_JAR_SHA256"]+"/") {
t.Errorf("%s_JAR_SHA256 is not the digest in %s_JAR_URL", name, name)
}
}
if !strings.Contains(lock["LIMBO_JAR_URL"], "-"+lock["MC_VERSION"]+".jar") {
t.Errorf("LIMBO_JAR_URL %s is not a Minecraft %s build", lock["LIMBO_JAR_URL"], lock["MC_VERSION"])
}
if !strings.Contains(lock["PAPER_JAR_URL"], "/paper-"+lock["MC_VERSION"]+"-") {
t.Errorf("PAPER_JAR_URL %s is not a Minecraft %s build; the lobby would not speak the login gate's protocol", lock["PAPER_JAR_URL"], lock["MC_VERSION"])
}
}
// Each downloaded jar's digest is a build-arg bootstrap.sh passes and the Dockerfile must
// both require and spend on the file it downloaded; docker only warns about an unknown
// --build-arg, so a renamed arg would ship an unchecked jar.
func TestGameStackDigestsReachTheImageBuilds(t *testing.T) {
script := BootstrapScript()
for _, c := range []struct{ dockerfile, arg, path string }{
{"deploy/limbo/Dockerfile", "LIMBO_JAR_SHA256", "/limbo/Limbo.jar"},
{"deploy/limbo/Dockerfile", "LIMBO_SCHEM_SHA256", "/limbo/spawn.schem"},
{"deploy/lobby/Dockerfile", "LUCKPERMS_JAR_SHA256", "/paper/plugins/LuckPerms.jar"},
} {
if !strings.Contains(script, "--build-arg "+c.arg+"=\"$"+c.arg+"\"") {
t.Errorf("bootstrap.sh never passes --build-arg %s", c.arg)
}
dockerfile := readGameStackFile(t, c.dockerfile)
if !strings.Contains(dockerfile, "ARG "+c.arg) {
t.Errorf("%s declares no ARG %s", c.dockerfile, c.arg)
}
if !strings.Contains(dockerfile, `echo "$`+c.arg+` `+c.path+`" | sha256sum -c`) {
t.Errorf("%s never verifies %s against %s", c.dockerfile, c.path, c.arg)
}
}
}
// A base image named by tag alone is whatever the tag points at on build day.
func TestDockerfileBaseImagesArePinnedByDigest(t *testing.T) {
root, err := os.ReadFile("Dockerfile")
if err != nil {
t.Fatal(err)
}
files := map[string]string{"Dockerfile": string(root)}
for _, name := range []string{"deploy/limbo/Dockerfile", "deploy/lobby/Dockerfile", "deploy/paper/Dockerfile"} {
files[name] = readGameStackFile(t, name)
}
pinned := regexp.MustCompile(`^FROM (--platform=\S+ )?\S+:\S+@sha256:[0-9a-f]{64}( AS \S+)?$`)
for name, body := range files {
n := 0
for line := range strings.SplitSeq(body, "\n") {
if !strings.HasPrefix(line, "FROM ") {
continue
}
n++
if !pinned.MatchString(line) {
t.Errorf("%s: %q is not pinned by digest", name, line)
}
}
if n == 0 {
t.Errorf("%s has no FROM line", name)
}
}
}
func readGameStackFile(t *testing.T, name string) string { func readGameStackFile(t *testing.T, name string) string {
t.Helper() t.Helper()
b, err := gameStackAssets.ReadFile(name) b, err := gameStackAssets.ReadFile(name)
+256 -15
View File
@@ -5,10 +5,12 @@ import (
"flag" "flag"
"fmt" "fmt"
"io" "io"
"log/slog"
"net/http" "net/http"
"os" "os"
"regexp" "regexp"
"strings" "strings"
"sync/atomic"
"time" "time"
"felis.lolicon.best/internal/api" "felis.lolicon.best/internal/api"
@@ -17,10 +19,15 @@ import (
"felis.lolicon.best/internal/build" "felis.lolicon.best/internal/build"
"felis.lolicon.best/internal/config" "felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/fileedit" "felis.lolicon.best/internal/fileedit"
"felis.lolicon.best/internal/imagepin"
"felis.lolicon.best/internal/mail" "felis.lolicon.best/internal/mail"
"felis.lolicon.best/internal/metrics"
"felis.lolicon.best/internal/naming"
"felis.lolicon.best/internal/panel" "felis.lolicon.best/internal/panel"
"felis.lolicon.best/internal/passkey" "felis.lolicon.best/internal/passkey"
"felis.lolicon.best/internal/platform" "felis.lolicon.best/internal/platform"
"felis.lolicon.best/internal/reaper"
"felis.lolicon.best/internal/registryprune"
"felis.lolicon.best/internal/restore" "felis.lolicon.best/internal/restore"
"felis.lolicon.best/internal/store" "felis.lolicon.best/internal/store"
"felis.lolicon.best/internal/submit" "felis.lolicon.best/internal/submit"
@@ -108,6 +115,8 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
return 1 return 1
} }
metrics.SetBuildInfo("api", resolvedVersion())
token := os.Getenv("FELIS_SERVICE_TOKEN") token := os.Getenv("FELIS_SERVICE_TOKEN")
if token == "" { if token == "" {
fmt.Fprintln(stderr, "felis api: warning: FELIS_SERVICE_TOKEN unset — internal face will reject all callers") fmt.Fprintln(stderr, "felis api: warning: FELIS_SERVICE_TOKEN unset — internal face will reject all callers")
@@ -142,11 +151,17 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
// build namespace and pushes to the internal registry. The build Pod never // build namespace and pushes to the internal registry. The build Pod never
// holds DB credentials — felis-api owns the PG store and admits scanned // holds DB credentials — felis-api owns the PG store and admits scanned
// images, so the Builder is constructed here with both bindings. // images, so the Builder is constructed here with both bindings.
buildCfg := buildConfig(cfg)
// The fetch initContainer runs THIS image's fetch-context entrypoint, so the
// build config carries the api's own image (the platform sets FELIS_IMAGE).
buildCfg.FelisImage = os.Getenv("FELIS_IMAGE")
buildJobs := build.NewK8sJobs(cl, buildCfg)
builder := &build.Builder{ builder := &build.Builder{
Store: build.NewPGStore(drv.DB()), Store: build.NewPGStore(drv.DB()),
Jobs: build.NewK8sJobs(cl, buildConfig(cfg)), Jobs: buildJobs,
Config: buildConfig(cfg), Config: buildCfg,
} }
go probeBuildUserNamespaces(ctx, buildJobs, buildCfg, stderr)
// User-modpack approval lane (user-directed extension over §16; see // User-modpack approval lane (user-directed extension over §16; see
// internal/submit). An ordinary user may only SUBMIT a // internal/submit). An ordinary user may only SUBMIT a
@@ -159,14 +174,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
// The blob upload transport is selected by the shape of user_uploads_context — // The blob upload transport is selected by the shape of user_uploads_context —
// the two backends the setup wizard chooses between. A local path wires // the two backends the setup wizard chooses between. A local path wires
// LocalContextStore (the mounted uploads PVC); an s3:// base wires // LocalContextStore (the mounted uploads PVC); an s3:// base wires
// S3ContextStore when its credentials resolve. Either way the store's target is // S3ContextStore when its credentials resolve. Anything else — or an s3:// base
// derived from the SAME config field the context ref uses, so the blob lands // with no credentials configured — leaves Blobs nil so POST
// exactly where Kaniko's --context points. Anything else — or an s3:// base with
// no credentials configured — leaves Blobs nil so POST
// /me/submissions/{id}/context returns 503, honest like the restore executor // /me/submissions/{id}/context returns 503, honest like the restore executor
// when its PVC is not supplied. (Letting the sandboxed Kaniko build Pod READ the // when its PVC is not supplied.
// context — PVC mount for local, creds+egress for S3 — is a separate deployment //
// integration.) // Reading the blob back is the API's job, not Kaniko's: the build Pod runs in
// another namespace and can neither mount the uploads PVC (a PVC does not cross
// namespaces) nor hold object-store credentials, so ContextBaseURL makes the
// derived context ref an internal-face URL that the build Job's fetch
// initContainer streams (cmd/felis fetch-context). The platform renders this
// address into the api Deployment (felis API base URL env); the fallback keeps
// a hand-rolled deployment working under the platform's default control
// namespace.
contextBase := cfg.Registry.UserUploadsContext contextBase := cfg.Registry.UserUploadsContext
var blobs submit.Blobs var blobs submit.Blobs
switch { switch {
@@ -186,11 +206,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
fmt.Fprintf(stderr, "felis api: user-uploads context %q is neither a local path nor an s3:// base — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", contextBase) fmt.Fprintf(stderr, "felis api: user-uploads context %q is neither a local path nor an s3:// base — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", contextBase)
} }
submissions := &submit.Manager{ submissions := &submit.Manager{
Store: submit.NewPGStore(drv.DB()), Store: submit.NewPGStore(drv.DB()),
Builds: builder, Builds: builder,
Registry: cfg.Registry.URL, Registry: cfg.Registry.URL,
ContextStore: contextBase, ContextStore: contextBase,
Blobs: blobs, ContextBaseURL: internalAPIBaseURL(),
Blobs: blobs,
}
if v := cfg.Registry.UserUploadsMaxBytes; v != "" {
if n, err := parseByteSize(v); err != nil || n <= 0 {
fmt.Fprintf(stderr, "felis api: [registry] user_uploads_max_bytes %q is not a positive size such as 4Gi; keeping the default\n", v)
} else {
submissions.MaxStoredBytesTotal = n
}
} }
// Restore subsystem (spec §7): the weak-SA restore Job mounts the target // Restore subsystem (spec §7): the weak-SA restore Job mounts the target
@@ -247,9 +275,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
// agree on what local auth knows. // agree on what local auth knows.
repo := api.NewPGRepo(drv.DB()) repo := api.NewPGRepo(drv.DB())
// The owner's on-demand backup levers come from [archive], the same keys the
// backup Job and the reaper read. A malformed key leaves the defaults in
// place here; the reaper Job fails on it and names it.
rcfg, err := reaperConfig(cfg)
if err != nil {
fmt.Fprintf(stderr, "felis api: %v; using the default backup limits\n", err)
rcfg = reaper.DefaultConfig()
}
cluster := api.NewK8sCluster(cl, cfg.K8s.Namespace)
a := &api.API{ a := &api.API{
Repo: repo, Repo: repo,
Cluster: api.NewK8sCluster(cl, cfg.K8s.Namespace), Cluster: cluster,
Console: api.NewK8sConsole(cl, cfg.K8s.Namespace), Console: api.NewK8sConsole(cl, cfg.K8s.Namespace),
Logs: api.NewK8sLogStreamer(clientset, cfg.K8s.Namespace), Logs: api.NewK8sLogStreamer(clientset, cfg.K8s.Namespace),
// Build-log stream (spec §16) is scoped to the BUILD namespace — the same // Build-log stream (spec §16) is scoped to the BUILD namespace — the same
@@ -257,8 +295,10 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
BuildLogs: api.NewK8sBuildLogStreamer(clientset, cfg.Registry.BuildNamespace), BuildLogs: api.NewK8sBuildLogStreamer(clientset, cfg.Registry.BuildNamespace),
Internal: api.BearerTokenAuth{Token: token}, Internal: api.BearerTokenAuth{Token: token},
Builder: builder, Builder: builder,
Images: imagePinner(cfg.Registry.URL),
Restorer: restorer, Restorer: restorer,
Backuper: backuper, Backuper: backuper,
JobStatus: api.NewK8sJobStatus(cl, cfg.K8s.Namespace),
Files: files, Files: files,
Submissions: submissions, Submissions: submissions,
Mailer: mailer, Mailer: mailer,
@@ -279,12 +319,33 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
AdminHostname: cfg.Auth.AdminHostname, AdminHostname: cfg.Auth.AdminHostname,
PanelHostname: cfg.Auth.PanelHostname, PanelHostname: cfg.Auth.PanelHostname,
WakeCooldown: 30 * time.Second, WakeCooldown: 30 * time.Second,
// An owner may start one backup per server per manual_cooldown, and none
// while the store is at max_local_bytes (data-durability-9).
BackupCooldown: rcfg.ManualCooldown,
BackupStoreCap: rcfg.MaxLocalBytes,
// The user-modpack lane's per-user throttles: a create spaces out
// review-queue rows, an upload spaces out (up to 1 GiB) context streams.
// Separate keys, so the normal create→upload sequence stays immediate.
SubmitCreateCooldown: 30 * time.Second,
SubmitUploadCooldown: 15 * time.Second,
// Bound concurrent console/build-log SSE streams per principal. Generous enough // Bound concurrent console/build-log SSE streams per principal. Generous enough
// for legitimate multi-tab / multi-server watching, while capping how many // for legitimate multi-tab / multi-server watching, while capping how many
// upstream follow connections a single caller can tie up if their streams stall. // upstream follow connections a single caller can tie up if their streams stall.
MaxStreamsPerPrincipal: 16, MaxStreamsPerPrincipal: 16,
// Public sign-in doors, per client address: a person signing in makes a
// handful of calls, so 20 at once refilled at 20 a minute never bites a
// real user and still turns a spray into a trickle. The client address
// is the edge's header when the install names one (config.AuthConfig).
AuthDoorLimit: api.RateLimit{Burst: 20, PerMinute: 20},
ClientIPHeader: cfg.Auth.EffectiveClientIPHeader(),
MailLimit: mailLimit(cfg.SMTP.MaxPerHour),
} }
fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)") fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)")
if a.ClientIPHeader != "" {
fmt.Fprintf(stderr, "felis api: sign-in rate limit keys on the %s header\n", a.ClientIPHeader)
} else {
fmt.Fprintln(stderr, "felis api: sign-in rate limit keys on the TCP peer ([auth] client_ip_header unset)")
}
// Felis-nano: the multi-source hasJoined multiplexer. Mojang leads as the code-owned // Felis-nano: the multi-source hasJoined multiplexer. Mojang leads as the code-owned
// identity anchor (正版优先); config can only append namespace-rewritten third-party // identity anchor (正版优先); config can only append namespace-rewritten third-party
@@ -348,6 +409,11 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
// reconciles it, but this loop converges builds nobody is polling. // reconciles it, but this loop converges builds nobody is polling.
go reconcileBuilds(ctx, builder, stderr) go reconcileBuilds(ctx, builder, stderr)
if pruner := registryPruner(cfg, builder.Store, cluster, stderr); pruner != nil {
go pruner.Loop(ctx, registryPruneInterval)
}
go reapRejectedContexts(ctx, submissions, stderr)
select { select {
case <-ctx.Done(): case <-ctx.Done():
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second) shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
@@ -401,9 +467,60 @@ func buildConfig(cfg *config.Config) build.Config {
return build.Config{ return build.Config{
Namespace: cfg.Registry.BuildNamespace, Namespace: cfg.Registry.BuildNamespace,
RegistryURL: cfg.Registry.URL, RegistryURL: cfg.Registry.URL,
// Empty overrides fall back to the registry's copies of the tools
// (build.Tools), which felis mirror-build-tools keeps current.
KanikoImage: cfg.Registry.KanikoImage,
TrivyImage: cfg.Registry.TrivyImage,
CPULimit: cfg.Registry.BuildCPULimit,
MemLimit: cfg.Registry.BuildMemLimit,
DiskLimit: cfg.Registry.BuildDiskLimit,
// "auto" follows the startup probe (see probeBuildUserNamespaces).
UserNamespaces: cfg.Registry.BuildUserNamespaces,
UserNamespacesProbe: new(atomic.Bool),
RuntimeClass: cfg.Registry.BuildRuntimeClass,
MaxConcurrent: cfg.Registry.MaxConcurrentBuilds,
TrivyDBRepository: cfg.Registry.TrivyDBRepository,
TrivyJavaDBRepository: cfg.Registry.TrivyJavaDBRepository,
// The submit lane's derived context URLs live here; the fetch step's
// service token goes nowhere else.
ContextOrigin: internalAPIBaseURL(),
} }
} }
// probeBuildUserNamespaces settles build_user_namespaces = "auto": one probe
// pod with hostUsers: false tells whether this node's kernel and runtime can run
// build pods in a user namespace. Builds submitted before it answers run without.
func probeBuildUserNamespaces(ctx context.Context, jobs *build.K8sJobs, cfg build.Config, stderr io.Writer) {
if mode := cfg.UserNamespaces; mode != "" && mode != build.UserNamespacesAuto {
return
}
if cfg.FelisImage == "" {
fmt.Fprintln(stderr, "felis api: FELIS_IMAGE unset — build pods run without a user namespace")
return
}
ok, err := jobs.ProbeUserNamespaces(ctx, cfg.FelisImage)
cfg.UserNamespacesProbe.Store(ok)
switch {
case ok:
fmt.Fprintln(stderr, "felis api: build pods run in a user namespace (hostUsers: false)")
case err != nil:
fmt.Fprintf(stderr, "felis api: build pods run without a user namespace: the probe failed: %v\n", err)
default:
fmt.Fprintln(stderr, "felis api: build pods run without a user namespace: this node cannot start a pod with hostUsers: false")
}
}
// internalAPIBaseURL resolves the platform's internal-face base URL: the address
// the platform rendered into this pod (felis API base URL env), or — for a
// hand-rolled deployment that set none — the platform default control namespace,
// the same fallback setup.go uses to hand the login gate its address.
func internalAPIBaseURL() string {
if base := os.Getenv(naming.EnvAPIBaseURL); base != "" {
return base
}
return platform.InternalAPIBaseURL(platform.DefaultControlNamespace)
}
// uploadsSchemeRE matches a leading URL scheme like "s3://" or "gs://". // uploadsSchemeRE matches a leading URL scheme like "s3://" or "gs://".
var uploadsSchemeRE = regexp.MustCompile(`^[a-zA-Z][a-zA-Z0-9+.-]*://`) var uploadsSchemeRE = regexp.MustCompile(`^[a-zA-Z][a-zA-Z0-9+.-]*://`)
@@ -506,3 +623,127 @@ func reconcileBuilds(ctx context.Context, b *build.Builder, stderr io.Writer) {
} }
} }
} }
// reapRejectedContexts deletes, once an hour, the uploaded contexts of
// submissions rejected more than submit.RejectedContextRetention ago. Without it a
// rejected modpack keeps its bytes on the uploads store (and against its
// submitter's budget) until an admin deletes the row.
func reapRejectedContexts(ctx context.Context, m *submit.Manager, stderr io.Writer) {
t := time.NewTicker(time.Hour)
defer t.Stop()
for {
n, err := m.ReapRejected(ctx, submit.RejectedContextRetention)
if err != nil {
fmt.Fprintf(stderr, "felis api: reap rejected uploads: %v\n", err)
}
if n > 0 {
fmt.Fprintf(stderr, "felis api: deleted the uploaded contexts of %d rejected submission(s)\n", n)
}
select {
case <-ctx.Done():
return
case <-t.C:
}
}
}
// registryPruneInterval spaces the registry pruner's runs. The registry-gc
// sidecar sweeps once a day, so pruning more often only changes which sweep frees
// a layer.
const registryPruneInterval = 6 * time.Hour
// registryPruner deletes the registry manifests nothing references
// (internal/registryprune); the registry-gc sidecar frees their layers on its next
// sweep. It acts as the gate's prune principal, whose token the api Deployment
// injects from felis-registry-auth. Without the token the registry only grows,
// which is said once here.
func registryPruner(cfg *config.Config, store imageRefStore, servers serverLister, stderr io.Writer) *registryprune.Pruner {
if cfg.Registry.URL == "" {
return nil
}
token := os.Getenv(platform.RegistryPruneTokenEnv)
if token == "" {
fmt.Fprintf(stderr, "felis api: registry pruner disabled (%s unset) — images nothing uses are never deleted from the registry\n", platform.RegistryPruneTokenEnv)
return nil
}
static := append([]string{os.Getenv("FELIS_IMAGE")}, buildConfig(cfg).ToolRefs()...)
return &registryprune.Pruner{
Registry: &registryprune.Client{Endpoint: "http://" + cfg.Registry.URL, Token: token},
Host: cfg.Registry.URL,
Refs: func(ctx context.Context) ([]string, error) {
return inUseImageRefs(ctx, store, servers, static)
},
Log: slog.New(slog.NewTextHandler(stderr, nil)),
}
}
type imageRefStore interface {
ListImages(ctx context.Context) ([]build.Image, error)
ListUnfinishedBuilds(ctx context.Context) ([]build.Build, error)
}
type serverLister interface {
ListServers(ctx context.Context) ([]api.ServerInfo, error)
PodImages(ctx context.Context) ([]string, error)
}
// inUseImageRefs lists every image reference the platform still depends on: the
// whitelist (disabled rows too, an admin may enable them again), every server's
// spec, the images the game pods run, builds still running, and the images the
// control plane and the build Jobs run. Any source failing fails the whole list,
// so the pruner never decides on a partial view.
//
// The pods matter for the felis image: a running server keeps the one it started
// with across platform upgrades (operator.PodTemplateAnnotation), which after a
// few releases is no longer among the newest tags the pruner keeps anyway, and
// the pod needs it again whenever it is recreated.
func inUseImageRefs(ctx context.Context, store imageRefStore, servers serverLister, static []string) ([]string, error) {
refs := append([]string(nil), static...)
images, err := store.ListImages(ctx)
if err != nil {
return nil, fmt.Errorf("image whitelist: %w", err)
}
for _, img := range images {
refs = append(refs, img.ImageRef)
}
srvs, err := servers.ListServers(ctx)
if err != nil {
return nil, fmt.Errorf("servers: %w", err)
}
for _, s := range srvs {
refs = append(refs, s.Image)
}
podImages, err := servers.PodImages(ctx)
if err != nil {
return nil, fmt.Errorf("game pods: %w", err)
}
refs = append(refs, podImages...)
builds, err := store.ListUnfinishedBuilds(ctx)
if err != nil {
return nil, fmt.Errorf("running builds: %w", err)
}
for _, b := range builds {
refs = append(refs, b.ImageRef)
}
return refs, nil
}
// mailLimit turns smtp.max_per_hour into the API's install-wide mail bucket:
// the hourly cap as the refill rate, with a quarter of it (at least 5) allowed
// at once so a burst of real sign-ins is not queued behind the average.
func mailLimit(perHour int) api.RateLimit {
if perHour <= 0 {
perHour = config.DefaultMailPerHour
}
return api.RateLimit{Burst: max(perHour/4, 5), PerMinute: float64(perHour) / 60}
}
// imagePinner resolves a new server's image against the platform registry
// through its in-cluster Service, the address its refs already spell. An install
// without a registry has no platform-built images to pin.
func imagePinner(registry string) api.ImagePinner {
if registry == "" {
return nil
}
return imagepin.Resolver{Registry: registry}
}
+90
View File
@@ -1,9 +1,14 @@
package main package main
import ( import (
"context"
"errors"
"fmt"
"net/http" "net/http"
"testing" "testing"
"felis.lolicon.best/internal/api"
"felis.lolicon.best/internal/build"
"felis.lolicon.best/internal/config" "felis.lolicon.best/internal/config"
) )
@@ -44,6 +49,32 @@ func TestAuthSourcesFromConfig(t *testing.T) {
} }
} }
// TestBuildConfig_ProjectsOverrides pins the [registry] overrides reaching the
// build subsystem: unset fields must stay EMPTY (the build package's compiled-in
// defaults apply there, not here), and set fields must pass through verbatim —
// an air-gapped install points these at its imported mirrors.
func TestBuildConfig_ProjectsOverrides(t *testing.T) {
empty := buildConfig(&config.Config{})
if empty.KanikoImage != "" || empty.TrivyImage != "" || empty.CPULimit != "" || empty.MemLimit != "" {
t.Errorf("empty registry config must project empty overrides (defaults live in internal/build), got %+v", empty)
}
full := buildConfig(&config.Config{Registry: config.RegistryConfig{
URL: "registry.felis.svc:5000",
BuildNamespace: "felis-build",
KanikoImage: "reg/kaniko:v1",
TrivyImage: "reg/trivy:v1",
BuildCPULimit: "1",
BuildMemLimit: "2Gi",
}})
if full.KanikoImage != "reg/kaniko:v1" || full.TrivyImage != "reg/trivy:v1" ||
full.CPULimit != "1" || full.MemLimit != "2Gi" {
t.Errorf("registry overrides did not reach build.Config: %+v", full)
}
if full.Namespace != "felis-build" || full.RegistryURL != "registry.felis.svc:5000" {
t.Errorf("namespace/registry url must keep projecting: %+v", full)
}
}
// TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every // TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every
// felis-api listener: the shared factory must bound the header and idle phases // felis-api listener: the shared factory must bound the header and idle phases
// (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because // (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because
@@ -65,3 +96,62 @@ func TestNewAPIServerSetsHardenedTimeouts(t *testing.T) {
t.Errorf("ReadTimeout = %v, want 0 (unset) so a slow SSE attach is not capped", srv.ReadTimeout) t.Errorf("ReadTimeout = %v, want 0 (unset) so a slow SSE attach is not capped", srv.ReadTimeout)
} }
} }
type fakeRefStore struct {
images []build.Image
builds []build.Build
err error
}
func (f fakeRefStore) ListImages(context.Context) ([]build.Image, error) { return f.images, f.err }
func (f fakeRefStore) ListUnfinishedBuilds(context.Context) ([]build.Build, error) {
return f.builds, nil
}
type fakeServers struct {
list []api.ServerInfo
pods []string
podsErr error
}
func (f fakeServers) ListServers(context.Context) ([]api.ServerInfo, error) { return f.list, nil }
func (f fakeServers) PodImages(context.Context) ([]string, error) { return f.pods, f.podsErr }
// The registry pruner deletes whatever this list does not name, so every source of
// a reference has to be in it, and a failing source must fail the list.
func TestInUseImageRefsCoversEverySource(t *testing.T) {
const reg = "registry.felis.svc:5000/"
store := fakeRefStore{
images: []build.Image{{ImageRef: reg + "modpacks/pack:*"}, {ImageRef: reg + "felis/paper:demo"}},
builds: []build.Build{{ImageRef: reg + "user-uploads/sub-9:latest"}},
}
servers := fakeServers{
list: []api.ServerInfo{{Name: "s1", Image: reg + "felis/paper:demo@sha256:" + fmt.Sprintf("%064d", 1)}},
pods: []string{reg + "felis/felis:v1.0.0"},
}
got, err := inUseImageRefs(context.Background(), store, servers, []string{reg + "felis/felis:b60"})
if err != nil {
t.Fatal(err)
}
want := []string{
reg + "felis/felis:b60",
reg + "modpacks/pack:*", reg + "felis/paper:demo",
reg + "felis/paper:demo@sha256:" + fmt.Sprintf("%064d", 1),
reg + "felis/felis:v1.0.0",
reg + "user-uploads/sub-9:latest",
}
if fmt.Sprint(got) != fmt.Sprint(want) {
t.Fatalf("refs = %v\nwant %v", got, want)
}
servers.podsErr = errors.New("apiserver down")
if _, err := inUseImageRefs(context.Background(), store, servers, nil); err == nil {
t.Fatal("a failing pod list produced a reference list")
}
servers.podsErr = nil
store.err = errors.New("db down")
if _, err := inUseImageRefs(context.Background(), store, servers, nil); err == nil {
t.Fatal("a failing whitelist read produced a reference list")
}
}
+13
View File
@@ -243,6 +243,19 @@ func buildMinecraftServerFromApplyRequest(req applyRequest, namespace string) (*
AutostartPolicy: policy, AutostartPolicy: policy,
Storage: v1alpha1.StorageSpec{Size: storageQ.String()}, Storage: v1alpha1.StorageSpec{Size: storageQ.String()},
Resources: corev1.ResourceRequirements{Limits: limits, Requests: requests}, Resources: corev1.ResourceRequirements{Limits: limits, Requests: requests},
// The rest matches what felis-api's create writes (K8sCluster.CreateServer):
// fall back to the login gate while stopped, RCON on (readiness, the
// player count and the console all ride it; the operator mints the
// password), and the default idle stop.
FallbackServer: naming.SystemLoginServer,
Rcon: v1alpha1.RconSpec{
Enabled: true,
SecretRef: v1alpha1.SecretKeyRef{
Name: naming.RconSecretName(req.Name),
Key: naming.RconSecretKey,
},
},
Idle: v1alpha1.DefaultIdle(),
}, },
}, nil }, nil
} }
+13
View File
@@ -6,6 +6,7 @@ import (
"testing" "testing"
"felis.lolicon.best/internal/apis/felis/v1alpha1" "felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/naming"
corev1 "k8s.io/api/core/v1" corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/api/resource" "k8s.io/apimachinery/pkg/api/resource"
) )
@@ -167,6 +168,18 @@ func TestBuildMinecraftServerFromApplyRequest_Valid(t *testing.T) {
if ms.Spec.Storage.Size != "20Gi" { if ms.Spec.Storage.Size != "20Gi" {
t.Errorf("Storage.Size = %q, want 20Gi", ms.Spec.Storage.Size) t.Errorf("Storage.Size = %q, want 20Gi", ms.Spec.Storage.Size)
} }
// Same operational defaults as the API create path: without RCON the server
// never reports players and the console answers 503; without spec.idle it
// never stops on its own.
if !ms.Spec.Rcon.Enabled || ms.Spec.Rcon.SecretRef.Name != naming.RconSecretName("test-server") {
t.Errorf("Rcon = %+v, want enabled with the operator-minted secret", ms.Spec.Rcon)
}
if ms.Spec.Idle != v1alpha1.DefaultIdle() {
t.Errorf("Idle = %+v, want the default %+v", ms.Spec.Idle, v1alpha1.DefaultIdle())
}
if ms.Spec.FallbackServer != naming.SystemLoginServer {
t.Errorf("FallbackServer = %q, want the login gate", ms.Spec.FallbackServer)
}
mem, ok := ms.Spec.Resources.Limits[corev1.ResourceMemory] mem, ok := ms.Spec.Resources.Limits[corev1.ResourceMemory]
if !ok { if !ok {
t.Fatal("memory limit missing") t.Fatal("memory limit missing")
+37 -4
View File
@@ -1,6 +1,7 @@
package main package main
import ( import (
"context"
"crypto/rand" "crypto/rand"
"encoding/hex" "encoding/hex"
"flag" "flag"
@@ -52,8 +53,8 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
fmt.Fprintf(stderr, "felis backup: archive store %q is not implemented in this build (only tarLocal)\n", cfg.Archive.Store) fmt.Fprintf(stderr, "felis backup: archive store %q is not implemented in this build (only tarLocal)\n", cfg.Archive.Store)
return 1 return 1
} }
// Reuse the reaper's retention derivation so an on-demand backup expires on the // The [archive] parse the reaper uses; an on-demand backup takes its
// same clock as an inactivity backup — one retention policy, not two. // manual_retention and manual_keep.
rcfg, err := reaperConfig(cfg) rcfg, err := reaperConfig(cfg)
if err != nil { if err != nil {
fmt.Fprintf(stderr, "felis backup: %v\n", err) fmt.Fprintf(stderr, "felis backup: %v\n", err)
@@ -72,6 +73,13 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
ctx := ctrl.SetupSignalHandler() ctx := ctrl.SetupSignalHandler()
// The archive store shares the node's disk with every world and the
// database: an owner's backup must not be what tips it into eviction.
if err := backup.CheckRoom(cfg.Archive.LocalPath, *worldsRoot, backup.MinFreeAfter); err != nil {
fmt.Fprintf(stderr, "felis backup: %v\n", err)
return 1
}
ref, size, err := archiver.Archive(ctx, *server, naming.WorldPVCName(*server)) ref, size, err := archiver.Archive(ctx, *server, naming.WorldPVCName(*server))
if err != nil { if err != nil {
fmt.Fprintf(stderr, "felis backup: archive: %v\n", err) fmt.Fprintf(stderr, "felis backup: archive: %v\n", err)
@@ -92,9 +100,10 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
BackupRef: string(ref), BackupRef: string(ref),
SizeBytes: size, SizeBytes: size,
Reason: "manual", Reason: "manual",
ExpiresAt: time.Now().Add(rcfg.Retention), ExpiresAt: time.Now().Add(rcfg.ManualRetention),
} }
if err := reaper.NewPGStore(drv.DB()).InsertBackup(ctx, rec); err != nil { st := reaper.NewPGStore(drv.DB())
if err := st.InsertBackup(ctx, rec); err != nil {
// The archive is written but unrecorded — an orphan the retention pass would // The archive is written but unrecorded — an orphan the retention pass would
// never expire. Delete it so a failed backup leaves no leaked bytes, mirroring // never expire. Delete it so a failed backup leaves no leaked bytes, mirroring
// the reaper's archive-then-record atomicity. // the reaper's archive-then-record atomicity.
@@ -107,9 +116,33 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
} }
fmt.Fprintf(stdout, "felis backup: server=%s archived %d bytes to %s (backup %s)\n", *server, size, ref, rec.ID) fmt.Fprintf(stdout, "felis backup: server=%s archived %d bytes to %s (backup %s)\n", *server, size, ref, rec.ID)
pruneManualBackups(ctx, st, archiver, *server, rcfg.ManualKeep, stdout, stderr)
return 0 return 0
} }
// pruneManualBackups keeps server's newest keep on-demand backups and removes
// the rest, oldest first, so repeated backups of one world cannot fill the
// shared archive store. The new backup is already recorded; a removal that
// fails is reported and retried after the next backup.
func pruneManualBackups(ctx context.Context, st *reaper.PGStore, archiver backup.WorldArchiver, server string, keep int, stdout, stderr io.Writer) {
excess, err := st.ExcessManualBackups(ctx, server, keep)
if err != nil {
fmt.Fprintf(stderr, "felis backup: list older backups of %s: %v\n", server, err)
return
}
for _, b := range excess {
if err := archiver.Delete(ctx, backup.ArchiveRef(b.BackupRef)); err != nil {
fmt.Fprintf(stderr, "felis backup: remove older backup %s: %v\n", b.ID, err)
continue
}
if err := st.MarkBackupDeleted(ctx, b.ID, time.Now()); err != nil {
fmt.Fprintf(stderr, "felis backup: record the removal of %s: %v\n", b.ID, err)
continue
}
fmt.Fprintf(stdout, "felis backup: removed older backup %s of %s (keeping the newest %d)\n", b.ID, server, keep)
}
}
// newBackupID mints a world_backups primary key, matching the reaper's "bk-"+hex // newBackupID mints a world_backups primary key, matching the reaper's "bk-"+hex
// scheme so a manual and an inactivity backup are indistinguishable downstream. // scheme so a manual and an inactivity backup are indistinguishable downstream.
func newBackupID() string { func newBackupID() string {
+33 -9
View File
@@ -86,23 +86,31 @@ func requestBackup(ctx context.Context, hc *http.Client, baseURL, token, name, o
// backupErrorFromResponse turns a non-202 into a human message. The well-known codes get // backupErrorFromResponse turns a non-202 into a human message. The well-known codes get
// an operator-facing explanation; anything else falls back to the API's // an operator-facing explanation; anything else falls back to the API's
// {"error":{message}} body, then the bare status code. // {"error":{code,message}} body, then the bare status code.
func backupErrorFromResponse(resp *http.Response) error { func backupErrorFromResponse(resp *http.Response) error {
switch resp.StatusCode {
case http.StatusConflict: // not_stopped
return fmt.Errorf("the server must be stopped before its world can be backed up — halt it first")
case http.StatusServiceUnavailable: // backup_unavailable
return fmt.Errorf("the backup subsystem is not configured on felis-api (FELIS_IMAGE / FELIS_BACKUP_PVC unset)")
case http.StatusNotFound:
return fmt.Errorf("no such server")
}
var e struct { var e struct {
Error struct { Error struct {
Code string `json:"code"`
Message string `json:"message"` Message string `json:"message"`
} `json:"error"` } `json:"error"`
} }
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<16)) raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<16))
_ = json.Unmarshal(raw, &e) _ = json.Unmarshal(raw, &e)
switch resp.StatusCode {
case http.StatusConflict:
// Two refusals share 409: the stopped gate and the missing-world-volume
// gate. The body's code distinguishes them; a code-less body reads as the
// stopped gate (the only 409 before the volume gate existed), and any other
// coded 409 falls through to the API's own operator text.
if e.Error.Code == "" || e.Error.Code == "not_stopped" {
return fmt.Errorf("the server must be stopped before its world can be backed up — halt it first")
}
case http.StatusServiceUnavailable: // backup_unavailable
return fmt.Errorf("the backup subsystem is not configured on felis-api (FELIS_IMAGE / FELIS_BACKUP_PVC unset)")
case http.StatusNotFound:
return fmt.Errorf("no such server")
}
if e.Error.Message != "" { if e.Error.Message != "" {
return fmt.Errorf("felis-api: %s", e.Error.Message) return fmt.Errorf("felis-api: %s", e.Error.Message)
} }
@@ -120,3 +128,19 @@ func performBackupNow(ctx context.Context, cl client.Client, controlNamespace, n
hc := &http.Client{Timeout: 10 * time.Second} hc := &http.Client{Timeout: 10 * time.Second}
return requestBackup(ctx, hc, baseURL, token, name, osUser) return requestBackup(ctx, hc, baseURL, token, name, osUser)
} }
// backupPickable narrows the backup picker to servers the backup API can accept.
// System servers (login/lobby) are excluded: they have no row in the servers
// table and carry reserved names, so every attempt dies in name validation —
// offering them would be a dead pick. The halt picker keeps them on purpose
// (break-glass retains full power over system servers); only the API-backed
// backup op cannot reach them.
func backupPickable(servers []haltableServer) []haltableServer {
out := make([]haltableServer, 0, len(servers))
for _, s := range servers {
if !s.system {
out = append(out, s)
}
}
return out
}
+32 -3
View File
@@ -114,16 +114,23 @@ func TestRequestBackup(t *testing.T) {
cases := []struct { cases := []struct {
name string name string
code int code int
body string // optional JSON error body
expect string expect string
}{ }{
{"409 not_stopped", http.StatusConflict, "must be stopped"}, {"409 not_stopped", http.StatusConflict, "", "must be stopped"},
{"503 backup_unavailable", http.StatusServiceUnavailable, "not configured"}, {"409 no_world_volume surfaces the API's own text", http.StatusConflict,
{"404 not found", http.StatusNotFound, "no such server"}, `{"error":{"code":"no_world_volume","message":"this server has no world volume yet — start it once to create it, then retry"}}`,
"no world volume yet"},
{"503 backup_unavailable", http.StatusServiceUnavailable, "", "not configured"},
{"404 not found", http.StatusNotFound, "", "no such server"},
} }
for _, tc := range cases { for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) { t.Run(tc.name, func(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(tc.code) w.WriteHeader(tc.code)
if tc.body != "" {
_, _ = io.WriteString(w, tc.body)
}
})) }))
defer srv.Close() defer srv.Close()
_, err := requestBackup(context.Background(), hc, srv.URL, "tok", "survival", "alice") _, err := requestBackup(context.Background(), hc, srv.URL, "tok", "survival", "alice")
@@ -142,3 +149,25 @@ func TestRequestBackup(t *testing.T) {
} }
}) })
} }
// The Sync picker must not offer system servers: the backup API validates names
// and resolves a servers-table row, so a lobby/login pick can only die in
// validation — a dead choice in an emergency console.
func TestBackupPickable(t *testing.T) {
got := backupPickable([]haltableServer{
{name: "lobby", phase: "Running", system: true},
{name: "login", phase: "Running", system: true},
{name: "test-one", phase: "Stopped"},
})
if len(got) != 1 || got[0].name != "test-one" || got[0].system {
t.Fatalf("backupPickable = %+v, want only the user server", got)
}
// Survivors keep their input order (the picker's cursor math depends on it).
got = backupPickable([]haltableServer{
{name: "alpha"}, {name: "login", system: true}, {name: "beta"},
})
if len(got) != 2 || got[0].name != "alpha" || got[1].name != "beta" {
t.Fatalf("backupPickable order = %+v, want [alpha beta]", got)
}
}
+43 -12
View File
@@ -76,12 +76,19 @@ type ownerStore interface {
AdminExists(ctx context.Context) (bool, error) AdminExists(ctx context.Context) (bool, error)
// UserByUsername loads a staff login projection. // UserByUsername loads a staff login projection.
UserByUsername(ctx context.Context, username string) (*api.StaffUser, error) UserByUsername(ctx context.Context, username string) (*api.StaffUser, error)
// OwnerUsername names the single active Owner seat, or "" when none exists.
// provisionOwner refuses to re-target anything but this username: with the
// seat occupied, a fresh name would mint a second owner row (UpsertOwner's
// insert arm) while the existing — possibly compromised — seat stays live,
// and no supported path can delete an owner row afterwards.
OwnerUsername(ctx context.Context) (string, error)
UpsertOwner(ctx context.Context, id, username, email string) error UpsertOwner(ctx context.Context, id, username, email string) error
// InsertOperator mints a NEW Operator staff account. Unlike UpsertOwner it is // InsertOperator mints a NEW Operator staff account. Unlike UpsertOwner it is
// insert-only: a username already taken is a conflict (api.ErrConflict), never a // insert-only: a username already taken is a conflict (api.ErrConflict), never a
// silent reset, so adding an Operator can never clobber the Owner or an existing // silent reset, so adding an Operator can never clobber the Owner or an existing
// Operator. The row is role=admin, identical in shape to the Owner — Felis has no // Operator. The row is role=admin — an Operator is staff BELOW the single
// separate operator DB role (migration 0003: staff = role=admin). // role=owner identity (migration 0011 adds that role); the two are the only
// staff roles.
InsertOperator(ctx context.Context, id, username, email string) error InsertOperator(ctx context.Context, id, username, email string) error
// CompleteOwnerSetup atomically consumes the in-game link code, creates or // CompleteOwnerSetup atomically consumes the in-game link code, creates or
// promotes the bound Owner, enables local auth, and stores the one-time setup // promotes the bound Owner, enables local auth, and stores the one-time setup
@@ -266,20 +273,31 @@ func authenticateAdmin(ctx context.Context, s ownerStore, username string) (matc
if err != nil { if err != nil {
return "", false, err return "", false, err
} }
if u.Role != "admin" { // Staff means admin OR owner: recovery attribution must accept the Owner (the
// primary break-glass identity), not just plain admins.
if u.Role != "admin" && u.Role != "owner" {
return "", false, nil return "", false, nil
} }
return u.Username, true, nil return u.Username, true, nil
} }
// provisionOwner mints or resets the single Owner account direct-to-Postgres, // provisionOwner mints or resets the single Owner account direct-to-Postgres,
// passwordless. The account is role=admin with no password — the Owner completes // passwordless. The account is role=owner with no password — the Owner completes
// passwordless login setup via the web setup-token flow after `felis setup`. // passwordless login setup via the web setup-token flow after `felis setup`.
// With a seat already occupied the reset must name that seat (ownerSeatTakenError
// otherwise): the upsert's insert arm would silently mint a SECOND owner, and
// every owner row is undeletable through the panel, so the tier could never
// converge back to one.
func provisionOwner(ctx context.Context, s ownerStore, username, email string) error { func provisionOwner(ctx context.Context, s ownerStore, username, email string) error {
username = strings.TrimSpace(username) username = strings.TrimSpace(username)
if username == "" { if username == "" {
return errors.New("owner username is required") return errors.New("owner username is required")
} }
if seat, err := s.OwnerUsername(ctx); err != nil {
return fmt.Errorf("check the owner seat: %w", err)
} else if seat != "" && seat != username {
return &ownerSeatTakenError{seat: seat}
}
id := newOwnerID() id := newOwnerID()
if id == "" { if id == "" {
return errors.New("generate owner id: entropy source failed") return errors.New("generate owner id: entropy source failed")
@@ -290,13 +308,26 @@ func provisionOwner(ctx context.Context, s ownerStore, username, email string) e
return nil return nil
} }
// provisionOperator mints a NEW Operator staff account direct-to-Postgres. Like the // ownerSeatTakenError refuses an Owner reset that names anything but the
// Owner it is role=admin and passwordless — Felis has no separate operator DB role, // occupied seat, naming it so the operator can retype. Is reports
// so an Operator is simply an additional staff admin (migration 0003). UNLIKE // api.ErrConflict so the TUI's recoverable-error branch (shared with the
// provisionOwner, which upserts the single Owner and resets it on a username // operator path's taken-name clash) routes back to the form instead of ending
// conflict, this is insert-only: a username already taken returns api.ErrConflict // the console.
// rather than overwriting a live account, so adding an Operator can never silently type ownerSeatTakenError struct{ seat string }
// clobber the Owner's or another Operator's account.
func (e *ownerSeatTakenError) Error() string {
return fmt.Sprintf("an Owner already exists as %q — enter that username to reset the Owner", e.seat)
}
func (e *ownerSeatTakenError) Is(target error) bool { return target == api.ErrConflict }
// provisionOperator mints a NEW Operator staff account direct-to-Postgres. It is
// role=admin and passwordless — an additional staff admin below the single
// role=owner identity (migrations 0003 + 0011). UNLIKE provisionOwner, which
// upserts the single Owner and resets it on a username conflict, this is
// insert-only: a username already taken returns api.ErrConflict rather than
// overwriting a live account, so adding an Operator can never silently clobber
// the Owner's or another Operator's account.
func provisionOperator(ctx context.Context, s ownerStore, username, email string) error { func provisionOperator(ctx context.Context, s ownerStore, username, email string) error {
username = strings.TrimSpace(username) username = strings.TrimSpace(username)
if username == "" { if username == "" {
@@ -387,7 +418,7 @@ func newSetupToken() (raw, hash string, err error) {
// performSetupMCBind is the `felis setup` Owner-establishment path: the operator // performSetupMCBind is the `felis setup` Owner-establishment path: the operator
// binds their Minecraft account via a one-time link code the login gate handed // binds their Minecraft account via a one-time link code the login gate handed
// them in-game, the bound user is promoted to role='admin' (passwordless Owner), // them in-game, the bound user is promoted to role='owner' (passwordless Owner),
// local auth is enabled, and a one-time setup URL is minted for the first web // local auth is enabled, and a one-time setup URL is minted for the first web
// login where the Owner verifies email / enrolls a passkey. adminHostname is the // login where the Owner verifies email / enrolls a passkey. adminHostname is the
// operator-console host the URL points at (op.console.<root>): the Owner is staff, // operator-console host the URL points at (op.console.<root>): the Owner is staff,
+57 -8
View File
@@ -19,14 +19,15 @@ import (
// terminal. The design is passwordless: accounts carry no credential, and the // terminal. The design is passwordless: accounts carry no credential, and the
// Owner completes first-login through the setup-token web flow. // Owner completes first-login through the setup-token web flow.
type fakeOwnerStore struct { type fakeOwnerStore struct {
upserts []upsertCall upserts []upsertCall
inserts []upsertCall inserts []upsertCall
settings map[string][]byte settings map[string][]byte
audits []api.AuditEntry audits []api.AuditEntry
tokens []setupTokenCall tokens []setupTokenCall
redeems []redeemCall redeems []redeemCall
users map[string]*api.StaffUser // keyed by username users map[string]*api.StaffUser // keyed by username
admins bool // AdminExists answer admins bool // AdminExists answer
ownerSeat string // OwnerUsername answer: the occupied seat, "" when none
// CompleteOwnerSetup's success result. redeemUserID defaults to the fresh id // CompleteOwnerSetup's success result. redeemUserID defaults to the fresh id
// the caller passes (the unlinked-UUID case) when left empty. // the caller passes (the unlinked-UUID case) when left empty.
@@ -40,6 +41,7 @@ type fakeOwnerStore struct {
auditErr error auditErr error
userErr error // non-not-found error from UserByUsername userErr error // non-not-found error from UserByUsername
adminErr error adminErr error
seatErr error
redeemErr error redeemErr error
createTokenErr error createTokenErr error
} }
@@ -80,6 +82,15 @@ func (f *fakeOwnerStore) UserByUsername(_ context.Context, username string) (*ap
return nil, api.ErrNotFound return nil, api.ErrNotFound
} }
// OwnerUsername reports the single active Owner seat. Tests set ownerSeat; the
// zero value models a fresh install where bootstrap is free to mint.
func (f *fakeOwnerStore) OwnerUsername(_ context.Context) (string, error) {
if f.seatErr != nil {
return "", f.seatErr
}
return f.ownerSeat, nil
}
func (f *fakeOwnerStore) UpsertOwner(_ context.Context, id, username, email string) error { func (f *fakeOwnerStore) UpsertOwner(_ context.Context, id, username, email string) error {
if f.upsertErr != nil { if f.upsertErr != nil {
return f.upsertErr return f.upsertErr
@@ -201,6 +212,34 @@ func TestProvisionOwner(t *testing.T) {
} }
}) })
t.Run("an occupied seat refuses any other username", func(t *testing.T) {
// The seat is the single owner row: upserting a fresh name would take the
// insert arm and mint a SECOND owner, while the existing seat — possibly the
// compromised account this reset was meant to replace — stays live, and no
// supported path can delete an owner row.
f := &fakeOwnerStore{ownerSeat: "seat-holder"}
err := provisionOwner(ctx, f, "someone-else", "")
if !errors.Is(err, api.ErrConflict) {
t.Fatalf("error = %v, want it to wrap api.ErrConflict so the TUI routes back to the form", err)
}
if !strings.Contains(err.Error(), `"seat-holder"`) {
t.Errorf("error = %q, want it to name the occupied seat", err)
}
if len(f.upserts) != 0 {
t.Errorf("want no write against an occupied seat, got %d", len(f.upserts))
}
})
t.Run("the occupied seat's own username still resets", func(t *testing.T) {
f := &fakeOwnerStore{ownerSeat: "seat-holder"}
if err := provisionOwner(ctx, f, "seat-holder", "[email protected]"); err != nil {
t.Fatalf("provisionOwner(reset): %v", err)
}
if len(f.upserts) != 1 || f.upserts[0].username != "seat-holder" || f.upserts[0].email != "[email protected]" {
t.Fatalf("want 1 reset upsert for the seat, got %+v", f.upserts)
}
})
t.Run("propagates a store error", func(t *testing.T) { t.Run("propagates a store error", func(t *testing.T) {
f := &fakeOwnerStore{upsertErr: errors.New("boom")} f := &fakeOwnerStore{upsertErr: errors.New("boom")}
if err := provisionOwner(ctx, f, "owner", ""); err == nil { if err := provisionOwner(ctx, f, "owner", ""); err == nil {
@@ -264,6 +303,16 @@ func TestAuthenticateAdmin(t *testing.T) {
} }
}) })
t.Run("the owner role attributes like an admin", func(t *testing.T) {
owner := mkAdmin("root")
owner.Role = "owner" // the platform owner is staff too (migration 0011)
f := &fakeOwnerStore{users: map[string]*api.StaffUser{"root": owner}}
matched, ok, err := authenticateAdmin(ctx, f, "root")
if err != nil || !ok || matched != "root" {
t.Fatalf("authenticateAdmin(owner) = (%q, %v, %v), want (root, true, nil)", matched, ok, err)
}
})
t.Run("an unknown user is a non-match, not an error", func(t *testing.T) { t.Run("an unknown user is a non-match, not an error", func(t *testing.T) {
f := &fakeOwnerStore{} f := &fakeOwnerStore{}
_, ok, err := authenticateAdmin(ctx, f, "nobody") _, ok, err := authenticateAdmin(ctx, f, "nobody")
+105
View File
@@ -0,0 +1,105 @@
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"os"
"strings"
"felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/platform"
"sigs.k8s.io/controller-runtime/pkg/client"
)
// cmdConverge is the explicit convergence pass over already-installed system
// servers (#1), plus the idle-stop default for user servers that predate it. Provisioning is create-if-absent, so a field the desired spec
// gained after an install (spec.rcon, spec.startup.healthHTTPPort, a derived env
// key) never reaches the existing CR — and nothing says so. This command fills
// exactly those zero-value fields; see convergeSystemServers for the full contract
// and why it is a separate, operator-timed step rather than part of setup.
//
// It reads the same host config as setup (the control plane's felis.toml) and
// talks to the cluster with the local kubeconfig, so it must run as root on the
// control-plane host.
func cmdConverge(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("converge", flag.ContinueOnError)
fs.SetOutput(stderr)
cfgPath := fs.String("config", defaultSetupConfigPath, "path to felis.toml")
if err := fs.Parse(args); err != nil {
if errors.Is(err, flag.ErrHelp) {
return 0
}
return 2
}
if os.Geteuid() != 0 {
fmt.Fprintln(stderr, "felis converge: refused — converging needs the cluster credentials, so it must run as root (try: sudo felis converge)")
return 1
}
cfg, err := config.Load(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis converge: %v\n", err)
fmt.Fprintln(stderr, "If this host was never installed, run `sudo felis setup` first.")
return 1
}
cl, err := buildSystemServerClient()
if err != nil {
fmt.Fprintf(stderr, "felis converge: %v\n", err)
return 1
}
controlNS := platform.DefaultControlNamespace
outcomes := convergeSystemServers(context.Background(), cl, cfg.K8s.Namespace,
cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage,
platform.InternalAPIBaseURL(controlNS), cfg.Server.RootDomain,
defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname))
outcomes = append(outcomes, convergeUserServerIdle(context.Background(), cl, cfg.K8s.Namespace)...)
fmt.Fprintln(stdout, "felis converge: filling fields an installed server predates (operator-set values are never overwritten):")
exit := 0
for _, o := range outcomes {
switch {
case o.err != nil:
fmt.Fprintf(stdout, " - %s: ERROR %v\n", o.name, o.err)
exit = 1
case len(o.changes) > 0:
fmt.Fprintf(stdout, " - %s: updated (%s)\n", o.name, strings.Join(o.changes, ", "))
default:
fmt.Fprintf(stdout, " - %s: %s\n", o.name, o.skipped)
}
}
return exit
}
// convergeUserServerIdle gives every user server that predates the idle default
// (spec.idle entirely unset) the default idle stop. A server whose idle stop was
// turned off keeps a duration on its spec, so it is not "unset" and is left
// alone; system servers never idle out and are skipped. Servers that already
// carry a value produce no line, so a converged fleet prints nothing here.
func convergeUserServerIdle(ctx context.Context, cl client.Client, namespace string) []systemServerOutcome {
var list v1alpha1.MinecraftServerList
if err := cl.List(ctx, &list, client.InNamespace(namespace)); err != nil {
return []systemServerOutcome{{name: "user servers", err: fmt.Errorf("list servers: %w", err)}}
}
var out []systemServerOutcome
for i := range list.Items {
ms := &list.Items[i]
if ms.Labels[v1alpha1.LabelSystemRole] != "" || ms.Spec.Idle != (v1alpha1.IdleSpec{}) {
continue
}
patch := client.MergeFrom(ms.DeepCopy())
ms.Spec.Idle = v1alpha1.DefaultIdle()
if err := cl.Patch(ctx, ms, patch); err != nil {
out = append(out, systemServerOutcome{name: ms.Name, err: fmt.Errorf("converge %s: %w", ms.Name, err)})
continue
}
out = append(out, systemServerOutcome{name: ms.Name, available: true, updated: true,
changes: []string{fmt.Sprintf("spec.idle (stop after %ds empty)", v1alpha1.DefaultEmptySecondsBeforeStop)}})
}
return out
}
+231
View File
@@ -0,0 +1,231 @@
package main
import (
"context"
"slices"
"strings"
"testing"
"felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/naming"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/client/fake"
)
// converge is the explicit pass over an installed system server whose CR predates
// a field the desired spec has since gained (#1). It must fill exactly the
// zero-valued whitelist fields and the derived env, and must not touch anything a
// non-zero value already occupies — that is the operator's.
func TestConvergeSystemServersFillsPredatedFields(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
// An old install: the lobby CR was created before the desired spec began
// rendering spec.rcon, and the login CR before the HTTP readiness gate existed.
// One derived env key is absent entirely (as if it were added later), and one
// hand-added env var plus a non-whitelisted spec field must survive.
lobby, err := lobbySystemServer("reg/lobby:1", "minecraft")
if err != nil {
t.Fatalf("build lobby: %v", err)
}
lobby.Spec.Rcon = v1alpha1.RconSpec{}
lobby.Spec.JavaMemory = "999Mi"
login, err := loginSystemServer("reg/limbo:1", "minecraft",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
if err != nil {
t.Fatalf("build login: %v", err)
}
login.Spec.Startup.HealthHTTPPort = 0
kept := login.Spec.Env
login.Spec.Env = nil
for _, e := range kept {
if e.Name != envPanelHostname {
login.Spec.Env = append(login.Spec.Env, e)
}
}
login.Spec.Env = append(login.Spec.Env, v1alpha1.EnvVar{Name: "OPERATOR_TUNING", Value: "keep-me"})
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(lobby, login).Build()
outcomes := convergeSystemServers(ctx, cl, "minecraft", "reg/limbo:1", "reg/lobby:1",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
byName := map[string]systemServerOutcome{}
for _, o := range outcomes {
if o.err != nil {
t.Fatalf("%s: unexpected error: %v", o.name, o.err)
}
byName[o.name] = o
}
lobbyOut := byName[naming.SystemLobbyServer]
if len(lobbyOut.changes) != 1 || lobbyOut.changes[0] != "spec.rcon" {
t.Errorf("lobby changes = %v, want [spec.rcon] (only the zero-valued field)", lobbyOut.changes)
}
loginOut := byName[naming.SystemLoginServer]
if !slices.Contains(loginOut.changes, "spec.startup.healthHTTPPort") || !slices.Contains(loginOut.changes, "env "+envPanelHostname) {
t.Errorf("login changes = %v, want the health port plus the missing derived env key", loginOut.changes)
}
var gotLobby v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLobbyServer}, &gotLobby); err != nil {
t.Fatalf("get lobby: %v", err)
}
if !gotLobby.Spec.Rcon.Enabled ||
gotLobby.Spec.Rcon.SecretRef.Name != naming.RconSecretName(naming.SystemLobbyServer) ||
gotLobby.Spec.Rcon.SecretRef.Key != naming.RconSecretKey {
t.Errorf("lobby rcon = %+v, want the desired block with the %s secret",
gotLobby.Spec.Rcon, naming.RconSecretName(naming.SystemLobbyServer))
}
if gotLobby.Spec.JavaMemory != "999Mi" {
t.Errorf("lobby javaMemory = %q, want 999Mi — converge fills new fields, it does not rewrite the spec", gotLobby.Spec.JavaMemory)
}
var gotLogin v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLoginServer}, &gotLogin); err != nil {
t.Fatalf("get login: %v", err)
}
if gotLogin.Spec.Startup.HealthHTTPPort != felisLimboHealthPort {
t.Errorf("login healthHTTPPort = %d, want %d", gotLogin.Spec.Startup.HealthHTTPPort, felisLimboHealthPort)
}
env := map[string]string{}
for _, e := range gotLogin.Spec.Env {
env[e.Name] = e.Value
}
if env[envPanelHostname] != "console.mc.example.net" {
t.Errorf("%s was not added back: %q", envPanelHostname, env[envPanelHostname])
}
if env["OPERATOR_TUNING"] != "keep-me" {
t.Error("a hand-added env var was dropped; converge only touches config-derived names")
}
}
// A field already holding a non-zero value belongs to the operator: converge must
// report "already converged" and write nothing.
func TestConvergeSystemServersLeavesNonZeroFieldsAlone(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
lobby, err := lobbySystemServer("reg/lobby:1", "minecraft")
if err != nil {
t.Fatalf("build lobby: %v", err)
}
lobby.Spec.Rcon = v1alpha1.RconSpec{
Enabled: true,
SecretRef: v1alpha1.SecretKeyRef{Name: "operator-rotated", Key: "password"},
}
login, err := loginSystemServer("reg/limbo:1", "minecraft",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
if err != nil {
t.Fatalf("build login: %v", err)
}
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(lobby, login).Build()
for _, o := range convergeSystemServers(ctx, cl, "minecraft", "reg/limbo:1", "reg/lobby:1",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net") {
if o.err != nil {
t.Fatalf("%s: unexpected error: %v", o.name, o.err)
}
if len(o.changes) != 0 || o.skipped != "already converged" {
t.Errorf("%s outcome = %+v, want already converged with no writes", o.name, o)
}
}
var got v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLobbyServer}, &got); err != nil {
t.Fatalf("get lobby: %v", err)
}
if got.Spec.Rcon.SecretRef.Name != "operator-rotated" {
t.Errorf("lobby rcon secretRef = %q — converge overwrote a field the operator had already set",
got.Spec.Rcon.SecretRef.Name)
}
}
// Guards: an absent CR is reported (creation is setup's job), a foreign CR is
// refused rather than adopted, and an unset image skips like the provisioner does.
func TestConvergeSystemServersGuards(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
run := func(cl client.Client, loginImage, lobbyImage string) []systemServerOutcome {
return convergeSystemServers(ctx, cl, "minecraft", loginImage, lobbyImage,
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
}
t.Run("absent CRs are reported, not created", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).Build()
for _, o := range run(cl, "reg/limbo:1", "reg/lobby:1") {
if o.err != nil {
t.Fatalf("%s: %v", o.name, o.err)
}
if o.created || !strings.Contains(o.skipped, "not present") {
t.Errorf("%s outcome = %+v, want a not-present skip", o.name, o)
}
}
})
t.Run("foreign CR is refused", func(t *testing.T) {
foreign := &v1alpha1.MinecraftServer{}
foreign.Name = naming.SystemLoginServer
foreign.Namespace = "minecraft"
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(foreign).Build()
out := run(cl, "reg/limbo:1", "")
if len(out) != 2 {
t.Fatalf("outcomes = %d, want 2", len(out))
}
if out[0].err == nil || !strings.Contains(out[0].err.Error(), "not marked") {
t.Fatalf("login error = %v, want an unmarked-name refusal", out[0].err)
}
})
t.Run("unset image skips", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).Build()
out := run(cl, "", "reg/lobby:1")
if out[0].skipped != "image not configured" {
t.Errorf("login skipped = %q, want %q", out[0].skipped, "image not configured")
}
})
}
// TestConvergeUserServerIdle fills the idle default only where spec.idle was
// never set: a server whose idle stop was turned off (duration kept), one with
// its own duration, and a system server all stay as they are.
func TestConvergeUserServerIdle(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
mk := func(name string, idle v1alpha1.IdleSpec, role string) *v1alpha1.MinecraftServer {
ms := &v1alpha1.MinecraftServer{}
ms.Name, ms.Namespace = name, "minecraft"
ms.Spec.Idle = idle
if role != "" {
ms.Labels = map[string]string{v1alpha1.LabelSystemRole: role}
}
return ms
}
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
mk("legacy", v1alpha1.IdleSpec{}, ""),
mk("off", v1alpha1.IdleSpec{EmptySecondsBeforeStop: 600}, ""),
mk("custom", v1alpha1.IdleSpec{AutoStopEnabled: true, EmptySecondsBeforeStop: 1800}, ""),
mk(naming.SystemLobbyServer, v1alpha1.IdleSpec{}, naming.SystemLobbyServer),
).Build()
outcomes := convergeUserServerIdle(ctx, cl, "minecraft")
if len(outcomes) != 1 || outcomes[0].name != "legacy" || outcomes[0].err != nil {
t.Fatalf("outcomes = %+v, want exactly one fill for legacy", outcomes)
}
want := map[string]v1alpha1.IdleSpec{
"legacy": v1alpha1.DefaultIdle(),
"off": {EmptySecondsBeforeStop: 600},
"custom": {AutoStopEnabled: true, EmptySecondsBeforeStop: 1800},
naming.SystemLobbyServer: {},
}
for name, idle := range want {
var ms v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: name}, &ms); err != nil {
t.Fatalf("get %s: %v", name, err)
}
if ms.Spec.Idle != idle {
t.Errorf("%s idle = %+v, want %+v", name, ms.Spec.Idle, idle)
}
}
if again := convergeUserServerIdle(ctx, cl, "minecraft"); len(again) != 0 {
t.Fatalf("second pass = %+v, want nothing to do", again)
}
}
+334
View File
@@ -0,0 +1,334 @@
package main
import (
"context"
"encoding/json"
"errors"
"flag"
"fmt"
"io"
"os"
"os/exec"
"path/filepath"
"strings"
"time"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/dbbackup"
)
const dbUsage = `usage:
felis db backup [-config path] [-dir dir] [-label daily|manual|...] [-keep n] [-state-dir dir]
[-no-servers] [-metrics-file path]
felis db restore [-config path] [-dir dir] [-yes] [-force] [-no-safety-backup] <bundle>
felis db verify [-dir dir] <bundle>
felis db list [-dir dir]
felis db check [-dir dir] [-max-age 26h]
`
// defaultKeep is how many bundles of a label a backup leaves behind. Manual
// bundles are the operator's own and are never pruned.
var defaultKeep = map[string]int{
dbbackup.LabelDaily: 14,
dbbackup.LabelPreMigrate: 10,
dbbackup.LabelPreRestore: 5,
}
// cmdDB implements `felis db`: logical backups of the control-plane database
// together with the host state a rebuild needs (internal/dbbackup). The verb
// comes first for the same reason as `felis migrate up`.
func cmdDB(args []string, stdout, stderr io.Writer) int {
if len(args) == 0 {
fmt.Fprint(stderr, dbUsage)
return 2
}
verb, rest := args[0], args[1:]
fs := flag.NewFlagSet("db "+verb, flag.ContinueOnError)
fs.SetOutput(stderr)
fs.Usage = func() { fmt.Fprint(stderr, dbUsage) }
dir := fs.String("dir", dbbackup.DefaultDir, "bundle directory")
switch verb {
case "backup":
return dbBackup(fs, dir, rest, stdout, stderr)
case "restore":
return dbRestore(fs, dir, rest, stdout, stderr)
case "verify":
return dbVerify(fs, dir, rest, stdout, stderr)
case "list":
return dbList(fs, dir, rest, stdout, stderr)
case "check":
return dbCheck(fs, dir, rest, stdout, stderr)
case "-h", "--help", "help":
fmt.Fprint(stdout, dbUsage)
return 0
}
fmt.Fprintf(stderr, "felis db: unknown verb %q\n%s", verb, dbUsage)
return 2
}
// parseWithArg parses flags that may sit on either side of one positional
// argument (`restore -yes x.tar` and `restore x.tar -yes` both work) and
// returns that argument.
func parseWithArg(fs *flag.FlagSet, args []string) (string, bool) {
if err := fs.Parse(args); err != nil {
return "", false
}
if fs.NArg() == 0 {
return "", true
}
arg := fs.Arg(0)
if err := fs.Parse(fs.Args()[1:]); err != nil {
return "", false
}
if fs.NArg() > 0 {
fmt.Fprintf(fs.Output(), "felis db: unexpected argument %q\n", fs.Arg(0))
return "", false
}
return arg, true
}
func dbDatabaseURL(path string) (string, error) {
cfg, err := config.Load(path)
if err != nil {
return "", err
}
return cfg.Database.URL, nil
}
func dbBackup(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
label := fs.String("label", dbbackup.LabelManual, "bundle label; daily/pre-migrate/pre-restore bundles are pruned, manual ones never")
keep := fs.Int("keep", -1, "bundles of this label to keep (default: daily 14, pre-migrate 10, pre-restore 5, manual all)")
stateDir := fs.String("state-dir", dbbackup.DefaultStateDir, `host state directory to bundle ("" for none)`)
noServers := fs.Bool("no-servers", false, "leave the MinecraftServer objects out of the bundle")
metrics := fs.String("metrics-file", "", "node-exporter textfile to rewrite on success (e.g. /var/lib/node_exporter/textfile_collector/felis_db_backup.prom)")
if err := fs.Parse(args); err != nil {
return 2
}
if fs.NArg() > 0 {
fmt.Fprint(stderr, dbUsage)
return 2
}
url, err := dbDatabaseURL(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis db backup: %v\n", err)
return 1
}
if *keep < 0 {
*keep = defaultKeep[*label]
}
o := dbbackup.BackupOptions{
DatabaseURL: url, Dir: *dir, Label: *label, Keep: *keep,
StateDir: *stateDir, Version: resolvedVersion(), Log: stderr,
MetricsFile: *metrics, Record: true,
}
if !*noServers {
o.ExportServers = exportMinecraftServers
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Minute)
defer cancel()
path, err := dbbackup.Backup(ctx, o)
if err != nil {
fmt.Fprintf(stderr, "felis db backup: %v\n", err)
return 1
}
fmt.Fprintf(stdout, "felis db backup: wrote %s\n", path)
return 0
}
// resolveBundle accepts a path, or a bare bundle name looked up in dir.
func resolveBundle(dir, arg string) string {
if strings.ContainsRune(arg, os.PathSeparator) {
return arg
}
if _, err := os.Stat(arg); err == nil {
return arg
}
return filepath.Join(dir, arg)
}
func dbRestore(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
yes := fs.Bool("yes", false, "replace the database's contents (required)")
force := fs.Bool("force", false, "restore even while other clients are connected")
noSafety := fs.Bool("no-safety-backup", false, "skip the bundle of the current database taken first")
stateDir := fs.String("state-dir", dbbackup.DefaultStateDir, "host state directory for the safety bundle")
arg, ok := parseWithArg(fs, args)
if !ok {
return 2
}
if arg == "" {
fmt.Fprint(stderr, dbUsage)
return 2
}
bundle := resolveBundle(*dir, arg)
m, err := dbbackup.Verify(bundle)
if err != nil {
fmt.Fprintf(stderr, "felis db restore: %v\n", err)
return 1
}
if !*yes {
fmt.Fprintf(stderr, "felis db restore: this replaces every table in the felis database with %s (%s, taken %s, schema %d).\n",
filepath.Base(bundle), m.Label, m.CreatedAt.Format(time.RFC3339), m.SchemaVersion)
fmt.Fprintln(stderr, "Scale felis-api and felis-operator to 0 first, then re-run with -yes.")
return 2
}
url, err := dbDatabaseURL(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis db restore: %v\n", err)
return 1
}
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Minute)
defer cancel()
_, safety, err := dbbackup.Restore(ctx, dbbackup.RestoreOptions{
DatabaseURL: url, Bundle: bundle, Dir: *dir, Force: *force, SkipSafetyBackup: *noSafety,
Safety: dbbackup.BackupOptions{Keep: defaultKeep[dbbackup.LabelPreRestore], StateDir: *stateDir,
Version: resolvedVersion(), ExportServers: exportMinecraftServers},
Log: stderr,
})
if err != nil {
fmt.Fprintf(stderr, "felis db restore: %v\n", err)
if errors.Is(err, dbbackup.ErrClientsConnected) {
fmt.Fprintln(stderr, " kubectl -n felis scale deployment felis-api felis-operator --replicas=0")
}
return 1
}
fmt.Fprintf(stdout, "felis db restore: restored %s (schema %d)\n", filepath.Base(bundle), m.SchemaVersion)
if safety != "" {
fmt.Fprintf(stdout, " the database as it was before is in %s\n", safety)
}
// Nothing migrates at startup, so a control plane newer than the bundle needs
// its migrations re-applied; rolling back to the release that wrote the bundle
// must skip that, or the rollback is undone.
fmt.Fprintf(stdout, " next: felis migrate up -config %s (skip it when rolling back to felis %s, which wrote this bundle)\n", *cfgPath, orUnknown(m.FelisVersion))
fmt.Fprintln(stdout, " kubectl -n felis scale deployment felis-api felis-operator --replicas=1")
return 0
}
func dbVerify(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
arg, ok := parseWithArg(fs, args)
if !ok {
return 2
}
if arg == "" {
fmt.Fprint(stderr, dbUsage)
return 2
}
bundle := resolveBundle(*dir, arg)
m, err := dbbackup.Verify(bundle)
if err != nil {
fmt.Fprintf(stderr, "felis db verify: %v\n", err)
return 1
}
fmt.Fprintf(stdout, "%s: ok\n taken %s (%s)\n felis %s\n schema %d\n %s\n",
filepath.Base(bundle), m.CreatedAt.Format(time.RFC3339), m.Label, orUnknown(m.FelisVersion), m.SchemaVersion, orUnknown(m.PGDumpVersion))
for _, f := range m.Files {
if f.Link != "" {
fmt.Fprintf(stdout, " %-40s -> %s\n", f.Name, f.Link)
continue
}
fmt.Fprintf(stdout, " %-40s %d bytes\n", f.Name, f.Size)
}
if m.ServersError != "" {
fmt.Fprintf(stdout, " (no MinecraftServer objects: %s)\n", m.ServersError)
}
return 0
}
func orUnknown(s string) string {
if s == "" {
return "unknown"
}
return s
}
func dbList(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
if err := fs.Parse(args); err != nil {
return 2
}
all, err := dbbackup.List(*dir)
if err != nil {
fmt.Fprintf(stderr, "felis db list: %v\n", err)
return 1
}
if len(all) == 0 {
fmt.Fprintf(stdout, "no database backups in %s\n", *dir)
return 0
}
now := time.Now()
for _, b := range all {
fmt.Fprintf(stdout, "%-50s %-12s %10s %s ago\n", b.Name, b.Label, humanBytes(b.Size), dbbackup.Age(now.Sub(b.Created)))
}
return 0
}
func humanBytes(n int64) string {
const unit = 1024
if n < unit {
return fmt.Sprintf("%d B", n)
}
div, exp := int64(unit), 0
for m := n / unit; m >= unit; m /= unit {
div *= unit
exp++
}
return fmt.Sprintf("%.1f %ciB", float64(n)/float64(div), "KMGTPE"[exp])
}
// dbCheck is the freshness probe: exit 1 when the newest bundle is missing or
// older than -max-age, for a monitor or the break-glass console to act on.
func dbCheck(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
maxAge := fs.Duration("max-age", dbbackup.StaleAfter, "oldest acceptable newest bundle")
if err := fs.Parse(args); err != nil {
return 2
}
b, err := dbbackup.Check(*dir, *maxAge, time.Now())
if err != nil {
fmt.Fprintf(stderr, "felis db check: %v\n", err)
return 1
}
fmt.Fprintf(stdout, "felis db check: ok, newest backup %s (%s ago)\n", b.Name, dbbackup.Age(time.Since(b.Created)))
return 0
}
// exportMinecraftServers reads every MinecraftServer through the host's k3s
// kubectl and strips what the API server owns, so the result can be fed back
// with `kubectl apply -f` on a rebuilt cluster.
func exportMinecraftServers(ctx context.Context) ([]byte, error) {
ctx, cancel := context.WithTimeout(ctx, 30*time.Second)
defer cancel()
// Output, not the CombinedOutput kubectlOutput uses: a deprecation warning
// on stderr must not end up inside the JSON.
cmd := exec.CommandContext(ctx, "k3s", "kubectl", "get", "minecraftservers.felis.lolicon.best", "-A", "-o", "json")
cmd.Env = append(os.Environ(), "KUBECONFIG="+hostBootstrapKubeconfigPath)
var errBuf strings.Builder
cmd.Stderr = &errBuf
out, err := cmd.Output()
if err != nil {
return nil, fmt.Errorf("k3s kubectl get minecraftservers: %w: %s", err, strings.TrimSpace(errBuf.String()))
}
return cleanServerList(out)
}
// cleanServerList drops status and the server-assigned metadata from a
// `kubectl get -o json` List.
func cleanServerList(raw []byte) ([]byte, error) {
var list struct {
Items []map[string]any `json:"items"`
}
if err := json.Unmarshal(raw, &list); err != nil {
return nil, fmt.Errorf("parse MinecraftServer list: %w", err)
}
for _, it := range list.Items {
delete(it, "status")
if md, ok := it["metadata"].(map[string]any); ok {
for _, k := range []string{"resourceVersion", "uid", "creationTimestamp", "generation", "managedFields", "selfLink"} {
delete(md, k)
}
}
}
if list.Items == nil {
list.Items = []map[string]any{}
}
return json.MarshalIndent(map[string]any{"apiVersion": "v1", "kind": "List", "items": list.Items}, "", " ")
}
+151
View File
@@ -0,0 +1,151 @@
package main
import (
"bytes"
"context"
"encoding/json"
"flag"
"io"
"strings"
"testing"
"felis.lolicon.best/internal/store"
)
func TestDBUsage(t *testing.T) {
for _, args := range [][]string{{"db"}, {"db", "frobnicate"}, {"db", "restore"}, {"db", "verify"}, {"db", "backup", "extra"}} {
var out, errBuf bytes.Buffer
if code := run(args, &out, &errBuf); code != 2 {
t.Errorf("%v: exit %d, want 2", args, code)
}
if !strings.Contains(errBuf.String(), "felis db restore") {
t.Errorf("%v: no usage on stderr: %q", args, errBuf.String())
}
}
}
func TestDBRestoreNeedsYes(t *testing.T) {
// A bundle that does not exist fails verification (1) before -yes matters;
// the -yes gate itself is exercised against a real bundle in internal/dbbackup
// and on the VM. Here: the refusal path never reaches the config or database.
var out, errBuf bytes.Buffer
if code := run([]string{"db", "restore", "-dir", t.TempDir(), "missing.tar"}, &out, &errBuf); code != 1 {
t.Fatalf("exit %d, stderr %q", code, errBuf.String())
}
}
func TestParseWithArg(t *testing.T) {
for _, args := range [][]string{{"-yes", "b.tar"}, {"b.tar", "-yes"}} {
fs := flag.NewFlagSet("t", flag.ContinueOnError)
fs.SetOutput(io.Discard)
yes := fs.Bool("yes", false, "")
arg, ok := parseWithArg(fs, args)
if !ok || arg != "b.tar" || !*yes {
t.Errorf("%v -> %q ok=%v yes=%v", args, arg, ok, *yes)
}
}
fs := flag.NewFlagSet("t", flag.ContinueOnError)
fs.SetOutput(io.Discard)
if _, ok := parseWithArg(fs, []string{"a.tar", "b.tar"}); ok {
t.Error("two positional arguments accepted")
}
}
func TestResolveBundle(t *testing.T) {
if got := resolveBundle("/var/lib/felis/db-backups", "felis-db-x.tar"); got != "/var/lib/felis/db-backups/felis-db-x.tar" {
t.Errorf("bare name -> %s", got)
}
if got := resolveBundle("/var/lib/felis/db-backups", "/root/copy.tar"); got != "/root/copy.tar" {
t.Errorf("path -> %s", got)
}
}
func TestCleanServerList(t *testing.T) {
raw := `{"apiVersion":"v1","kind":"List","metadata":{"resourceVersion":""},"items":[{
"apiVersion":"felis.lolicon.best/v1alpha1","kind":"MinecraftServer",
"metadata":{"name":"survival","namespace":"minecraft","uid":"u","resourceVersion":"42","generation":3,
"creationTimestamp":"2026-09-01T00:00:00Z","managedFields":[{}],"labels":{"a":"b"}},
"spec":{"desiredState":"Running"},"status":{"phase":"Running"}}]}`
out, err := cleanServerList([]byte(raw))
if err != nil {
t.Fatal(err)
}
var got struct {
Kind string `json:"kind"`
Items []map[string]any `json:"items"`
}
if err := json.Unmarshal(out, &got); err != nil {
t.Fatal(err)
}
if got.Kind != "List" || len(got.Items) != 1 {
t.Fatalf("got %s", out)
}
it := got.Items[0]
if _, ok := it["status"]; ok {
t.Error("status kept")
}
md := it["metadata"].(map[string]any)
for _, k := range []string{"uid", "resourceVersion", "generation", "creationTimestamp", "managedFields"} {
if _, ok := md[k]; ok {
t.Errorf("metadata.%s kept", k)
}
}
if md["name"] != "survival" || md["namespace"] != "minecraft" || md["labels"] == nil {
t.Errorf("identity lost: %v", md)
}
if it["spec"].(map[string]any)["desiredState"] != "Running" {
t.Error("spec lost")
}
empty, err := cleanServerList([]byte(`{"items":null}`))
if err != nil || !strings.Contains(string(empty), `"items": []`) {
t.Errorf("empty list -> %s, %v", empty, err)
}
if _, err := cleanServerList([]byte("Warning: x\n{")); err == nil {
t.Error("garbage parsed")
}
}
func TestHasPending(t *testing.T) {
ms := []store.Migration{{Version: 1}, {Version: 2}, {Version: 3}}
if hasPending(map[int]struct{}{1: {}, 2: {}, 3: {}}, ms) {
t.Error("fully applied reported pending")
}
if !hasPending(map[int]struct{}{1: {}, 2: {}}, ms) {
t.Error("missing 3 not reported")
}
}
type appliedDriver struct {
store.Driver
done map[int]struct{}
}
func (d appliedDriver) EnsureVersionTable(context.Context) error { return nil }
func (d appliedDriver) AppliedVersions(context.Context) (map[int]struct{}, error) {
return d.done, nil
}
func TestPreMigrateBackupOnlyGuardsAPopulatedDatabase(t *testing.T) {
ms := []store.Migration{{Version: 1}, {Version: 2}}
// An unusable URL makes an attempted backup observable as an error without
// any PostgreSQL tooling.
const badURL = "not-a-url"
for _, tc := range []struct {
name string
done map[int]struct{}
attempt bool
}{
{"fresh database", map[int]struct{}{}, false},
{"up to date", map[int]struct{}{1: {}, 2: {}}, false},
{"pending on a populated database", map[int]struct{}{1: {}}, true},
} {
path, err := preMigrateBackup(context.Background(), appliedDriver{done: tc.done}, ms, badURL, t.TempDir(), io.Discard)
if attempted := err != nil; attempted != tc.attempt {
t.Errorf("%s: attempted = %v (err %v), want %v", tc.name, attempted, err, tc.attempt)
}
if path != "" {
t.Errorf("%s: path = %q", tc.name, path)
}
}
}
+66
View File
@@ -0,0 +1,66 @@
package main
import (
"flag"
"fmt"
"io"
"net"
"os"
"time"
)
// Vars so tests can shrink them. A dial that neither connects nor is refused
// within egressDialTimeout counts as blocked: a policy that drops packets looks
// exactly like that.
var (
egressDialTimeout = 500 * time.Millisecond
egressPollInterval = 200 * time.Millisecond
)
// cmdEgressGate is the first initContainer of every build pod. The pod's
// NetworkPolicy is programmed asynchronously after the pod starts (live on k3s:
// a build-labelled pod reached the internet and the Kubernetes API for its first
// ~0.7 s), so the gate dials a destination the policy denies until it stops
// answering, and only then lets the pod's next container, eventually the
// untrusted Dockerfile, start.
//
// The default probe is the Kubernetes API Service, which the kubelet names in
// every pod's environment and the build policy never admits. A probe that still
// answers after --wait means the policy is not enforced at all (a CNI without
// NetworkPolicy support, or k3s run with --disable-network-policy), and the
// build fails closed.
func cmdEgressGate(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("egress-gate", flag.ContinueOnError)
fs.SetOutput(stderr)
probe := fs.String("probe", "", "host:port the build NetworkPolicy denies (default: the Kubernetes API Service from KUBERNETES_SERVICE_HOST/PORT)")
wait := fs.Duration("wait", 2*time.Minute, "how long the probe may keep answering before the build is refused")
if err := fs.Parse(args); err != nil {
return 2
}
if *probe == "" {
host, port := os.Getenv("KUBERNETES_SERVICE_HOST"), os.Getenv("KUBERNETES_SERVICE_PORT")
if host == "" || port == "" {
fmt.Fprintln(stderr, "felis egress-gate: no --probe and no KUBERNETES_SERVICE_HOST/PORT to default to")
return 2
}
*probe = net.JoinHostPort(host, port)
}
start := time.Now()
for {
conn, err := net.DialTimeout("tcp", *probe, egressDialTimeout)
if err != nil {
fmt.Fprintf(stdout, "felis egress-gate: %s is unreachable after %s (%v); the egress lock is in effect\n",
*probe, time.Since(start).Round(time.Millisecond), err)
return 0
}
_ = conn.Close()
if time.Since(start) >= *wait {
fmt.Fprintf(stderr, "felis egress-gate: %s still answers after %s: the build namespace's NetworkPolicy is not enforced "+
"(a CNI without NetworkPolicy support, or k3s started with --disable-network-policy); refusing to run the build\n",
*probe, *wait)
return 1
}
time.Sleep(egressPollInterval)
}
}
+98
View File
@@ -0,0 +1,98 @@
package main
import (
"bytes"
"net"
"strings"
"testing"
"time"
)
func shrinkEgressGate(t *testing.T) {
t.Helper()
dial, poll := egressDialTimeout, egressPollInterval
egressDialTimeout, egressPollInterval = 200*time.Millisecond, 10*time.Millisecond
t.Cleanup(func() { egressDialTimeout, egressPollInterval = dial, poll })
}
// The gate holds while the probe answers and lets the pod go on once the policy
// lands, which the test plays by closing the listener.
func TestEgressGateWaitsForTheLock(t *testing.T) {
shrinkEgressGate(t)
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
accepted := make(chan struct{}, 100)
go func() {
for {
c, err := ln.Accept()
if err != nil {
return
}
_ = c.Close()
accepted <- struct{}{}
}
}()
go func() {
for i := 0; i < 3; i++ {
<-accepted
}
_ = ln.Close()
}()
var out, errb bytes.Buffer
if code := cmdEgressGate([]string{"--probe", ln.Addr().String(), "--wait", "10s"}, &out, &errb); code != 0 {
t.Fatalf("exit %d: %s", code, errb.String())
}
if !strings.Contains(out.String(), "egress lock is in effect") {
t.Errorf("stdout = %q", out.String())
}
}
// A probe that keeps answering means no policy is enforced: the build must not run.
func TestEgressGateRefusesAnOpenNetwork(t *testing.T) {
shrinkEgressGate(t)
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
defer ln.Close()
go func() {
for {
c, err := ln.Accept()
if err != nil {
return
}
_ = c.Close()
}
}()
var out, errb bytes.Buffer
if code := cmdEgressGate([]string{"--probe", ln.Addr().String(), "--wait", "100ms"}, &out, &errb); code != 1 {
t.Fatalf("exit %d, want 1", code)
}
if !strings.Contains(errb.String(), "not enforced") {
t.Errorf("stderr = %q", errb.String())
}
}
func TestEgressGateDefaultsToTheKubernetesService(t *testing.T) {
shrinkEgressGate(t)
t.Setenv("KUBERNETES_SERVICE_HOST", "")
t.Setenv("KUBERNETES_SERVICE_PORT", "")
var out, errb bytes.Buffer
if code := cmdEgressGate(nil, &out, &errb); code != 2 {
t.Fatalf("exit %d without a probe, want 2", code)
}
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
host, port, _ := net.SplitHostPort(ln.Addr().String())
_ = ln.Close() // closed: the lock reads as in effect at once
t.Setenv("KUBERNETES_SERVICE_HOST", host)
t.Setenv("KUBERNETES_SERVICE_PORT", port)
out.Reset()
if code := cmdEgressGate(nil, &out, &errb); code != 0 || !strings.Contains(out.String(), ln.Addr().String()) {
t.Fatalf("exit %d, stdout %q", code, out.String())
}
}
+259
View File
@@ -0,0 +1,259 @@
package main
import (
"archive/tar"
"compress/gzip"
"context"
"crypto/sha256"
"encoding/hex"
"errors"
"flag"
"fmt"
"io"
"net/http"
"os"
"os/signal"
"path/filepath"
"strings"
"syscall"
"time"
"felis.lolicon.best/internal/build"
)
// cmdFetchContext is the in-Pod entrypoint the build Job's context-fetch
// initContainer runs. It reads the blob the platform stored for a submission
// from the felis-api INTERNAL face (with a bounded retry — see
// fetchContextWithRetry) and extracts it into the shared emptyDir the Kaniko
// container then builds from.
//
// Why this exists: the build Pod runs in the build namespace, where it can neither
// mount the control-plane uploads PVC (a PVC does not cross namespaces) nor hold
// object-store credentials, so the API that WROTE the blob is the transport. The
// route is service-token-gated; the token arrives through a namespace-local Secret
// mounted only into this initContainer, never into Kaniko's — so the untrusted
// Dockerfile's build steps have no credential to read (their containers share no
// environment, no PID namespace, and Kaniko itself mounts the context read-only).
//
// The extraction is deliberately paranoid: the tarball is attacker-controlled
// input, so absolute paths, ".." escapes, links, and special files are refused
// rather than sanitized. Kaniko treats the extracted tree as hostile regardless
// (spec §16), but the pod's own filesystem still must not be written outside the
// context directory it was given.
func cmdFetchContext(args []string, _, stderr io.Writer) int {
fs := flag.NewFlagSet("fetch-context", flag.ContinueOnError)
fs.SetOutput(stderr)
url := fs.String("url", "", "internal-face URL of the submission's build-context tarball")
out := fs.String("out", "/context", "directory to extract the build context into")
want := fs.String("sha256", "", "refuse the context unless the tarball's sha256 is this lowercase hex digest")
if err := fs.Parse(args); err != nil {
return 2
}
if *want != "" && !build.IsSHA256Hex(*want) {
fmt.Fprintf(stderr, "felis fetch-context: --sha256 %q is not a lowercase hex sha256\n", *want)
return 2
}
if *url == "" {
fmt.Fprintln(stderr, "felis fetch-context: --url is required")
return 2
}
token := os.Getenv("FELIS_SERVICE_TOKEN")
if token == "" {
fmt.Fprintln(stderr, "felis fetch-context: FELIS_SERVICE_TOKEN is empty — the internal face rejects anonymous reads")
return 2
}
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
// Validate the URL once up front: a bad one is a usage error (2), not
// something to sit in the retry loop.
if _, err := http.NewRequest(http.MethodGet, *url, nil); err != nil {
fmt.Fprintf(stderr, "felis fetch-context: bad --url: %v\n", err)
return 2
}
// No overall client timeout: a legitimate modpack context can be large and the
// Job's activeDeadlineSeconds is the real bound. The header timeout catches a
// wedged endpoint without capping a healthy download.
// Redirects are refused: the request carries the service token, and the
// internal face never redirects, so a 3xx is someone steering the token.
client := &http.Client{
Transport: &http.Transport{ResponseHeaderTimeout: time.Minute},
CheckRedirect: func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse },
}
resp, err := fetchContextWithRetry(ctx, client, *url, token, stderr)
if err != nil {
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
return 1
}
defer resp.Body.Close()
h := sha256.New()
body := io.TeeReader(resp.Body, h)
if err := extractTarGz(body, *out); err != nil {
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
return 1
}
if *want == "" {
return 0
}
// The tar end marker comes before the gzip trailer and whatever follows it,
// so read to EOF: the digest must cover every byte the blob holds. The blob
// itself is size-capped at upload, which bounds this read.
if _, err := io.Copy(io.Discard, io.LimitReader(body, maxContextBytes)); err != nil {
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
return 1
}
if got := hex.EncodeToString(h.Sum(nil)); got != *want {
// The init container failing is what keeps Kaniko from ever starting on
// the extracted tree.
fmt.Fprintf(stderr, "felis fetch-context: the context's sha256 is %s, the approved digest is %s: it changed after approval; refusing to build\n", got, *want)
return 1
}
return 0
}
// fetchRetryInterval/fetchRetryWindow bound how long the fetch waits out a
// control-plane blip before giving up. The api pod being replaced is a normal
// event (rollout, eviction, a chaos drill), and without a retry one refused
// dial turns it into a failed build: BackoffLimit=0 gives the Job no second
// Pod, so the terminal verdict costs a manual re-approval — the live drill hit
// exactly this (context-fetch exit 1 on `connect: connection refused` while
// the api pod rolled; the new pod was serving 11 seconds later and the same
// 198-byte blob). The window is tiny next to the Job's 30-minute
// activeDeadline; a 4xx (missing blob, rejected token) still fails fast.
//
// Vars, not consts, so tests can shrink the window.
var (
fetchRetryInterval = 3 * time.Second
fetchRetryWindow = 45 * time.Second
)
// fetchContextWithRetry GETs the context tarball, retrying transport failures
// and 5xx responses until fetchRetryWindow runs out. A 4xx is an answer, not a
// blip — retrying it only delays the honest error.
func fetchContextWithRetry(ctx context.Context, client *http.Client, url, token string, stderr io.Writer) (*http.Response, error) {
deadline := time.Now().Add(fetchRetryWindow)
for {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, fmt.Errorf("bad --url: %w", err)
}
req.Header.Set("Authorization", "Bearer "+token)
resp, err := client.Do(req)
if err == nil && resp.StatusCode == http.StatusOK {
return resp, nil
}
if err == nil {
status := resp.Status
_ = resp.Body.Close()
err = fmt.Errorf("GET returned %s", status)
if resp.StatusCode < 500 {
return nil, err
}
}
if ctx.Err() != nil {
return nil, fmt.Errorf("GET failed: %w", err)
}
if time.Now().After(deadline) {
return nil, fmt.Errorf("GET failed (retried for %s): %w", fetchRetryWindow, err)
}
fmt.Fprintf(stderr, "felis fetch-context: %v; retrying (the internal face may be restarting)\n", err)
select {
case <-ctx.Done():
return nil, fmt.Errorf("GET failed: %w", err)
case <-time.After(fetchRetryInterval):
}
}
}
// maxContextBytes / maxContextEntries bound what one context may expand to. The
// compressed upload is capped at 1 GiB, but gzip turns that into hundreds of GiB
// or millions of empty files, and the emptyDir's 4 GiB sizeLimit is only
// enforced by the kubelet's periodic sweep, after the disk has filled. The byte
// cap matches that sizeLimit; the entry cap is far above any real modpack (a
// large one is a few thousand files) and far below an inode exhaustion.
//
// Vars, not consts, so tests can shrink them.
var (
maxContextBytes int64 = 4 << 30
maxContextEntries = 200_000
)
// extractTarGz streams a gzip'd tarball into root, creating directories as
// needed. Every entry is vetted BEFORE anything is written: a path that is
// absolute or escapes root (via ".."), a link (symlink or hardlink), or any
// special file kind aborts the whole extraction. Refusing rather than skipping is
// deliberate — a context that needs one of those constructs is not a context this
// transport carries, and silently dropping entries would build from a corpus the
// submitter did not upload. The whole extraction is also bounded by
// maxContextBytes and maxContextEntries.
func extractTarGz(r io.Reader, root string) error {
if err := os.MkdirAll(root, 0o755); err != nil {
return fmt.Errorf("create context dir: %w", err)
}
zr, err := gzip.NewReader(r)
if err != nil {
return fmt.Errorf("context is not a valid gzip tarball: %w", err)
}
defer zr.Close()
tr := tar.NewReader(zr)
var written int64
entries := 0
for {
hdr, err := tr.Next()
if errors.Is(err, io.EOF) {
return nil
}
if err != nil {
return fmt.Errorf("read context tarball: %w", err)
}
if entries++; entries > maxContextEntries {
return fmt.Errorf("the build context has more than %d entries", maxContextEntries)
}
name := filepath.Clean(hdr.Name)
if name == "." {
continue
}
// The zip-slip guard: reject, never rewrite. filepath.Clean collapses any
// "a/../../b", so these two checks are sufficient once Clean has run.
if filepath.IsAbs(name) || name == ".." || strings.HasPrefix(name, ".."+string(filepath.Separator)) {
return fmt.Errorf("context entry %q escapes the context directory", hdr.Name)
}
target := filepath.Join(root, name)
switch hdr.Typeflag {
case tar.TypeDir:
if err := os.MkdirAll(target, 0o755); err != nil {
return fmt.Errorf("create %q: %w", name, err)
}
case tar.TypeReg:
if err := os.MkdirAll(filepath.Dir(target), 0o755); err != nil {
return fmt.Errorf("create parent of %q: %w", name, err)
}
mode := os.FileMode(0o644)
if hdr.FileInfo().Mode()&0o111 != 0 {
mode = 0o755 // preserve executability (entrypoint scripts), nothing else
}
f, err := os.OpenFile(target, os.O_CREATE|os.O_WRONLY|os.O_TRUNC, mode)
if err != nil {
return fmt.Errorf("create %q: %w", name, err)
}
n, err := io.Copy(f, io.LimitReader(tr, maxContextBytes-written+1))
written += n
if err != nil {
_ = f.Close()
return fmt.Errorf("write %q: %w", name, err)
}
if written > maxContextBytes {
_ = f.Close()
return fmt.Errorf("the build context expands past %d bytes", maxContextBytes)
}
if err := f.Close(); err != nil {
return fmt.Errorf("close %q: %w", name, err)
}
default:
return fmt.Errorf("context entry %q has unsupported type %q (links and special files are refused)", hdr.Name, string(hdr.Typeflag))
}
}
}
+404
View File
@@ -0,0 +1,404 @@
package main
import (
"archive/tar"
"bytes"
"compress/gzip"
"crypto/sha256"
"encoding/hex"
"io"
"net"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strings"
"sync/atomic"
"testing"
"time"
)
type tarEntry struct {
name string
body string
mode int64
typ byte
linkname string
}
// tgzBody builds an in-memory .tar.gz from entries, preserving each entry's type
// and mode so the tests can exercise the guards with exactly the bytes an
// attacker could upload.
func tgzBody(t *testing.T, entries ...tarEntry) []byte {
t.Helper()
var buf bytes.Buffer
zw := gzip.NewWriter(&buf)
tw := tar.NewWriter(zw)
for _, e := range entries {
typ := e.typ
if typ == 0 {
typ = tar.TypeReg
}
mode := e.mode
if mode == 0 {
mode = 0o644
}
hdr := &tar.Header{Name: e.name, Typeflag: typ, Mode: mode, Size: int64(len(e.body))}
if typ == tar.TypeSymlink {
hdr.Linkname = e.linkname
hdr.Size = 0
}
if err := tw.WriteHeader(hdr); err != nil {
t.Fatalf("write header %q: %v", e.name, err)
}
if hdr.Size > 0 {
if _, err := tw.Write([]byte(e.body)); err != nil {
t.Fatalf("write body %q: %v", e.name, err)
}
}
}
if err := tw.Close(); err != nil {
t.Fatalf("close tar: %v", err)
}
if err := zw.Close(); err != nil {
t.Fatalf("close gzip: %v", err)
}
return buf.Bytes()
}
// A normal context extracts with its tree intact, and the executable bit that
// modpack entrypoints rely on survives.
func TestExtractTarGzRoundTrip(t *testing.T) {
dir := t.TempDir()
body := tgzBody(t,
tarEntry{name: "Dockerfile", body: "FROM scratch\n"},
tarEntry{name: "mods/example.jar", body: "jar-bytes"},
tarEntry{name: "start.sh", body: "#!/bin/sh\n", mode: 0o755},
tarEntry{name: "mods/", typ: tar.TypeDir, mode: 0o755},
)
if err := extractTarGz(bytes.NewReader(body), dir); err != nil {
t.Fatalf("extract: %v", err)
}
for name, want := range map[string]string{
"Dockerfile": "FROM scratch\n",
"mods/example.jar": "jar-bytes",
} {
got, err := os.ReadFile(filepath.Join(dir, name))
if err != nil || string(got) != want {
t.Fatalf("%s = (%q, %v), want %q", name, got, err, want)
}
}
fi, err := os.Stat(filepath.Join(dir, "start.sh"))
if err != nil || fi.Mode()&0o111 == 0 {
t.Fatalf("entrypoint script lost its exec bit: %v (%v)", fi, err)
}
}
// The guards: "..", absolute paths, symlinks, and special files are refused whole
// — nothing escapes, and nothing is silently skipped.
func TestExtractTarGzRefusesEscapes(t *testing.T) {
cases := []struct {
name string
entries []tarEntry
}{
{"dotdot", []tarEntry{{name: "../outside", body: "x"}}},
{"nested dotdot", []tarEntry{{name: "a/../../outside", body: "x"}}},
{"absolute", []tarEntry{{name: "/etc/outside", body: "x"}}},
{"symlink", []tarEntry{{name: "link", typ: tar.TypeSymlink, linkname: "/etc"}}},
{"hardlink", []tarEntry{{name: "hard", typ: tar.TypeLink, linkname: "somewhere"}}},
{"device", []tarEntry{{name: "dev", typ: tar.TypeChar}}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
dir := t.TempDir()
if err := extractTarGz(bytes.NewReader(tgzBody(t, tc.entries...)), dir); err == nil {
t.Fatal("extract accepted a hostile entry, want an error")
}
// Nothing may have been written outside the target (or at all).
entries, _ := os.ReadDir(dir)
if len(entries) != 0 {
t.Fatalf("hostile archive left %d entries behind", len(entries))
}
})
}
}
// A context that expands past the byte or entry cap is refused, however small
// it was compressed: gzip bombs and inode floods stop at the cap.
func TestExtractTarGzCapsExpansion(t *testing.T) {
bytesCap, entriesCap := maxContextBytes, maxContextEntries
t.Cleanup(func() { maxContextBytes, maxContextEntries = bytesCap, entriesCap })
maxContextBytes, maxContextEntries = 1000, 5
fits := tgzBody(t, tarEntry{name: "a", body: strings.Repeat("x", 600)}, tarEntry{name: "b", body: strings.Repeat("y", 400)})
if err := extractTarGz(bytes.NewReader(fits), t.TempDir()); err != nil {
t.Fatalf("a context exactly at the byte cap: %v", err)
}
big := tgzBody(t, tarEntry{name: "a", body: strings.Repeat("x", 600)}, tarEntry{name: "b", body: strings.Repeat("y", 401)})
if err := extractTarGz(bytes.NewReader(big), t.TempDir()); err == nil || !strings.Contains(err.Error(), "expands past") {
t.Fatalf("one byte over the cap: err = %v", err)
}
var many []tarEntry
for i := 0; i < 6; i++ {
many = append(many, tarEntry{name: "d" + string(rune('0'+i)) + "/", typ: tar.TypeDir})
}
if err := extractTarGz(bytes.NewReader(tgzBody(t, many...)), t.TempDir()); err == nil || !strings.Contains(err.Error(), "entries") {
t.Fatalf("six entries over a cap of five: err = %v", err)
}
}
// The command end to end: it dials the URL with the bearer token from the
// environment, and refuses to run without it (the internal face would 401
// anyway; failing at parse time is the honest earlier error).
func TestCmdFetchContextFetchAndExtract(t *testing.T) {
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Header.Get("Authorization") != "Bearer test-token" {
w.WriteHeader(http.StatusUnauthorized)
return
}
w.Header().Set("Content-Type", "application/gzip")
_, _ = w.Write(body)
}))
defer srv.Close()
dir := t.TempDir()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + dir}, io.Discard, io.Discard); code != 0 {
t.Fatalf("cmdFetchContext exit = %d, want 0", code)
}
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
}
// No token: refuse before dialing.
t.Setenv("FELIS_SERVICE_TOKEN", "")
var stderr bytes.Buffer
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 2 {
t.Fatalf("missing token exit = %d, want 2 (stderr %q)", code, stderr.String())
}
// A non-200 answer (e.g. the route's 404 for a never-uploaded context) fails.
srv404 := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusNotFound)
}))
defer srv404.Close()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
if code := cmdFetchContext([]string{"--url=" + srv404.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, io.Discard); code != 1 {
t.Fatalf("404 exit = %d, want 1", code)
}
}
// With --sha256 the fetch refuses any bytes but the approved ones, including
// a tarball that extracts cleanly: that is exactly the context an uploader
// swapped in after the review.
func TestCmdFetchContextChecksDigest(t *testing.T) {
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write(body)
}))
defer srv.Close()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
sum := sha256.Sum256(body)
good := hex.EncodeToString(sum[:])
args := func(digest string) []string {
return []string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir(), "--sha256=" + digest}
}
if code := cmdFetchContext(args(good), io.Discard, io.Discard); code != 0 {
t.Fatalf("matching digest exit = %d, want 0", code)
}
var stderr bytes.Buffer
other := strings.Repeat("0", 64)
if code := cmdFetchContext(args(other), io.Discard, &stderr); code != 1 || !strings.Contains(stderr.String(), "changed after approval") {
t.Fatalf("mismatched digest exit = %d, stderr %q; want 1 naming the change", code, stderr.String())
}
stderr.Reset()
if code := cmdFetchContext(args("ABC"), io.Discard, &stderr); code != 2 {
t.Fatalf("malformed digest exit = %d, want 2 (stderr %q)", code, stderr.String())
}
// Bytes after the tar end marker still count: appending to an approved blob
// must change what the fetch accepts.
padded := append(append([]byte{}, body...), "trailing"...)
srvPadded := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write(padded)
}))
defer srvPadded.Close()
if code := cmdFetchContext([]string{"--url=" + srvPadded.URL + "/c", "--out=" + t.TempDir(), "--sha256=" + good}, io.Discard, io.Discard); code != 1 {
t.Fatalf("padded blob exit = %d, want 1", code)
}
}
// The request carries the service token, so a redirect is a failure: the token
// never follows it to another host (build-supply-chain-13).
func TestCmdFetchContextRefusesRedirects(t *testing.T) {
var leaked bool
elsewhere := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
leaked = true
}))
defer elsewhere.Close()
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, elsewhere.URL+"/steal", http.StatusFound)
}))
defer srv.Close()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/c", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
t.Fatalf("redirect exit = %d, want 1 (stderr %q)", code, stderr.String())
}
if leaked {
t.Fatal("the fetch followed the redirect")
}
}
// A body that is not a gzip tarball must fail the extraction rather than produce
// an empty (or partial) context Kaniko would then try to build.
func TestExtractTarGzRejectsNonGzip(t *testing.T) {
dir := t.TempDir()
err := extractTarGz(strings.NewReader("not a tarball"), dir)
if err == nil || !strings.Contains(err.Error(), "gzip") {
t.Fatalf("err = %v, want a gzip complaint", err)
}
}
// shrinkFetchWindow swaps the retry knobs for a faster test and restores them
// afterwards, so no test leaks a tiny window into another.
func shrinkFetchWindow(t *testing.T, interval, window time.Duration) {
t.Helper()
oldInterval, oldWindow := fetchRetryInterval, fetchRetryWindow
fetchRetryInterval, fetchRetryWindow = interval, window
t.Cleanup(func() { fetchRetryInterval, fetchRetryWindow = oldInterval, oldWindow })
}
// A control-plane blip mid-fetch is survived: a 5xx on the first attempt is
// retried and the second attempt's tarball extracts. This walks back the live
// drill's failure, where the api pod rolled mid-fetch and the single attempt
// died, failing the build Job.
func TestFetchContextRetriesThroughBlip(t *testing.T) {
shrinkFetchWindow(t, 10*time.Millisecond, time.Second)
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
var calls int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
if atomic.AddInt32(&calls, 1) == 1 {
w.WriteHeader(http.StatusBadGateway) // the port is up, the API is not
return
}
_, _ = w.Write(body)
}))
defer srv.Close()
dir := t.TempDir()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + dir}, io.Discard, &stderr); code != 0 {
t.Fatalf("exit = %d, want 0 (stderr %q)", code, stderr.String())
}
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
}
if !strings.Contains(stderr.String(), "retrying") {
t.Fatalf("stderr %q does not mention the retry", stderr.String())
}
}
// The live drill's exact shape: the dial itself is refused (the api pod is
// gone and no endpoint answers). A refused dial is retried like any other
// transport failure, and once the face is back the fetch completes.
func TestFetchContextRetriesRefusedDial(t *testing.T) {
shrinkFetchWindow(t, 10*time.Millisecond, 5*time.Second)
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
// Borrow a listen address, then close it: the first attempts dial into a
// refused connection, exactly like a restarting control plane.
probe := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
addr := strings.TrimPrefix(probe.URL, "http://")
probe.Close()
dir := t.TempDir()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
// Start the fetch; while the retry loop burns refused dials, bring the same
// address back.
result := make(chan int, 1)
go func() {
result <- cmdFetchContext([]string{"--url=http://" + addr + "/sub-1/context", "--out=" + dir}, io.Discard, &stderr)
}()
time.Sleep(100 * time.Millisecond) // let a handful of dials be refused
ln, err := net.Listen("tcp", addr)
if err != nil {
t.Fatalf("rebind %s: %v", addr, err)
}
back := &http.Server{Handler: http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Header.Get("Authorization") != "Bearer test-token" {
w.WriteHeader(http.StatusUnauthorized)
return
}
_, _ = w.Write(body)
})}
defer back.Close()
go func() { _ = back.Serve(ln) }()
code := <-result
if code != 0 {
t.Fatalf("exit = %d, want 0 (stderr %q)", code, stderr.String())
}
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
}
if !strings.Contains(stderr.String(), "retrying") {
t.Fatalf("stderr %q does not mention the retry", stderr.String())
}
}
// A 4xx is an answer, not a blip: a missing/never-uploaded context fails
// immediately — no retry loop burns the build's deadline on a terminal error.
func TestFetchContextDoesNotRetry4xx(t *testing.T) {
shrinkFetchWindow(t, 5*time.Millisecond, 200*time.Millisecond)
var calls int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
atomic.AddInt32(&calls, 1)
w.WriteHeader(http.StatusNotFound)
}))
defer srv.Close()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
t.Fatalf("exit = %d, want 1 (stderr %q)", code, stderr.String())
}
if got := atomic.LoadInt32(&calls); got != 1 {
t.Fatalf("server saw %d attempts, want exactly 1", got)
}
if strings.Contains(stderr.String(), "retrying") {
t.Fatalf("stderr %q mentions a retry for a terminal 4xx", stderr.String())
}
}
// The retry is bounded: an internal face that stays down does not hang the
// build pod; the window runs out and the fetch reports the exhausted retries.
func TestFetchContextGivesUpAfterWindow(t *testing.T) {
shrinkFetchWindow(t, 5*time.Millisecond, 60*time.Millisecond)
var calls int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
atomic.AddInt32(&calls, 1)
w.WriteHeader(http.StatusServiceUnavailable)
}))
defer srv.Close() // the face is up but never healthy: 503 forever
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
start := time.Now()
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
t.Fatalf("exit = %d, want 1 (stderr %q)", code, stderr.String())
}
if elapsed := time.Since(start); elapsed > 5*time.Second {
t.Fatalf("gave up after %v; the window is supposed to bound it", elapsed)
}
if got := atomic.LoadInt32(&calls); got < 2 {
t.Fatalf("server saw %d attempts, want at least one retry", got)
}
if !strings.Contains(stderr.String(), "retried for") {
t.Fatalf("stderr %q does not report the exhausted retry window", stderr.String())
}
}
+11 -16
View File
@@ -20,18 +20,14 @@ const forwardingSecretEnv = "FELIS_FORWARDING_SECRET"
// config/ and server.properties live under it. // config/ and server.properties live under it.
const defaultForwardingDataDir = "/data" const defaultForwardingDataDir = "/data"
// fwd*Mode make the written config readable AND rewritable by the main server // fwd*Mode are the modes the written config lands with. The initContainer runs as
// container, whose UID we do not control (an arbitrary user image). The // the same uid as the server container (naming.GameUID, pinned by the operator in
// initContainer runs as root (see buildStatefulSet) so it can write into a data // the pod securityContext) after the prepare-data initContainer has handed the
// volume of unknown ownership; 0666/0777 then let a non-root Paper rewrite the // whole volume to that uid, so owner read/write is all the server needs to rewrite
// same files on boot. // these files on boot and nothing else on the node gets write access to them.
//
// This relies on the initContainer running as root to write into a volume of
// unknown ownership; that is how the operator schedules it. If that ever changes,
// give the server pod an fsGroup so the shared volume is group-writable instead.
const ( const (
fwdFileMode os.FileMode = 0o666 fwdFileMode os.FileMode = 0o644
fwdDirMode os.FileMode = 0o777 fwdDirMode os.FileMode = 0o755
) )
// cmdInitForwarding is the felis-image initContainer entrypoint that makes an // cmdInitForwarding is the felis-image initContainer entrypoint that makes an
@@ -96,9 +92,8 @@ func writePaperGlobal(dataDir, secret string) error {
if err := os.MkdirAll(dir, fwdDirMode); err != nil { if err := os.MkdirAll(dir, fwdDirMode); err != nil {
return fmt.Errorf("create %s: %w", dir, err) return fmt.Errorf("create %s: %w", dir, err)
} }
// MkdirAll honours the process umask (root's is typically 022 → 0755); chmod // MkdirAll honours the process umask; chmod does not, so a directory an older
// does not, and a non-root main container must be able to place/replace the // release left at 0777 is brought back to fwdDirMode here.
// file in this directory on boot.
if err := os.Chmod(dir, fwdDirMode); err != nil { if err := os.Chmod(dir, fwdDirMode); err != nil {
return fmt.Errorf("chmod %s: %w", dir, err) return fmt.Errorf("chmod %s: %w", dir, err)
} }
@@ -188,8 +183,8 @@ func upsertProperty(content []byte, key, value string) []byte {
} }
// writeFileMode writes data then forces the mode, since WriteFile honours the // writeFileMode writes data then forces the mode, since WriteFile honours the
// umask (root's is typically 022 → 0644) but a non-root main container must be // umask and leaves an existing file's mode alone: a file an older release wrote
// able to rewrite these files on boot. // world-writable (0666) is tightened back to fwdFileMode on the next boot.
func writeFileMode(path string, data []byte) error { func writeFileMode(path string, data []byte) error {
if err := os.WriteFile(path, data, fwdFileMode); err != nil { if err := os.WriteFile(path, data, fwdFileMode); err != nil {
return fmt.Errorf("write %s: %w", path, err) return fmt.Errorf("write %s: %w", path, err)
+3 -2
View File
@@ -164,8 +164,9 @@ func TestUpsertPropertyAppends(t *testing.T) {
} }
} }
// The written files must be group/world writable so a non-root main container can // The written files land at fwdFileMode: owner-writable for the game uid the init
// rewrite them. chmod semantics are POSIX-only, so this asserts on non-Windows. // shares with the server container, and no longer world-writable. chmod semantics
// are POSIX-only, so this asserts on non-Windows.
func TestWriteForwardingFileModes(t *testing.T) { func TestWriteForwardingFileModes(t *testing.T) {
if runtime.GOOS == "windows" { if runtime.GOOS == "windows" {
t.Skip("POSIX file modes not represented on Windows") t.Skip("POSIX file modes not represented on Windows")
+117
View File
@@ -0,0 +1,117 @@
package main
import (
"errors"
"flag"
"fmt"
"io"
"io/fs"
"os"
"syscall"
"felis.lolicon.best/internal/naming"
)
// cmdInitVolume is the felis-image `prepare-data` initContainer entrypoint: it
// hands every entry of a server's world volume to the game uid/gid before the
// server container starts. The operator runs the server itself as naming.GameUID,
// so a world written by an earlier release (whose server ran as root), a restore
// Job (which extracts as root), or a storage provisioner that creates the volume
// root-owned would otherwise leave files the server cannot write — a world that
// boots and then fails every save.
//
// fsGroup covers only part of this: kubelet applies it to volume types that
// support ownership management, and a k3s local-path PV is a hostPath underneath,
// which it skips. A walk from inside the pod works for every volume type.
//
// Only mismatched entries are touched, so a volume already owned by the game uid
// costs one lstat per entry and no writes. The walk runs inside an os.Root at the
// data dir and uses lchown, so a symlink a plugin planted is re-owned as a link
// and never followed out of the volume.
//
// A single entry that cannot be chowned is reported and skipped: failing the pod
// over one odd file would keep the whole server down, while the server itself
// reports the one file it cannot write. Only an unreadable data dir fails.
func cmdInitVolume(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("init-volume", flag.ContinueOnError)
fs.SetOutput(stderr)
dataDir := fs.String("data", defaultForwardingDataDir, "world volume mount to hand to the game uid")
uid := fs.Int64("uid", naming.GameUID, "owner uid for every entry")
gid := fs.Int64("gid", naming.GameGID, "owner gid for every entry")
if err := fs.Parse(args); err != nil {
return 2
}
root, err := os.OpenRoot(*dataDir)
if err != nil {
fmt.Fprintf(stderr, "felis init-volume: open %s: %v\n", *dataDir, err)
return 1
}
defer root.Close()
res, err := chownTree(root, int(*uid), int(*gid), root.Lchown)
if err != nil {
fmt.Fprintf(stderr, "felis init-volume: %v\n", err)
return 1
}
for _, f := range res.failures {
fmt.Fprintf(stderr, "felis init-volume: %s\n", f)
}
fmt.Fprintf(stdout, "felis init-volume: %d entries checked, %d handed to %d:%d, %d failed\n",
res.checked, res.changed, *uid, *gid, len(res.failures))
return 0
}
// chownResult tallies one walk; failures is capped so a volume of thousands of
// unownable files cannot flood the pod log.
type chownResult struct {
checked int
changed int
failures []string
}
const maxReportedChownFailures = 20
// chownTree walks root and calls chown on every entry (the root dir included)
// whose owner is not uid:gid. It returns an error only when the root itself
// cannot be read; per-entry failures are collected in the result.
func chownTree(root *os.Root, uid, gid int, chown func(name string, uid, gid int) error) (chownResult, error) {
var res chownResult
fail := func(name string, err error) {
if len(res.failures) < maxReportedChownFailures {
res.failures = append(res.failures, fmt.Sprintf("%s: %v", name, err))
} else if len(res.failures) == maxReportedChownFailures {
res.failures = append(res.failures, "further failures not listed")
}
}
err := fs.WalkDir(root.FS(), ".", func(name string, d fs.DirEntry, walkErr error) error {
if walkErr != nil {
if name == "." {
return walkErr
}
fail(name, walkErr)
// A directory that cannot be listed is skipped as a whole; a file
// error has nothing below it to skip.
if d != nil && d.IsDir() {
return fs.SkipDir
}
return nil
}
res.checked++
info, err := d.Info()
if err != nil {
fail(name, err)
return nil
}
if st, ok := info.Sys().(*syscall.Stat_t); ok && int(st.Uid) == uid && int(st.Gid) == gid {
return nil
}
if err := chown(name, uid, gid); err != nil {
if !errors.Is(err, fs.ErrNotExist) { // gone mid-walk: nothing left to own
fail(name, err)
}
return nil
}
res.changed++
return nil
})
return res, err
}
+124
View File
@@ -0,0 +1,124 @@
package main
import (
"bytes"
"os"
"path/filepath"
"runtime"
"slices"
"strconv"
"testing"
)
// openTree builds a small world under a temp dir: nested dirs, a file, and a
// symlink pointing out of the volume that the walk must not follow.
func openTree(t *testing.T) *os.Root {
t.Helper()
dir := t.TempDir()
for _, d := range []string{"world/region", "plugins"} {
if err := os.MkdirAll(filepath.Join(dir, d), 0o755); err != nil {
t.Fatal(err)
}
}
for _, f := range []string{"level.dat", "world/region/r.0.0.mca"} {
if err := os.WriteFile(filepath.Join(dir, f), []byte("x"), 0o600); err != nil {
t.Fatal(err)
}
}
outside := t.TempDir()
if err := os.Symlink(outside, filepath.Join(dir, "plugins", "escape")); err != nil {
t.Fatal(err)
}
root, err := os.OpenRoot(dir)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { root.Close() })
return root
}
// Every entry owned by someone else is handed over, the root dir included, and a
// symlink is re-owned as a link rather than walked into.
func TestChownTreeHandsOverMismatchedEntries(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("POSIX ownership not represented on Windows")
}
root := openTree(t)
var got []string
res, err := chownTree(root, os.Getuid()+1, os.Getgid(), func(name string, uid, gid int) error {
got = append(got, name)
return nil
})
if err != nil {
t.Fatalf("chownTree: %v", err)
}
want := []string{".", "level.dat", "plugins", "plugins/escape", "world", "world/region", "world/region/r.0.0.mca"}
slices.Sort(got)
if !slices.Equal(got, want) {
t.Errorf("chowned %v, want %v", got, want)
}
if res.changed != len(want) || res.checked != len(want) || len(res.failures) != 0 {
t.Errorf("result = %+v, want %d checked and changed, no failures", res, len(want))
}
}
// A volume already owned by the game uid costs no chown at all: this is the steady
// state every restart after the first one hits.
func TestChownTreeSkipsMatchingOwner(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("POSIX ownership not represented on Windows")
}
root := openTree(t)
calls := 0
res, err := chownTree(root, os.Getuid(), os.Getgid(), func(string, int, int) error {
calls++
return nil
})
if err != nil {
t.Fatalf("chownTree: %v", err)
}
if calls != 0 || res.changed != 0 {
t.Errorf("chown called %d times on an already-owned tree (result %+v)", calls, res)
}
}
// One entry that refuses the chown is reported and the walk carries on: a single
// odd file must not keep the whole server from starting.
func TestChownTreeContinuesPastFailures(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("POSIX ownership not represented on Windows")
}
root := openTree(t)
res, err := chownTree(root, os.Getuid()+1, os.Getgid(), func(name string, uid, gid int) error {
if name == "level.dat" {
return os.ErrPermission
}
return nil
})
if err != nil {
t.Fatalf("chownTree: %v", err)
}
if len(res.failures) != 1 || res.changed != 6 {
t.Errorf("result = %+v, want 1 failure and 6 changed", res)
}
}
// Against a real directory the owner already matches, so the command succeeds
// without needing CAP_CHOWN — the path every test runner (non-root) can take.
func TestCmdInitVolumeOwnedTree(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("POSIX ownership not represented on Windows")
}
dir := t.TempDir()
if err := os.WriteFile(filepath.Join(dir, "server.properties"), []byte("x"), 0o644); err != nil {
t.Fatal(err)
}
var out, errb bytes.Buffer
code := cmdInitVolume([]string{"--data", dir, "--uid", strconv.Itoa(os.Getuid()), "--gid", strconv.Itoa(os.Getgid())}, &out, &errb)
if code != 0 {
t.Fatalf("exit %d, stderr %q", code, errb.String())
}
if code := cmdInitVolume([]string{"--data", filepath.Join(dir, "missing")}, &out, &errb); code != 1 {
t.Errorf("missing data dir exit = %d, want 1", code)
}
}
+70 -28
View File
@@ -26,8 +26,10 @@ func (m *multiFlag) Set(v string) error {
// felis-reaper identity only when the retention reaper is enabled, gated with // felis-reaper identity only when the retention reaper is enabled, gated with
// its CronJob), the weak build/restore Job SAs, the build/minecraft // its CronJob), the weak build/restore Job SAs, the build/minecraft
// NetworkPolicies, and the running control-plane workloads (felis-api/operator // NetworkPolicies, and the running control-plane workloads (felis-api/operator
// Deployments + the in-cluster registry Deployment/Service/PVC) — as a single // Deployments + the in-cluster registry Deployment/Service/PVC + the
// multi-document YAML stream on stdout, ready for `kubectl apply -f -`. // world-archive PVC that backs backup/restore, unless --backup-pvc is emptied)
// — as a single multi-document YAML stream on stdout, ready for
// `kubectl apply -f -`.
// //
// It is a pure renderer: it never contacts a cluster and holds no credentials. // It is a pure renderer: it never contacts a cluster and holds no credentials.
// --velocity-cidr records the proxy host addresses allowed by the game NetworkPolicy. // --velocity-cidr records the proxy host addresses allowed by the game NetworkPolicy.
@@ -43,14 +45,22 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
registryPort := fs.Int("registry-port", 5000, "port the in-cluster registry listens on") registryPort := fs.Int("registry-port", 5000, "port the in-cluster registry listens on")
panelNodePort := fs.Int("panel-node-port", int(platform.DefaultPanelNodePort), "NodePort that exposes the built-in HTTPS panel/API origin") panelNodePort := fs.Int("panel-node-port", int(platform.DefaultPanelNodePort), "NodePort that exposes the built-in HTTPS panel/API origin")
felisImage := fs.String("felis-image", "", "container image the felis-api/operator Deployments run, also passed through as FELIS_IMAGE (REQUIRED)") felisImage := fs.String("felis-image", "", "container image the felis-api/operator Deployments run, also passed through as FELIS_IMAGE (REQUIRED)")
registryImage := fs.String("registry-image", "", "in-cluster registry image (default: registry:2)") registryImage := fs.String("registry-image", "", "in-cluster registry image (default: registry 2.8.3, pinned by digest)")
backupPVC := fs.String("backup-pvc", "", "name of the backup PVC advertised to the restore executor via FELIS_BACKUP_PVC (default none = restore endpoint returns 503)") backupPVC := fs.String("backup-pvc", "felis-backups", "name of the world-archive PVC this bundle renders in the Minecraft namespace and advertises to the backup/restore executors via FELIS_BACKUP_PVC (default: felis-backups; pass an empty value to render none, leaving backup/restore answering 503)")
worldsHostPath := fs.String("worlds-host-path", "", "node directory under which each world PVC is visible as <path>/<pvc>; enables the reaper CronJob (requires --backup-pvc and --archive-local-path)") worldsHostPath := fs.String("worlds-host-path", "", "node directory the reaper reads worlds from: each world PVC resolves as <path>/<pvc>, or as the stock local-path directory <path>/<pv-name>_<ns>_<pvc-name> (k3s storage root: /var/lib/rancher/k3s/storage); enables the reaper CronJob (requires --archive-local-path and a non-empty --backup-pvc)")
archiveLocalPath := fs.String("archive-local-path", "", "path the backup PVC is mounted at in the reaper CronJob; MUST equal felis.toml [archive] local_path") archiveLocalPath := fs.String("archive-local-path", "", "path the backup PVC is mounted at in the reaper CronJob; MUST equal felis.toml [archive] local_path")
registryStorage := fs.String("registry-storage", "", "capacity the registry PVC requests (default 10Gi; k3s local-path does not enforce it)")
uploadsStorage := fs.String("uploads-storage", "", "capacity the uploads PVC requests (default 5Gi; k3s local-path does not enforce it)")
backupStorage := fs.String("backup-storage", "", "capacity the world-archive PVC requests (default 10Gi; k3s local-path does not enforce it)")
reaperNode := fs.String("reaper-node", "", "node that holds --worlds-host-path: pins the reaper CronJob's pod there via nodeSelector kubernetes.io/hostname (multi-node clusters need this, or the reaper may schedule where the hostPath is empty)")
var velocityCIDRs multiFlag var velocityCIDRs multiFlag
fs.Var(&velocityCIDRs, "velocity-cidr", "CIDR of a Velocity proxy host allowed to reach game port 25565 (repeatable, REQUIRED)") fs.Var(&velocityCIDRs, "velocity-cidr", "CIDR of a Velocity proxy host allowed to reach game port 25565 (repeatable, REQUIRED)")
var packageCIDRs multiFlag var packageCIDRs multiFlag
fs.Var(&packageCIDRs, "package-cidr", "CIDR of a package mirror build Pods may reach (repeatable; default none = no internet egress)") fs.Var(&packageCIDRs, "package-cidr", "CIDR of a package mirror build Pods may reach (repeatable; default none = no internet egress)")
var serverDenyCIDRs multiFlag
fs.Var(&serverDenyCIDRs, "server-egress-deny-cidr", "extra CIDR game server pods may never reach, e.g. the node's public address (repeatable)")
var serverAllowCIDRs multiFlag
fs.Var(&serverAllowCIDRs, "server-egress-allow-cidr", "private CIDR game server pods may reach despite the private-range block, e.g. a LAN database (repeatable)")
if err := fs.Parse(args); err != nil { if err := fs.Parse(args); err != nil {
return 2 return 2
} }
@@ -71,7 +81,9 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
"(the felis-api/operator Deployments run it and it is passed through as FELIS_IMAGE, e.g. --felis-image registry.felis.svc:5000/felis:v1)") "(the felis-api/operator Deployments run it and it is passed through as FELIS_IMAGE, e.g. --felis-image registry.felis.svc:5000/felis:v1)")
return 2 return 2
} }
for _, cidr := range append(append([]string{}, velocityCIDRs...), packageCIDRs...) { allCIDRs := append(append([]string{}, velocityCIDRs...), packageCIDRs...)
allCIDRs = append(append(allCIDRs, serverDenyCIDRs...), serverAllowCIDRs...)
for _, cidr := range allCIDRs {
if _, _, err := net.ParseCIDR(cidr); err != nil { if _, _, err := net.ParseCIDR(cidr); err != nil {
fmt.Fprintf(stderr, "felis manifests: invalid CIDR %q: %v\n", cidr, err) fmt.Fprintf(stderr, "felis manifests: invalid CIDR %q: %v\n", cidr, err)
return 2 return 2
@@ -81,39 +93,56 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
fmt.Fprintf(stderr, "felis manifests: --panel-node-port must be in Kubernetes NodePort range 30000-32767 (got %d)\n", *panelNodePort) fmt.Fprintf(stderr, "felis manifests: --panel-node-port must be in Kubernetes NodePort range 30000-32767 (got %d)\n", *panelNodePort)
return 2 return 2
} }
// The node pin exists only for the reaper's hostPath: naming a node without the
// worlds root would be silently dropped (no CronJob renders), so fail loud like
// the storage-trio check below.
if *reaperNode != "" && *worldsHostPath == "" {
fmt.Fprintln(stderr, "felis manifests: --reaper-node requires --worlds-host-path "+
"(it pins the reaper CronJob, which renders only with the retention storage trio)")
return 2
}
// Retention/reaper rendering is opt-in and needs all three storage coordinates // Retention/reaper rendering is opt-in and needs a storage topology together:
// together: where worlds live (to read+archive them), the backup PVC (to write // where worlds live (to read+archive them), a backup PVC (to write archives
// archives into), and the path it is mounted at (which MUST equal felis.toml // into — rendered from --backup-pvc), and the path it is mounted at (which MUST
// [archive] local_path so tarLocal's absolute archive refs resolve). A partial // equal felis.toml [archive] local_path so tarLocal's absolute archive refs
// configuration is almost certainly an operator mistake, so fail loud rather than // resolve). A partial configuration is almost certainly an operator mistake, so
// silently drop retention. Asking for it without the other two is rejected; an // fail loud rather than silently drop retention or render a reaper with nowhere
// empty trio renders the bundle WITHOUT the reaper and says so. // to write. The backup PVC itself defaults to felis-backups (it is what makes a
// default install's backup endpoint work at all); retention additionally needs
// --worlds-host-path.
if *worldsHostPath != "" { if *worldsHostPath != "" {
if *backupPVC == "" || *archiveLocalPath == "" { if *backupPVC == "" || *archiveLocalPath == "" {
fmt.Fprintln(stderr, "felis manifests: --worlds-host-path enables the reaper CronJob and requires "+ fmt.Fprintln(stderr, "felis manifests: --worlds-host-path enables the reaper CronJob and requires "+
"--backup-pvc and --archive-local-path too (--archive-local-path must equal felis.toml [archive] local_path)") "--archive-local-path (must equal felis.toml [archive] local_path) and a non-empty --backup-pvc "+
"(the archive store; default felis-backups)")
return 2 return 2
} }
// The reaper WILL render. Two deployment preconditions this generator cannot // The reaper WILL render. Two deployment facts this generator cannot check
// check would SILENTLY turn retention into a no-op if unmet — surface them as // would silently turn retention into a no-op if unmet — surface them as
// loudly as the fail-closed cases above, so an operator is never left with a // loudly as the fail-closed cases above, so an operator is never left with a
// reaper that reaps an empty directory. (Both are also in the WorldsHostPath // reaper that reaps nothing. (Both are also in the WorldsHostPath flag/field
// flag/field docs, but nobody deploying from stdout reads those.) // docs, but nobody deploying from stdout reads those.)
pin := "the CronJob sets NO nodeSelector: a single-node starter pins it to the worlds implicitly, but on a " +
"multi-node cluster you MUST pass --reaper-node <name> (or add a nodeSelector) for the node holding the " +
"worlds, or the reaper may schedule where the hostPath is empty"
if *reaperNode != "" {
pin = fmt.Sprintf("the CronJob and its worlds-root PV are pinned to node %q via kubernetes.io/hostname — "+
"keep this pointed at the node that actually holds the world volumes", *reaperNode)
}
fmt.Fprintf(stderr, "felis manifests: note: rendering the retention reaper CronJob (worlds hostPath %q). "+ fmt.Fprintf(stderr, "felis manifests: note: rendering the retention reaper CronJob (worlds hostPath %q). "+
"Two preconditions are NOT verified here:\n"+ "These points are NOT verified here:\n"+
" - each world PVC must be visible at %s/<pvc> on the node: a stock local-path-provisioner lays "+ " - the node's world volumes must actually live below %s: the reaper resolves a world as "+
"volumes under PV-name paths (.../pvc-<uuid>_<ns>_<pvc>/), so unless the worlds StorageClass is "+ "%s/<pvc>, then as the stock local-path directory <path>/<pv-name>_<ns>_<pvc-name> (what k3s "+
"arranged to expose <path>/<pvc>, the reaper tars an empty directory;\n"+ "writes under /var/lib/rancher/k3s/storage). Any other provisioner needs its volumes exposed as "+
" - the CronJob sets NO nodeSelector: a single-node starter pins it to the worlds implicitly, but "+ "<path>/<pvc>, or each candidate's archive fails and the world is preserved;\n"+
"on a multi-node cluster you MUST add a nodeSelector for the node holding the worlds, or the reaper "+ " - %s.\n", *worldsHostPath, *worldsHostPath, *worldsHostPath, pin)
"may schedule where the hostPath is empty.\n", *worldsHostPath, *worldsHostPath)
} else { } else {
fmt.Fprintln(stderr, "felis manifests: note: retention reaper CronJob not rendered "+ fmt.Fprintln(stderr, "felis manifests: note: retention reaper CronJob not rendered "+
"(pass --worlds-host-path, --backup-pvc and --archive-local-path to enable it)") "(pass --worlds-host-path and --archive-local-path — the archive PVC defaults to felis-backups — to enable it)")
} }
out, err := platform.RenderYAML(platform.Params{ params := platform.Params{
ControlNamespace: *controlNS, ControlNamespace: *controlNS,
MinecraftNamespace: *minecraftNS, MinecraftNamespace: *minecraftNS,
BuildNamespace: *buildNS, BuildNamespace: *buildNS,
@@ -124,10 +153,23 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
RegistryImage: *registryImage, RegistryImage: *registryImage,
BackupPVC: *backupPVC, BackupPVC: *backupPVC,
WorldsHostPath: *worldsHostPath, WorldsHostPath: *worldsHostPath,
ReaperNode: *reaperNode,
ArchiveLocalPath: *archiveLocalPath, ArchiveLocalPath: *archiveLocalPath,
VelocityCIDRs: []string(velocityCIDRs), VelocityCIDRs: []string(velocityCIDRs),
PackageSourceCIDRs: []string(packageCIDRs), PackageSourceCIDRs: []string(packageCIDRs),
})
ServerEgressDenyCIDRs: []string(serverDenyCIDRs),
ServerEgressAllowCIDRs: []string(serverAllowCIDRs),
RegistryStorage: *registryStorage,
UploadsStorage: *uploadsStorage,
BackupStorage: *backupStorage,
}
if err := params.Validate(); err != nil {
fmt.Fprintf(stderr, "felis manifests: %v\n", err)
return 2
}
out, err := platform.RenderYAML(params)
if err != nil { if err != nil {
fmt.Fprintf(stderr, "felis manifests: render: %v\n", err) fmt.Fprintf(stderr, "felis manifests: render: %v\n", err)
return 1 return 1
+95 -8
View File
@@ -77,6 +77,10 @@ func TestManifestsRendersBundle(t *testing.T) {
"10.0.0.5/32", "10.0.0.5/32",
// The felis image flows through to the Deployments. // The felis image flows through to the Deployments.
"registry.felis.svc:5000/felis:v1", "registry.felis.svc:5000/felis:v1",
// Backup works out of the box: the archive PVC renders and the api gets
// the env that wires the backup/restore executors to it.
"name: felis-backups",
"name: FELIS_BACKUP_PVC",
} { } {
if !strings.Contains(text, want) { if !strings.Contains(text, want) {
t.Errorf("rendered bundle missing %q", want) t.Errorf("rendered bundle missing %q", want)
@@ -95,15 +99,19 @@ func TestManifestsRendersBundle(t *testing.T) {
} }
} }
// TestManifestsReaperRequiresTrio proves --worlds-host-path is a fail-loud opt-in: // TestManifestsReaperRequiresStorage proves --worlds-host-path is a fail-loud
// asking for the reaper without the backup PVC and its mount path (which must equal // opt-in: asking for the reaper without a writable archive store (the backup PVC,
// [archive] local_path) is rejected rather than silently dropping retention. // which defaults to felis-backups but can be emptied) and its mount path (which
func TestManifestsReaperRequiresTrio(t *testing.T) { // must equal [archive] local_path) is rejected rather than silently dropping
// retention or deleting worlds it could not archive first.
func TestManifestsReaperRequiresStorage(t *testing.T) {
base := []string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32", "--worlds-host-path", "/var/lib/felis/worlds"} base := []string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32", "--worlds-host-path", "/var/lib/felis/worlds"}
for _, extra := range [][]string{ for _, extra := range [][]string{
{}, // neither backup-pvc nor archive-local-path {}, // missing archive-local-path (backup-pvc defaults)
{"--backup-pvc", "felis-backups"}, // missing archive-local-path {"--backup-pvc", "other"}, // still missing archive-local-path
{"--archive-local-path", "/backups"}, // missing backup-pvc // A reaper with no archive store would have nowhere to write the archive
// it must verify before deleting a world; emptying the PVC is rejected.
{"--archive-local-path", "/backups", "--backup-pvc="},
} { } {
var out, errBuf bytes.Buffer var out, errBuf bytes.Buffer
code := run(append(append([]string{}, base...), extra...), &out, &errBuf) code := run(append(append([]string{}, base...), extra...), &out, &errBuf)
@@ -119,6 +127,23 @@ func TestManifestsReaperRequiresTrio(t *testing.T) {
} }
} }
// TestManifestsBackupPVCOptOut proves --backup-pvc= renders a bundle with no
// archive store at all: no PVC and no FELIS_BACKUP_PVC env, so backup/restore
// answer 503 instead of pointing Jobs at a claim nobody provisions.
func TestManifestsBackupPVCOptOut(t *testing.T) {
var out, errBuf bytes.Buffer
code := run([]string{"manifests", "--felis-image", "reg/felis:test",
"--velocity-cidr", "10.0.0.5/32", "--backup-pvc="}, &out, &errBuf)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, errBuf.String())
}
for _, absent := range []string{"felis-backups", "FELIS_BACKUP_PVC"} {
if strings.Contains(out.String(), absent) {
t.Errorf("--backup-pvc= bundle must not contain %q", absent)
}
}
}
// TestManifestsRendersReaper proves the happy path with the full retention trio: // TestManifestsRendersReaper proves the happy path with the full retention trio:
// a batch/v1 CronJob is emitted, named felis-reaper, mounting the backup PVC at the // a batch/v1 CronJob is emitted, named felis-reaper, mounting the backup PVC at the
// supplied archive path. // supplied archive path.
@@ -150,9 +175,71 @@ func TestManifestsRendersReaper(t *testing.T) {
// this generator cannot verify (else a misarranged hostPath silently no-ops // this generator cannot verify (else a misarranged hostPath silently no-ops
// retention): the <path>/<pvc> arrangement-dependency and the multi-node // retention): the <path>/<pvc> arrangement-dependency and the multi-node
// nodeSelector hazard. // nodeSelector hazard.
for _, want := range []string{"local-path-provisioner", "nodeSelector"} { for _, want := range []string{"local-path", "nodeSelector"} {
if !strings.Contains(errBuf.String(), want) { if !strings.Contains(errBuf.String(), want) {
t.Errorf("reaper render must warn operators about %q on stderr, got %q", want, errBuf.String()) t.Errorf("reaper render must warn operators about %q on stderr, got %q", want, errBuf.String())
} }
} }
} }
// TestManifestsReaperNodePin: --reaper-node pins the rendered CronJob's pod via
// kubernetes.io/hostname and replaces the "no nodeSelector" hazard note with the
// pin confirmation; using it without the worlds root is a fail-loud 2.
func TestManifestsReaperNodePin(t *testing.T) {
var out, errBuf bytes.Buffer
code := run([]string{
"manifests",
"--felis-image", "registry.felis.svc:5000/felis:v1",
"--velocity-cidr", "10.0.0.5/32",
"--worlds-host-path", "/var/lib/felis/worlds",
"--archive-local-path", "/backups",
"--reaper-node", "node-a",
}, &out, &errBuf)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, errBuf.String())
}
for _, want := range []string{
"kubernetes.io/hostname: node-a",
} {
if !strings.Contains(out.String(), want) {
t.Errorf("pinned render missing %q", want)
}
}
if !strings.Contains(errBuf.String(), "node-a") {
t.Errorf("stderr must confirm the pin, got %q", errBuf.String())
}
var out2, err2 bytes.Buffer
if code := run([]string{
"manifests",
"--felis-image", "registry.felis.svc:5000/felis:v1",
"--velocity-cidr", "10.0.0.5/32",
"--reaper-node", "node-a",
}, &out2, &err2); code != 2 {
t.Errorf("--reaper-node without --worlds-host-path: exit = %d, want 2", code)
}
}
// TestManifestsStorageSizes proves the PVC size flags reach the rendered claims
// and a size the API server would reject fails before anything is applied.
func TestManifestsStorageSizes(t *testing.T) {
var out, errBuf bytes.Buffer
code := run([]string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32",
"--registry-storage", "40Gi", "--uploads-storage", "8Gi"}, &out, &errBuf)
if code != 0 {
t.Fatalf("exit code = %d, stderr = %s", code, errBuf.String())
}
for _, want := range []string{"storage: 40Gi", "storage: 8Gi"} {
if !strings.Contains(out.String(), want) {
t.Errorf("bundle lacks %q", want)
}
}
out.Reset()
errBuf.Reset()
code = run([]string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32",
"--registry-storage", "lots"}, &out, &errBuf)
if code == 0 || !strings.Contains(errBuf.String(), "registry storage") {
t.Errorf("--registry-storage lots: exit %d, stderr %q; want a refusal naming the flag", code, errBuf.String())
}
}
+51
View File
@@ -7,15 +7,24 @@ import (
"io" "io"
"felis.lolicon.best/internal/config" "felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/dbbackup"
"felis.lolicon.best/internal/store" "felis.lolicon.best/internal/store"
) )
// cmdMigrate implements `felis migrate up`: load config, open the database, and // cmdMigrate implements `felis migrate up`: load config, open the database, and
// apply every pending embedded migration under the advisory lock (spec §6). // apply every pending embedded migration under the advisory lock (spec §6).
//
// Migrations only roll forward, and some drop data (0017_drop_password), so a
// database that already holds a schema and has migrations pending is bundled
// first (internal/dbbackup, label pre-migrate). A failed snapshot stops the
// upgrade; -no-backup is the explicit way past it, e.g. for an external
// database whose server is newer than the host's pg_dump.
func cmdMigrate(args []string, stdout, stderr io.Writer) int { func cmdMigrate(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("migrate", flag.ContinueOnError) fs := flag.NewFlagSet("migrate", flag.ContinueOnError)
fs.SetOutput(stderr) fs.SetOutput(stderr)
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml") cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
backupDir := fs.String("backup-dir", dbbackup.DefaultDir, "where the pre-migration snapshot goes")
noBackup := fs.Bool("no-backup", false, "apply pending migrations without snapshotting the database first")
// The "up" verb precedes any flags (felis migrate up -config path). Go's // The "up" verb precedes any flags (felis migrate up -config path). Go's
// flag.Parse stops at the first non-flag token and would never see a flag // flag.Parse stops at the first non-flag token and would never see a flag
// placed after "up", silently falling back to the default -config. Pull the // placed after "up", silently falling back to the default -config. Pull the
@@ -48,6 +57,18 @@ func cmdMigrate(args []string, stdout, stderr io.Writer) int {
return 1 return 1
} }
if !*noBackup {
path, err := preMigrateBackup(ctx, drv, migrations, cfg.Database.URL, *backupDir, stderr)
if err != nil {
fmt.Fprintf(stderr, "felis migrate: pre-migration backup failed, nothing applied: %v\n", err)
fmt.Fprintln(stderr, " fix the backup, or re-run with -no-backup to migrate without one")
return 1
}
if path != "" {
fmt.Fprintf(stdout, "felis migrate: database snapshot %s\n", path)
}
}
applied, err := store.Up(ctx, drv, migrations) applied, err := store.Up(ctx, drv, migrations)
if err != nil { if err != nil {
fmt.Fprintf(stderr, "felis migrate: %v\n", err) fmt.Fprintf(stderr, "felis migrate: %v\n", err)
@@ -60,3 +81,33 @@ func cmdMigrate(args []string, stdout, stderr io.Writer) int {
} }
return 0 return 0
} }
// preMigrateBackup bundles the database when it already carries a schema and
// some of migrations are not applied yet, and returns the bundle's path ("" when
// there was nothing to protect: a fresh database, or nothing pending).
func preMigrateBackup(ctx context.Context, drv store.Driver, migrations []store.Migration, dbURL, dir string, log io.Writer) (string, error) {
if err := drv.EnsureVersionTable(ctx); err != nil {
return "", fmt.Errorf("ensure version table: %w", err)
}
done, err := drv.AppliedVersions(ctx)
if err != nil {
return "", fmt.Errorf("read applied versions: %w", err)
}
if len(done) == 0 || !hasPending(done, migrations) {
return "", nil
}
return dbbackup.Backup(ctx, dbbackup.BackupOptions{
DatabaseURL: dbURL, Dir: dir, Label: dbbackup.LabelPreMigrate,
Keep: defaultKeep[dbbackup.LabelPreMigrate], StateDir: dbbackup.DefaultStateDir,
Version: resolvedVersion(), Log: log, Record: true,
})
}
func hasPending(done map[int]struct{}, migrations []store.Migration) bool {
for _, m := range migrations {
if _, ok := done[m.Version]; !ok {
return true
}
}
return false
}
+125
View File
@@ -0,0 +1,125 @@
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"os"
"os/signal"
"strings"
"syscall"
"time"
"felis.lolicon.best/internal/build"
"felis.lolicon.best/internal/imagepush"
"felis.lolicon.best/internal/registrygate"
)
// defaultBuildToolsStatus is where mirror-build-tools records its last run; the
// watchdog reads it to tell a vulnerability DB that stopped refreshing.
const defaultBuildToolsStatus = "/var/lib/felis/build-tools/status.json"
// cmdMirrorBuildTools copies the build lane's tools (build.Tools: the kaniko and
// trivy images, Trivy's vulnerability and Java DBs) from upstream into the
// platform registry, where build Jobs pull them. deploy/bootstrap.sh runs it at
// install and from felis-build-tools.timer twice a day, which is what keeps the
// DBs fresh; a root shell can run it the same way to refresh now.
//
// It writes as the platform principal through the node's loopback hostPort, the
// same way the installer pushes, reading the token from the environment or from
// /etc/felis/secrets.env.
func cmdMirrorBuildTools(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("mirror-build-tools", flag.ContinueOnError)
fs.SetOutput(stderr)
endpoint := fs.String("endpoint", "127.0.0.1:5000", "host[:port] of the registry to write to (plain HTTP)")
only := fs.String("only", "", "comma-separated tool names to copy (default: all of "+toolNames()+")")
status := fs.String("status", defaultBuildToolsStatus, `file to record the run in ("" records nothing)`)
secrets := fs.String("secrets-env", "/etc/felis/secrets.env", "installer secrets file holding REGISTRY_PLATFORM_TOKEN, read when FELIS_REGISTRY_PASSWORD is unset")
platformFlag := fs.String("platform", "", "os/arch of the images to copy (default: this machine's)")
if err := fs.Parse(args); err != nil {
if errors.Is(err, flag.ErrHelp) {
return 0
}
return 2
}
tools, err := selectTools(*only)
if err != nil {
fmt.Fprintf(stderr, "felis mirror-build-tools: %v\n", err)
return 2
}
if err := loadEnvFile(*secrets); err != nil {
fmt.Fprintf(stderr, "felis mirror-build-tools: read %s: %v\n", *secrets, err)
return 1
}
user, pass := os.Getenv("FELIS_REGISTRY_USERNAME"), os.Getenv("FELIS_REGISTRY_PASSWORD")
if pass == "" {
user, pass = registrygate.PrincipalPlatform, os.Getenv("REGISTRY_PLATFORM_TOKEN")
}
if pass == "" {
fmt.Fprintln(stderr, "felis mirror-build-tools: no registry credential: set FELIS_REGISTRY_PASSWORD or run as root on the node (REGISTRY_PLATFORM_TOKEN in /etc/felis/secrets.env)")
return 2
}
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
p := &imagepush.Pusher{Scheme: "http", Username: user, Password: pass, Log: stdout}
src := &imagepush.Source{Platform: *platformFlag}
started := time.Now()
var failed []string
for _, t := range tools {
dst := strings.TrimSuffix(*endpoint, "/") + "/" + t.Mirror
if _, err := p.Mirror(ctx, src, t.Source, dst); err != nil {
fmt.Fprintf(stderr, "felis mirror-build-tools: %s: %v\n", t.Name, err)
failed = append(failed, t.Name+": "+err.Error())
}
}
if *status != "" {
st, err := imagepush.ReadMirrorStatus(*status)
if err != nil || st == nil {
st = &imagepush.MirrorStatus{}
}
st.LastAttempt = started
st.LastError = strings.Join(failed, "; ")
if len(failed) == 0 {
st.LastSuccess = started
}
if err := imagepush.WriteMirrorStatus(*status, *st); err != nil {
fmt.Fprintf(stderr, "felis mirror-build-tools: record %s: %v\n", *status, err)
}
}
if len(failed) > 0 {
return 1
}
return 0
}
func toolNames() string {
var names []string
for _, t := range build.Tools {
names = append(names, t.Name)
}
return strings.Join(names, ",")
}
func selectTools(only string) ([]build.Tool, error) {
if only == "" {
return build.Tools, nil
}
var out []build.Tool
for _, name := range strings.Split(only, ",") {
name = strings.TrimSpace(name)
found := false
for _, t := range build.Tools {
if t.Name == name {
out = append(out, t)
found = true
}
}
if !found {
return nil, fmt.Errorf("unknown tool %q (known: %s)", name, toolNames())
}
}
return out, nil
}
+556
View File
@@ -0,0 +1,556 @@
package main
import (
"bufio"
"context"
"errors"
"flag"
"fmt"
"io"
"os"
"path/filepath"
"strings"
"time"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/dbbackup"
"felis.lolicon.best/internal/offsite"
"felis.lolicon.best/internal/platform"
"felis.lolicon.best/internal/store"
appsv1 "k8s.io/api/apps/v1"
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/client"
)
const offsiteUsage = `usage:
felis offsite sync [-config path] [-archive-dir dir] [-db-dir dir] [-status-file path]
felis offsite status [-config path] [-status-file path]
felis offsite list [-config path]
felis offsite fetch-db [-config path | -endpoint url -bucket name [-region r] [-prefix p]]
[-dir dir] latest|<bundle>
felis offsite fetch-worlds [-config path] [-archive-dir dir]
felis offsite keygen
Every verb but keygen reads the bucket credentials and the encryption key from
the variables [offsite] names (default FELIS_OFFSITE_ACCESS_KEY,
FELIS_OFFSITE_SECRET_KEY, FELIS_OFFSITE_KEY), taking any that are unset from
-env-file (default /etc/felis/offsite.env).
`
// defaultOffsiteEnvFile is where bootstrap keeps the [offsite] secrets; the
// felis-offsite.service unit loads it as its EnvironmentFile.
const defaultOffsiteEnvFile = "/etc/felis/offsite.env"
// cmdOffsite implements `felis offsite`: the off-site copy of the world
// archives and the database bundles (internal/offsite). felis-offsite.timer
// runs `sync` hourly on the host; the fetch verbs are the way back after the
// node is lost (docs/troubleshooting.md §16).
func cmdOffsite(args []string, stdout, stderr io.Writer) int {
if len(args) == 0 {
fmt.Fprint(stderr, offsiteUsage)
return 2
}
verb, rest := args[0], args[1:]
fs := flag.NewFlagSet("offsite "+verb, flag.ContinueOnError)
fs.SetOutput(stderr)
fs.Usage = func() { fmt.Fprint(stderr, offsiteUsage) }
switch verb {
case "sync":
return offsiteSync(fs, rest, stdout, stderr)
case "status":
return offsiteStatus(fs, rest, stdout, stderr)
case "list":
return offsiteList(fs, rest, stdout, stderr)
case "fetch-db":
return offsiteFetchDB(fs, rest, stdout, stderr)
case "fetch-worlds":
return offsiteFetchWorlds(fs, rest, stdout, stderr)
case "keygen":
k, err := offsite.NewKey()
if err != nil {
fmt.Fprintf(stderr, "felis offsite keygen: %v\n", err)
return 1
}
fmt.Fprintln(stdout, k)
return 0
case "-h", "--help", "help":
fmt.Fprint(stdout, offsiteUsage)
return 0
}
fmt.Fprintf(stderr, "felis offsite: unknown verb %q\n%s", verb, offsiteUsage)
return 2
}
// offsiteEnv is the resolved [offsite] binding: the bucket and the key.
type offsiteEnv struct {
cfg config.OffsiteConfig
bucket *offsite.S3
key []byte
}
// loadEnvFile sets each KEY=VALUE of path that is not already in the
// environment, so a root shell runs a command the same way its unit does.
// A missing file is not an error.
func loadEnvFile(path string) error {
if path == "" {
return nil
}
f, err := os.Open(path)
if errors.Is(err, os.ErrNotExist) {
return nil
}
if err != nil {
return err
}
defer f.Close()
sc := bufio.NewScanner(f)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
k, v, ok := strings.Cut(line, "=")
if !ok {
continue
}
k = strings.TrimSpace(strings.TrimPrefix(k, "export "))
v = strings.TrimSpace(v)
if len(v) >= 2 && (v[0] == '"' || v[0] == '\'') && v[len(v)-1] == v[0] {
v = v[1 : len(v)-1]
}
if os.Getenv(k) == "" {
os.Setenv(k, v)
}
}
return sc.Err()
}
// resolveOffsite builds the bucket client and parses the key for c.
func resolveOffsite(c config.OffsiteConfig) (*offsiteEnv, error) {
if !c.Enabled() {
return nil, errors.New("no [offsite] bucket is configured (docs/troubleshooting.md §16, \"Keep a copy somewhere else\")")
}
need := func(ref, what string) (string, error) {
v := os.Getenv(ref)
if v == "" {
return "", fmt.Errorf("%s: environment variable %s is empty (set it, or put it in %s)", what, ref, defaultOffsiteEnvFile)
}
return v, nil
}
ak, err := need(c.AccessKeyRef, "access key")
if err != nil {
return nil, err
}
sk, err := need(c.SecretKeyRef, "secret key")
if err != nil {
return nil, err
}
rawKey, err := need(c.KeyRef, "encryption key")
if err != nil {
return nil, err
}
key, err := offsite.ParseKey(rawKey)
if err != nil {
return nil, err
}
b, err := offsite.NewS3(offsite.S3Config{
Endpoint: c.Endpoint, Region: c.Region, Bucket: c.Bucket, Prefix: c.Prefix,
AccessKey: ak, SecretKey: sk,
})
if err != nil {
return nil, err
}
return &offsiteEnv{cfg: c, bucket: b, key: key}, nil
}
// loadOffsite loads felis.toml and the env file and resolves [offsite].
func loadOffsite(cfgPath, envFile string) (*config.Config, *offsiteEnv, error) {
if err := loadEnvFile(envFile); err != nil {
return nil, nil, fmt.Errorf("read %s: %w", envFile, err)
}
cfg, err := config.Load(cfgPath)
if err != nil {
return nil, nil, err
}
env, err := resolveOffsite(cfg.Offsite)
if err != nil {
return cfg, nil, err
}
return cfg, env, nil
}
func offsiteSync(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (the host copy, which reaches PostgreSQL on 127.0.0.1)")
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
archiveDir := fs.String("archive-dir", "", "host directory of the world archive volume (default: resolved from the backup PVC through the cluster)")
backupPVC := fs.String("backup-pvc", "felis-backups", `the world archive PVC, in the [k8s] namespace ("" when backups are off)`)
dbDir := fs.String("db-dir", dbbackup.DefaultDir, `database bundle directory ("" copies no bundles)`)
statusFile := fs.String("status-file", offsite.DefaultStatusFile, "where the result of this run is recorded for the watchdog and `status`")
if err := fs.Parse(args); err != nil {
return 2
}
cfg, env, err := loadOffsite(*cfgPath, *envFile)
if err != nil {
fmt.Fprintf(stderr, "felis offsite sync: %v\n", err)
return 1
}
st := offsite.Status{
LastAttempt: time.Now().UTC(), Endpoint: env.cfg.Endpoint, Bucket: env.cfg.Bucket,
Prefix: env.cfg.Prefix, KeyID: offsite.KeyID(env.key),
}
if prev, _ := offsite.ReadStatus(*statusFile); prev != nil {
st.LastSuccess = prev.LastSuccess
}
res, err := runOffsiteSync(cfg, env, *archiveDir, *backupPVC, *dbDir, stderr)
st.Result = res
if err != nil {
st.LastError = err.Error()
} else {
st.LastSuccess = st.LastAttempt
}
if werr := offsite.WriteStatus(*statusFile, st); werr != nil {
fmt.Fprintf(stderr, "felis offsite sync: record status: %v\n", werr)
}
fmt.Fprintf(stdout, "felis offsite sync: worlds copied=%d pending=%d missing=%d expired=%d; bundles copied=%d pruned=%d; bucket holds %d worlds (%s) and %d bundles\n",
res.WorldsUploaded, res.WorldsPending, len(res.WorldsMissing), res.WorldsExpired,
res.DBUploaded, res.DBPruned, res.RemoteWorlds, offsite.HumanBytes(res.RemoteBytes), res.RemoteDB)
for _, m := range res.WorldsMissing {
fmt.Fprintf(stderr, "felis offsite sync: recorded archive not on the volume, nothing to copy: %s\n", m)
}
if err != nil {
fmt.Fprintf(stderr, "felis offsite sync: %v\n", err)
return 1
}
return 0
}
func runOffsiteSync(cfg *config.Config, env *offsiteEnv, archiveDir, backupPVC, dbDir string, log io.Writer) (offsite.Result, error) {
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Minute)
defer cancel()
checkCtx, checkCancel := context.WithTimeout(ctx, 30*time.Second)
err := env.bucket.Check(checkCtx)
checkCancel()
if err != nil {
return offsite.Result{}, err
}
if archiveDir == "" && backupPVC != "" {
dir, err := resolveArchiveDir(ctx, cfg.K8s.Namespace, backupPVC, false, log)
if err != nil {
return offsite.Result{}, err
}
archiveDir = dir
}
drv, err := store.Open(ctx, cfg.Database.URL)
if err != nil {
return offsite.Result{}, fmt.Errorf("open database: %w", err)
}
defer drv.Close()
s := &offsite.Syncer{
Bucket: env.bucket, Catalog: offsite.PGCatalog{DB: drv.DB()}, Key: env.key,
ArchiveDir: archiveDir, DBDir: dbDir, DBKeep: env.cfg.DBKeep, Log: log,
}
return s.Run(ctx)
}
// resolveArchiveDir finds the host directory behind the world archive PVC: a
// local-path volume is a directory on this node. A PVC still waiting for its
// first consumer holds nothing yet: without bind that is "" (no archives),
// with bind it is bound first, for fetch-worlds to write into.
func resolveArchiveDir(ctx context.Context, ns, pvcName string, bind bool, log io.Writer) (string, error) {
if ns == "" {
ns = platform.DefaultMinecraftNamespace
}
cl, err := buildSystemServerClient()
if err != nil {
return "", fmt.Errorf("reach the cluster to find the archive volume (or pass -archive-dir): %w", err)
}
var pvc corev1.PersistentVolumeClaim
if err := cl.Get(ctx, types.NamespacedName{Namespace: ns, Name: pvcName}, &pvc); err != nil {
return "", fmt.Errorf("archive volume %s/%s: %w", ns, pvcName, err)
}
if pvc.Spec.VolumeName == "" {
if !bind {
fmt.Fprintf(log, "felis offsite: archive volume %s/%s is not bound yet; no world has been archived\n", ns, pvcName)
return "", nil
}
if err := bindVolume(ctx, cl, ns, pvcName, log); err != nil {
return "", err
}
if err := cl.Get(ctx, types.NamespacedName{Namespace: ns, Name: pvcName}, &pvc); err != nil {
return "", err
}
}
var pv corev1.PersistentVolume
if err := cl.Get(ctx, types.NamespacedName{Name: pvc.Spec.VolumeName}, &pv); err != nil {
return "", fmt.Errorf("archive volume %s: %w", pvc.Spec.VolumeName, err)
}
var dir string
switch {
case pv.Spec.Local != nil:
dir = pv.Spec.Local.Path
case pv.Spec.HostPath != nil:
dir = pv.Spec.HostPath.Path
default:
return "", fmt.Errorf("archive volume %s is not a directory on a node (local or hostPath); pass -archive-dir with where it is mounted on this host", pv.Name)
}
if fi, err := os.Stat(dir); err != nil || !fi.IsDir() {
return "", fmt.Errorf("archive volume %s is %s on its node, which is not a directory here; run this on the node that holds it, or pass -archive-dir", pv.Name, dir)
}
return dir, nil
}
// bindVolume runs a pod that mounts the PVC and exits, which is what makes a
// WaitForFirstConsumer volume (k3s local-path) get provisioned. The pod uses
// the control plane's own image, which every install already has.
func bindVolume(ctx context.Context, cl client.Client, ns, pvcName string, log io.Writer) error {
var api appsv1.Deployment
if err := cl.Get(ctx, types.NamespacedName{Namespace: platform.DefaultControlNamespace, Name: "felis-api"}, &api); err != nil {
return fmt.Errorf("find the felis image to bind the archive volume with: %w", err)
}
if len(api.Spec.Template.Spec.Containers) == 0 {
return errors.New("felis-api has no container to take the image from")
}
image := api.Spec.Template.Spec.Containers[0].Image
pod := platform.VolumeBinderPod(ns, pvcName, image)
if err := cl.Create(ctx, pod); err != nil {
return fmt.Errorf("start a pod to bind the archive volume: %w", err)
}
fmt.Fprintf(log, "felis offsite: binding the archive volume %s/%s (pod %s)\n", ns, pvcName, pod.Name)
defer func() {
_ = cl.Delete(context.Background(), pod, client.PropagationPolicy(metav1.DeletePropagationBackground))
}()
deadline := time.Now().Add(3 * time.Minute)
for time.Now().Before(deadline) {
var pvc corev1.PersistentVolumeClaim
if err := cl.Get(ctx, types.NamespacedName{Namespace: ns, Name: pvcName}, &pvc); err == nil && pvc.Spec.VolumeName != "" && pvc.Status.Phase == corev1.ClaimBound {
return nil
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(2 * time.Second):
}
}
return fmt.Errorf("the archive volume %s/%s did not bind within 3 minutes; see kubectl -n %s describe pod %s", ns, pvcName, ns, pod.Name)
}
func offsiteStatus(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
statusFile := fs.String("status-file", offsite.DefaultStatusFile, "the record `sync` writes")
if err := fs.Parse(args); err != nil {
return 2
}
cfg, err := config.Load(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis offsite status: %v\n", err)
return 1
}
if !cfg.Offsite.Enabled() {
fmt.Fprintln(stdout, "off-site copy: not configured. World archives and database bundles exist on this machine only.")
fmt.Fprintln(stdout, "See docs/troubleshooting.md §16, \"Keep a copy somewhere else\".")
return 1
}
o := cfg.Offsite
fmt.Fprintf(stdout, "bucket: %s at %s", o.Bucket, o.Endpoint)
if o.Prefix != "" {
fmt.Fprintf(stdout, ", prefix %s", o.Prefix)
}
fmt.Fprintln(stdout)
st, err := offsite.ReadStatus(*statusFile)
if err != nil {
fmt.Fprintf(stderr, "felis offsite status: %v\n", err)
return 1
}
if st == nil {
fmt.Fprintln(stdout, "last sync: never (sudo systemctl start felis-offsite.service)")
return 1
}
now := time.Now()
fmt.Fprintf(stdout, "key id: %s\n", st.KeyID)
fmt.Fprintf(stdout, "last attempt: %s (%s ago)\n", st.LastAttempt.Local().Format(time.DateTime), dbbackup.Age(now.Sub(st.LastAttempt)))
if st.LastSuccess.IsZero() {
fmt.Fprintln(stdout, "last success: never")
} else {
fmt.Fprintf(stdout, "last success: %s (%s ago)\n", st.LastSuccess.Local().Format(time.DateTime), dbbackup.Age(now.Sub(st.LastSuccess)))
}
if st.LastError != "" {
fmt.Fprintf(stdout, "last error: %s\n", st.LastError)
}
r := st.Result
fmt.Fprintf(stdout, "bucket holds: %d world archives (%s), %d database bundles, newest %s\n",
r.RemoteWorlds, offsite.HumanBytes(r.RemoteBytes), r.RemoteDB, orNone(r.NewestDB))
fmt.Fprintf(stdout, "waiting: %d world archives not yet copied\n", r.WorldsPending)
for _, m := range r.WorldsMissing {
fmt.Fprintf(stdout, "missing: %s is recorded but not on the volume\n", m)
}
if st.LastSuccess.IsZero() || now.Sub(st.LastSuccess) > offsite.StaleAfter {
fmt.Fprintf(stdout, "\nThe last successful sync is older than %s: journalctl -u felis-offsite -n 50\n", dbbackup.Age(offsite.StaleAfter))
return 1
}
return 0
}
func orNone(s string) string {
if s == "" {
return "none"
}
return s
}
func offsiteList(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
if err := fs.Parse(args); err != nil {
return 2
}
_, env, err := loadOffsite(*cfgPath, *envFile)
if err != nil {
fmt.Fprintf(stderr, "felis offsite list: %v\n", err)
return 1
}
return printOffsiteList(env, stdout, stderr)
}
func printOffsiteList(env *offsiteEnv, stdout, stderr io.Writer) int {
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
bundles, err := offsite.ListDB(ctx, env.bucket)
if err != nil {
fmt.Fprintf(stderr, "felis offsite list: %v\n", err)
return 1
}
worlds, err := env.bucket.List(ctx, "worlds/")
if err != nil {
fmt.Fprintf(stderr, "felis offsite list: %v\n", err)
return 1
}
fmt.Fprintf(stdout, "database bundles (%d, newest first):\n", len(bundles))
for _, b := range bundles {
fmt.Fprintf(stdout, " %s %s\n", b.Key, offsite.HumanBytes(b.Size))
}
var total int64
for _, w := range worlds {
total += w.Size
}
fmt.Fprintf(stdout, "world archives: %d (%s)\n", len(worlds), offsite.HumanBytes(total))
return 0
}
func offsiteFetchDB(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml; on a host with no install yet, give -endpoint and -bucket instead")
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
endpoint := fs.String("endpoint", "", "bucket endpoint, when there is no felis.toml")
bucket := fs.String("bucket", "", "bucket name, when there is no felis.toml")
region := fs.String("region", "", "bucket region, when there is no felis.toml")
prefix := fs.String("prefix", "", "key prefix, when there is no felis.toml")
dir := fs.String("dir", dbbackup.DefaultDir, "directory to write the bundle to")
arg, ok := parseWithArg(fs, args)
if !ok {
return 2
}
if arg == "" {
fmt.Fprint(stderr, offsiteUsage)
return 2
}
if err := loadEnvFile(*envFile); err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-db: read %s: %v\n", *envFile, err)
return 1
}
var oc config.OffsiteConfig
if *bucket != "" {
oc = config.OffsiteConfig{
Endpoint: *endpoint, Bucket: *bucket, Region: *region, Prefix: *prefix,
AccessKeyRef: config.DefaultOffsiteAccessKeyEnv, SecretKeyRef: config.DefaultOffsiteSecretKeyEnv,
KeyRef: config.DefaultOffsiteKeyEnv,
}
} else {
cfg, err := config.Load(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-db: %v (on a host with no install yet, pass -endpoint and -bucket)\n", err)
return 1
}
oc = cfg.Offsite
}
env, err := resolveOffsite(oc)
if err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
return 1
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Minute)
defer cancel()
name := arg
if name == "latest" {
bundles, err := offsite.ListDB(ctx, env.bucket)
if err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
return 1
}
if len(bundles) == 0 {
fmt.Fprintln(stderr, "felis offsite fetch-db: the bucket holds no database bundle")
return 1
}
name = bundles[0].Key
}
if _, _, ok := dbbackup.ParseBundleName(name); !ok {
fmt.Fprintf(stderr, "felis offsite fetch-db: %q is not a bundle name (felis-db-<stamp>-<label>.tar); see `felis offsite list`\n", name)
return 2
}
if err := os.MkdirAll(*dir, 0o700); err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
return 1
}
dst := filepath.Join(*dir, name)
if err := offsite.FetchObject(ctx, env.bucket, env.key, offsite.DBKey(name), dst, 0o600); err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
return 1
}
if _, err := dbbackup.Verify(dst); err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-db: fetched %s but it does not verify: %v\n", dst, err)
return 1
}
fmt.Fprintf(stdout, "felis offsite fetch-db: wrote %s (verified)\n", dst)
return 0
}
func offsiteFetchWorlds(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (the host copy)")
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
archiveDir := fs.String("archive-dir", "", "host directory of the world archive volume (default: resolved from the backup PVC, binding it if needed)")
backupPVC := fs.String("backup-pvc", "felis-backups", "the world archive PVC, in the [k8s] namespace")
if err := fs.Parse(args); err != nil {
return 2
}
cfg, env, err := loadOffsite(*cfgPath, *envFile)
if err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-worlds: %v\n", err)
return 1
}
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
defer cancel()
dir := *archiveDir
if dir == "" {
if dir, err = resolveArchiveDir(ctx, cfg.K8s.Namespace, *backupPVC, true, stderr); err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-worlds: %v\n", err)
return 1
}
}
drv, err := store.Open(ctx, cfg.Database.URL)
if err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-worlds: open database: %v\n", err)
return 1
}
defer drv.Close()
res, err := offsite.FetchWorlds(ctx, env.bucket, offsite.PGCatalog{DB: drv.DB()}, env.key, dir, stderr)
fmt.Fprintf(stdout, "felis offsite fetch-worlds: %d recorded archives, %d fetched into %s, %d with no copy in the bucket\n",
res.Present, len(res.Fetched), dir, len(res.Missing))
for _, m := range res.Missing {
fmt.Fprintf(stdout, " no off-site copy: %s\n", m)
}
if err != nil {
fmt.Fprintf(stderr, "felis offsite fetch-worlds: %v\n", err)
return 1
}
return 0
}
+96
View File
@@ -0,0 +1,96 @@
package main
import (
"bytes"
"os"
"path/filepath"
"strings"
"testing"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/offsite"
)
func TestLoadOffsiteEnvFile(t *testing.T) {
path := filepath.Join(t.TempDir(), "offsite.env")
body := `# written by bootstrap
FELIS_OFFSITE_ACCESS_KEY=AKIA123
export FELIS_OFFSITE_SECRET_KEY="se=cret"
FELIS_OFFSITE_KEY='k'
not a line
`
if err := os.WriteFile(path, []byte(body), 0o600); err != nil {
t.Fatal(err)
}
t.Setenv("FELIS_OFFSITE_ACCESS_KEY", "from-the-shell")
t.Setenv("FELIS_OFFSITE_SECRET_KEY", "")
t.Setenv("FELIS_OFFSITE_KEY", "")
if err := loadEnvFile(path); err != nil {
t.Fatal(err)
}
for k, want := range map[string]string{
"FELIS_OFFSITE_ACCESS_KEY": "from-the-shell", // the environment wins
"FELIS_OFFSITE_SECRET_KEY": "se=cret",
"FELIS_OFFSITE_KEY": "k",
} {
if got := os.Getenv(k); got != want {
t.Errorf("%s = %q, want %q", k, got, want)
}
}
if err := loadEnvFile(filepath.Join(t.TempDir(), "absent")); err != nil {
t.Errorf("a missing env file is not an error: %v", err)
}
}
func TestResolveOffsiteNamesTheMissingVariable(t *testing.T) {
c := config.OffsiteConfig{
Endpoint: "https://s3.example", Bucket: "b",
AccessKeyRef: "T_AK", SecretKeyRef: "T_SK", KeyRef: "T_KEY",
}
t.Setenv("T_AK", "ak")
t.Setenv("T_SK", "sk")
t.Setenv("T_KEY", "")
if _, err := resolveOffsite(c); err == nil || !strings.Contains(err.Error(), "T_KEY") {
t.Fatalf("err = %v, want it to name T_KEY", err)
}
t.Setenv("T_KEY", "not base64 at all")
if _, err := resolveOffsite(c); err == nil {
t.Fatal("a malformed key was accepted")
}
key, _ := offsite.NewKey()
t.Setenv("T_KEY", key)
env, err := resolveOffsite(c)
if err != nil {
t.Fatal(err)
}
if len(env.key) != offsite.KeySize {
t.Fatalf("key is %d bytes", len(env.key))
}
if _, err := resolveOffsite(config.OffsiteConfig{}); err == nil {
t.Fatal("an unconfigured [offsite] resolved")
}
}
func TestOffsiteKeygen(t *testing.T) {
var out, errb bytes.Buffer
if code := cmdOffsite([]string{"keygen"}, &out, &errb); code != 0 {
t.Fatalf("exit %d: %s", code, errb.String())
}
if _, err := offsite.ParseKey(strings.TrimSpace(out.String())); err != nil {
t.Fatalf("keygen printed %q: %v", out.String(), err)
}
}
func TestOffsiteFetchDBRejectsOddNames(t *testing.T) {
key, _ := offsite.NewKey()
t.Setenv("FELIS_OFFSITE_ACCESS_KEY", "ak")
t.Setenv("FELIS_OFFSITE_SECRET_KEY", "sk")
t.Setenv("FELIS_OFFSITE_KEY", key)
var out, errb bytes.Buffer
code := cmdOffsite([]string{"fetch-db", "-env-file", "", "-endpoint", "http://127.0.0.1:1", "-bucket", "b",
"-dir", t.TempDir(), "../../etc/shadow"}, &out, &errb)
if code != 2 || !strings.Contains(errb.String(), "not a bundle name") {
t.Fatalf("exit %d: %s", code, errb.String())
}
}
+58 -2
View File
@@ -1,19 +1,26 @@
package main package main
import ( import (
"context"
"errors"
"flag" "flag"
"fmt" "fmt"
"io" "io"
"log/slog"
"net/http"
"os" "os"
"time"
"felis.lolicon.best/internal/apis/felis/v1alpha1" "felis.lolicon.best/internal/apis/felis/v1alpha1"
felismetrics "felis.lolicon.best/internal/metrics" felismetrics "felis.lolicon.best/internal/metrics"
"felis.lolicon.best/internal/operator" "felis.lolicon.best/internal/operator"
"github.com/go-logr/logr"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
utilruntime "k8s.io/apimachinery/pkg/util/runtime" utilruntime "k8s.io/apimachinery/pkg/util/runtime"
clientgoscheme "k8s.io/client-go/kubernetes/scheme" clientgoscheme "k8s.io/client-go/kubernetes/scheme"
ctrl "sigs.k8s.io/controller-runtime" ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/cache" "sigs.k8s.io/controller-runtime/pkg/cache"
"sigs.k8s.io/controller-runtime/pkg/healthz"
ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics"
metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server" metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server"
) )
@@ -25,6 +32,11 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
fs := flag.NewFlagSet("operator", flag.ContinueOnError) fs := flag.NewFlagSet("operator", flag.ContinueOnError)
fs.SetOutput(stderr) fs.SetOutput(stderr)
metricsAddr := fs.String("metrics-bind-address", ":8080", "address the metric endpoint binds to") metricsAddr := fs.String("metrics-bind-address", ":8080", "address the metric endpoint binds to")
// healthAddr serves the manager's health endpoints (/healthz, /readyz) that the
// Deployment's probes dial. Without it the operator pod would carry no probe at
// all, and a wedged manager would keep its endpoint forever. It must differ from
// metricsAddr: the metrics server owns :8080.
healthAddr := fs.String("health-probe-bind-address", ":8081", "address the health probe endpoint binds to")
// namespace MUST equal the [k8s] namespace felis-api is configured with, and // namespace MUST equal the [k8s] namespace felis-api is configured with, and
// the deployment manifests (felis manifests) render both from one value. It // the deployment manifests (felis manifests) render both from one value. It
// scopes the manager's cache (informers) to a single namespace so the operator // scopes the manager's cache (informers) to a single namespace so the operator
@@ -42,9 +54,16 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
utilruntime.Must(clientgoscheme.AddToScheme(scheme)) utilruntime.Must(clientgoscheme.AddToScheme(scheme))
utilruntime.Must(v1alpha1.AddToScheme(scheme)) utilruntime.Must(v1alpha1.AddToScheme(scheme))
// controller-runtime logs through its own logr sink; without one, its first
// reconcile prints "log.SetLogger(...) was never called" ATTACHED TO A FULL
// GOROUTINE STACK — pure noise, not signal. Route it to slog's default handler
// so its messages appear as ordinary stderr lines.
ctrl.SetLogger(logr.FromSlogHandler(slog.Default().Handler()))
mgr, err := ctrl.NewManager(ctrl.GetConfigOrDie(), ctrl.Options{ mgr, err := ctrl.NewManager(ctrl.GetConfigOrDie(), ctrl.Options{
Scheme: scheme, Scheme: scheme,
Metrics: metricsserver.Options{BindAddress: *metricsAddr}, Metrics: metricsserver.Options{BindAddress: *metricsAddr},
HealthProbeBindAddress: *healthAddr,
// Scope every informer to the single watched namespace. Without this the // Scope every informer to the single watched namespace. Without this the
// cached client (mgr.GetClient) would LIST/WATCH cluster-wide, which a // cached client (mgr.GetClient) would LIST/WATCH cluster-wide, which a
// namespaced Role cannot grant — the operator would fail closed at runtime // namespaced Role cannot grant — the operator would fail closed at runtime
@@ -60,6 +79,23 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
} }
fmt.Fprintf(stderr, "felis operator: watching namespace %q\n", *namespace) fmt.Fprintf(stderr, "felis operator: watching namespace %q\n", *namespace)
// /healthz fails while a reconcile pass has been stuck past its limit, so the
// liveness probe restarts an operator whose workers are wedged (a Pod whose
// process answers but no server starts or stops). /readyz waits for the
// informer caches: until they sync the operator acts on nothing, and one that
// never syncs (lost RBAC, an unreachable API) never reports Available.
// A dependency hiccup fails neither: the caches ride through API blips, and
// each pass is bounded well inside the stuck limit.
watch := &operator.ReconcileWatch{}
if err := mgr.AddHealthzCheck("reconcile", watch.Check); err != nil {
fmt.Fprintf(stderr, "felis operator: register healthz check: %v\n", err)
return 1
}
if err := mgr.AddReadyzCheck("informers", cacheSynced(mgr.GetCache())); err != nil {
fmt.Fprintf(stderr, "felis operator: register readyz check: %v\n", err)
return 1
}
// Publish the named felis_* metrics (spec §23) on the endpoint the manager // Publish the named felis_* metrics (spec §23) on the endpoint the manager
// already serves (metricsAddr). controller-runtime's metrics server exposes // already serves (metricsAddr). controller-runtime's metrics server exposes
// its global Registry, so registering into it is all that is needed for // its global Registry, so registering into it is all that is needed for
@@ -69,6 +105,7 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
fmt.Fprintf(stderr, "felis operator: register metrics: %v\n", err) fmt.Fprintf(stderr, "felis operator: register metrics: %v\n", err)
return 1 return 1
} }
felismetrics.SetBuildInfo("operator", resolvedVersion())
r := &operator.Reconciler{ r := &operator.Reconciler{
Client: mgr.GetClient(), Client: mgr.GetClient(),
@@ -78,6 +115,10 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
// injects into user servers. The Deployment passes it as FELIS_IMAGE (see // injects into user servers. The Deployment passes it as FELIS_IMAGE (see
// platform.OperatorDeployment); absent, that injection is simply skipped. // platform.OperatorDeployment); absent, that injection is simply skipped.
FelisImage: os.Getenv("FELIS_IMAGE"), FelisImage: os.Getenv("FELIS_IMAGE"),
// Uncached: the maintenance-lock check lists Jobs only when a server is
// about to start, which does not justify a namespace-wide Job informer.
Jobs: mgr.GetAPIReader(),
Watch: watch,
} }
if err := r.SetupWithManager(mgr); err != nil { if err := r.SetupWithManager(mgr); err != nil {
fmt.Fprintf(stderr, "felis operator: setup controller: %v\n", err) fmt.Fprintf(stderr, "felis operator: setup controller: %v\n", err)
@@ -99,3 +140,18 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
} }
return 0 return 0
} }
// cacheSynced is a readyz check that passes once every informer the manager
// started has synced. It waits at most a second, well inside the probe timeout.
func cacheSynced(c interface {
WaitForCacheSync(ctx context.Context) bool
}) healthz.Checker {
return func(req *http.Request) error {
ctx, cancel := context.WithTimeout(req.Context(), time.Second)
defer cancel()
if !c.WaitForCacheSync(ctx) {
return errors.New("informer caches not synced")
}
return nil
}
}
+28
View File
@@ -0,0 +1,28 @@
package main
import (
"context"
"net/http/httptest"
"testing"
)
type fakeCache bool
func (f fakeCache) WaitForCacheSync(ctx context.Context) bool {
if !f {
<-ctx.Done()
}
return bool(f)
}
// TestCacheSynced: the operator reports ready only once its informers synced,
// and a check against caches that never sync returns within its own deadline.
func TestCacheSynced(t *testing.T) {
req := httptest.NewRequest("GET", "/readyz", nil)
if err := cacheSynced(fakeCache(true))(req); err != nil {
t.Errorf("synced: %v", err)
}
if err := cacheSynced(fakeCache(false))(req); err == nil {
t.Error("unsynced caches reported ready")
}
}
+126
View File
@@ -0,0 +1,126 @@
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"strings"
"time"
"felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/imagepin"
"felis.lolicon.best/internal/platform"
"k8s.io/apimachinery/pkg/api/meta"
"sigs.k8s.io/controller-runtime/pkg/client"
)
// defaultRegistryURL is the [registry] url every install uses; deploy/bootstrap.sh
// spells the same value as REGISTRY_URL.
const defaultRegistryURL = "registry.felis.svc:5000"
// cmdPinImages pins every user server whose spec.image still names a mutable tag
// in the platform registry to the digest that tag names now (internal/imagepin).
// felis-api pins on create, so this covers the servers created before it did.
//
// deploy/bootstrap.sh runs it before it rebuilds the game images and pushes them
// over the same tags: run after the push, it would pin those servers to the new
// build, which is exactly the silent Minecraft upgrade pinning exists to stop.
// It reaches the registry through the node's loopback hostPort, the same way the
// installer pushes.
func cmdPinImages(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("pin-images", flag.ContinueOnError)
fs.SetOutput(stderr)
namespace := fs.String("namespace", platform.DefaultMinecraftNamespace, "namespace the MinecraftServers live in")
registry := fs.String("registry", defaultRegistryURL, "registry host[:port] the image refs spell")
endpoint := fs.String("endpoint", "", "host[:port] to reach the registry at (default: 127.0.0.1 on the registry's port, its hostPort on this node)")
if err := fs.Parse(args); err != nil {
if errors.Is(err, flag.ErrHelp) {
return 0
}
return 2
}
if *endpoint == "" {
*endpoint = loopbackEndpoint(*registry)
}
cl, err := buildSystemServerClient()
if err != nil {
fmt.Fprintf(stderr, "felis pin-images: %v\n", err)
return 1
}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel()
outcomes, err := pinUserServerImages(ctx, cl, *namespace, imagepin.Resolver{Registry: *registry, Endpoint: *endpoint})
if meta.IsNoMatchError(err) {
fmt.Fprintln(stdout, "felis pin-images: no MinecraftServer CRD yet, so no server to pin")
return 0
}
if err != nil {
fmt.Fprintf(stderr, "felis pin-images: %v\n", err)
return 1
}
if len(outcomes) == 0 {
fmt.Fprintln(stdout, "felis pin-images: every user server already runs a pinned image")
return 0
}
fmt.Fprintln(stdout, "felis pin-images: pinning user servers to the build their tag names now:")
exit := 0
for _, o := range outcomes {
if o.err != nil {
fmt.Fprintf(stdout, " - %s: ERROR %v\n", o.name, o.err)
exit = 1
continue
}
fmt.Fprintf(stdout, " - %s: %s\n", o.name, strings.Join(o.changes, ", "))
}
return exit
}
// loopbackEndpoint is the registry's port on 127.0.0.1: the registry Deployment
// binds it as a hostPort, and containerd's mirror and the installer's pushes use
// the same address.
func loopbackEndpoint(registry string) string {
if i := strings.LastIndex(registry, ":"); i >= 0 {
return "127.0.0.1" + registry[i:]
}
return "127.0.0.1"
}
// pinUserServerImages patches spec.image of every user server whose image the
// resolver covers and is not yet pinned. System servers are left on their tags:
// the installer rebuilds and restarts them on purpose (restart_existing_system_servers).
// A server that is already pinned, or runs an image from elsewhere, produces no
// outcome, so a pinned fleet reports nothing. A running server restarts once as
// the operator rolls its StatefulSet onto the pinned ref, which is the build it
// already runs.
func pinUserServerImages(ctx context.Context, cl client.Client, namespace string, r imagepin.Resolver) ([]systemServerOutcome, error) {
var list v1alpha1.MinecraftServerList
if err := cl.List(ctx, &list, client.InNamespace(namespace)); err != nil {
return nil, fmt.Errorf("list servers: %w", err)
}
var out []systemServerOutcome
for i := range list.Items {
ms := &list.Items[i]
if ms.Labels[v1alpha1.LabelSystemRole] != "" || imagepin.Pinned(ms.Spec.Image) || !r.Covers(ms.Spec.Image) {
continue
}
pinned, err := r.Pin(ctx, ms.Spec.Image)
if errors.Is(err, imagepin.ErrNotFound) {
err = fmt.Errorf("%s is not in the registry, so there is no build to pin it to; left unpinned: %w", ms.Spec.Image, err)
}
if err != nil {
out = append(out, systemServerOutcome{name: ms.Name, err: err})
continue
}
patch := client.MergeFrom(ms.DeepCopy())
ms.Spec.Image = pinned
if err := cl.Patch(ctx, ms, patch); err != nil {
out = append(out, systemServerOutcome{name: ms.Name, err: fmt.Errorf("patch %s: %w", ms.Name, err)})
continue
}
out = append(out, systemServerOutcome{name: ms.Name, available: true, updated: true,
changes: []string{"spec.image pinned to " + pinned}})
}
return out, nil
}
+112
View File
@@ -0,0 +1,112 @@
package main
import (
"context"
"net/http"
"net/http/httptest"
"strings"
"testing"
"felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/imagepin"
"felis.lolicon.best/internal/naming"
"k8s.io/apimachinery/pkg/api/meta"
"k8s.io/apimachinery/pkg/runtime/schema"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/client/fake"
"sigs.k8s.io/controller-runtime/pkg/client/interceptor"
)
const pinTestDigest = "sha256:2222222222222222222222222222222222222222222222222222222222222222"
// TestPinUserServerImages pins exactly the user servers still on a platform tag,
// reports a tag the registry lost as an error without touching that server, and
// has nothing left to do on a second pass.
func TestPinUserServerImages(t *testing.T) {
reg := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/v2/felis/paper/manifests/demo" {
http.NotFound(w, r)
return
}
w.Header().Set("Docker-Content-Digest", pinTestDigest)
}))
defer reg.Close()
res := imagepin.Resolver{Registry: defaultRegistryURL, Endpoint: strings.TrimPrefix(reg.URL, "http://")}
paper := defaultRegistryURL + "/felis/paper:demo"
mk := func(name, image, role string) *v1alpha1.MinecraftServer {
ms := &v1alpha1.MinecraftServer{}
ms.Name, ms.Namespace = name, "minecraft"
ms.Spec.Image = image
if role != "" {
ms.Labels = map[string]string{v1alpha1.LabelSystemRole: role}
}
return ms
}
cl := fake.NewClientBuilder().WithScheme(newSystemServerScheme(t)).WithObjects(
mk("legacy", paper, ""),
mk("pinned", paper+"@sha256:"+strings.Repeat("3", 64), ""),
mk("external", "docker.io/itzg/minecraft-server:java21", ""),
mk("gone", defaultRegistryURL+"/felis/paper:old", ""),
mk(naming.SystemLobbyServer, defaultRegistryURL+"/felis/felis-lobby:demo", naming.SystemLobbyServer),
).Build()
ctx := context.Background()
outcomes, err := pinUserServerImages(ctx, cl, "minecraft", res)
if err != nil {
t.Fatalf("pinUserServerImages: %v", err)
}
byName := map[string]systemServerOutcome{}
for _, o := range outcomes {
byName[o.name] = o
}
if len(outcomes) != 2 || byName["legacy"].err != nil || byName["gone"].err == nil {
t.Fatalf("outcomes = %+v, want legacy pinned and gone reported", outcomes)
}
want := map[string]string{
"legacy": paper + "@" + pinTestDigest,
"pinned": paper + "@sha256:" + strings.Repeat("3", 64),
"external": "docker.io/itzg/minecraft-server:java21",
"gone": defaultRegistryURL + "/felis/paper:old",
naming.SystemLobbyServer: defaultRegistryURL + "/felis/felis-lobby:demo",
}
for name, image := range want {
var ms v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: name}, &ms); err != nil {
t.Fatalf("get %s: %v", name, err)
}
if ms.Spec.Image != image {
t.Errorf("%s image = %q, want %q", name, ms.Spec.Image, image)
}
}
again, err := pinUserServerImages(ctx, cl, "minecraft", res)
if err != nil || len(again) != 1 || again[0].name != "gone" {
t.Fatalf("second pass = %+v, %v; want only the unresolvable server again", again, err)
}
}
// A fresh install has no CRD yet; the command must read that as nothing to pin.
func TestPinUserServerImagesNoCRD(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(newSystemServerScheme(t)).WithInterceptorFuncs(interceptor.Funcs{
List: func(context.Context, client.WithWatch, client.ObjectList, ...client.ListOption) error {
return &meta.NoKindMatchError{GroupKind: schema.GroupKind{Group: "felis.lolicon.best", Kind: "MinecraftServer"}}
},
}).Build()
_, err := pinUserServerImages(context.Background(), cl, "minecraft", imagepin.Resolver{Registry: defaultRegistryURL})
if !meta.IsNoMatchError(err) {
t.Fatalf("err = %v, want a NoMatch error the command can recognise", err)
}
}
func TestLoopbackEndpoint(t *testing.T) {
for in, want := range map[string]string{
"registry.felis.svc:5000": "127.0.0.1:5000",
"registry.example": "127.0.0.1",
} {
if got := loopbackEndpoint(in); got != want {
t.Errorf("loopbackEndpoint(%q) = %q, want %q", in, got, want)
}
}
}
+176 -25
View File
@@ -1,9 +1,13 @@
package main package main
import ( import (
"context"
"database/sql"
"errors"
"flag" "flag"
"fmt" "fmt"
"io" "io"
"os"
"path/filepath" "path/filepath"
"strconv" "strconv"
"strings" "strings"
@@ -12,8 +16,11 @@ import (
"felis.lolicon.best/internal/apis/felis/v1alpha1" "felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/backup" "felis.lolicon.best/internal/backup"
"felis.lolicon.best/internal/config" "felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/mail"
"felis.lolicon.best/internal/platform"
"felis.lolicon.best/internal/reaper" "felis.lolicon.best/internal/reaper"
"felis.lolicon.best/internal/store" "felis.lolicon.best/internal/store"
corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
utilruntime "k8s.io/apimachinery/pkg/util/runtime" utilruntime "k8s.io/apimachinery/pkg/util/runtime"
clientgoscheme "k8s.io/client-go/kubernetes/scheme" clientgoscheme "k8s.io/client-go/kubernetes/scheme"
@@ -30,7 +37,7 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("reaper", flag.ContinueOnError) fs := flag.NewFlagSet("reaper", flag.ContinueOnError)
fs.SetOutput(stderr) fs.SetOutput(stderr)
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml") cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
worldsRoot := fs.String("worlds-root", "/worlds", "mount root under which world PVCs are visible (tarLocal: <root>/<pvc>)") worldsRoot := fs.String("worlds-root", "/worlds", "mount root under which world PVCs are visible (tarLocal: <root>/<pvc>, else the stock local-path <root>/<pv-name>_<ns>_<pvc-name>)")
if err := fs.Parse(args); err != nil { if err := fs.Parse(args); err != nil {
return 2 return 2
} }
@@ -47,21 +54,8 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
return 1 return 1
} }
archiver, err := buildArchiver(cfg, *worldsRoot)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: %v\n", err)
return 1
}
ctx := ctrl.SetupSignalHandler() ctx := ctrl.SetupSignalHandler()
drv, err := store.Open(ctx, cfg.Database.URL)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: open database: %v\n", err)
return 1
}
defer drv.Close()
scheme := runtime.NewScheme() scheme := runtime.NewScheme()
utilruntime.Must(clientgoscheme.AddToScheme(scheme)) utilruntime.Must(clientgoscheme.AddToScheme(scheme))
utilruntime.Must(v1alpha1.AddToScheme(scheme)) utilruntime.Must(v1alpha1.AddToScheme(scheme))
@@ -71,6 +65,19 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
return 1 return 1
} }
archiver, err := buildArchiver(ctx, cfg, *worldsRoot, cl)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: %v\n", err)
return 1
}
drv, err := store.Open(ctx, cfg.Database.URL)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: open database: %v\n", err)
return 1
}
defer drv.Close()
r := &reaper.Reaper{ r := &reaper.Reaper{
Cfg: rcfg, Cfg: rcfg,
Store: reaper.NewPGStore(drv.DB()), Store: reaper.NewPGStore(drv.DB()),
@@ -78,19 +85,107 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
Archiver: archiver, Archiver: archiver,
} }
// Pre-reap warnings go out by email when [smtp] is configured (the same
// relay and password_ref convention felis-api uses); without it the channel
// stays nil and the reaper logs each suppressed warning instead of stamping
// it, so a later SMTP setup still gets to warn. The owner must have a
// VERIFIED address — that flag is what proves the mailbox.
if cfg.SMTP.Host != "" {
passRef := cfg.SMTP.PasswordRef
if passRef == "" {
passRef = platform.SMTPPasswordEnv
}
password := os.Getenv(passRef)
if cfg.SMTP.Username != "" && password == "" {
fmt.Fprintf(stderr, "felis reaper: warning: [smtp] username is set but credentials env %s is empty — warning emails will fail AUTH\n", passRef)
}
db := drv.DB()
r.Warner = &mailWarner{
lookupEmail: func(ctx context.Context, ownerID string) (string, error) {
var email string
switch err := db.QueryRowContext(ctx,
`SELECT email FROM users
WHERE id = $1 AND email_verified = true AND COALESCE(email, '') <> ''`,
ownerID).Scan(&email); {
case errors.Is(err, sql.ErrNoRows):
return "", fmt.Errorf("owner %s has no verified email", ownerID)
case err != nil:
return "", err
}
return email, nil
},
notifier: &mail.SMTP{
Host: cfg.SMTP.Host,
Port: cfg.SMTP.Port,
From: cfg.SMTP.From,
Username: cfg.SMTP.Username,
Password: password,
},
}
} else {
fmt.Fprintln(stderr, "felis reaper: [smtp] not configured — pre-reap warnings are logged and NOT marked sent")
}
sum, err := r.RunOnce(ctx) sum, err := r.RunOnce(ctx)
if err != nil { if err != nil {
fmt.Fprintf(stderr, "felis reaper: %v\n", err) fmt.Fprintf(stderr, "felis reaper: %v\n", err)
return 1 return 1
} }
fmt.Fprintf(stdout, "felis reaper: evaluated=%d reaped=%d warned=%d skipped=%d evicted=%d expired=%d\n", return reportReaperRun(sum, stdout, stderr)
sum.Evaluated, sum.WorldsReaped, sum.Warned, sum.Skipped, sum.EvictedEarly, sum.BackupsExpired) }
return 0
// reportReaperRun prints the run's tally and turns a run that left work undone
// into exit 1, so the Job fails and the watchdog's job-failed check (and the
// FelisWorldJobFailed rule) reach the operator: a world that cannot be archived
// is kept, and without this nobody would learn that it is never reaped.
func reportReaperRun(sum reaper.Summary, stdout, stderr io.Writer) int {
fmt.Fprintf(stdout, "felis reaper: evaluated=%d reaped=%d awaiting_offsite=%d warned=%d skipped=%d store_full=%d evicted=%d expired=%d expire_failed=%d\n",
sum.Evaluated, sum.WorldsReaped, sum.AwaitingOffsite, sum.Warned, sum.Skipped, sum.StoreFull,
sum.EvictedEarly, sum.BackupsExpired, sum.ExpireFailed)
if !sum.Failed() {
return 0
}
fmt.Fprintf(stderr, "felis reaper: %d servers failed (%d kept because the backup store is full) and %d expired backups were not removed; the errors are above, and each is retried next run\n",
sum.Skipped, sum.StoreFull, sum.ExpireFailed)
return 1
}
// mailWarner delivers a pre-reap notice to the owner's verified email — the
// only channel this build can reach. Unowned owners and owners who never proved
// a mailbox yield an error; the reaper retries such notices on its next run and
// never lets them block the reap (red line ⑤).
type mailWarner struct {
lookupEmail func(ctx context.Context, ownerID string) (string, error)
notifier noticeNotifier
}
// noticeNotifier is the slice of mail.SMTP the warner needs (injected in tests).
type noticeNotifier interface {
SendNotice(ctx context.Context, email, subject, body string) error
}
func (w *mailWarner) Warn(ctx context.Context, ownerID, server, remaining string) error {
email, err := w.lookupEmail(ctx, ownerID)
if err != nil {
return fmt.Errorf("resolve owner email: %w", err)
}
subject := fmt.Sprintf("Felis: 服务器 %s 将在 %s 后回收 · server reaped in %s", server, remaining, remaining)
body := fmt.Sprintf(
"Felis 世界回收提醒 / world-reaper notice\r\n"+
"\r\n"+
"服务器 / Server: %s\r\n"+
"距回收 / Time left: %s\r\n"+
"\r\n"+
"闲置的服务器会先自动备份,再释放世界;有人加入游戏即可重置倒计时。\r\n"+
"Idle servers are backed up and then released; any join resets the countdown.\r\n",
server, remaining)
return w.notifier.SendNotice(ctx, email, subject, body)
} }
// reaperConfig derives the reaper's retention windows from felis.toml. The 15d // reaperConfig derives the reaper's retention windows from felis.toml. The 15d
// idle deadline is fixed by §18; only the warning offsets, retention, and the // idle deadline is fixed by §18; only the warning offsets, retention, the
// store soft-cap are configurable (§24). // store soft-cap and the on-demand backup bounds are configurable (§24). The
// backup Job and felis-api read the manual_* bounds through it too.
func reaperConfig(cfg *config.Config) (reaper.Config, error) { func reaperConfig(cfg *config.Config) (reaper.Config, error) {
rc := reaper.DefaultConfig() rc := reaper.DefaultConfig()
if v := cfg.Archive.Retention; v != "" { if v := cfg.Archive.Retention; v != "" {
@@ -118,26 +213,82 @@ func reaperConfig(cfg *config.Config) (reaper.Config, error) {
} }
rc.MaxLocalBytes = b rc.MaxLocalBytes = b
} }
if v := cfg.Archive.ManualRetention; v != "" {
d, err := parseSpanDuration(v)
if err != nil || d <= 0 {
return rc, fmt.Errorf("[archive] manual_retention %q: want a positive span such as 30d", v)
}
rc.ManualRetention = d
}
switch n := cfg.Archive.ManualKeep; {
case n < 0:
return rc, fmt.Errorf("[archive] manual_keep %d: want 1 or more", n)
case n > 0:
rc.ManualKeep = n
}
if v := cfg.Archive.ManualCooldown; v != "" {
d, err := parseSpanDuration(v)
if err != nil || d < 0 {
return rc, fmt.Errorf("[archive] manual_cooldown %q: want a span such as 10m (0s for none)", v)
}
rc.ManualCooldown = d
}
rc.RequireOffsite = cfg.Offsite.Enabled()
return rc, nil return rc, nil
} }
// buildArchiver constructs the WorldArchiver. Only tarLocal is implemented in // buildArchiver constructs the WorldArchiver. Only tarLocal is implemented in
// this build; the resolver maps each world PVC to <worldsRoot>/<pvc>, the mount // this build; the resolver maps each world PVC to its directory under worldsRoot
// convention the reaper Job is deployed with. // (resolveWorldDir).
func buildArchiver(cfg *config.Config, worldsRoot string) (backup.WorldArchiver, error) { func buildArchiver(ctx context.Context, cfg *config.Config, worldsRoot string, cl client.Client) (backup.WorldArchiver, error) {
switch cfg.Archive.Store { switch cfg.Archive.Store {
case "tarLocal": case "tarLocal":
return &backup.TarLocal{ return &backup.TarLocal{
BackupRoot: cfg.Archive.LocalPath, BackupRoot: cfg.Archive.LocalPath,
Resolve: func(pvc string) (string, error) { Resolve: resolveWorldDir(ctx, cl, cfg.K8s.Namespace, worldsRoot),
return filepath.Join(worldsRoot, pvc), nil
},
}, nil }, nil
default: default:
return nil, fmt.Errorf("[archive] store %q is not implemented in this build (only tarLocal)", cfg.Archive.Store) return nil, fmt.Errorf("[archive] store %q is not implemented in this build (only tarLocal)", cfg.Archive.Store)
} }
} }
// resolveWorldDir maps a world PVC to its directory under worldsRoot, supporting
// the two layouts a Felis host actually has:
//
// 1. <root>/<pvc> — the reaper's documented arrangement (worlds exposed by PVC
// name, e.g. via mounting each volume or a crafted storage class).
// 2. <root>/<pv-name>_<namespace>_<pvc-name> — what a stock k3s install gets:
// local-path-provisioner stores every volume under its storage root as that
// exact directory name. Without this arm, retention on a default install could
// only ever fail to find a world (a no-op reaper, or worse an operator
// arranging paths by hand).
//
// The second path is derived EXACTLY from the live PVC's spec.volumeName, never
// from a glob: a leftover directory of an old, deleted PV must never be mistaken
// for the world the PVC currently binds, because the reaper archives the resolved
// directory and then deletes that PVC — archiving stale bytes and deleting the
// real world would be data loss. When neither path exists the first is returned,
// so the archive walk fails loudly against the documented path.
func resolveWorldDir(ctx context.Context, cl client.Client, namespace, worldsRoot string) backup.PVCResolver {
return func(pvc string) (string, error) {
direct := filepath.Join(worldsRoot, pvc)
if _, err := os.Stat(direct); err == nil {
return direct, nil
}
var claim corev1.PersistentVolumeClaim
if err := cl.Get(ctx, client.ObjectKey{Namespace: namespace, Name: pvc}, &claim); err != nil {
return "", fmt.Errorf("resolve world PVC %s: %w", pvc, err)
}
if pv := claim.Spec.VolumeName; pv != "" {
volDir := filepath.Join(worldsRoot, fmt.Sprintf("%s_%s_%s", pv, claim.Namespace, claim.Name))
if _, err := os.Stat(volDir); err == nil {
return volDir, nil
}
}
return direct, nil
}
}
// parseSpanDuration parses the human spans used in felis.toml's [archive] table: // parseSpanDuration parses the human spans used in felis.toml's [archive] table:
// "3mo" (months≈30d), "15d" (days), or any time.ParseDuration unit ("12h"). // "3mo" (months≈30d), "15d" (days), or any time.ParseDuration unit ("12h").
func parseSpanDuration(s string) (time.Duration, error) { func parseSpanDuration(s string) (time.Duration, error) {
+186
View File
@@ -0,0 +1,186 @@
package main
import (
"bytes"
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"time"
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client/fake"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/reaper"
)
// TestReportReaperRunFailsTheJob: a run that could not process a server, or
// could not remove an expired backup, exits 1 so the Job shows as failed.
func TestReportReaperRunFailsTheJob(t *testing.T) {
for _, tc := range []struct {
name string
sum reaper.Summary
want int
}{
{"clean", reaper.Summary{Evaluated: 3, WorldsReaped: 1, AwaitingOffsite: 1}, 0},
{"server failed", reaper.Summary{Evaluated: 3, Skipped: 1}, 1},
{"store full", reaper.Summary{Evaluated: 3, Skipped: 1, StoreFull: 1}, 1},
{"expiry failed", reaper.Summary{Evaluated: 3, ExpireFailed: 2}, 1},
} {
var out, errb bytes.Buffer
if got := reportReaperRun(tc.sum, &out, &errb); got != tc.want {
t.Errorf("%s: exit %d, want %d", tc.name, got, tc.want)
}
if !strings.Contains(out.String(), "skipped=") || !strings.Contains(out.String(), "expire_failed=") {
t.Errorf("%s: summary line = %q", tc.name, out.String())
}
if (tc.want == 1) != (errb.Len() > 0) {
t.Errorf("%s: stderr = %q", tc.name, errb.String())
}
}
}
// TestResolveWorldDir pins the two world layouts the reaper must find, and the
// fail-closed miss. The stock local-path arm is derived from the live PVC's
// volumeName — a name-based guess (glob) could tar a stale deleted PV's bytes and
// then delete the current world, which is why it is read from the API instead.
// TestReaperConfigManualKeys: the on-demand backup keys default to 30 days,
// five per server and a ten-minute cooldown, accept overrides, and refuse
// values that would keep nothing or throttle backwards.
func TestReaperConfigManualKeys(t *testing.T) {
rc, err := reaperConfig(&config.Config{})
if err != nil {
t.Fatal(err)
}
if rc.ManualRetention != 30*reaper.Day || rc.ManualKeep != 5 || rc.ManualCooldown != 10*time.Minute {
t.Fatalf("defaults = %v / %d / %v", rc.ManualRetention, rc.ManualKeep, rc.ManualCooldown)
}
rc, err = reaperConfig(&config.Config{Archive: config.ArchiveConfig{
ManualRetention: "7d", ManualKeep: 2, ManualCooldown: "0s"}})
if err != nil {
t.Fatal(err)
}
if rc.ManualRetention != 7*reaper.Day || rc.ManualKeep != 2 || rc.ManualCooldown != 0 {
t.Fatalf("overrides = %v / %d / %v", rc.ManualRetention, rc.ManualKeep, rc.ManualCooldown)
}
for _, bad := range []config.ArchiveConfig{
{ManualRetention: "0d"},
{ManualRetention: "soon"},
{ManualKeep: -1},
{ManualCooldown: "-5m"},
{ManualCooldown: "often"},
} {
if _, err := reaperConfig(&config.Config{Archive: bad}); err == nil {
t.Errorf("%+v was accepted", bad)
}
}
}
func TestResolveWorldDir(t *testing.T) {
ctx := context.Background()
root := t.TempDir()
// Arrange a world under the documented <root>/<pvc> layout.
named := filepath.Join(root, "world-named-0")
if err := os.MkdirAll(named, 0o750); err != nil {
t.Fatal(err)
}
// Arrange a second world the way k3s local-path stores it.
pvDir := filepath.Join(root, "pvc-11111111-2222-3333-4444-555555555555_minecraft_world-live-0")
if err := os.MkdirAll(pvDir, 0o750); err != nil {
t.Fatal(err)
}
claim := &corev1.PersistentVolumeClaim{
ObjectMeta: metav1.ObjectMeta{Name: "world-live-0", Namespace: "minecraft"},
Spec: corev1.PersistentVolumeClaimSpec{
VolumeName: "pvc-11111111-2222-3333-4444-555555555555",
},
}
cl := fake.NewClientBuilder().WithScheme(haltScheme(t)).WithObjects(claim).Build()
resolve := resolveWorldDir(ctx, cl, "minecraft", root)
t.Run("documented name layout wins", func(t *testing.T) {
got, err := resolve("world-named-0")
if err != nil || got != named {
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, named)
}
})
t.Run("stock local-path layout resolves exactly", func(t *testing.T) {
got, err := resolve("world-live-0")
if err != nil || got != pvDir {
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, pvDir)
}
})
t.Run("neither layout present falls back to the documented path", func(t *testing.T) {
// The claim exists but its directory does not: return the documented path so
// the archive walk fails there, and the reaper preserves the world.
missing := &corev1.PersistentVolumeClaim{
ObjectMeta: metav1.ObjectMeta{Name: "world-gone-0", Namespace: "minecraft"},
Spec: corev1.PersistentVolumeClaimSpec{VolumeName: "pvc-99999999-0000-0000-0000-000000000000"},
}
cl := fake.NewClientBuilder().WithScheme(haltScheme(t)).WithObjects(missing).Build()
got, err := resolveWorldDir(ctx, cl, "minecraft", root)("world-gone-0")
if err != nil || got != filepath.Join(root, "world-gone-0") {
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, filepath.Join(root, "world-gone-0"))
}
})
t.Run("unknown pvc is an error, not a guess", func(t *testing.T) {
_, err := resolve("world-unknown-0")
if err == nil || !strings.Contains(err.Error(), "resolve world PVC world-unknown-0") {
t.Fatalf("err = %v, want a resolve-world-PVC error", err)
}
})
}
// The pre-reap warner resolves the owner's VERIFIED email and hands the notice
// to the mailer. Every failure (no verified address, relay refusal) returns an
// error so the reaper retries on its next run instead of stamping a notice
// nobody received.
func TestMailWarner(t *testing.T) {
lookup := func(email string, err error) func(context.Context, string) (string, error) {
return func(context.Context, string) (string, error) { return email, err }
}
n := &captureNotifier{}
w := &mailWarner{lookupEmail: lookup("[email protected]", nil), notifier: n}
if err := w.Warn(context.Background(), "u1", "survival", "3d"); err != nil {
t.Fatalf("Warn: %v", err)
}
if n.email != "[email protected]" || !strings.Contains(n.subject, "survival") || !strings.Contains(n.subject, "3d") {
t.Fatalf("notice envelope = (%q, %q)", n.email, n.subject)
}
if !strings.Contains(n.body, "survival") || !strings.Contains(n.body, "3d") {
t.Fatalf("body missing server/remaining:\n%s", n.body)
}
w = &mailWarner{lookupEmail: lookup("", errors.New("owner u2 has no verified email")), notifier: n}
if err := w.Warn(context.Background(), "u2", "survival", "3d"); err == nil || !strings.Contains(err.Error(), "verified email") {
t.Fatalf("unverified owner = %v, want the lookup error surfaced", err)
}
w = &mailWarner{lookupEmail: lookup("[email protected]", nil), notifier: &captureNotifier{err: errors.New("relay down")}}
if err := w.Warn(context.Background(), "u1", "survival", "3d"); err == nil || !strings.Contains(err.Error(), "relay down") {
t.Fatalf("relay failure = %v, want it surfaced", err)
}
}
type captureNotifier struct {
email, subject, body string
err error
}
func (n *captureNotifier) SendNotice(_ context.Context, email, subject, body string) error {
if n.err != nil {
return n.err
}
n.email, n.subject, n.body = email, subject, body
return nil
}
+164
View File
@@ -0,0 +1,164 @@
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"log/slog"
"net"
"net/http"
"net/url"
"os"
"os/signal"
"path/filepath"
"strings"
"syscall"
"time"
"felis.lolicon.best/internal/imagepush"
"felis.lolicon.best/internal/registrygate"
)
// cmdRegistryGate is the sidecar entrypoint in the registry pod: it owns the
// registry port (and the loopback hostPort containerd pulls through), lets reads
// through anonymously, and forwards writes to the loopback-only registry:2 only
// for an authenticated principal allowed to write that repository. See
// internal/registrygate for the policy.
//
// Tokens are files under --auth-dir, one per principal (platform, build, prune),
// mounted from the registry-auth Secret. A missing file disables that principal:
// writes fail closed while every pull keeps working, which is the right way round
// for a registry the running workloads depend on.
//
// --maint-listen is the GC sidecar's read-only handshake (registrygate.MaintHandler).
// It has no authentication, so it must name a loopback address; --maint-dir keeps
// an open window across a gate restart.
func cmdRegistryGate(args []string, _, stderr io.Writer) int {
fs := flag.NewFlagSet("registry-gate", flag.ContinueOnError)
fs.SetOutput(stderr)
listen := fs.String("listen", ":5000", "address the gate serves the registry API on")
upstream := fs.String("upstream", "http://127.0.0.1:5001", "the loopback registry the gate forwards to")
authDir := fs.String("auth-dir", "/etc/felis-registry-auth", "directory holding one token file per principal")
maintListen := fs.String("maint-listen", "", "loopback address for the GC sidecar's read-only handshake (empty disables it)")
maintDir := fs.String("maint-dir", "", "directory that keeps an open read-only window across a gate restart")
quiet := fs.Duration("maint-quiet", registrygate.DefaultQuiet, "how long writes must be idle before a read-only window is granted")
dataDir := fs.String("data-dir", "", "the registry's storage root, mounted read-only, for the manifest index (empty disables it)")
if err := fs.Parse(args); err != nil {
return 2
}
if *maintListen != "" && !loopbackAddr(*maintListen) {
fmt.Fprintf(stderr, "felis registry-gate: --maint-listen %q must be a loopback address: the handshake has no authentication\n", *maintListen)
return 2
}
target, err := url.Parse(*upstream)
if err != nil || target.Scheme == "" || target.Host == "" {
fmt.Fprintf(stderr, "felis registry-gate: bad --upstream %q\n", *upstream)
return 2
}
log := slog.New(slog.NewTextHandler(stderr, nil))
tokens := map[string]string{}
for _, p := range registrygate.Principals {
b, err := os.ReadFile(filepath.Join(*authDir, p))
tok := strings.TrimSpace(string(b))
if err != nil || tok == "" {
log.Warn("registry principal disabled: no token", "principal", p, "dir", *authDir)
continue
}
tokens[p] = tok
}
gate := registrygate.New(target, tokens, log)
gate.SetQuiet(*quiet)
gate.DataDir = *dataDir
if *maintDir != "" {
if err := gate.SetMaintenanceState(registrygate.MaintStatePath(*maintDir)); err != nil {
// A corrupt file must not keep the registry from serving pulls.
log.Warn("ignoring the saved read-only window", "err", err)
}
}
srv := &http.Server{
Addr: *listen,
Handler: gate,
ReadHeaderTimeout: 10 * time.Second,
}
var maint *http.Server
if *maintListen != "" {
maint = &http.Server{Addr: *maintListen, Handler: gate.MaintHandler(), ReadHeaderTimeout: 10 * time.Second}
go func() {
if err := maint.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Error("maintenance listener stopped; garbage collection cannot get a read-only window", "err", err)
}
}()
}
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
go func() {
<-ctx.Done()
shutdown, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
_ = srv.Shutdown(shutdown)
if maint != nil {
_ = maint.Shutdown(shutdown)
}
}()
log.Info("registry gate listening", "addr", *listen, "upstream", target.String(), "principals", len(tokens))
if err := srv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
fmt.Fprintf(stderr, "felis registry-gate: %v\n", err)
return 1
}
return 0
}
// loopbackAddr reports whether a host:port listen address binds loopback only.
func loopbackAddr(addr string) bool {
host, _, err := net.SplitHostPort(addr)
if err != nil {
return false
}
if host == "localhost" {
return true
}
ip := net.ParseIP(host)
return ip != nil && ip.IsLoopback()
}
// cmdPushImage is the build Job's publish step. It runs after Kaniko built the
// image into a tarball (--no-push) and Trivy passed that tarball, and it is the
// only container of the build pod that holds the registry credential — the one
// executing the untrusted Dockerfile never sees it.
func cmdPushImage(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("push-image", flag.ContinueOnError)
fs.SetOutput(stderr)
tarPath := fs.String("tar", "", "image tarball Kaniko wrote with --tar-path")
ref := fs.String("ref", "", "host/repository:tag to publish it as")
scheme := fs.String("scheme", "http", "registry scheme: http for the in-cluster registry, https otherwise")
if err := fs.Parse(args); err != nil {
return 2
}
if *tarPath == "" || *ref == "" {
fmt.Fprintln(stderr, "felis push-image: --tar and --ref are required")
return 2
}
if *scheme != "http" && *scheme != "https" {
fmt.Fprintf(stderr, "felis push-image: bad --scheme %q\n", *scheme)
return 2
}
user := os.Getenv("FELIS_REGISTRY_USERNAME")
pass := os.Getenv("FELIS_REGISTRY_PASSWORD")
if user == "" || pass == "" {
fmt.Fprintln(stderr, "felis push-image: FELIS_REGISTRY_USERNAME/FELIS_REGISTRY_PASSWORD are empty — the registry refuses anonymous writes")
return 2
}
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
p := &imagepush.Pusher{Scheme: *scheme, Username: user, Password: pass, Log: stderr}
digest, err := p.Push(ctx, *tarPath, *ref)
if err != nil {
fmt.Fprintf(stderr, "felis push-image: %v\n", err)
return 1
}
fmt.Fprintln(stdout, digest)
return 0
}
+37 -17
View File
@@ -11,7 +11,9 @@ Usage:
felis <command> [flags] felis <command> [flags]
Commands: Commands:
migrate up Apply embedded database migrations under an advisory lock migrate up Apply embedded database migrations under an advisory lock (snapshots the database first)
db Back up, verify, list and restore the control-plane database (backup|restore|verify|list|check)
offsite Copy world archives and database bundles to an off-site bucket, and fetch them back (sync|status|list|fetch-db|fetch-worlds|keygen)
operator Run the MinecraftServer controller-manager operator Run the MinecraftServer controller-manager
api Run the felis-api HTTP server api Run the felis-api HTTP server
nano Run the Felis-nano hasJoined multiplexer (multi-Yggdrasil, no control plane) nano Run the Felis-nano hasJoined multiplexer (multi-Yggdrasil, no control plane)
@@ -19,9 +21,16 @@ Commands:
restore Extract a world archive into a world volume (internal Job entrypoint) restore Extract a world archive into a world volume (internal Job entrypoint)
backup Archive a world into the backup store and record it (internal Job entrypoint) backup Archive a world into the backup store and record it (internal Job entrypoint)
files List/read/write one file in a stopped server's world (internal Job entrypoint) files List/read/write one file in a stopped server's world (internal Job entrypoint)
egress-gate Hold a build pod until its egress NetworkPolicy is enforced (internal Job entrypoint)
fetch-context Fetch and extract a submission's build context (internal Job entrypoint)
push-image Push a scanned image tarball to the registry (internal Job entrypoint)
mirror-build-tools Copy kaniko, trivy and Trivy's DBs into the registry (run by felis-build-tools.timer)
registry-gate Authorize registry writes in front of registry:2 (internal sidecar entrypoint)
manifests Render the control-plane RBAC + NetworkPolicy install bundle as YAML manifests Render the control-plane RBAC + NetworkPolicy install bundle as YAML
apply Create a MinecraftServer CRD (direct K8s write; use -f server.json) apply Create a MinecraftServer CRD (direct K8s write; use -f server.json)
setup Run host bootstrap + first-run setup console (TUI; requires root/sudo) setup Run host bootstrap + first-run setup console (TUI; requires root/sudo)
converge Fill in fields a newer desired spec added to already-installed system servers
watchdog Check the platform once and mail the owners what has gone wrong (run by felis-watchdog.timer)
version Print the build stamp of this binary version Print the build stamp of this binary
update Report which platform components have updates available update Report which platform components have updates available
breakGlass Open the local break-glass emergency console (TUI; requires root/sudo) breakGlass Open the local break-glass emergency console (TUI; requires root/sudo)
@@ -39,22 +48,33 @@ Run "felis <command> -h" for command-specific flags.
// The help aliases are deliberately NOT entries: they print usage rather than run a // The help aliases are deliberately NOT entries: they print usage rather than run a
// subcommand, and listing them would make the table disagree with the command list. // subcommand, and listing them would make the table disagree with the command list.
var commands = map[string]func(args []string, stdout, stderr io.Writer) int{ var commands = map[string]func(args []string, stdout, stderr io.Writer) int{
"migrate": cmdMigrate, "migrate": cmdMigrate,
"operator": cmdOperator, "db": cmdDB,
"api": cmdAPI, "offsite": cmdOffsite,
"nano": cmdNano, "operator": cmdOperator,
"reaper": cmdReaper, "api": cmdAPI,
"restore": cmdRestore, "nano": cmdNano,
"backup": cmdBackup, "reaper": cmdReaper,
"files": cmdFiles, "restore": cmdRestore,
"manifests": cmdManifests, "backup": cmdBackup,
"apply": cmdApply, "files": cmdFiles,
"setup": cmdSetup, "egress-gate": cmdEgressGate,
"breakGlass": cmdBreakGlass, "fetch-context": cmdFetchContext,
"bootstrap-assets": cmdBootstrapAssets, "push-image": cmdPushImage,
"init-forwarding": cmdInitForwarding, "mirror-build-tools": cmdMirrorBuildTools,
"version": cmdVersion, "registry-gate": cmdRegistryGate,
"update": cmdUpdate, "manifests": cmdManifests,
"apply": cmdApply,
"setup": cmdSetup,
"converge": cmdConverge,
"breakGlass": cmdBreakGlass,
"bootstrap-assets": cmdBootstrapAssets,
"init-forwarding": cmdInitForwarding,
"init-volume": cmdInitVolume,
"pin-images": cmdPinImages,
"version": cmdVersion,
"update": cmdUpdate,
"watchdog": cmdWatchdog,
} }
// run dispatches a subcommand. It is separate from main so the router is // run dispatches a subcommand. It is separate from main so the router is
+7 -3
View File
@@ -38,9 +38,13 @@ func TestRunUnknownCommand(t *testing.T) {
} }
// undocumentedCommands are routable on purpose but kept out of the usage text: they // undocumentedCommands are routable on purpose but kept out of the usage text: they
// are called by deploy/bootstrap.sh, not by a human at a prompt. Listing them here is // are called by deploy/bootstrap.sh or the operator's initContainers, not by a human
// what makes their absence from usage a deliberate decision rather than an oversight. // at a prompt. Listing them here is what makes their absence from usage a deliberate
var undocumentedCommands = map[string]bool{"bootstrap-assets": true, "init-forwarding": true} // decision rather than an oversight.
var undocumentedCommands = map[string]bool{
"bootstrap-assets": true, "init-forwarding": true, "init-volume": true,
"pin-images": true,
}
// The usage text and the dispatch table must describe the same set of commands. // The usage text and the dispatch table must describe the same set of commands.
// //
+28 -3
View File
@@ -233,13 +233,36 @@ func provisionSystemServers(ctx context.Context, cfg *config.Config, out io.Writ
// (the on-demand BACKUP Job runs in the minecraft namespace and mounts it to // (the on-demand BACKUP Job runs in the minecraft namespace and mounts it to
// self-record its world_backups row; without the replica the Job's volume // self-record its world_backups row; without the replica the Job's volume
// mount fails and every backup request strands in the cluster). // mount fails and every backup request strands in the cluster).
// An empty build_namespace means the build system's compiled-in default; the
// replica must target the namespace the Jobs actually run in.
buildNS := cfg.Registry.BuildNamespace
if buildNS == "" {
buildNS = platform.DefaultBuildNamespace
}
secretOutcomes := []systemServerOutcome{ secretOutcomes := []systemServerOutcome{
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace, ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token"), naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "minecraft ns", false),
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace, ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
naming.ForwardingSecretName, naming.ForwardingSecretKey, "forwarding-secret"), naming.ForwardingSecretName, naming.ForwardingSecretKey, "forwarding-secret", "minecraft ns", false),
// refresh=true: felis-config is the rendered config, not a credential. The
// backup/restore/fileedit Jobs and the reaper mount this copy, so a re-run
// must update it when the control plane's render has moved on (a stale copy
// e.g. keeps an old database URL after a credential rotation).
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace, ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
"felis-config", "felis.toml", "config"), "felis-config", "felis.toml", "config", "minecraft ns", true),
// The reaper's pre-reap warning emails authenticate with the same relay
// password felis-api uses; the reaper pod runs in the minecraft namespace,
// where a secretKeyRef resolves only against a local mirror. Skipped while
// the relay is not configured yet — the "configure email" screen refreshes
// both mirrors when it applies.
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
"felis-smtp", "password", "smtp", "minecraft ns", false),
// The build namespace needs the same token: the build Job's fetch
// initContainer reads the submission context from the internal face. Best
// effort — a deployment that only installs the control plane simply never
// builds a user submission.
ensureSecretReplica(ctx, cl, controlNS, buildNS,
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "felis-build ns", false),
} }
outcomes := ensureSystemServers(ctx, cl, cfg.K8s.Namespace, cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage, apiBaseURL, cfg.Server.RootDomain, defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname)) outcomes := ensureSystemServers(ctx, cl, cfg.K8s.Namespace, cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage, apiBaseURL, cfg.Server.RootDomain, defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname))
outcomes = append(secretOutcomes, outcomes...) outcomes = append(secretOutcomes, outcomes...)
@@ -250,6 +273,8 @@ func provisionSystemServers(ctx context.Context, cfg *config.Config, out io.Writ
fmt.Fprintf(out, " - %s: ERROR %v\n", o.name, o.err) fmt.Fprintf(out, " - %s: ERROR %v\n", o.name, o.err)
case o.created: case o.created:
fmt.Fprintf(out, " - %s: created (DesiredState=Running)\n", o.name) fmt.Fprintf(out, " - %s: created (DesiredState=Running)\n", o.name)
case o.updated:
fmt.Fprintf(out, " - %s: refreshed from the control namespace\n", o.name)
default: default:
fmt.Fprintf(out, " - %s: skipped (%s)\n", o.name, o.skipped) fmt.Fprintf(out, " - %s: skipped (%s)\n", o.name, o.skipped)
} }
+194 -35
View File
@@ -1,6 +1,7 @@
package main package main
import ( import (
"bytes"
"context" "context"
"fmt" "fmt"
"time" "time"
@@ -95,7 +96,7 @@ const felisLimboHealthPort int32 = 8080
// fail-safes to readiness-only, so a login pod that has the URL/domain but not yet // fail-safes to readiness-only, so a login pod that has the URL/domain but not yet
// the token is safe (it simply does not authenticate) rather than broken. // the token is safe (it simply does not authenticate) rather than broken.
const ( const (
envAPIBaseURL = "FELIS_API_BASE_URL" envAPIBaseURL = naming.EnvAPIBaseURL
envRootDomain = "FELIS_ROOT_DOMAIN" envRootDomain = "FELIS_ROOT_DOMAIN"
envPanelHostname = "FELIS_PANEL_HOSTNAME" envPanelHostname = "FELIS_PANEL_HOSTNAME"
envLobbyServer = "FELIS_LOBBY_SERVER" envLobbyServer = "FELIS_LOBBY_SERVER"
@@ -255,10 +256,31 @@ func buildSystemServerClient() (client.Client, error) {
// setup can report it without the provisioner deciding on the output format. // setup can report it without the provisioner deciding on the output format.
type systemServerOutcome struct { type systemServerOutcome struct {
name string name string
created bool // true = we created it this run created bool // true = we created it this run
available bool // true = the required object now exists updated bool // true = we refreshed an existing replica from the source
skipped string // non-empty = why it was skipped (image unset / already exists) available bool // true = the required object now exists
err error // non-nil = create failed skipped string // non-empty = why it was skipped (image unset / already exists)
err error // non-nil = create failed
changes []string // converge only: the fields this pass filled
}
// systemServerPlan is one system service in the provisioner's table: its name,
// the image config gives it, and the pure builder for its desired CR.
type systemServerPlan struct {
name string
image string
build func(image, namespace string) (*v1alpha1.MinecraftServer, error)
}
// systemServerPlans is the single description of the login+lobby pair, shared by
// ensureSystemServers (create-if-absent) and convergeSystemServers (field fill).
func systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerPlan {
return []systemServerPlan{
{name: naming.SystemLoginServer, image: loginImage, build: func(image, ns string) (*v1alpha1.MinecraftServer, error) {
return loginSystemServer(image, ns, apiBaseURL, rootDomain, panelHostname)
}},
{name: naming.SystemLobbyServer, image: lobbyImage, build: lobbySystemServer},
}
} }
// ensureSystemServers idempotently creates the login and lobby system services. // ensureSystemServers idempotently creates the login and lobby system services.
@@ -269,17 +291,7 @@ type systemServerOutcome struct {
// K8s client and namespace; this function performs no signal-handler or client // K8s client and namespace; this function performs no signal-handler or client
// setup of its own. // setup of its own.
func ensureSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome { func ensureSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome {
type plan struct { plans := systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname)
name string
image string
build func(image, namespace string) (*v1alpha1.MinecraftServer, error)
}
plans := []plan{
{name: naming.SystemLoginServer, image: loginImage, build: func(image, ns string) (*v1alpha1.MinecraftServer, error) {
return loginSystemServer(image, ns, apiBaseURL, rootDomain, panelHostname)
}},
{name: naming.SystemLobbyServer, image: lobbyImage, build: lobbySystemServer},
}
outcomes := make([]systemServerOutcome, 0, len(plans)) outcomes := make([]systemServerOutcome, 0, len(plans))
for _, p := range plans { for _, p := range plans {
@@ -379,12 +391,7 @@ var derivedSystemEnv = map[string]bool{
// deliberate removal is indistinguishable from drift and re-adding it would fight the // deliberate removal is indistinguishable from drift and re-adding it would fight the
// operator every run. // operator every run.
func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired *v1alpha1.MinecraftServer) (bool, error) { func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired *v1alpha1.MinecraftServer) (bool, error) {
want := make(map[string]string, len(derivedSystemEnv)) want := derivedEnvWanted(desired)
for _, e := range desired.Spec.Env {
if derivedSystemEnv[e.Name] {
want[e.Name] = e.Value
}
}
changed := false changed := false
for i, e := range existing.Spec.Env { for i, e := range existing.Spec.Env {
@@ -402,6 +409,116 @@ func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired
return true, nil return true, nil
} }
// derivedEnvWanted maps the derived env keys of desired onto their values.
func derivedEnvWanted(desired *v1alpha1.MinecraftServer) map[string]string {
want := make(map[string]string, len(derivedSystemEnv))
for _, e := range desired.Spec.Env {
if derivedSystemEnv[e.Name] {
want[e.Name] = e.Value
}
}
return want
}
// convergeSystemServers is the explicit convergence pass over already-installed
// system servers (#1). ensureSystemServers is create-if-absent by design — an
// existing CR is left alone so a re-run cannot clobber an operator's edits — and
// that leaves no path for a field the DESIRED spec gained after the install:
// spec.rcon (the lobby's write channel), spec.startup.healthHTTPPort (the login
// gate's readiness probe), or a config-derived env key that did not exist yet.
// Such fields sit at their zero value forever while re-running setup reports
// success, which is exactly the reported "configuration updates never reach an
// installed deployment" symptom.
//
// This pass fills exactly those zero-value fields and the config-derived env keys,
// and nothing else: a field already holding a non-zero value is the operator's and
// is never overwritten. It is an explicit command rather than an implicit step of
// setup because some fills need an ordering only the operator knows — enabling
// RCON or the HTTP readiness gate on a server whose image predates the listener
// would hold that server in Starting until it was marked Failed. Rebuild (or
// upgrade) the images first, then run this.
func convergeSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome {
outcomes := make([]systemServerOutcome, 0, 2)
for _, p := range systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname) {
if p.image == "" {
outcomes = append(outcomes, systemServerOutcome{name: p.name, skipped: "image not configured"})
continue
}
desired, err := p.build(p.image, namespace)
if err != nil {
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: err})
continue
}
var existing v1alpha1.MinecraftServer
switch err := cl.Get(ctx, client.ObjectKeyFromObject(desired), &existing); {
case apierrors.IsNotFound(err):
outcomes = append(outcomes, systemServerOutcome{name: p.name,
skipped: "not present — run `sudo felis setup` first"})
continue
case err != nil:
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: err})
continue
}
if existing.Labels[v1alpha1.LabelSystemRole] != p.name {
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: fmt.Errorf(
"existing MinecraftServer %s/%s is not marked as the Felis %q system role; refusing to converge it",
namespace, p.name, p.name,
)})
continue
}
var changes []string
if existing.Spec.Rcon == (v1alpha1.RconSpec{}) && desired.Spec.Rcon != (v1alpha1.RconSpec{}) {
existing.Spec.Rcon = desired.Spec.Rcon
changes = append(changes, "spec.rcon")
}
if existing.Spec.Startup.HealthHTTPPort == 0 && desired.Spec.Startup.HealthHTTPPort != 0 {
existing.Spec.Startup.HealthHTTPPort = desired.Spec.Startup.HealthHTTPPort
changes = append(changes, "spec.startup.healthHTTPPort")
}
changes = append(changes, convergeDerivedEnv(&existing, desired)...)
if len(changes) == 0 {
outcomes = append(outcomes, systemServerOutcome{name: p.name, available: true, skipped: "already converged"})
continue
}
if err := cl.Update(ctx, &existing); err != nil {
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: fmt.Errorf("converge %s: %w", p.name, err)})
continue
}
outcomes = append(outcomes, systemServerOutcome{name: p.name, available: true, updated: true, changes: changes})
}
return outcomes
}
// convergeDerivedEnv makes the config-derived env match the desired values: a key
// whose value drifted is overwritten, and a key missing entirely is added. This is
// the wider half of the same explicit pass — refreshDerivedEnv's present-only loop
// can never introduce a NEW key, which is how a derived key added after an install
// never reached it at all.
func convergeDerivedEnv(existing, desired *v1alpha1.MinecraftServer) []string {
want := derivedEnvWanted(desired)
var changes []string
present := make(map[string]bool, len(existing.Spec.Env))
for i := range existing.Spec.Env {
e := &existing.Spec.Env[i]
present[e.Name] = true
if v, ok := want[e.Name]; ok && v != e.Value {
e.Value = v
changes = append(changes, "env "+e.Name)
}
}
for _, e := range desired.Spec.Env {
if !derivedSystemEnv[e.Name] || present[e.Name] {
continue
}
existing.Spec.Env = append(existing.Spec.Env, e)
changes = append(changes, "env "+e.Name)
}
return changes
}
// The login gate is a hard prerequisite of the Owner bind, so setup waits for it // The login gate is a hard prerequisite of the Owner bind, so setup waits for it
// rather than racing it. The ceiling covers a cold image pull on a fresh node; // rather than racing it. The ceiling covers a cold image pull on a fresh node;
// the poll is fast enough that a warm start feels immediate. // the poll is fast enough that a warm start feels immediate.
@@ -471,26 +588,35 @@ func phaseOrPending(p v1alpha1.Phase) string {
return string(p) return string(p)
} }
// ensureSecretReplica copies one Secret from the control namespace into the minecraft // ensureSecretReplica copies one Secret from the control namespace into a workload
// namespace so a backend pod can mount it via secretKeyRef. A secretKeyRef is // namespace (minecraft — or the build namespace, whose fetch initContainer reads the
// namespace-local, but the backends run in the minecraft namespace while the sources // context from the felis-api internal face with the same token) so a pod can mount it
// of truth live beside the control plane — so without this replica the operator's // via secretKeyRef. A secretKeyRef is namespace-local, but those workloads do not run
// injected secretKeyRef would dangle and wedge the pod in CreateContainerConfigError. // beside the control plane — so without this replica the secretKeyRef would dangle and
// wedge the pod in CreateContainerConfigError.
// //
// Two Secrets need it, for different reasons: the service token (login only — it // Three Secrets need it, for different reasons: the service token (the login limbo and
// authenticates the limbo plugin to the felis-api internal face) and the Velocity // the build Pod's context fetch — both authenticate to the felis-api internal face),
// modern-forwarding secret (every backend — it is how a backend knows a login really // the Velocity modern-forwarding secret (every backend — it is how a backend knows
// came from the proxy, and so that the player's UUID is Mojang-verified rather than // a login really came from the proxy, and so that the player's UUID is Mojang-verified
// offline-derived). // rather than offline-derived), and the SMTP relay password (the reaper's pre-reap
// warning emails; the felis-config mirror is what carries [smtp] into its pod).
// //
// It is create-if-absent: an existing replica is left untouched so a hand-rotated // It is create-if-absent: an existing replica is left untouched so a hand-rotated
// value in the minecraft namespace is never clobbered (to rotate, delete the replica // value in the workload namespace is never clobbered (to rotate, delete the replica
// and re-run setup). Best-effort like the rest of the provisioner: a missing source or // and re-run setup). Best-effort like the rest of the provisioner: a missing source or
// a create failure degrades to a reported outcome, never a hard setup failure. It // a create failure degrades to a reported outcome, never a hard setup failure. It
// copies only Type and Data — never labels/annotations/ownerRefs — so the replica // copies only Type and Data — never labels/annotations/ownerRefs — so the replica
// carries no accidental GC owner or managed-by lineage. // carries no accidental GC owner or managed-by lineage.
func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace, minecraftNamespace, secretName, secretKey, label string) systemServerOutcome { //
name := label + " (minecraft ns)" // refreshExisting switches the felis-config mirror to refresh-in-place: that Secret is
// a rendered config, never a hand-rotated credential, and the workload Jobs that mount
// it (backup/restore/fileedit) plus the reaper silently misbehave on a stale copy —
// e.g. after a database credential rotation the control plane moves on while every
// backup Job keeps failing auth. Credential Secrets keep the never-overwrite rule so a
// rotated value survives; to rotate those, delete the replica and re-run setup.
func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace, minecraftNamespace, secretName, secretKey, label, where string, refreshExisting bool) systemServerOutcome {
name := label + " (" + where + ")"
validate := func(secret *corev1.Secret, location, skipped string) systemServerOutcome { validate := func(secret *corev1.Secret, location, skipped string) systemServerOutcome {
if len(secret.Data[secretKey]) == 0 { if len(secret.Data[secretKey]) == 0 {
return systemServerOutcome{name: name, skipped: fmt.Sprintf( return systemServerOutcome{name: name, skipped: fmt.Sprintf(
@@ -498,6 +624,33 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
} }
return systemServerOutcome{name: name, available: true, skipped: skipped} return systemServerOutcome{name: name, available: true, skipped: skipped}
} }
// refreshFromControl updates an existing replica from the control-namespace source
// when the rendered key differs. Only the felis-config mirror opts in.
refreshFromControl := func(existing *corev1.Secret) systemServerOutcome {
var src corev1.Secret
if err := cl.Get(ctx, client.ObjectKey{Namespace: controlNamespace, Name: secretName}, &src); err != nil {
if apierrors.IsNotFound(err) {
return systemServerOutcome{name: name, skipped: fmt.Sprintf(
"source Secret %s/%s not found — provision it (deploy/bootstrap.sh), then re-run setup",
controlNamespace, secretName)}
}
return systemServerOutcome{name: name, err: err}
}
if out := validate(&src, controlNamespace, ""); !out.available {
return out
}
if bytes.Equal(existing.Data[secretKey], src.Data[secretKey]) {
return validate(existing, minecraftNamespace, "already current")
}
if existing.Data == nil {
existing.Data = map[string][]byte{}
}
existing.Data[secretKey] = src.Data[secretKey]
if err := cl.Update(ctx, existing); err != nil {
return systemServerOutcome{name: name, err: err}
}
return systemServerOutcome{name: name, updated: true, available: true}
}
if controlNamespace == minecraftNamespace { if controlNamespace == minecraftNamespace {
// Same namespace needs no replica, but the source still has to exist. // Same namespace needs no replica, but the source still has to exist.
var existing corev1.Secret var existing corev1.Secret
@@ -516,6 +669,9 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
var existing corev1.Secret var existing corev1.Secret
getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing) getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing)
if getErr == nil { if getErr == nil {
if refreshExisting {
return refreshFromControl(&existing)
}
return validate(&existing, minecraftNamespace, "already exists") return validate(&existing, minecraftNamespace, "already exists")
} }
if !apierrors.IsNotFound(getErr) { if !apierrors.IsNotFound(getErr) {
@@ -544,6 +700,9 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
if getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing); getErr != nil { if getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing); getErr != nil {
return systemServerOutcome{name: name, err: getErr} return systemServerOutcome{name: name, err: getErr}
} }
if refreshExisting {
return refreshFromControl(&existing)
}
return validate(&existing, minecraftNamespace, "already exists") return validate(&existing, minecraftNamespace, "already exists")
} }
return systemServerOutcome{name: name, err: err} return systemServerOutcome{name: name, err: err}
+82 -1
View File
@@ -156,7 +156,7 @@ func TestEnsureSecretReplica(t *testing.T) {
} }
replicate := func(cl client.Client, controlNS, mcNS string) systemServerOutcome { replicate := func(cl client.Client, controlNS, mcNS string) systemServerOutcome {
return ensureSecretReplica(ctx, cl, controlNS, mcNS, return ensureSecretReplica(ctx, cl, controlNS, mcNS,
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token") naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "minecraft ns", false)
} }
t.Run("replicates when absent", func(t *testing.T) { t.Run("replicates when absent", func(t *testing.T) {
@@ -251,6 +251,87 @@ func TestEnsureSecretReplica(t *testing.T) {
}) })
} }
// The felis-config mirror is the one replica that must refresh: it is a rendered
// config, and a stale workload-side copy (backup/restore/fileedit Jobs, the reaper)
// misbehaves silently — a rotated database credential keeps the control plane moving
// while every backup Job keeps failing auth. Credential Secrets keep create-if-absent.
func TestEnsureSecretReplicaRefresh(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
configSecret := func(ns, body string) *corev1.Secret {
return &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: "felis-config", Namespace: ns},
Type: corev1.SecretTypeOpaque,
Data: map[string][]byte{"felis.toml": []byte(body)},
}
}
refresh := func(cl client.Client) systemServerOutcome {
return ensureSecretReplica(ctx, cl, "felis", "minecraft",
"felis-config", "felis.toml", "config", "minecraft ns", true)
}
replicaBody := func(t *testing.T, cl client.Client) string {
t.Helper()
var got corev1.Secret
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: "felis-config"}, &got); err != nil {
t.Fatalf("get replica: %v", err)
}
return string(got.Data["felis.toml"])
}
t.Run("refreshes a stale config replica", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("felis", "current"),
configSecret("minecraft", "stale"),
).Build()
out := refresh(cl)
if out.err != nil || !out.updated || !out.available {
t.Fatalf("outcome = %+v, want refreshed", out)
}
if got := replicaBody(t, cl); got != "current" {
t.Errorf("replica = %q, want current", got)
}
})
t.Run("leaves a current config replica alone", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("felis", "same"),
configSecret("minecraft", "same"),
).Build()
out := refresh(cl)
if out.err != nil || out.updated || !out.available || out.skipped != "already current" {
t.Fatalf("outcome = %+v, want already current", out)
}
})
t.Run("fills an empty-key replica", func(t *testing.T) {
empty := &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: "felis-config", Namespace: "minecraft"},
Data: map[string][]byte{},
}
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("felis", "current"), empty).Build()
out := refresh(cl)
if out.err != nil || !out.updated {
t.Fatalf("outcome = %+v, want refreshed", out)
}
if got := replicaBody(t, cl); got != "current" {
t.Errorf("replica = %q, want current", got)
}
})
t.Run("missing source degrades to a skip", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("minecraft", "stale")).Build()
out := refresh(cl)
if out.err != nil || out.updated || out.available || out.skipped == "" {
t.Fatalf("outcome = %+v, want skipped (source missing)", out)
}
if got := replicaBody(t, cl); got != "stale" {
t.Errorf("replica = %q, want untouched stale", got)
}
})
}
func TestRequiredProvisioningError(t *testing.T) { func TestRequiredProvisioningError(t *testing.T) {
ready := []systemServerOutcome{ ready := []systemServerOutcome{
{name: "service-token (minecraft ns)", available: true}, {name: "service-token (minecraft ns)", available: true},
+1 -1
View File
@@ -86,7 +86,7 @@ func (m *backupModel) loadCmd() tea.Cmd {
if err != nil { if err != nil {
return backupListMsg{err: fmt.Errorf("list servers: %w", err)} return backupListMsg{err: fmt.Errorf("list servers: %w", err)}
} }
return backupListMsg{cl: cl, servers: servers} return backupListMsg{cl: cl, servers: backupPickable(servers)}
} }
} }
+4
View File
@@ -437,6 +437,10 @@ func (m *edgeModel) errorView() string {
if m.lastErr != nil { if m.lastErr != nil {
b.WriteString(tuiHint.Render(m.lastErr.Error()) + "\n") b.WriteString(tuiHint.Render(m.lastErr.Error()) + "\n")
} }
// Nothing done above is rolled back, and nothing needs to be: every step finds what an
// earlier attempt created (the tunnel, its DNS route, the Access app and policy) and
// carries on from it.
b.WriteString("\n" + tuiHint.Render("Retrying is safe: it reuses the tunnel, DNS record and Access app created so far instead of making duplicates.") + "\n")
b.WriteString("\n" + tuiAction("enter", "retry", "esc", "edit")) b.WriteString("\n" + tuiAction("enter", "retry", "esc", "edit"))
return b.String() return b.String()
} }
+50 -6
View File
@@ -32,7 +32,9 @@ func applyCloudflareEdge(ctx context.Context, result *cfsetup.Result, panelHost,
if adminHost == "" { if adminHost == "" {
return fmt.Errorf("admin hostname is required") return fmt.Errorf("admin hostname is required")
} }
if err := writeConnectionConfig(panelHost, adminHost, result.AccessAud); err != nil { // cloudflared is the only way in once the NodePort is fenced, so the
// visitor address it writes can key the sign-in rate limit.
if err := writeConnectionConfig(panelHost, adminHost, result.AccessAud, "CF-Connecting-IP"); err != nil {
return err return err
} }
if err := applyFelisConfigSecret(ctx); err != nil { if err := applyFelisConfigSecret(ctx); err != nil {
@@ -78,7 +80,8 @@ func applyReverseProxy(ctx context.Context, panelHost, adminHost string) error {
if adminHost == "" { if adminHost == "" {
return fmt.Errorf("admin hostname is required") return fmt.Errorf("admin hostname is required")
} }
if err := writeConnectionConfig(panelHost, adminHost, ""); err != nil { // Caddy, nginx and Traefik all append the peer they saw to X-Forwarded-For.
if err := writeConnectionConfig(panelHost, adminHost, "", "X-Forwarded-For"); err != nil {
return err return err
} }
if err := applyFelisConfigSecret(ctx); err != nil { if err := applyFelisConfigSecret(ctx); err != nil {
@@ -93,16 +96,17 @@ func applyReverseProxy(ctx context.Context, panelHost, adminHost string) error {
// writeConnectionConfig stamps the chosen hostnames (and optional Access audience) // writeConnectionConfig stamps the chosen hostnames (and optional Access audience)
// into both the host and pod config files. An empty aud clears any prior // into both the host and pod config files. An empty aud clears any prior
// Cloudflare audience, which is correct when switching to a non-Access front. // Cloudflare audience, which is correct when switching to a non-Access front.
func writeConnectionConfig(panelHost, adminHost, aud string) error { // clientIPHeader is the header that front writes the visitor address into.
func writeConnectionConfig(panelHost, adminHost, aud, clientIPHeader string) error {
for _, path := range []string{hostSetupConfigPath, podSetupConfigPath} { for _, path := range []string{hostSetupConfigPath, podSetupConfigPath} {
if err := updateAuthConfig(path, panelHost, adminHost, aud); err != nil { if err := updateAuthConfig(path, panelHost, adminHost, aud, clientIPHeader); err != nil {
return err return err
} }
} }
return nil return nil
} }
func updateAuthConfig(path, panelHost, adminHost, aud string) error { func updateAuthConfig(path, panelHost, adminHost, aud, clientIPHeader string) error {
cfg, err := config.Load(path) cfg, err := config.Load(path)
if err != nil { if err != nil {
return err return err
@@ -112,6 +116,7 @@ func updateAuthConfig(path, panelHost, adminHost, aud string) error {
} }
cfg.Auth.AdminHostname = adminHost cfg.Auth.AdminHostname = adminHost
cfg.Auth.AccessJWTAud = aud cfg.Auth.AccessJWTAud = aud
cfg.Auth.ClientIPHeader = clientIPHeader
return writeConfig(path, cfg) return writeConfig(path, cfg)
} }
@@ -133,6 +138,11 @@ func writeConfig(path string, cfg *config.Config) error {
return os.Rename(tmpPath, path) return os.Rename(tmpPath, path)
} }
// applyFelisConfigSecret applies the rendered config to the control namespace and
// then converges the workload-namespace mirror best-effort. The mirror feeds the
// backup/restore/fileedit Jobs and the reaper; without this refresh a reconfigure
// here would leave those readers on the previous render until the next `felis
// setup` run (startup pass) or installer re-run.
func applyFelisConfigSecret(ctx context.Context) error { func applyFelisConfigSecret(ctx context.Context) error {
out, err := kubectlOutput(ctx, out, err := kubectlOutput(ctx,
"-n", "felis", "create", "secret", "generic", "felis-config", "-n", "felis", "create", "secret", "generic", "felis-config",
@@ -142,7 +152,41 @@ func applyFelisConfigSecret(ctx context.Context) error {
if err != nil { if err != nil {
return err return err
} }
return kubectlWithInput(ctx, out, "apply", "-f", "-") if err := kubectlWithInput(ctx, out, "apply", "-f", "-"); err != nil {
return err
}
if err := replicateFelisConfigToWorkloadNamespace(ctx); err != nil {
fmt.Fprintf(os.Stderr, "felis setup: warning: the control-plane config is applied, but the workload-namespace mirror could not be refreshed (%v); re-run felis setup once that is fixed\n", err)
}
return nil
}
// replicateFelisConfigToWorkloadNamespace overwrites the workload-namespace
// felis-config mirror with the freshly rendered pod config. Deliberately a full
// replace, not create-if-absent: a stale mirror is exactly what silently hands
// the Jobs that mount it old settings after a reconfigure. No-op when the
// workload namespace is unset or is the control namespace itself.
func replicateFelisConfigToWorkloadNamespace(ctx context.Context) error {
cfg, err := config.Load(hostSetupConfigPath)
if err != nil {
return err
}
ns := cfg.K8s.Namespace
if ns == "" || ns == "felis" {
return nil
}
manifest, err := kubectlOutput(ctx,
"-n", ns, "create", "secret", "generic", "felis-config",
"--from-file=felis.toml="+podSetupConfigPath,
"--dry-run=client", "-o", "yaml",
)
if err != nil {
return fmt.Errorf("render felis-config for %s: %w", ns, err)
}
if err := kubectlWithInput(ctx, manifest, "-n", ns, "apply", "-f", "-"); err != nil {
return fmt.Errorf("replicate felis-config to %s: %w", ns, err)
}
return nil
} }
func installCloudflaredService(ctx context.Context, cloudflaredBin, configPath string) error { func installCloudflaredService(ctx context.Context, cloudflaredBin, configPath string) error {
+5 -1
View File
@@ -207,7 +207,11 @@ func (m *mcBindModel) doneView() string {
if box.Len() > 0 { if box.Len() > 0 {
box.WriteString("\n") box.WriteString("\n")
} }
box.WriteString(tuiLabel.Render("setup URL ") + "\n" + tuiPassword.Render(m.setupTokenURL) + "\n\n") box.WriteString(tuiLabel.Render("setup URL ") + "\n")
for _, line := range wrapDisplayURL(m.setupTokenURL, 70) {
box.WriteString(tuiPassword.Render(line) + "\n")
}
box.WriteString("\n")
box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup.\nIt is shown only once.")) box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup.\nIt is shown only once."))
} }
if m.auditWarning != "" { if m.auditWarning != "" {
+34 -7
View File
@@ -151,19 +151,46 @@ func TestOwnerModelProvisionErrorRouting(t *testing.T) {
} }
}) })
t.Run("a conflict on the Owner path is not a retry", func(t *testing.T) { t.Run("the Owner seat refusal returns to the form naming the seat", func(t *testing.T) {
// Defensive: the Owner upserts and so never conflicts, but were one ever to // Upserting a fresh username while a seat is occupied would mint a second
// surface it must end the session rather than loop the form — only the // owner, so provisionOwner refuses with ownerSeatTakenError (Is
// insert-only operator path is retryable. // api.ErrConflict) and the console must route back for a retype — the same
// recoverable contract as the operator clash, and the only Owner-path
// conflict there is.
m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true) m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true)
seatErr := &ownerSeatTakenError{seat: "seat-holder"}
next, cmd := m.Update(owProvisionMsg{err: conflict}) next, cmd := m.Update(owProvisionMsg{err: seatErr})
om := next.(*ownerModel)
if om.step != owProvision {
t.Fatalf("step = %v, want owProvision — the seat refusal is recoverable", om.step)
}
if om.provisionErr == nil || !errors.Is(om.provisionErr, api.ErrConflict) || !strings.Contains(om.provisionErr.Error(), "seat-holder") {
t.Errorf("provisionErr = %v, want the seat refusal naming the seat", om.provisionErr)
}
// Feed the rebuilt form's init message back through so its view renders;
// then the note must carry the seat name (the operator's retype cue).
if cmd != nil {
if msg := cmd(); msg != nil {
if n2, _ := om.Update(msg); n2 != nil {
om = n2.(*ownerModel)
}
}
}
if view := om.form.View(); !strings.Contains(view, "seat-holder") {
t.Errorf("the provision form must surface the seat refusal:\n%s", view)
}
})
t.Run("a generic Owner-path fault still tears the console down", func(t *testing.T) {
m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true)
next, cmd := m.Update(owProvisionMsg{err: errors.New("boom")})
om := next.(*ownerModel) om := next.(*ownerModel)
if om.provisionErr != nil { if om.provisionErr != nil {
t.Error("the Owner path recorded a retryable conflict; only the operator path retries") t.Error("a generic fault must not be treated as a retryable refusal")
} }
if res, ok := cmd().(ownerResultMsg); !ok || res.err == nil { if res, ok := cmd().(ownerResultMsg); !ok || res.err == nil {
t.Error("an Owner-path conflict should tear down via an error result") t.Error("a generic Owner-path fault should tear down via an error result")
} }
}) })
} }
+8 -1
View File
@@ -3,8 +3,10 @@ package main
import ( import (
"context" "context"
"fmt" "fmt"
"io"
"time" "time"
"felis.lolicon.best/internal/dbbackup"
"felis.lolicon.best/internal/store" "felis.lolicon.best/internal/store"
) )
@@ -12,7 +14,7 @@ import (
// applied count. Used by the preflight stage to self-heal a freshly bootstrapped // applied count. Used by the preflight stage to self-heal a freshly bootstrapped
// (or upgraded) database. // (or upgraded) database.
func applyMigrations(dbURL string) (int, error) { func applyMigrations(dbURL string) (int, error) {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel() defer cancel()
drv, err := store.Open(ctx, dbURL) drv, err := store.Open(ctx, dbURL)
if err != nil { if err != nil {
@@ -23,6 +25,11 @@ func applyMigrations(dbURL string) (int, error) {
if err != nil { if err != nil {
return 0, err return 0, err
} }
// Same guard as `felis migrate up`: never roll a populated database forward
// without a snapshot to roll back to.
if _, err := preMigrateBackup(ctx, drv, migrations, dbURL, dbbackup.DefaultDir, io.Discard); err != nil {
return 0, fmt.Errorf("pre-migration backup: %w", err)
}
if _, err := store.Up(ctx, drv, migrations); err != nil { if _, err := store.Up(ctx, drv, migrations); err != nil {
return 0, err return 0, err
} }
+21 -10
View File
@@ -174,12 +174,13 @@ func (m *ownerModel) Update(msg tea.Msg) (tea.Model, tea.Cmd) {
case owProvisionMsg: case owProvisionMsg:
if msg.err != nil { if msg.err != nil {
// A taken Operator username is the expected, recoverable outcome of the // api.ErrConflict marks the two recoverable refusals: a taken Operator
// insert-only operator path (refusing the clash is the whole reason it is // username (insert-only clash) and an Owner reset naming anything but the
// insert-only, not an upsert). Route back to the form with a note so the // occupied seat (ownerSeatTakenError Is ErrConflict). Route back to the
// operator can pick another name, rather than tearing down the console — // form with a note so the operator can retype, rather than tearing down
// any other error is a genuine fault and still ends the session. // the console — any other error is a genuine fault and still ends the
if m.operation == bgAddOperator && errors.Is(msg.err, api.ErrConflict) { // session.
if errors.Is(msg.err, api.ErrConflict) {
m.provisionErr = msg.err m.provisionErr = msg.err
m.step = owProvision m.step = owProvision
m.form = m.sized(m.buildProvisionForm()) m.form = m.sized(m.buildProvisionForm())
@@ -364,9 +365,15 @@ func (m *ownerModel) buildProvisionForm() *huh.Form {
} }
} }
if m.provisionErr != nil { if m.provisionErr != nil {
// The only error routed back to this form is a username clash on the insert-only // Recoverable refusals routed back here: the seat refusal already names the
// operator path; show a concrete prompt to choose another name. // username to enter, so show it verbatim; the operator-name clash gets the
desc = "That username is already taken — choose a different one.\n\n" + desc // generic retry prompt.
note := "That username is already taken — choose a different one."
var seatErr *ownerSeatTakenError
if errors.As(m.provisionErr, &seatErr) {
note = seatErr.Error()
}
desc = note + "\n\n" + desc
} }
fields := []huh.Field{ fields := []huh.Field{
@@ -420,7 +427,11 @@ func (m *ownerModel) doneView() string {
var box strings.Builder var box strings.Builder
box.WriteString(tuiLabel.Render("username ") + m.username + "\n") box.WriteString(tuiLabel.Render("username ") + m.username + "\n")
if m.setupTokenURL != "" { if m.setupTokenURL != "" {
box.WriteString("\n" + tuiLabel.Render("setup URL ") + "\n" + tuiPassword.Render(m.setupTokenURL) + "\n\n") box.WriteString("\n" + tuiLabel.Render("setup URL ") + "\n")
for _, line := range wrapDisplayURL(m.setupTokenURL, 70) {
box.WriteString(tuiPassword.Render(line) + "\n")
}
box.WriteString("\n")
box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup. It is shown only once.")) box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup. It is shown only once."))
} }
if m.auditWarning != "" { if m.auditWarning != "" {
+1
View File
@@ -553,6 +553,7 @@ func (m *rootModel) showSummary() (tea.Model, tea.Cmd) {
storageLabel: m.result.storageDetail, storageLabel: m.result.storageDetail,
routedHosts: routed, routedHosts: routed,
localHint: m.result.connectMethod == connectLocal, localHint: m.result.connectMethod == connectLocal,
alreadySetUp: m.result.alreadySetUp,
}) })
} }
+28
View File
@@ -249,6 +249,34 @@ func TestRootReconfigureSMTP(t *testing.T) {
} }
} }
// TestRootReconfigureStorageKeepsStatusFraming locks the same rule for the
// "change storage" path: on a re-run, completing it must land back on the
// alreadySetUp status framing (with the updated recap), not "Setup complete."
func TestRootReconfigureStorageKeepsStatusFraming(t *testing.T) {
m := newTestRoot(true, consoleModeSetup, "")
m = drive(t, m, preflightDoneMsg{})
if _, ok := m.screen.(*summaryModel); !ok {
t.Fatalf("re-run after preflight, screen = %T, want *summaryModel", m.screen)
}
m = drive(t, m, reconfigureStorageMsg{})
if _, ok := m.screen.(*storageChooserModel); !ok {
t.Fatalf("reconfigure-storage screen = %T, want *storageChooserModel", m.screen)
}
m = drive(t, m, storageResultMsg{method: storageLocal, detail: "local disk · /var/lib/felis/uploads"})
sum, ok := m.screen.(*summaryModel)
if !ok {
t.Fatalf("after reconfigure-storage, screen = %T, want *summaryModel", m.screen)
}
if !sum.alreadySetUp {
t.Fatalf("after reconfigure-storage, summary should keep the alreadySetUp framing")
}
if sum.storageLabel != "local disk · /var/lib/felis/uploads" {
t.Fatalf("storageLabel = %q, want the updated recap", sum.storageLabel)
}
}
func TestRootRerunLandsOnStatus(t *testing.T) { func TestRootRerunLandsOnStatus(t *testing.T) {
// adminExists at start of a setup run = re-run: preflight should skip straight // adminExists at start of a setup run = re-run: preflight should skip straight
// to the "manage in panel" status screen, never touching owner/connect. // to the "manage in panel" status screen, never touching owner/connect.
+51 -5
View File
@@ -4,6 +4,7 @@ import (
"context" "context"
"errors" "errors"
"fmt" "fmt"
"os"
"strconv" "strconv"
"strings" "strings"
@@ -338,20 +339,30 @@ func applySMTPConfig(ctx context.Context, in smtpInputs) error {
if err := applyFelisConfigSecret(ctx); err != nil { if err := applyFelisConfigSecret(ctx); err != nil {
return err return err
} }
// Refresh the workload-namespace copies too (the reaper's warning path): the
// OTP path is already live in the control namespace, so a replica miss is
// reported but not fatal.
if err := replicateSMTPToWorkloadNamespace(ctx, in.password); err != nil {
fmt.Fprintf(os.Stderr, "felis setup: warning: email is configured, but refreshing the workload copies failed (pre-reap warning emails may stay suppressed): %v\n", err)
}
if err := kubectl(ctx, "-n", "felis", "rollout", "restart", "deployment/felis-api"); err != nil { if err := kubectl(ctx, "-n", "felis", "rollout", "restart", "deployment/felis-api"); err != nil {
return err return err
} }
return kubectl(ctx, "-n", "felis", "rollout", "status", "deployment/felis-api", "--timeout=180s") return kubectl(ctx, "-n", "felis", "rollout", "status", "deployment/felis-api", "--timeout=180s")
} }
// applySMTPSecret creates (or replaces) the felis-smtp Secret the felis-api // smtpSecretManifest renders the felis-smtp Secret for the given namespace, the
// Deployment injects the relay password from. Rendered in-process and piped to // one the receiving Deployment/CronJob resolves its secretKeyRef against (felis
// for felis-api, the workload namespace for the reaper's mirror). The namespace
// must be IN the manifest: kubectl rejects a manifest whose namespace conflicts
// with -n, so leaving the control namespace hardcoded made every workload-ns
// replica fail before it started. Rendered in-process and piped to
// `kubectl apply` — the password is never a command-line arg, so it never // `kubectl apply` — the password is never a command-line arg, so it never
// appears in the host process table. // appears in the host process table.
func applySMTPSecret(ctx context.Context, password string) error { func smtpSecretManifest(password, namespace string) ([]byte, error) {
secret := &corev1.Secret{ secret := &corev1.Secret{
TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Secret"}, TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Secret"},
ObjectMeta: metav1.ObjectMeta{Name: platform.SMTPSecretName, Namespace: "felis"}, ObjectMeta: metav1.ObjectMeta{Name: platform.SMTPSecretName, Namespace: namespace},
Type: corev1.SecretTypeOpaque, Type: corev1.SecretTypeOpaque,
StringData: map[string]string{ StringData: map[string]string{
platform.SMTPSecretPasswordKey: password, platform.SMTPSecretPasswordKey: password,
@@ -359,7 +370,42 @@ func applySMTPSecret(ctx context.Context, password string) error {
} }
manifest, err := yaml.Marshal(secret) manifest, err := yaml.Marshal(secret)
if err != nil { if err != nil {
return fmt.Errorf("render smtp secret: %w", err) return nil, fmt.Errorf("render smtp secret: %w", err)
}
return manifest, nil
}
func applySMTPSecret(ctx context.Context, password string) error {
manifest, err := smtpSecretManifest(password, "felis")
if err != nil {
return err
} }
return kubectlWithInput(ctx, manifest, "apply", "-f", "-") return kubectlWithInput(ctx, manifest, "apply", "-f", "-")
} }
// replicateSMTPToWorkloadNamespace refreshes the workload-namespace (minecraft)
// copy of felis-smtp after email is reconfigured. The reaper's CronJob runs
// there and resolves the password by local reference — a secretKeyRef is
// namespace-local — so without this refresh a later SMTP change would never
// reach the pre-reap warning emails. Deliberately OVERWRITES: this is a mirror
// of the control-namespace source, and a stale mirror is exactly the failure
// this closes. The felis-config mirror rides along in applyFelisConfigSecret,
// which every apply path refreshes.
func replicateSMTPToWorkloadNamespace(ctx context.Context, password string) error {
cfg, err := config.Load(hostSetupConfigPath)
if err != nil {
return err
}
ns := cfg.K8s.Namespace
if ns == "" || ns == "felis" {
return nil
}
smtpManifest, err := smtpSecretManifest(password, ns)
if err != nil {
return err
}
if err := kubectlWithInput(ctx, smtpManifest, "-n", ns, "apply", "-f", "-"); err != nil {
return fmt.Errorf("replicate %s to %s: %w", platform.SMTPSecretName, ns, err)
}
return nil
}
+37
View File
@@ -0,0 +1,37 @@
package main
import (
"strings"
"testing"
"sigs.k8s.io/yaml"
)
// TestSMTPSecretManifestCarriesTargetNamespace pins the fix for the
// workload-namespace replica: kubectl refuses a manifest whose namespace
// conflicts with -n ("the namespace from the provided object ... does not
// match"), so the mirror must render felis-smtp with the TARGET namespace —
// otherwise the "configure email" refresh fails on the first apply and the
// felis-config mirror never runs at all.
func TestSMTPSecretManifestCarriesTargetNamespace(t *testing.T) {
for _, ns := range []string{"felis", "minecraft"} {
b, err := smtpSecretManifest("pw", ns)
if err != nil {
t.Fatalf("render for %s: %v", ns, err)
}
var got struct {
Metadata struct {
Namespace string `json:"namespace"`
} `json:"metadata"`
}
if err := yaml.Unmarshal(b, &got); err != nil {
t.Fatalf("unmarshal for %s: %v", ns, err)
}
if got.Metadata.Namespace != ns {
t.Fatalf("manifest namespace = %q, want %q", got.Metadata.Namespace, ns)
}
if !strings.Contains(string(b), "name: felis-smtp") {
t.Fatalf("manifest must still name felis-smtp: %s", b)
}
}
}
+27
View File
@@ -29,6 +29,33 @@ func tuiSeparator() string {
return tuiHint.Render(strings.Repeat("─", 70)) return tuiHint.Render(strings.Repeat("─", 70))
} }
// wrapDisplayURL breaks a long URL into lines no wider than width so the TUI
// renderer never truncates it on a narrow terminal — the one-time setup URL
// carries a 43-char token and overruns 80 columns. It prefers breaking right
// after a '=' or '/' inside the window (the token then lands on its own line)
// and hard-wraps only when no boundary is available. Lines concatenate back to
// the original string.
func wrapDisplayURL(u string, width int) []string {
if width <= 0 {
width = 70
}
var lines []string
for len(u) > width {
cut := width
if i := strings.LastIndexByte(u[:width], '='); i >= 0 && i >= width/2 {
cut = i + 1
} else if i := strings.LastIndexByte(u[:width], '/'); i >= 0 && i >= width/2 {
cut = i + 1
}
lines = append(lines, u[:cut])
u = u[cut:]
}
if u != "" {
lines = append(lines, u)
}
return lines
}
// tuiStepRail renders a breadcrumb of wizard stages. Steps before `current` // tuiStepRail renders a breadcrumb of wizard stages. Steps before `current`
// render as done, `current` is highlighted, and later steps are dimmed. // render as done, `current` is highlighted, and later steps are dimmed.
func tuiStepRail(steps []string, current int) string { func tuiStepRail(steps []string, current int) string {
+26
View File
@@ -0,0 +1,26 @@
package main
import (
"strings"
"testing"
)
func TestWrapDisplayURL(t *testing.T) {
u := "https://op.console.example.net/setup?token=" + strings.Repeat("A", 43)
lines := wrapDisplayURL(u, 70)
if got := strings.Join(lines, ""); got != u {
t.Fatalf("concatenated lines = %q, want the original URL back", got)
}
for i, l := range lines {
if len(l) > 70 {
t.Errorf("line %d is %d cols wide: %q", i, len(l), l)
}
}
if len(lines) < 2 || !strings.HasSuffix(lines[0], "token=") {
t.Fatalf("want the first line to end at the 'token=' boundary, got %q", lines)
}
short := "https://a/b"
if got := wrapDisplayURL(short, 70); len(got) != 1 || got[0] != short {
t.Errorf("short URL should pass through unsplit, got %q", got)
}
}
+62 -27
View File
@@ -24,8 +24,9 @@ const updateTimeout = 60 * time.Second
// The apply side is deliberately NOT implemented in this command. Every component // The apply side is deliberately NOT implemented in this command. Every component
// here is installed by deploy/bootstrap.sh, which is idempotent, already handles the // here is installed by deploy/bootstrap.sh, which is idempotent, already handles the
// parts that are easy to get wrong (Velocity's pinned MINOR, the atomic jar install, // parts that are easy to get wrong (Velocity's pinned MINOR, the atomic jar install,
// the k3s image re-import that a byte-identical StatefulSet template will not // the image re-import + registry push that a byte-identical StatefulSet template
// trigger on its own), and is the path that gets exercised on every install. A // will not trigger on its own), and is the path that gets exercised on every
// install. A
// second installer living in this file would duplicate that policy, could drift from // second installer living in this file would duplicate that policy, could drift from
// it silently, and would be reachable only on a live node where a mistake takes the // it silently, and would be reachable only on a live node where a mistake takes the
// proxy or the control plane down. So `felis update` reports, and hands the operator // proxy or the control plane down. So `felis update` reports, and hands the operator
@@ -42,6 +43,21 @@ type updateTarget struct {
command string command string
} }
// installerRerun is the tested apply path for every planner-backed selector: re-run the
// installer. It is idempotent, and it is the only path that fetches a newer version --
// `felis setup` skips its host-bootstrap phase on a completed install (all four install
// markers already exist), so there it opens the config console and moves no component,
// and even on the bootstrap path it re-images felis-api from the binary setup is already
// running (FELIS_BOOTSTRAP_BINARY), which looks like an update and changes nothing.
//
// The URL is the one-liner both READMEs hand out, read at a tag rather than main: the
// script's release channel installs the newest release's binary, and main can carry
// installer changes that binary was never tested with. installerRef picks the tag and
// renderApplyGuidance substitutes it for {ref}. While the repo is private the URL
// answers 404 (raw.githubusercontent.com hides private repos), which is why the trailer
// below points at the README's token'd form for that case.
const installerRerun = "curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/{ref}/deploy/bootstrap.sh | sudo bash"
// updateTargets is the selector table. panel and plugins both resolve to felis-api // updateTargets is the selector table. panel and plugins both resolve to felis-api
// because they are not separately versioned: the panel is compiled into the felis // because they are not separately versioned: the panel is compiled into the felis
// binary with //go:embed, and the plugin jars are built from this same repo in the // binary with //go:embed, and the plugin jars are built from this same repo in the
@@ -51,19 +67,19 @@ var updateTargets = []updateTarget{
selector: "panel", selector: "panel",
component: "felis-api", component: "felis-api",
note: "the panel is embedded in the felis binary (//go:embed), so updating it means rebuilding the felis image and rolling felis-api", note: "the panel is embedded in the felis binary (//go:embed), so updating it means rebuilding the felis image and rolling felis-api",
command: "sudo felis setup", command: installerRerun,
}, },
{ {
selector: "velocity", selector: "velocity",
component: "velocity", component: "velocity",
note: "re-runs install_velocity: newest BUILD of the pinned minor (FELIS_VELOCITY_VERSION), atomic jar install, then restarts felis-velocity", note: "re-runs install_velocity: the build the release pins in deploy/game-stack.lock (FELIS_VELOCITY_VERSION=<minor> takes that minor's newest build instead), sha256-checked, atomic jar install, then restarts felis-velocity only if the jar or its config changed",
command: "sudo felis setup", command: installerRerun,
}, },
{ {
selector: "plugins", selector: "plugins",
component: "felis-api", component: "felis-api",
note: "felis-velocity.jar is a host-file swap, but felis-paper.jar and felis-limbo.jar are baked into the lobby/limbo images and need a rebuild + k3s image re-import", note: "felis-velocity.jar is a host-file swap, but felis-paper.jar and felis-limbo.jar are baked into the lobby/limbo images and need a rebuild + re-mirror into the in-cluster registry (the installer re-run does both)",
command: "sudo felis setup", command: installerRerun,
}, },
{ {
selector: "mc", selector: "mc",
@@ -219,7 +235,6 @@ func renderApplyGuidance(res updater.Result, selected map[string]bool, force boo
var b strings.Builder var b strings.Builder
var offeredCommand bool var offeredCommand bool
var offeredFelisAPI bool
for _, t := range updateTargets { for _, t := range updateTargets {
if !selected[t.selector] { if !selected[t.selector] {
continue continue
@@ -246,30 +261,50 @@ func renderApplyGuidance(res updater.Result, selected map[string]bool, force boo
// component, and reinstalling the current release is a valid repair action. // component, and reinstalling the current release is a valid repair action.
fmt.Fprintf(&b, " note: cannot tell whether %s is current — its latest version could not be discovered (see above); this reinstalls it either way\n", t.component) fmt.Fprintf(&b, " note: cannot tell whether %s is current — its latest version could not be discovered (see above); this reinstalls it either way\n", t.component)
} }
fmt.Fprintf(&b, " run: %s\n", t.command) fmt.Fprintf(&b, " run: %s\n", strings.ReplaceAll(t.command, "{ref}", installerRef(byComponent)))
offeredCommand = true offeredCommand = true
offeredFelisAPI = offeredFelisAPI || t.component == "felis-api"
} }
// Only explain the command when one was actually offered; a --mc-only run has // Only explain the command when one was actually offered; a --mc-only run has
// nothing to run and the trailer would be a non-sequitur. // nothing to run and the trailer would be a non-sequitur.
if offeredCommand {
b.WriteString("\nfelis setup is idempotent and re-runs the installer that owns these components;\nit does not reinstall what is already current. Restart game servers afterwards.\n")
}
// Scoped to felis-api because it is the only component setup cannot move forward.
// velocity is fine: install_velocity re-resolves the newest build of the pinned minor
// on every run. But setup hands deploy/bootstrap.sh the binary it is itself running
// (FELIS_BOOTSTRAP_BINARY), and that arm skips the release lookup entirely, so it
// rebuilds the image and rolls the deployment from the SAME binary -- a run that looks
// like a successful update and leaves the version unchanged.
// //
// The installer is the only thing that moves felis-api. It is safe to point at now // One trailer serves every selector now: setup is not an apply path at all on a
// that detect_node_ip reuses the installed root domain, so what is left to warn about // completed install (shouldRunHostBootstrapBeforeConfig only enters the host
// is the channel: FELIS_VERSION_BOOTSTRAP is not persisted anywhere and defaults to // bootstrap while an install marker is missing), so the installer re-run is the one
// release, so a bare re-run on a host tracking main quietly moves it onto releases. // worked path for all three components and there is no per-component exception left
// That is a channel change, not a broken install, which is why it is one clause and // to scope. Two caveats stay because following the advice without them bites real
// not a paragraph. // hosts: the channel is not persisted anywhere (a bare re-run on a main host quietly
if offeredFelisAPI { // moves it onto releases), and the private repo's one-liner needs the read token
b.WriteString("\nfelis-api (panel, plugins) is the exception: setup re-images it from the felis binary\nalready on this host, so it cannot install a NEWER felis-api. Re-run the bootstrap\ninstaller for that -- it keeps this install's root domain. It does default to the\nrelease channel, so pass FELIS_VERSION_BOOTSTRAP=dev if this host tracks main.\n") // back in the environment before it can resolve anything.
if offeredCommand {
b.WriteString("\nRe-running the installer applies everything above: it fetches the newest version on\nthe channel in effect and re-applies the bundle (release is the default). The channel\nis not persisted, so pass FELIS_VERSION_BOOTSTRAP=dev if this host tracks main. While\nthis repo is private, the one-liner above 404s without a token; the README's install\nsection has the token'd form that works. felis setup is not this path: on a completed\ninstall it opens the config console and installs nothing newer. Restart game servers\nafterwards.\n")
} }
return b.String() return b.String()
} }
// installerRef is the git ref the installer re-run reads bootstrap.sh from: the newest
// stable felis release when the feed answered, which is the release that script then
// installs; else the release this host runs; main only when neither is a release tag.
func installerRef(byComponent map[string]updates.Action) string {
a, ok := byComponent["felis-api"]
if !ok {
return "main"
}
if a.LatestKnown && isReleaseTag(a.Latest) {
return a.Latest.String()
}
if isReleaseTag(a.Current) {
return a.Current.String()
}
return "main"
}
// isReleaseTag reports whether v was read from a stable vX.Y.Z tag, the only refs
// release.yml publishes a binary for.
func isReleaseTag(v updates.Version) bool {
s := v.String()
if !strings.HasPrefix(s, "v") || v.IsPrerelease() {
return false
}
_, err := updates.Parse(s)
return err == nil
}
+52 -25
View File
@@ -107,7 +107,7 @@ func TestApplyGuidanceMinecraftOffersNoCommand(t *testing.T) {
if !strings.Contains(out, "pinned by policy") { if !strings.Contains(out, "pinned by policy") {
t.Fatalf("want the pin explained:\n%s", out) t.Fatalf("want the pin explained:\n%s", out)
} }
if strings.Contains(out, "run:") || strings.Contains(out, "felis setup is idempotent") { if strings.Contains(out, "run:") || strings.Contains(out, "Re-running the installer") {
t.Fatalf("--mc must offer no command and no command trailer:\n%s", out) t.Fatalf("--mc must offer no command and no command trailer:\n%s", out)
} }
} }
@@ -133,42 +133,69 @@ func TestUpdateTargetsMatchTopology(t *testing.T) {
} }
} }
// `sudo felis setup` is the right answer for velocity and the wrong one for felis-api, // Re-running the installer is the one apply path this table may hand out. setup is NOT an
// so the caveat has to be scoped rather than appended to every run. setup hands // updater on a completed install -- its host-bootstrap phase only runs while an install
// bootstrap the binary it is already running, and that arm skips the release lookup: // marker is missing, so it opens the config console and moves no component -- and even on
// the run rebuilds the image and rolls the deployment off the SAME binary, which looks // the bootstrap path it re-images felis-api from the binary setup is already running. The
// like a successful update and changes nothing. install_velocity, by contrast, really // table used to answer with "sudo felis setup" and scope a felis-api-only exception; both
// does re-resolve the newest build on every run. // taught a model that does not survive contact with an installed host.
func TestApplyGuidanceScopesTheFelisAPICaveat(t *testing.T) { func TestApplyGuidancePointsEveryComponentAtTheInstaller(t *testing.T) {
const caveat = "cannot install a NEWER felis-api"
api := renderApplyGuidance( api := renderApplyGuidance(
planResult([]updates.Action{{Component: "felis-api", Kind: updates.ActionNotify, LatestKnown: true}}), planResult([]updates.Action{{Component: "felis-api", Kind: updates.ActionNotify, LatestKnown: true}}),
map[string]bool{"panel": true}, false) map[string]bool{"panel": true}, false)
if !strings.Contains(api, caveat) { for _, want := range []string{"deploy/bootstrap.sh", "FELIS_VERSION_BOOTSTRAP=dev", "felis setup is not this path"} {
t.Fatalf("--panel resolves to felis-api and must carry the caveat:\n%s", api) if !strings.Contains(api, want) {
t.Fatalf("--panel guidance missing %q:\n%s", want, api)
}
} }
// Naming the installer obliges us to name what a bare re-run still changes. The domain if strings.Contains(api, "run: sudo felis setup") {
// is handled -- detect_node_ip reuses the installed one -- but the channel is not t.Fatalf("setup must never be offered as the apply command:\n%s", api)
// persisted at all and defaults to release, so a host tracking main gets moved onto
// releases by following this advice.
if !strings.Contains(api, "FELIS_VERSION_BOOTSTRAP=dev") {
t.Fatalf("pointing at the installer without the channel caveat misleads a dev host:\n%s", api)
} }
// The same path serves velocity; a scoped caveat would re-teach the old model that
// setup fixes velocity.
vel := renderApplyGuidance( vel := renderApplyGuidance(
planResult([]updates.Action{{Component: "velocity", Kind: updates.ActionNotify, LatestKnown: true}}), planResult([]updates.Action{{Component: "velocity", Kind: updates.ActionNotify, LatestKnown: true}}),
map[string]bool{"velocity": true}, false) map[string]bool{"velocity": true}, false)
if strings.Contains(vel, caveat) { if !strings.Contains(vel, "run: curl -fsSL") || !strings.Contains(vel, "felis setup is not this path") {
t.Fatalf("velocity IS fixed by setup; the caveat would misdirect the operator:\n%s", vel) t.Fatalf("velocity gets the same installer path:\n%s", vel)
}
if !strings.Contains(vel, "felis setup is idempotent") {
t.Fatalf("velocity still wants the ordinary trailer:\n%s", vel)
} }
// --mc offers no command at all, so neither trailer belongs. // --mc offers no command at all, so neither trailer belongs.
mc := renderApplyGuidance(planResult(nil), map[string]bool{"mc": true}, true) mc := renderApplyGuidance(planResult(nil), map[string]bool{"mc": true}, true)
if strings.Contains(mc, caveat) { if strings.Contains(mc, "deploy/bootstrap.sh") || strings.Contains(mc, "FELIS_VERSION_BOOTSTRAP") {
t.Fatalf("--mc offers no command; the caveat is a non-sequitur:\n%s", mc) t.Fatalf("--mc offers no command; the trailer is a non-sequitur:\n%s", mc)
}
}
// The re-run reads bootstrap.sh at the tag whose binary it installs. main can carry
// installer changes no release was tested with.
func TestApplyGuidanceReadsTheInstallerAtTheReleaseTag(t *testing.T) {
v := func(s string) updates.Version {
t.Helper()
out, err := updates.Parse(s)
if err != nil {
t.Fatal(err)
}
return out
}
cases := []struct {
name string
api []updates.Action
want string
}{
{"latest known", []updates.Action{{Component: "felis-api", Kind: updates.ActionNotify, Current: v("v1.3.0"), Latest: v("v1.4.0"), LatestKnown: true}}, "/FelisMC/Felis/v1.4.0/deploy/bootstrap.sh"},
{"latest unknown", []updates.Action{{Component: "felis-api", Kind: updates.ActionNone, Current: v("v1.3.0")}}, "/FelisMC/Felis/v1.3.0/deploy/bootstrap.sh"},
{"prerelease latest", []updates.Action{{Component: "felis-api", Kind: updates.ActionNone, Current: v("v1.3.0"), Latest: v("v1.4.0-rc.1"), LatestKnown: true}}, "/FelisMC/Felis/v1.3.0/deploy/bootstrap.sh"},
{"nothing known", nil, "/FelisMC/Felis/main/deploy/bootstrap.sh"},
}
for _, c := range cases {
out := renderApplyGuidance(planResult(c.api), map[string]bool{"velocity": true, "panel": true}, true)
if !strings.Contains(out, c.want) {
t.Errorf("%s: want %q in:\n%s", c.name, c.want, out)
}
if strings.Contains(out, "{ref}") {
t.Errorf("%s: placeholder left in:\n%s", c.name, out)
}
} }
} }
+269
View File
@@ -0,0 +1,269 @@
package main
import (
"context"
"database/sql"
"errors"
"flag"
"fmt"
"io"
"net"
"os"
"strings"
"time"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/mail"
"felis.lolicon.best/internal/offsite"
"felis.lolicon.best/internal/platform"
"felis.lolicon.best/internal/store"
"felis.lolicon.best/internal/watchdog"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"sigs.k8s.io/controller-runtime/pkg/client"
)
// proxyFor is how long the game proxy may refuse connections before it is
// mailed: a restart takes seconds.
const proxyFor = 3 * time.Minute
// cmdWatchdog runs one pass of the platform watchdog (internal/watchdog): it
// checks the cluster, PostgreSQL, the game proxy, the database backups and the
// host, prints every finding, and mails the platform owners what came due.
// deploy/bootstrap.sh runs it every two minutes from felis-watchdog.timer.
func cmdWatchdog(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("watchdog", flag.ContinueOnError)
fs.SetOutput(stderr)
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (the host copy, which reaches PostgreSQL on 127.0.0.1)")
statePath := fs.String("state", "/var/lib/felis/watchdog/state.json", "state kept between runs (root only: it caches the relay password)")
quietPath := fs.String("quiet-file", "/run/felis/watchdog-quiet-until", "Unix time before which nothing is mailed; the installer writes it while it restarts things on purpose")
backupDir := fs.String("backup-dir", "/var/lib/felis/db-backups", `control-plane database backups to check for freshness ("" skips the check)`)
diskPaths := fs.String("disk-paths", "/,/var/lib/rancher/k3s,/var/lib/postgresql,/var/lib/felis", "comma-separated paths whose filesystems must keep free space")
proxyAddr := fs.String("proxy-addr", "", `game proxy address to dial, e.g. 127.0.0.1:25565 ("" skips the check)`)
controlNS := fs.String("control-namespace", platform.DefaultControlNamespace, "namespace of the control plane")
offsiteStatus := fs.String("offsite-status", offsite.DefaultStatusFile, "the record `felis offsite sync` leaves, checked when [offsite] is configured")
toolsStatus := fs.String("build-tools-status", defaultBuildToolsStatus, "the record `felis mirror-build-tools` leaves, checked when builds scan against the registry's DB copy")
dryRun := fs.Bool("dry-run", false, "print every finding and the mail that is due; send nothing and keep the state as it was")
if err := fs.Parse(args); err != nil {
if errors.Is(err, flag.ErrHelp) {
return 0
}
return 2
}
cfg, err := config.Load(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis watchdog: %v\n", err)
return 1
}
state, err := watchdog.LoadState(*statePath)
if err != nil {
fmt.Fprintf(stderr, "felis watchdog: %v\n", err)
return 1
}
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
now := time.Now()
var report watchdog.Report
add := func(f *watchdog.Finding) {
if f != nil {
report.Findings = append(report.Findings, *f)
}
}
// The cluster: one unreachable API server stands in for every check behind it.
minecraftNS := cfg.K8s.Namespace
if minecraftNS == "" {
minecraftNS = platform.DefaultMinecraftNamespace
}
cl, err := buildSystemServerClient()
var found []watchdog.Finding
if err == nil {
found, err = watchdog.Cluster{Client: cl, ControlNamespace: *controlNS, MinecraftNamespace: minecraftNS}.Check(ctx, now)
}
if err != nil {
f := watchdog.KubeAPIDown(err)
add(&f)
report.Unknown = append(report.Unknown, watchdog.ClusterPrefixes...)
} else {
report.Findings = append(report.Findings, found...)
if cfg.SMTP.Host != "" {
refreshSMTPPassword(ctx, cl, *controlNS, state, stderr)
}
}
if recipients, err := ownerEmails(ctx, cfg.Database.URL); err != nil {
f := watchdog.PostgresDown(err)
add(&f)
} else {
state.Recipients = recipients
}
if *proxyAddr != "" {
add(proxyFinding(ctx, *proxyAddr))
}
if *backupDir != "" {
add(watchdog.BackupFinding(*backupDir, now))
}
if cfg.Offsite.Enabled() {
add(watchdog.OffsiteFinding(*offsiteStatus, now))
}
if usesMirroredScanDB(cfg) {
add(watchdog.ScanDBFinding(*toolsStatus, now))
}
report.Findings = append(report.Findings, watchdog.DiskFindings(splitList(*diskPaths))...)
add(watchdog.MemoryFinding("/proc/meminfo"))
if len(report.Findings) == 0 {
fmt.Fprintln(stdout, "felis watchdog: every check passed")
}
for _, f := range report.Findings {
fmt.Fprintf(stdout, "felis watchdog: [%s] %s: %s\n", f.Severity, f.Key, f.SummaryEN)
}
plan := state.Observe(report, now)
host, _ := os.Hostname()
subject, body := plan.Message(host, now)
if *dryRun {
if plan.Empty() {
fmt.Fprintln(stdout, "felis watchdog: nothing is due to be mailed")
} else {
fmt.Fprintf(stdout, "felis watchdog: due to be mailed to %s:\nSubject: %s\n\n%s", strings.Join(state.Recipients, ", "), subject, strings.ReplaceAll(body, "\r\n", "\n"))
}
return 0
}
save := func() int {
if err := watchdog.SaveState(*statePath, state); err != nil {
fmt.Fprintf(stderr, "felis watchdog: save state: %v\n", err)
return 1
}
return 0
}
if plan.Empty() {
return save()
}
if until := watchdog.QuietUntil(*quietPath); now.Before(until) {
fmt.Fprintf(stdout, "felis watchdog: quiet until %s (installer running); holding this mail: %s\n", until.UTC().Format(time.RFC3339), subject)
return save()
}
switch {
case cfg.SMTP.Host == "":
fmt.Fprintf(stdout, "felis watchdog: no [smtp] relay configured, so this is logged only: %s\n", subject)
case len(state.Recipients) == 0:
fmt.Fprintf(stdout, "felis watchdog: no owner account has a verified email, so this is logged only: %s\n", subject)
default:
if err := sendAlert(ctx, cfg, state, subject, body); err != nil {
// Not committed: the same alerts come due again next run.
fmt.Fprintf(stderr, "felis watchdog: mail %q: %v\n", subject, err)
save()
return 1
}
fmt.Fprintf(stdout, "felis watchdog: mailed %s: %s\n", strings.Join(state.Recipients, ", "), subject)
}
state.Commit(plan, now)
return save()
}
// usesMirroredScanDB reports whether build scans read the vulnerability DB copy
// felis mirror-build-tools keeps in the platform registry: the default, or an
// explicit trivy_db_repository under the registry's mirror/.
func usesMirroredScanDB(cfg *config.Config) bool {
if cfg.Registry.URL == "" {
return false
}
repo := cfg.Registry.TrivyDBRepository
return repo == "" || strings.HasPrefix(repo, cfg.Registry.URL+"/mirror/")
}
// refreshSMTPPassword caches the relay password from the felis-smtp Secret, or
// forgets it when the Secret is gone (a relay without AUTH). An env var named by
// [smtp] password_ref, when set, wins at send time instead.
func refreshSMTPPassword(ctx context.Context, cl client.Client, ns string, state *watchdog.State, stderr io.Writer) {
var sec corev1.Secret
err := cl.Get(ctx, client.ObjectKey{Namespace: ns, Name: platform.SMTPSecretName}, &sec)
switch {
case apierrors.IsNotFound(err):
state.SMTPPassword = ""
case err != nil:
fmt.Fprintf(stderr, "felis watchdog: read %s/%s (keeping the cached relay password): %v\n", ns, platform.SMTPSecretName, err)
default:
state.SMTPPassword = string(sec.Data[platform.SMTPSecretPasswordKey])
}
}
// ownerEmails pings PostgreSQL and returns the verified addresses of the
// enabled owner accounts, the people who can act on an alert.
func ownerEmails(ctx context.Context, url string) ([]string, error) {
ctx, cancel := context.WithTimeout(ctx, 15*time.Second)
defer cancel()
drv, err := store.Open(ctx, url)
if err != nil {
return nil, err
}
defer drv.Close()
rows, err := drv.DB().QueryContext(ctx,
`SELECT email FROM users
WHERE role = 'owner' AND email_verified AND COALESCE(email, '') <> ''
AND NOT disabled AND deleted_at IS NULL
ORDER BY email`)
if err != nil {
return nil, err
}
defer rows.Close()
var out []string
for rows.Next() {
var email sql.NullString
if err := rows.Scan(&email); err != nil {
return nil, err
}
out = append(out, email.String)
}
return out, rows.Err()
}
// proxyFinding dials the game proxy; players reach every server through it.
func proxyFinding(ctx context.Context, addr string) *watchdog.Finding {
d := net.Dialer{Timeout: 5 * time.Second}
conn, err := d.DialContext(ctx, "tcp", addr)
if err == nil {
conn.Close()
return nil
}
return &watchdog.Finding{
Key: "proxy", Severity: watchdog.Critical, For: proxyFor,
Summary: fmt.Sprintf("游戏代理 %s 无法连接:玩家进不了任何服务器", addr),
SummaryEN: fmt.Sprintf("the game proxy at %s refuses connections: players cannot reach any server", addr),
Hint: fmt.Sprintf("systemctl status felis-velocity; journalctl -u felis-velocity -n 200 (%v)", err),
}
}
// sendAlert mails subject/body to every recipient; it fails only when no
// recipient got it.
func sendAlert(ctx context.Context, cfg *config.Config, state *watchdog.State, subject, body string) error {
password := state.SMTPPassword
if ref := cfg.SMTP.PasswordRef; ref != "" && os.Getenv(ref) != "" {
password = os.Getenv(ref)
}
relay := &mail.SMTP{Host: cfg.SMTP.Host, Port: cfg.SMTP.Port, From: cfg.SMTP.From, Username: cfg.SMTP.Username, Password: password}
var errs []error
for _, to := range state.Recipients {
if err := relay.SendNotice(ctx, to, subject, body); err != nil {
errs = append(errs, fmt.Errorf("%s: %w", to, err))
}
}
if len(errs) == len(state.Recipients) {
return errors.Join(errs...)
}
return nil
}
func splitList(s string) []string {
var out []string
for _, p := range strings.Split(s, ",") {
if p = strings.TrimSpace(p); p != "" {
out = append(out, p)
}
}
return out
}
+42
View File
@@ -0,0 +1,42 @@
package main
import (
"context"
"net"
"strings"
"testing"
)
// TestProxyFinding: a listening proxy is healthy; a closed port is the critical
// "players cannot reach any server" finding.
func TestProxyFinding(t *testing.T) {
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
addr := ln.Addr().String()
go func() {
for {
c, err := ln.Accept()
if err != nil {
return
}
c.Close()
}
}()
if f := proxyFinding(context.Background(), addr); f != nil {
t.Fatalf("listening proxy reported: %+v", f)
}
ln.Close()
f := proxyFinding(context.Background(), addr)
if f == nil || f.Key != "proxy" || !strings.Contains(f.SummaryEN, addr) {
t.Fatalf("closed proxy = %+v, want the proxy finding", f)
}
}
func TestSplitList(t *testing.T) {
got := splitList(" /, /var/lib/felis ,,")
if strings.Join(got, "|") != "/|/var/lib/felis" {
t.Fatalf("splitList = %q", got)
}
}
+284
View File
@@ -0,0 +1,284 @@
# Felis alert rules — plain Prometheus format (also the promtool-tested source
# for felis-prometheusrule.yaml). See docs/troubleshooting.md §14 for scraping
# and loading instructions. These are for a deployment that brings its own
# Prometheus; every install already runs `felis watchdog` on the host, which
# checks the same conditions without one and mails the owners (§14).
#
# felis_* series come from these processes:
# - felis-operator-metrics Service :8080 → felis_servers_total, felis_server_phase,
# felis_start_duration_seconds,
# felis_build_info{component="operator"},
# controller_runtime_*, workqueue_*
# - felis-api-internal Service :8081 → felis_image_build_failures_total,
# felis_mail_total, felis_rate_limited_total,
# felis_auth_otp_lockouts_total,
# felis_auth_failures_total,
# felis_audit_write_failures_total,
# felis_build_info{component="api"}
# - node-exporter textfile collector → felis_db_backup_* (felis-db-backup.timer)
# node_* / kube_* series come from node-exporter / kube-state-metrics.
groups:
- name: felis.rules
rules:
- alert: FelisImageBuildFailures
expr: increase(felis_image_build_failures_total[6h]) > 0
for: 5m
labels:
severity: warning
annotations:
summary: "modpack/image build failed in the last 6h"
description: >-
felis_image_build_failures_total increased. Inspect the failed build Job
(kubectl logs -n felis-build job/<build-job>); the same error text is on
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
- alert: FelisSlowServerStarts
expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300
for: 15m
labels:
severity: warning
annotations:
summary: "p90 server start time exceeds 5 minutes"
description: >-
Starts regularly take over five minutes (felis_start_duration_seconds,
observed when readiness is first reached). A start that never completes
records nothing — cross-check desiredState=Running servers with no ready
phase (troubleshooting §1).
- name: felis.platform.rules
rules:
- alert: FelisOperatorDown
expr: absent(felis_build_info{component="operator"})
for: 10m
labels:
severity: critical
annotations:
summary: "felis-operator is down or not scraped"
description: >-
No felis_build_info{component="operator"} series for 10 minutes. Without
the operator no server starts, stops or recovers. Check
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
healthy, the felis-operator-metrics Service is not being scraped
(troubleshooting §14).
- alert: FelisAPIDown
expr: absent(felis_build_info{component="api"})
for: 10m
labels:
severity: critical
annotations:
summary: "felis-api is down or not scraped"
description: >-
No felis_build_info{component="api"} series for 10 minutes. The panel,
sign-in and the proxy's player lookups all go through felis-api. Check
`kubectl -n felis get deploy felis-api` and its log; if the pod is
healthy, the felis-api-internal Service is not being scraped
(troubleshooting §14).
- alert: FelisLoginGateDown
expr: felis_server_phase{role="login",desired="Running",phase!="Running"} == 1
for: 10m
labels:
severity: critical
annotations:
summary: "the login gate {{ $labels.server }} is {{ $labels.phase }}"
description: >-
Every player connection passes through the login server first, so no one
can join. The MinecraftServer's conditions carry the reason:
`kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
(troubleshooting §1, §2).
- alert: FelisSystemServerDown
expr: felis_server_phase{role!="",role!="login",desired="Running",phase!="Running"} == 1
for: 10m
labels:
severity: warning
annotations:
summary: "system server {{ $labels.server }} ({{ $labels.role }}) is {{ $labels.phase }}"
description: >-
Players who sign in are sent to the lobby; while it is down they stay at
the gate. `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
shows the reason (troubleshooting §1, §2).
- alert: FelisServerFailed
expr: felis_server_phase{role="",phase="Failed"} == 1
for: 5m
labels:
severity: warning
annotations:
summary: "server {{ $labels.server }} is Failed"
description: >-
The operator gave up on this server (a crash loop, an image that will not
pull, a world volume that will not mount). Its conditions carry the
reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
(troubleshooting §2).
- alert: FelisReconcileErrors
expr: sum(increase(controller_runtime_reconcile_errors_total{controller="minecraftserver"}[15m])) > 10
for: 5m
labels:
severity: warning
annotations:
summary: "the operator failed over 10 reconciles in 15 minutes"
description: >-
Server changes are being retried instead of applied. The felis-operator
log names each failing server and its error.
- alert: FelisReconcileStuck
expr: max(workqueue_longest_running_processor_seconds{name="minecraftserver"}) > 300
for: 5m
labels:
severity: critical
annotations:
summary: "an operator reconcile has been running for over 5 minutes"
description: >-
Each reconcile is bounded at 3 minutes, so this one is ignoring its
deadline and holding a worker. The liveness probe restarts the operator
once a pass passes 10 minutes; the log from before the restart shows
where it hung.
- name: felis.jobs.rules
rules:
- alert: FelisWorldJobFailed
expr: kube_job_failed{namespace="minecraft",condition="true"} == 1
labels:
severity: warning
annotations:
summary: "Job {{ $labels.job_name }} failed"
description: >-
A world backup, restore or reaper run failed; after a failed backup that
world's newest archive is older than planned.
`kubectl -n minecraft logs job/{{ $labels.job_name }}` has the error
(troubleshooting §10).
- alert: FelisReaperStale
expr: time() - kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"} > 26 * 3600
for: 10m
labels:
severity: warning
annotations:
summary: "the world reaper has not succeeded in over 26h"
description: >-
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
lists its runs, and the newest one's log
shows why (troubleshooting §10).
- name: felis.node.rules
rules:
- alert: FelisNodeDiskSpaceLow
expr: >-
node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
/ node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15
for: 15m
labels:
severity: warning
annotations:
summary: "node filesystem {{ $labels.mountpoint }} below 15% available"
description: >-
Sustained disk pressure evicts game pods and garbage-collects images
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
- alert: FelisNodeDiskPressure
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
for: 5m
labels:
severity: critical
annotations:
summary: "kubelet reports DiskPressure on {{ $labels.node }}"
description: >-
The eviction chain is in progress: control-plane pods hold
system-cluster-critical and survive, game pods do not. Free disk now
(troubleshooting §13b).
- alert: FelisNodeMemoryLow
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
for: 15m
labels:
severity: warning
annotations:
summary: "node memory available below 10% for 15m"
description: >-
PostgreSQL, the control plane, the registry and game servers share one
node; sustained memory pressure risks OOM kills.
- name: felis.backup.rules
rules:
- alert: FelisDBBackupStale
expr: time() - max(felis_db_backup_last_success_timestamp_seconds) > 26 * 3600
for: 10m
labels:
severity: critical
annotations:
summary: "no control-plane database backup in over 26h"
description: >-
felis-db-backup.timer runs daily; the newest bundle is more than a day
old. Read `journalctl -u felis-db-backup` on the host, then take one now
with `sudo felis db backup` (troubleshooting §16).
- alert: FelisDBBackupMetricMissing
expr: absent(felis_db_backup_last_success_timestamp_seconds)
for: 2h
labels:
severity: warning
annotations:
summary: "database backup freshness is not being scraped"
description: >-
No felis_db_backup_last_success_timestamp_seconds series, so
FelisDBBackupStale cannot fire. Point node-exporter's
--collector.textfile.directory at the directory of
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
(troubleshooting §16).
- name: felis.auth.rules
rules:
- alert: FelisMailBudgetExhausted
expr: sum(increase(felis_mail_total{result="throttled"}[15m])) > 0
labels:
severity: warning
annotations:
summary: "the install-wide mail budget refused mail"
description: >-
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
spent, and every sign-in code is refused with 429 mail_rate_limited until
it refills. Check felis_rate_limited_total for a flood before raising the
budget (troubleshooting §17).
- alert: FelisMailDeliveryFailing
expr: sum(increase(felis_mail_total{result="failed"}[15m])) > 0
labels:
severity: warning
annotations:
summary: "the SMTP relay refused mail in the last 15m"
description: >-
felis_mail_total{result="failed"} increased: sign-in codes are not being
delivered (502 mail_undeliverable). The relay's reason is in the
felis-api log (troubleshooting §17).
- alert: FelisSignInFlood
expr: sum(rate(felis_rate_limited_total{scope="auth_door"}[5m])) * 60 > 10
for: 10m
labels:
severity: warning
annotations:
summary: "sign-in doors refusing over 10 requests a minute"
description: >-
The per-address sign-in limit has been refusing callers for 10 minutes.
A script is hammering the auth doors; if real users report rate_limited
at once instead, [auth] client_ip_header is missing and everyone shares
the proxy's address (troubleshooting §17).
- alert: FelisOTPAccountLocked
expr: sum by (purpose) (increase(felis_auth_otp_lockouts_total[1h])) > 0
labels:
severity: warning
annotations:
summary: "an account's email-code sign-in locked after 10 wrong codes"
description: >-
Someone entered 10 wrong codes for one account within 24h ({{ $labels.purpose }}).
The audit log names the account (action auth.otp.locked); the owner was
mailed. Unless they fumbled codes, someone is guessing at it
(troubleshooting §17).
- alert: FelisSignInFailures
expr: sum(increase(felis_auth_failures_total[15m])) > 30
for: 5m
labels:
severity: warning
annotations:
summary: "over 30 refused sign-ins in 15 minutes"
description: >-
Wrong codes, unknown addresses or bad passkey assertions well above people
mistyping: someone is guessing or enumerating. `sum by (door, reason)
(increase(felis_auth_failures_total[15m]))` shows where; the audit rows
(action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
- alert: FelisAuditWriteFailing
expr: increase(felis_audit_write_failures_total[15m]) > 0
labels:
severity: warning
annotations:
summary: "felis-api failed to write audit rows"
description: >-
The actions went through but their audit rows were lost. The felis-api log
names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
being unreachable or out of disk.
+509
View File
@@ -0,0 +1,509 @@
# promtool unit tests: `promtool test rules felis-alerts_test.yml`
# Proves every shipped rule actually fires on its target condition (and stays
# silent before it).
rule_files:
- felis-alerts.yaml
evaluation_interval: 1m
tests:
- name: build failure and slow starts
interval: 1m
input_series:
# counter: quiet for 5m, then one failure per step.
- series: 'felis_image_build_failures_total'
values: '0x5 1x15'
# histogram: all observations land in the (300,600] bucket.
- series: 'felis_start_duration_seconds_bucket{le="120"}'
values: '0x22'
- series: 'felis_start_duration_seconds_bucket{le="300"}'
values: '0x22'
- series: 'felis_start_duration_seconds_bucket{le="600"}'
values: '0+10x21'
- series: 'felis_start_duration_seconds_bucket{le="+Inf"}'
values: '0+10x21'
alert_rule_test:
- eval_time: 2m
alertname: FelisImageBuildFailures
exp_alerts: []
- eval_time: 20m
alertname: FelisImageBuildFailures
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "modpack/image build failed in the last 6h"
description: >-
felis_image_build_failures_total increased. Inspect the failed build Job
(kubectl logs -n felis-build job/<build-job>); the same error text is on
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
- eval_time: 20m
alertname: FelisSlowServerStarts
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "p90 server start time exceeds 5 minutes"
description: >-
Starts regularly take over five minutes (felis_start_duration_seconds,
observed when readiness is first reached). A start that never completes
records nothing — cross-check desiredState=Running servers with no ready
phase (troubleshooting §1).
- name: node disk and memory thresholds
interval: 1m
input_series:
- series: 'node_filesystem_avail_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
values: '10x26'
- series: 'node_filesystem_size_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
values: '100x26'
- series: 'kube_node_status_condition{condition="DiskPressure",node="n1",status="true"}'
values: '0x4 1x22'
- series: 'node_memory_MemAvailable_bytes{instance="node1",job="node-exporter"}'
values: '5x26'
- series: 'node_memory_MemTotal_bytes{instance="node1",job="node-exporter"}'
values: '100x26'
alert_rule_test:
- eval_time: 2m
alertname: FelisNodeDiskPressure
exp_alerts: []
- eval_time: 20m
alertname: FelisNodeDiskSpaceLow
exp_alerts:
- exp_labels:
device: /dev/vda1
fstype: xfs
instance: node1
job: node-exporter
mountpoint: /
severity: warning
exp_annotations:
summary: "node filesystem / below 15% available"
description: >-
Sustained disk pressure evicts game pods and garbage-collects images
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
- eval_time: 20m
alertname: FelisNodeDiskPressure
exp_alerts:
- exp_labels:
condition: DiskPressure
node: n1
status: "true"
severity: critical
exp_annotations:
summary: "kubelet reports DiskPressure on n1"
description: >-
The eviction chain is in progress: control-plane pods hold
system-cluster-critical and survive, game pods do not. Free disk now
(troubleshooting §13b).
- eval_time: 20m
alertname: FelisNodeMemoryLow
exp_alerts:
- exp_labels:
instance: node1
job: node-exporter
severity: warning
exp_annotations:
summary: "node memory available below 10% for 15m"
description: >-
PostgreSQL, the control plane, the registry and game servers share one
node; sustained memory pressure risks OOM kills.
- name: database backup freshness
interval: 1m
input_series:
# The newest bundle was taken at t=0 and none since.
- series: 'felis_db_backup_last_success_timestamp_seconds{instance="node1",job="node-exporter",label="daily"}'
values: '0x1630'
alert_rule_test:
- eval_time: 25h
alertname: FelisDBBackupStale
exp_alerts: []
- eval_time: 27h
alertname: FelisDBBackupStale
exp_alerts:
- exp_labels:
severity: critical
exp_annotations:
summary: "no control-plane database backup in over 26h"
description: >-
felis-db-backup.timer runs daily; the newest bundle is more than a day
old. Read `journalctl -u felis-db-backup` on the host, then take one now
with `sudo felis db backup` (troubleshooting §16).
- eval_time: 27h
alertname: FelisDBBackupMetricMissing
exp_alerts: []
- name: database backup freshness not scraped
interval: 1m
input_series:
- series: 'up{job="node-exporter"}'
values: '1x200'
alert_rule_test:
- eval_time: 1h
alertname: FelisDBBackupMetricMissing
exp_alerts: []
- eval_time: 3h
alertname: FelisDBBackupMetricMissing
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "database backup freshness is not being scraped"
description: >-
No felis_db_backup_last_success_timestamp_seconds series, so
FelisDBBackupStale cannot fire. Point node-exporter's
--collector.textfile.directory at the directory of
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
(troubleshooting §16).
- name: sign-in mail budget and relay
interval: 1m
input_series:
# Created at zero on start; the budget refuses one mail at t=3m.
- series: 'felis_mail_total{kind="otp",result="throttled",job="felis-api"}'
values: '0 0 0 1x30'
- series: 'felis_mail_total{kind="otp",result="failed",job="felis-api"}'
values: '0x33'
alert_rule_test:
- eval_time: 2m
alertname: FelisMailBudgetExhausted
exp_alerts: []
- eval_time: 5m
alertname: FelisMailBudgetExhausted
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "the install-wide mail budget refused mail"
description: >-
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
spent, and every sign-in code is refused with 429 mail_rate_limited until
it refills. Check felis_rate_limited_total for a flood before raising the
budget (troubleshooting §17).
- eval_time: 5m
alertname: FelisMailDeliveryFailing
exp_alerts: []
- name: sign-in flood
interval: 1m
input_series:
# 30 refusals a minute from t=0; a lone refused script at 2/min stays quiet.
- series: 'felis_rate_limited_total{scope="auth_door",job="felis-api"}'
values: '0+30x40'
alert_rule_test:
- eval_time: 10m
alertname: FelisSignInFlood
exp_alerts: []
- eval_time: 20m
alertname: FelisSignInFlood
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "sign-in doors refusing over 10 requests a minute"
description: >-
The per-address sign-in limit has been refusing callers for 10 minutes.
A script is hammering the auth doors; if real users report rate_limited
at once instead, [auth] client_ip_header is missing and everyone shares
the proxy's address (troubleshooting §17).
- name: sign-in trickle stays quiet
interval: 1m
input_series:
- series: 'felis_rate_limited_total{scope="auth_door",job="felis-api"}'
values: '0+2x40'
alert_rule_test:
- eval_time: 30m
alertname: FelisSignInFlood
exp_alerts: []
- name: account email-code lock
interval: 1m
input_series:
- series: 'felis_auth_otp_lockouts_total{purpose="login_email",job="felis-api"}'
values: '0 0 1x90'
- series: 'felis_auth_otp_lockouts_total{purpose="op_login",job="felis-api"}'
values: '0x92'
alert_rule_test:
- eval_time: 1m
alertname: FelisOTPAccountLocked
exp_alerts: []
- eval_time: 10m
alertname: FelisOTPAccountLocked
exp_alerts:
- exp_labels:
severity: warning
purpose: login_email
exp_annotations:
summary: "an account's email-code sign-in locked after 10 wrong codes"
description: >-
Someone entered 10 wrong codes for one account within 24h (login_email).
The audit log names the account (action auth.otp.locked); the owner was
mailed. Unless they fumbled codes, someone is guessing at it
(troubleshooting §17).
- eval_time: 90m
alertname: FelisOTPAccountLocked
exp_alerts: []
- name: sign-in failure rate
interval: 1m
input_series:
# Two doors failing at 3/min between them from t=0.
- series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}'
values: '0+2x40'
- series: 'felis_auth_failures_total{door="op_login",reason="no_account",job="felis-api"}'
values: '0+1x40'
alert_rule_test:
- eval_time: 8m
alertname: FelisSignInFailures
exp_alerts: []
- eval_time: 25m
alertname: FelisSignInFailures
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "over 30 refused sign-ins in 15 minutes"
description: >-
Wrong codes, unknown addresses or bad passkey assertions well above people
mistyping: someone is guessing or enumerating. `sum by (door, reason)
(increase(felis_auth_failures_total[15m]))` shows where; the audit rows
(action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
- name: people mistyping stays quiet
interval: 1m
input_series:
- series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}'
values: '0 0 1 1 2 2 3 3 4 4 5x30'
alert_rule_test:
- eval_time: 30m
alertname: FelisSignInFailures
exp_alerts: []
- name: audit rows lost
interval: 1m
input_series:
- series: 'felis_audit_write_failures_total{job="felis-api",instance="api-0"}'
values: '0 0 0 2x20'
alert_rule_test:
- eval_time: 2m
alertname: FelisAuditWriteFailing
exp_alerts: []
- eval_time: 5m
alertname: FelisAuditWriteFailing
exp_alerts:
- exp_labels:
severity: warning
job: felis-api
instance: api-0
exp_annotations:
summary: "felis-api failed to write audit rows"
description: >-
The actions went through but their audit rows were lost. The felis-api log
names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
being unreachable or out of disk.
- name: operator and api presence
interval: 1m
input_series:
- series: 'felis_build_info{component="operator",version="v1",job="felis-operator",instance="op-0"}'
values: '1x30'
# felis-api stops being scraped after 5m; the series goes stale 5m later.
- series: 'felis_build_info{component="api",version="v1",job="felis-api",instance="api-0"}'
values: '1x5'
alert_rule_test:
- eval_time: 25m
alertname: FelisOperatorDown
exp_alerts: []
- eval_time: 15m
alertname: FelisAPIDown
exp_alerts: []
- eval_time: 25m
alertname: FelisAPIDown
exp_alerts:
- exp_labels:
severity: critical
component: api
exp_annotations:
summary: "felis-api is down or not scraped"
description: >-
No felis_build_info{component="api"} series for 10 minutes. The panel,
sign-in and the proxy's player lookups all go through felis-api. Check
`kubectl -n felis get deploy felis-api` and its log; if the pod is
healthy, the felis-api-internal Service is not being scraped
(troubleshooting §14).
- name: operator never scraped
interval: 1m
input_series:
- series: 'felis_build_info{component="api",version="v1",job="felis-api",instance="api-0"}'
values: '1x30'
alert_rule_test:
- eval_time: 5m
alertname: FelisOperatorDown
exp_alerts: []
- eval_time: 15m
alertname: FelisOperatorDown
exp_alerts:
- exp_labels:
severity: critical
component: operator
exp_annotations:
summary: "felis-operator is down or not scraped"
description: >-
No felis_build_info{component="operator"} series for 10 minutes. Without
the operator no server starts, stops or recovers. Check
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
healthy, the felis-operator-metrics Service is not being scraped
(troubleshooting §14).
- name: system and user servers down
interval: 1m
input_series:
- series: 'felis_server_phase{server="login",role="login",phase="Starting",desired="Running"}'
values: '1x20'
- series: 'felis_server_phase{server="lobby",role="lobby",phase="Failed",desired="Running"}'
values: '1x20'
# Stopped on purpose: not an outage.
- series: 'felis_server_phase{server="lobby2",role="lobby",phase="Stopped",desired="Stopped"}'
values: '1x20'
# A user server carries no role label (the operator publishes role="").
- series: 'felis_server_phase{server="survival",phase="Failed",desired="Running"}'
values: '1x20'
- series: 'felis_server_phase{server="creative",phase="Running",desired="Running"}'
values: '1x20'
alert_rule_test:
- eval_time: 9m
alertname: FelisLoginGateDown
exp_alerts: []
- eval_time: 11m
alertname: FelisLoginGateDown
exp_alerts:
- exp_labels:
severity: critical
server: login
role: login
phase: Starting
desired: Running
exp_annotations:
summary: "the login gate login is Starting"
description: >-
Every player connection passes through the login server first, so no one
can join. The MinecraftServer's conditions carry the reason:
`kubectl -n minecraft describe minecraftserver login`
(troubleshooting §1, §2).
- eval_time: 11m
alertname: FelisSystemServerDown
exp_alerts:
- exp_labels:
severity: warning
server: lobby
role: lobby
phase: Failed
desired: Running
exp_annotations:
summary: "system server lobby (lobby) is Failed"
description: >-
Players who sign in are sent to the lobby; while it is down they stay at
the gate. `kubectl -n minecraft describe minecraftserver lobby`
shows the reason (troubleshooting §1, §2).
- eval_time: 3m
alertname: FelisServerFailed
exp_alerts: []
- eval_time: 6m
alertname: FelisServerFailed
exp_alerts:
- exp_labels:
severity: warning
server: survival
phase: Failed
desired: Running
exp_annotations:
summary: "server survival is Failed"
description: >-
The operator gave up on this server (a crash loop, an image that will not
pull, a world volume that will not mount). Its conditions carry the
reason: `kubectl -n minecraft describe minecraftserver survival`
(troubleshooting §2).
- name: operator reconcile errors and a stuck pass
interval: 1m
input_series:
# Two failed reconciles a minute from 6m on.
- series: 'controller_runtime_reconcile_errors_total{controller="minecraftserver",job="felis-operator"}'
values: '0x5 0+2x20'
# One pass that started at 5m and never returns.
- series: 'workqueue_longest_running_processor_seconds{name="minecraftserver",controller="minecraftserver",job="felis-operator"}'
values: '0x5 60+60x20'
alert_rule_test:
- eval_time: 5m
alertname: FelisReconcileErrors
exp_alerts: []
- eval_time: 25m
alertname: FelisReconcileErrors
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "the operator failed over 10 reconciles in 15 minutes"
description: >-
Server changes are being retried instead of applied. The felis-operator
log names each failing server and its error.
- eval_time: 12m
alertname: FelisReconcileStuck
exp_alerts: []
- eval_time: 20m
alertname: FelisReconcileStuck
exp_alerts:
- exp_labels:
severity: critical
exp_annotations:
summary: "an operator reconcile has been running for over 5 minutes"
description: >-
Each reconcile is bounded at 3 minutes, so this one is ignoring its
deadline and holding a worker. The liveness probe restarts the operator
once a pass passes 10 minutes; the log from before the restart shows
where it hung.
- name: world job failures and a late reaper
interval: 1m
input_series:
- series: 'kube_job_failed{namespace="minecraft",job_name="backup-survival-abc",condition="true"}'
values: '0x2 1x10'
- series: 'kube_job_failed{namespace="minecraft",job_name="backup-survival-abc",condition="false"}'
values: '1x2 0x10'
# A build Job in another namespace is FelisImageBuildFailures' business.
- series: 'kube_job_failed{namespace="felis-build",job_name="build-x",condition="true"}'
values: '1x12'
# Evaluation starts at the epoch, so "over a day ago" is a negative timestamp.
- series: 'kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"}'
values: '-100000x30'
alert_rule_test:
- eval_time: 1m
alertname: FelisWorldJobFailed
exp_alerts: []
- eval_time: 5m
alertname: FelisWorldJobFailed
exp_alerts:
- exp_labels:
severity: warning
namespace: minecraft
job_name: backup-survival-abc
condition: "true"
exp_annotations:
summary: "Job backup-survival-abc failed"
description: >-
A world backup, restore or reaper run failed; after a failed backup that
world's newest archive is older than planned.
`kubectl -n minecraft logs job/backup-survival-abc` has the error
(troubleshooting §10).
- eval_time: 5m
alertname: FelisReaperStale
exp_alerts: []
- eval_time: 15m
alertname: FelisReaperStale
exp_alerts:
- exp_labels:
severity: warning
namespace: minecraft
cronjob: felis-reaper
exp_annotations:
summary: "the world reaper has not succeeded in over 26h"
description: >-
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
lists its runs, and the newest one's log
shows why (troubleshooting §10).
- name: a reaper that ran yesterday stays quiet
interval: 1m
input_series:
- series: 'kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"}'
values: '-50000x30'
alert_rule_test:
- eval_time: 25m
alertname: FelisReaperStale
exp_alerts: []
+278
View File
@@ -0,0 +1,278 @@
# prometheus-operator twin of felis-alerts.yaml (kube-prometheus-stack loads
# rules through the PrometheusRule CRD, not rule_files). The plain file is the
# promtool-tested source; keep the groups in sync.
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: felis-alerts
namespace: monitoring
labels:
# Change to match your stack's ruleSelector (kube-prometheus-stack's
# default selects on the Helm release name).
release: kube-prometheus-stack
spec:
groups:
- name: felis.rules
rules:
- alert: FelisImageBuildFailures
expr: increase(felis_image_build_failures_total[6h]) > 0
for: 5m
labels:
severity: warning
annotations:
summary: "modpack/image build failed in the last 6h"
description: >-
felis_image_build_failures_total increased. Inspect the failed build Job
(kubectl logs -n felis-build job/<build-job>); the same error text is on
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
- alert: FelisSlowServerStarts
expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300
for: 15m
labels:
severity: warning
annotations:
summary: "p90 server start time exceeds 5 minutes"
description: >-
Starts regularly take over five minutes (felis_start_duration_seconds,
observed when readiness is first reached). A start that never completes
records nothing — cross-check desiredState=Running servers with no ready
phase (troubleshooting §1).
- name: felis.platform.rules
rules:
- alert: FelisOperatorDown
expr: absent(felis_build_info{component="operator"})
for: 10m
labels:
severity: critical
annotations:
summary: "felis-operator is down or not scraped"
description: >-
No felis_build_info{component="operator"} series for 10 minutes. Without
the operator no server starts, stops or recovers. Check
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
healthy, the felis-operator-metrics Service is not being scraped
(troubleshooting §14).
- alert: FelisAPIDown
expr: absent(felis_build_info{component="api"})
for: 10m
labels:
severity: critical
annotations:
summary: "felis-api is down or not scraped"
description: >-
No felis_build_info{component="api"} series for 10 minutes. The panel,
sign-in and the proxy's player lookups all go through felis-api. Check
`kubectl -n felis get deploy felis-api` and its log; if the pod is
healthy, the felis-api-internal Service is not being scraped
(troubleshooting §14).
- alert: FelisLoginGateDown
expr: felis_server_phase{role="login",desired="Running",phase!="Running"} == 1
for: 10m
labels:
severity: critical
annotations:
summary: "the login gate {{ $labels.server }} is {{ $labels.phase }}"
description: >-
Every player connection passes through the login server first, so no one
can join. The MinecraftServer's conditions carry the reason:
`kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
(troubleshooting §1, §2).
- alert: FelisSystemServerDown
expr: felis_server_phase{role!="",role!="login",desired="Running",phase!="Running"} == 1
for: 10m
labels:
severity: warning
annotations:
summary: "system server {{ $labels.server }} ({{ $labels.role }}) is {{ $labels.phase }}"
description: >-
Players who sign in are sent to the lobby; while it is down they stay at
the gate. `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
shows the reason (troubleshooting §1, §2).
- alert: FelisServerFailed
expr: felis_server_phase{role="",phase="Failed"} == 1
for: 5m
labels:
severity: warning
annotations:
summary: "server {{ $labels.server }} is Failed"
description: >-
The operator gave up on this server (a crash loop, an image that will not
pull, a world volume that will not mount). Its conditions carry the
reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
(troubleshooting §2).
- alert: FelisReconcileErrors
expr: sum(increase(controller_runtime_reconcile_errors_total{controller="minecraftserver"}[15m])) > 10
for: 5m
labels:
severity: warning
annotations:
summary: "the operator failed over 10 reconciles in 15 minutes"
description: >-
Server changes are being retried instead of applied. The felis-operator
log names each failing server and its error.
- alert: FelisReconcileStuck
expr: max(workqueue_longest_running_processor_seconds{name="minecraftserver"}) > 300
for: 5m
labels:
severity: critical
annotations:
summary: "an operator reconcile has been running for over 5 minutes"
description: >-
Each reconcile is bounded at 3 minutes, so this one is ignoring its
deadline and holding a worker. The liveness probe restarts the operator
once a pass passes 10 minutes; the log from before the restart shows
where it hung.
- name: felis.jobs.rules
rules:
- alert: FelisWorldJobFailed
expr: kube_job_failed{namespace="minecraft",condition="true"} == 1
labels:
severity: warning
annotations:
summary: "Job {{ $labels.job_name }} failed"
description: >-
A world backup, restore or reaper run failed; after a failed backup that
world's newest archive is older than planned.
`kubectl -n minecraft logs job/{{ $labels.job_name }}` has the error
(troubleshooting §10).
- alert: FelisReaperStale
expr: time() - kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"} > 26 * 3600
for: 10m
labels:
severity: warning
annotations:
summary: "the world reaper has not succeeded in over 26h"
description: >-
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
lists its runs, and the newest one's log
shows why (troubleshooting §10).
- name: felis.node.rules
rules:
- alert: FelisNodeDiskSpaceLow
expr: >-
node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
/ node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15
for: 15m
labels:
severity: warning
annotations:
summary: "node filesystem {{ $labels.mountpoint }} below 15% available"
description: >-
Sustained disk pressure evicts game pods and garbage-collects images
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
- alert: FelisNodeDiskPressure
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
for: 5m
labels:
severity: critical
annotations:
summary: "kubelet reports DiskPressure on {{ $labels.node }}"
description: >-
The eviction chain is in progress: control-plane pods hold
system-cluster-critical and survive, game pods do not. Free disk now
(troubleshooting §13b).
- alert: FelisNodeMemoryLow
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
for: 15m
labels:
severity: warning
annotations:
summary: "node memory available below 10% for 15m"
description: >-
PostgreSQL, the control plane, the registry and game servers share one
node; sustained memory pressure risks OOM kills.
- name: felis.backup.rules
rules:
- alert: FelisDBBackupStale
expr: time() - max(felis_db_backup_last_success_timestamp_seconds) > 26 * 3600
for: 10m
labels:
severity: critical
annotations:
summary: "no control-plane database backup in over 26h"
description: >-
felis-db-backup.timer runs daily; the newest bundle is more than a day
old. Read `journalctl -u felis-db-backup` on the host, then take one now
with `sudo felis db backup` (troubleshooting §16).
- alert: FelisDBBackupMetricMissing
expr: absent(felis_db_backup_last_success_timestamp_seconds)
for: 2h
labels:
severity: warning
annotations:
summary: "database backup freshness is not being scraped"
description: >-
No felis_db_backup_last_success_timestamp_seconds series, so
FelisDBBackupStale cannot fire. Point node-exporter's
--collector.textfile.directory at the directory of
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
(troubleshooting §16).
- name: felis.auth.rules
rules:
- alert: FelisMailBudgetExhausted
expr: sum(increase(felis_mail_total{result="throttled"}[15m])) > 0
labels:
severity: warning
annotations:
summary: "the install-wide mail budget refused mail"
description: >-
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
spent, and every sign-in code is refused with 429 mail_rate_limited until
it refills. Check felis_rate_limited_total for a flood before raising the
budget (troubleshooting §17).
- alert: FelisMailDeliveryFailing
expr: sum(increase(felis_mail_total{result="failed"}[15m])) > 0
labels:
severity: warning
annotations:
summary: "the SMTP relay refused mail in the last 15m"
description: >-
felis_mail_total{result="failed"} increased: sign-in codes are not being
delivered (502 mail_undeliverable). The relay's reason is in the
felis-api log (troubleshooting §17).
- alert: FelisSignInFlood
expr: sum(rate(felis_rate_limited_total{scope="auth_door"}[5m])) * 60 > 10
for: 10m
labels:
severity: warning
annotations:
summary: "sign-in doors refusing over 10 requests a minute"
description: >-
The per-address sign-in limit has been refusing callers for 10 minutes.
A script is hammering the auth doors; if real users report rate_limited
at once instead, [auth] client_ip_header is missing and everyone shares
the proxy's address (troubleshooting §17).
- alert: FelisOTPAccountLocked
expr: sum by (purpose) (increase(felis_auth_otp_lockouts_total[1h])) > 0
labels:
severity: warning
annotations:
summary: "an account's email-code sign-in locked after 10 wrong codes"
description: >-
Someone entered 10 wrong codes for one account within 24h ({{ $labels.purpose }}).
The audit log names the account (action auth.otp.locked); the owner was
mailed. Unless they fumbled codes, someone is guessing at it
(troubleshooting §17).
- alert: FelisSignInFailures
expr: sum(increase(felis_auth_failures_total[15m])) > 30
for: 5m
labels:
severity: warning
annotations:
summary: "over 30 refused sign-ins in 15 minutes"
description: >-
Wrong codes, unknown addresses or bad passkey assertions well above people
mistyping: someone is guessing or enumerating. `sum by (door, reason)
(increase(felis_auth_failures_total[15m]))` shows where; the audit rows
(action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
- alert: FelisAuditWriteFailing
expr: increase(felis_audit_write_failures_total[15m]) > 0
labels:
severity: warning
annotations:
summary: "felis-api failed to write audit rows"
description: >-
The actions went through but their audit rows were lost. The felis-api log
names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
being unreachable or out of disk.
+1517 -132
View File
File diff suppressed because it is too large. Load diff
+1356 -15
View File
File diff suppressed because it is too large. Load diff
@@ -123,10 +123,6 @@ spec:
lifecycle: lifecycle:
description: Lifecycle tunes graceful shutdown (spec §7). description: Lifecycle tunes graceful shutdown (spec §7).
properties: properties:
preStopSaveAndStop:
description: PreStopSaveAndStop enables the operator-injected
RCON save+stop preStop.
type: boolean
terminationGracePeriodSeconds: terminationGracePeriodSeconds:
description: TerminationGracePeriodSeconds is the pod grace period description: TerminationGracePeriodSeconds is the pod grace period
(default 300). (default 300).
@@ -347,6 +343,14 @@ spec:
- type - type
type: object type: object
type: array type: array
emptySince:
description: |-
EmptySince is when the operator first observed 0 online players during
a Running phase (spec §8 idle auto-stop). It is reset when a player joins
or the server stops, so the empty-duration counter starts fresh each time
the server becomes unoccupied.
format: date-time
type: string
endpoint: endpoint:
description: Endpoint is where the proxy should route traffic. description: Endpoint is where the proxy should route traffic.
properties: properties:
+25 -91
View File
@@ -1,30 +1,28 @@
#!/bin/bash #!/bin/bash
# demo-up.sh — one-shot Felis demo bring-up. # demo-up.sh — one-shot Felis demo bring-up.
# #
# Collapses the four manual steps (bootstrap -> build/import limbo+lobby images ->
# edit felis.toml -> felis setup) into a single command:
#
# sudo bash deploy/demo-up.sh # sudo bash deploy/demo-up.sh
# #
# It ends by exec'ing the interactive `felis setup` TUI (create the Owner account) — # Every piece a demo box needs — base platform (k3s + felis + docker + cloudflared +
# that human step is the only thing this script cannot do for you. # control plane), the limbo/lobby/paper images, the felis-velocity proxy plugin, and
# the [velocity] wiring in felis.host.toml — is built by deploy/bootstrap.sh. This
# wrapper adds only the one step the installer cannot do: the interactive
# `felis setup` TUI that creates the Owner account.
# #
# Image source, in order of preference: # It used to rebuild the game images here with its own copy of that logic, written
# 1. Prebuilt tars at deploy/images/felis-limbo.tar + felis-lobby.tar (imported as-is). # before bootstrap grew the job. The copy drifted: it pinned Paper 1.21.8 while the
# 2. Otherwise built on this host with docker, resolving the LOOHP/Limbo CI jar and # installer derives one version from the Limbo login gate (both hops of a login must
# the latest stable Paper jar automatically. Override any of: # speak one protocol), never built felis-velocity.jar (so the proxy it wired had
# LIMBO_JAR_URL LIMBO_SCHEM_URL LIMBO_VERSION PAPER_JAR_URL PAPER_JAR_SHA256 # nowhere to route), and left docker running. The installer is the single origin.
# PAPER_MC_VERSION
# #
# Toggles: SKIP_BOOTSTRAP=1 (base already up), SKIP_SETUP=1 (stop before the TUI). # Toggles: SKIP_BOOTSTRAP=1 (base + game stack already installed by a full bootstrap),
# SKIP_SETUP=1 (stop before the TUI).
set -Eeuo pipefail set -Eeuo pipefail
SRC_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" SRC_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
STATE_DIR=/etc/felis STATE_DIR=/etc/felis
HOST_TOML="$STATE_DIR/felis.host.toml" HOST_TOML="$STATE_DIR/felis.host.toml"
IMG_DIR="$SRC_DIR/deploy/images" PLUGIN_JAR=/opt/felis/velocity/plugins/felis-velocity.jar
LIMBO_IMAGE="felis-limbo:demo"
LOBBY_IMAGE="felis-lobby:demo"
K3S=/usr/local/bin/k3s K3S=/usr/local/bin/k3s
FELIS=/usr/local/bin/felis FELIS=/usr/local/bin/felis
@@ -33,11 +31,11 @@ die() { printf '\033[1;31mERROR: %s\033[0m\n' "$*" >&2; exit 1; }
[ "$(id -u)" -eq 0 ] || die "run as root (sudo bash deploy/demo-up.sh)" [ "$(id -u)" -eq 0 ] || die "run as root (sudo bash deploy/demo-up.sh)"
# 1. base platform (k3s + felis + docker + cloudflared + control plane) ---------- # 1. base platform + full game stack (deploy/bootstrap.sh) -----------------------
if [ "${SKIP_BOOTSTRAP:-0}" = 1 ]; then if [ "${SKIP_BOOTSTRAP:-0}" = 1 ]; then
log "SKIP_BOOTSTRAP=1 — assuming the base platform is already up" log "SKIP_BOOTSTRAP=1 — assuming the base platform is already up"
else else
log "bringing up the base platform (deploy/bootstrap.sh)" log "bringing up the base platform and the game stack (deploy/bootstrap.sh)"
bash "$SRC_DIR/deploy/bootstrap.sh" bash "$SRC_DIR/deploy/bootstrap.sh"
fi fi
@@ -46,83 +44,19 @@ command -v "$K3S" >/dev/null 2>&1 || K3S=k3s
command -v "$K3S" >/dev/null 2>&1 || die "k3s not found — did bootstrap complete?" command -v "$K3S" >/dev/null 2>&1 || die "k3s not found — did bootstrap complete?"
command -v "$FELIS" >/dev/null 2>&1 || die "felis not found — did bootstrap complete?" command -v "$FELIS" >/dev/null 2>&1 || die "felis not found — did bootstrap complete?"
# 2. get the two game images into k3s containerd -------------------------------- # SKIP_BOOTSTRAP=1 trusts an earlier run to be complete. Check that it actually left
if [ -f "$IMG_DIR/felis-limbo.tar" ] && [ -f "$IMG_DIR/felis-lobby.tar" ]; then # the full stack behind: a base from before the game-stack installer, or one whose
log "importing prebuilt image tars from $IMG_DIR" # pieces were pruned by hand, must fail here with a pointer — not present as a proxy
"$K3S" ctr images import "$IMG_DIR/felis-limbo.tar" # that accepts logins and routes nowhere, with nothing in any log to say why.
"$K3S" ctr images import "$IMG_DIR/felis-lobby.tar"
# Optional: the plain-Paper recommended base, if a tar was staged for it.
[ -f "$IMG_DIR/felis-paper.tar" ] && "$K3S" ctr images import "$IMG_DIR/felis-paper.tar"
else
log "no prebuilt tars in $IMG_DIR — building on this host with docker"
command -v docker >/dev/null 2>&1 || die "docker not found; cannot build images"
rel=$(curl -fsSL --max-time 30 "https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/api/json" \
| grep -oE 'target/Limbo-[0-9][^"]+\.jar' | head -1) || true
: "${LIMBO_JAR_URL:=https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/artifact/$rel}"
: "${LIMBO_SCHEM_URL:=https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/artifact/spawn.schem}"
: "${LIMBO_VERSION:=$(basename "$rel" | sed -E 's/^Limbo-//; s/\.jar$//; s/-[0-9]+\.[0-9]+$//')}"
[ -n "$rel" ] || [ -n "${LIMBO_JAR_URL##*artifact/}" ] || die "could not resolve the Limbo jar; set LIMBO_JAR_URL"
log "building $LIMBO_IMAGE (Limbo $LIMBO_VERSION)"
docker build -f "$SRC_DIR/deploy/limbo/Dockerfile" \
--build-arg LIMBO_JAR_URL="$LIMBO_JAR_URL" \
--build-arg LIMBO_SCHEM_URL="$LIMBO_SCHEM_URL" \
--build-arg LIMBO_VERSION="$LIMBO_VERSION" \
-t "$LIMBO_IMAGE" "$SRC_DIR"
docker save "$LIMBO_IMAGE" | "$K3S" ctr images import -
: "${PAPER_MC_VERSION:=1.21.8}"
: "${PAPER_JAR_URL:=$(curl -fsSL --max-time 30 "https://fill.papermc.io/v3/projects/paper/versions/${PAPER_MC_VERSION}/builds/latest" | grep -oE 'https://fill-data\.papermc\.io/[^"]+\.jar' | head -1)}"
[ -n "$PAPER_JAR_URL" ] || die "could not resolve the Paper jar; set PAPER_JAR_URL"
# Both Dockerfiles require the jar's digest. The fill-data URL is content-addressed
# (the objects/ path segment IS the sha256), so it is derived rather than asked for;
# a mirror override carries no such segment and must bring its own digest.
if [ -z "${PAPER_JAR_SHA256:-}" ]; then
sha="${PAPER_JAR_URL#*/objects/}"
sha="${sha%%/*}"
case "$sha" in
*[!0-9a-f]*|"") sha="" ;;
esac
if [ "${#sha}" -ne 64 ]; then
die "cannot derive the Paper jar sha256 from PAPER_JAR_URL (not a content-addressed fill-data URL); set PAPER_JAR_SHA256"
fi
PAPER_JAR_SHA256="$sha"
fi
log "building $LOBBY_IMAGE (Paper $PAPER_MC_VERSION)"
docker build -f "$SRC_DIR/deploy/lobby/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
--build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
-t "$LOBBY_IMAGE" "$SRC_DIR"
docker save "$LOBBY_IMAGE" | "$K3S" ctr images import -
# Plain Paper recommended base — same PAPER_JAR_URL, no plugins, no secret gate.
: "${PAPER_IMAGE:=felis-paper:demo}"
log "building $PAPER_IMAGE (plain Paper $PAPER_MC_VERSION, forwarding via the operator initContainer)"
docker build -f "$SRC_DIR/deploy/paper/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
--build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
-t "$PAPER_IMAGE" "$SRC_DIR"
docker save "$PAPER_IMAGE" | "$K3S" ctr images import -
fi
# 3. wire the images into the config `felis setup` reads ------------------------
log "wiring [velocity] images into $HOST_TOML"
[ -f "$HOST_TOML" ] || die "missing $HOST_TOML — did bootstrap run?" [ -f "$HOST_TOML" ] || die "missing $HOST_TOML — did bootstrap run?"
if grep -q '^\[velocity\]' "$HOST_TOML"; then grep -q '^\[velocity\]' "$HOST_TOML" \
echo " [velocity] table already present — leaving it untouched" || die "$HOST_TOML has no [velocity] section — re-run the installer without SKIP_BOOTSTRAP so the system servers get wired"
else [ -f "$PLUGIN_JAR" ] \
cat >> "$HOST_TOML" <<EOF || die "$PLUGIN_JAR missing — this base did not finish the full installer, and a proxy without it silently routes nothing; re-run the installer without SKIP_BOOTSTRAP"
[velocity] # 2. interactive Owner creation + system-server provisioning --------------------
login_image = "$LIMBO_IMAGE"
lobby_image = "$LOBBY_IMAGE"
EOF
echo " appended login_image=$LIMBO_IMAGE / lobby_image=$LOBBY_IMAGE"
fi
# 4. interactive Owner creation + system-server provisioning -------------------
if [ "${SKIP_SETUP:-0}" = 1 ]; then if [ "${SKIP_SETUP:-0}" = 1 ]; then
log "SKIP_SETUP=1 — base + images + config ready. Finish with: sudo felis setup" log "SKIP_SETUP=1 — base + game stack ready. Finish with: sudo felis setup"
else else
log "launching 'felis setup' — create the Owner account (this is the only interactive step)" log "launching 'felis setup' — create the Owner account (this is the only interactive step)"
exec "$FELIS" setup exec "$FELIS" setup
+23
View File
@@ -0,0 +1,23 @@
# The upstream builds this release installs. deploy/bootstrap.sh reads this file (the
# default FELIS_GAME_STACK=pinned), downloads exactly these artifacts and refuses any whose
# sha256 differs, so every host installing one release gets the same login gate, lobby,
# plain-Paper image and proxy, and a rerun rebuilds nothing that did not change.
#
# MC_VERSION is the protocol the whole stack speaks: Limbo speaks exactly one, and Paper
# follows it so a client that passes the login gate can also reach the lobby.
#
# Refresh with deploy/update-game-stack-lock.sh, which resolves upstream's newest builds and
# hashes them. Plain KEY=value lines only; bootstrap reads it without evaluating it.
MC_VERSION=26.3
LIMBO_VERSION=2026.0.3-ALPHA
LIMBO_JAR_URL=https://ci.loohpjames.com/job/Limbo/76/artifact/target/Limbo-2026.0.3-ALPHA-26.3.jar
LIMBO_JAR_SHA256=a2de91fcaa2255aed8a111b8786f370213423c7798eee778c00d016d83aa65d2
LIMBO_SCHEM_URL=https://ci.loohpjames.com/job/Limbo/76/artifact/spawn.schem
LIMBO_SCHEM_SHA256=70c85dae2db157971ef513e318820c5a5e10a96b813b70e21f8af2b632cbcbc7
PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/49399919246cbf443efc8507447dc948eb7477c41be560b0e87e2a455aff824a/paper-26.3-40.jar
PAPER_JAR_SHA256=49399919246cbf443efc8507447dc948eb7477c41be560b0e87e2a455aff824a
LUCKPERMS_JAR_URL=https://download.luckperms.net/1672/bukkit/loader/LuckPerms-Bukkit-5.5.85.jar
LUCKPERMS_JAR_SHA256=dc637ce18f48d3b75a7ffd1784b85be16a627359090dfe5adf15dab6d145dc7d
VELOCITY_VERSION=3.5.1
VELOCITY_JAR_URL=https://fill-data.papermc.io/v1/objects/b4e3164df5377346854dc6cb9e6a78022b1946ff69e89676313f5f6f1c6f0fb3/velocity-3.5.1-615.jar
VELOCITY_JAR_SHA256=b4e3164df5377346854dc6cb9e6a78022b1946ff69e89676313f5f6f1c6f0fb3
+27 -7
View File
@@ -10,10 +10,11 @@
# There is no bundled server.properties — Limbo writes a default on first run. # There is no bundled server.properties — Limbo writes a default on first run.
# So the runtime is assembled from those two URLs (not a zip) via --build-arg: # So the runtime is assembled from those two URLs (not a zip) via --build-arg:
# #
# . <(grep '^LIMBO_' deploy/game-stack.lock)
# docker build -f deploy/limbo/Dockerfile \ # docker build -f deploy/limbo/Dockerfile \
# --build-arg LIMBO_JAR_URL=https://ci.loohpjames.com/job/Limbo/<n>/artifact/target/Limbo-<ver>.jar \ # --build-arg LIMBO_JAR_URL="$LIMBO_JAR_URL" --build-arg LIMBO_JAR_SHA256="$LIMBO_JAR_SHA256" \
# --build-arg LIMBO_SCHEM_URL=https://ci.loohpjames.com/job/Limbo/<n>/artifact/spawn.schem \ # --build-arg LIMBO_SCHEM_URL="$LIMBO_SCHEM_URL" --build-arg LIMBO_SCHEM_SHA256="$LIMBO_SCHEM_SHA256" \
# --build-arg LIMBO_VERSION=<maven-api-version> \ # --build-arg LIMBO_VERSION="$LIMBO_VERSION" \
# -t felis-limbo:demo . # -t felis-limbo:demo .
# #
# Note LIMBO_VERSION (the maven API version the plugin compiles against, e.g. # Note LIMBO_VERSION (the maven API version the plugin compiles against, e.g.
@@ -33,7 +34,7 @@
# JDK 17 fails to read them with "wrong version 65.0, should be 61.0". The image # JDK 17 fails to read them with "wrong version 65.0, should be 61.0". The image
# also provides the `gradle` binary (this tree vendors no Gradle wrapper). # also provides the `gradle` binary (this tree vendors no Gradle wrapper).
# build.gradle still targets release 17 bytecode so the plugin loads on Java 17+. # build.gradle still targets release 17 bytecode so the plugin loads on Java 17+.
FROM gradle:8.14-jdk21 AS plugin FROM gradle:8.14-jdk21@sha256:5c4c0c4284de4a19951e82ac78f86dbcda2e136644bbfe159beba7ea3420cc80 AS plugin
WORKDIR /src WORKDIR /src
# Copy what the limbo module needs: its own tree plus the shared link core it # Copy what the limbo module needs: its own tree plus the shared link core it
# srcDir-includes (../shared/src/main/java → /src/plugins/shared/src/main/java), so # srcDir-includes (../shared/src/main/java → /src/plugins/shared/src/main/java), so
@@ -50,21 +51,33 @@ RUN cd plugins/limbo \
# 21-jre: the Limbo jar is Java 21 bytecode (class-file major 65), so a Java 17 # 21-jre: the Limbo jar is Java 21 bytecode (class-file major 65), so a Java 17
# JRE cannot run it (UnsupportedClassVersionError). A 21 JRE also runs the # JRE cannot run it (UnsupportedClassVersionError). A 21 JRE also runs the
# plugin's release-17 bytecode fine. # plugin's release-17 bytecode fine.
FROM eclipse-temurin:21-jre FROM eclipse-temurin:21-jre@sha256:49e21e16e3c86eb7816a44a67549910ed090fbeb40c29c525d58bf5e02e91b0f
ARG LIMBO_JAR_URL ARG LIMBO_JAR_URL
ARG LIMBO_JAR_SHA256
ARG LIMBO_SCHEM_URL ARG LIMBO_SCHEM_URL
ARG LIMBO_SCHEM_SHA256
WORKDIR /limbo WORKDIR /limbo
# Pull the two loose LOOHP/Limbo CI artifacts: the server jar (required, saved as # Pull the two loose LOOHP/Limbo CI artifacts: the server jar (required, saved as
# Limbo.jar) and the default spawn schematic (optional). Fail loudly if the jar # Limbo.jar) and the default spawn schematic (optional). Each is checked against the
# URL was not supplied. # digest deploy/game-stack.lock names (bootstrap.sh passes it): the login gate is the
# first thing every player's connection reaches, and Limbo's CI publishes no digest of
# its own.
RUN set -eu; \ RUN set -eu; \
if [ -z "${LIMBO_JAR_URL:-}" ]; then \ if [ -z "${LIMBO_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg LIMBO_JAR_URL=<Limbo server jar> is required" >&2; exit 1; \ echo "ERROR: --build-arg LIMBO_JAR_URL=<Limbo server jar> is required" >&2; exit 1; \
fi; \ fi; \
if [ -z "${LIMBO_JAR_SHA256:-}" ]; then \
echo "ERROR: --build-arg LIMBO_JAR_SHA256=<Limbo jar sha256> is required" >&2; exit 1; \
fi; \
if [ -n "${LIMBO_SCHEM_URL:-}" ] && [ -z "${LIMBO_SCHEM_SHA256:-}" ]; then \
echo "ERROR: --build-arg LIMBO_SCHEM_SHA256=<spawn.schem sha256> is required with LIMBO_SCHEM_URL" >&2; exit 1; \
fi; \
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \ apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
curl -fSL "$LIMBO_JAR_URL" -o /limbo/Limbo.jar; \ curl -fSL "$LIMBO_JAR_URL" -o /limbo/Limbo.jar; \
echo "$LIMBO_JAR_SHA256 /limbo/Limbo.jar" | sha256sum -c; \
if [ -n "${LIMBO_SCHEM_URL:-}" ]; then \ if [ -n "${LIMBO_SCHEM_URL:-}" ]; then \
curl -fSL "$LIMBO_SCHEM_URL" -o /limbo/spawn.schem; \ curl -fSL "$LIMBO_SCHEM_URL" -o /limbo/spawn.schem; \
echo "$LIMBO_SCHEM_SHA256 /limbo/spawn.schem" | sha256sum -c; \
fi; \ fi; \
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \ apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \
mkdir -p /limbo/plugins mkdir -p /limbo/plugins
@@ -78,6 +91,13 @@ COPY deploy/limbo/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
# The operator mounts the world PVC at /data. Runtime state lives there; /limbo # The operator mounts the world PVC at /data. Runtime state lives there; /limbo
# remains the immutable image seed copied into the volume by the entrypoint. # remains the immutable image seed copied into the volume by the entrypoint.
WORKDIR /data WORKDIR /data
# Run as the game uid (naming.GameUID in the Go tree). The operator pins the same uid in
# the pod securityContext whatever USER an image declares; declaring it here as well
# keeps a plain `docker run` of this image off root, and chowning the empty /data seed
# lets that run write its world. The jar seed above stays root-owned and read-only to
# the server.
RUN chown 1000:1000 /data
USER 1000:1000
ENV FELIS_HEALTH_PORT=8080 ENV FELIS_HEALTH_PORT=8080
# FELIS_GAME_PORT is the port the entrypoint pins Limbo to; it MUST equal the operator's # FELIS_GAME_PORT is the port the entrypoint pins Limbo to; it MUST equal the operator's
+18 -8
View File
@@ -90,14 +90,21 @@ docker build -f deploy/limbo/Dockerfile \
version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the
maven repo). maven repo).
Import into k3s and point config at it: Publish it into the cluster's registry and point config at it. On the node
itself (docker treats `127.0.0.1` as insecure by default):
``` ```
docker save felis-limbo:demo | sudo k3s ctr images import - docker tag felis-limbo:demo 127.0.0.1:5000/felis/limbo:demo
# felis.toml → [velocity] login_image = "felis-limbo:demo" docker push 127.0.0.1:5000/felis/limbo:demo
# felis.toml → [velocity] login_image = "registry.felis.svc:5000/felis/limbo:demo"
sudo felis setup sudo felis setup
``` ```
The registry keys a repository by the path after the host, so pushing through a
`kubectl -n felis port-forward svc/registry 5000:5000` from another machine is
equivalent. Hosting the image in the registry (rather than only importing it
into containerd) is what lets kubelet re-pull it after an image GC.
## Ports (handled for you) ## Ports (handled for you)
The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to
@@ -136,11 +143,14 @@ set them by hand:
pod's internal port 8081. That Service is deliberately separate from the external pod's internal port 8081. That Service is deliberately separate from the external
NodePort `felis-api` (443) so the no-Zero-Trust internal face is never published on NodePort `felis-api` (443) so the no-Zero-Trust internal face is never published on
a node's external IP. a node's external IP.
- **NetworkPolicy:** none is required today — neither the minecraft-namespace egress - **NetworkPolicy:** the minecraft namespace is egress-locked
nor the control-namespace ingress is policy-locked, so the login pod's call to the (`felis-server-egress`: DNS plus the public internet, every private range
API internal port is reachable. If a future deployment adds a minecraft egress lock excluded), so the internal API is unreachable from a game server by default.
or a control-namespace ingress fence, it must also open the login-pod → `felis-login-to-internal-api` opens exactly the login pod → felis-api (8081) path,
felis-api-internal (8081) path. selecting on the reserved `login` name AND the setup-owned
`felis.lolicon.best/system-role=login` label the operator copies onto the pod — the
same pair that decides who receives `FELIS_SERVICE_TOKEN`, so a user server cannot
match it by picking a name.
The Velocity gate/lobby wiring is printed by `felis setup` and enforces the The Velocity gate/lobby wiring is printed by `felis setup` and enforces the
invariant: fresh connections hit `login` first, and only an authenticated release invariant: fresh connections hit `login` first, and only an authenticated release
+24 -11
View File
@@ -5,13 +5,13 @@
# the POST-auth /menu hub: it is reached only when the login gate transfers an # the POST-auth /menu hub: it is reached only when the login gate transfers an
# authenticated player onward, and it must never be a fallback target. # authenticated player onward, and it must never be a fallback target.
# #
# Build (deploy/bootstrap.sh does this for you; the PAPER_JAR_URL comes from PaperMC's # Build (deploy/bootstrap.sh does this for you, with the URLs and digests
# Fill v3 API — api.papermc.io v2 has returned HTTP 410 since 2026-07-01): # deploy/game-stack.lock names):
# . <(grep -E '^(PAPER|LUCKPERMS)_' deploy/game-stack.lock)
# docker build -f deploy/lobby/Dockerfile \ # docker build -f deploy/lobby/Dockerfile \
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-26.2-<build>.jar \ # --build-arg PAPER_JAR_URL="$PAPER_JAR_URL" --build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
# --build-arg PAPER_JAR_SHA256=<that same sha — the objects/ path segment> \ # --build-arg LUCKPERMS_JAR_URL="$LUCKPERMS_JAR_URL" \
# --build-arg LUCKPERMS_JAR_URL="$(curl -fsSL https://metadata.luckperms.net/data/all \ # --build-arg LUCKPERMS_JAR_SHA256="$LUCKPERMS_JAR_SHA256" \
# | grep -o 'https://download.luckperms.net/[^"]*/bukkit/loader/[^"]*\.jar')" \
# -t felis-lobby:demo . # -t felis-lobby:demo .
# docker save felis-lobby:demo | sudo k3s ctr images import - # docker save felis-lobby:demo | sudo k3s ctr images import -
# # felis.toml → [velocity] lobby_image = "felis-lobby:demo" # # felis.toml → [velocity] lobby_image = "felis-lobby:demo"
@@ -28,7 +28,7 @@
# ---- build the felis-paper plugin jar (Paper API is Java 21) ---- # ---- build the felis-paper plugin jar (Paper API is Java 21) ----
# gradle:8.14-jdk21 — an official Gradle image on JDK 21 (this tree vendors no Gradle # gradle:8.14-jdk21 — an official Gradle image on JDK 21 (this tree vendors no Gradle
# wrapper, and a bare JDK image ships no `gradle`). JDK 21 matches the Paper API. # wrapper, and a bare JDK image ships no `gradle`). JDK 21 matches the Paper API.
FROM gradle:8.14-jdk21 AS plugin FROM gradle:8.14-jdk21@sha256:5c4c0c4284de4a19951e82ac78f86dbcda2e136644bbfe159beba7ea3420cc80 AS plugin
WORKDIR /src WORKDIR /src
COPY plugins/paper/ ./plugins/paper/ COPY plugins/paper/ ./plugins/paper/
COPY plugins/shared/ ./plugins/shared/ COPY plugins/shared/ ./plugins/shared/
@@ -41,7 +41,7 @@ RUN cd plugins/paper \
# 25-jre, not 21: Paper 26.2 declares `java.version.minimum = 25` (PaperMC Fill v3, # 25-jre, not 21: Paper 26.2 declares `java.version.minimum = 25` (PaperMC Fill v3,
# GET /v3/projects/paper/versions/26.2) and refuses to boot on anything older. A 25 JRE # GET /v3/projects/paper/versions/26.2) and refuses to boot on anything older. A 25 JRE
# also runs the plugin's Java-21 bytecode, so only the runtime moves. # also runs the plugin's Java-21 bytecode, so only the runtime moves.
FROM eclipse-temurin:25-jre FROM eclipse-temurin:25-jre@sha256:bb036ed6cfdc57e3da7c22634d15f1b840d2caf76183861c80e81ca4b5104abb
ARG PAPER_JAR_URL ARG PAPER_JAR_URL
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces # Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
# that shape at build time. Checking the digest after the download turns a truncated or # that shape at build time. Checking the digest after the download turns a truncated or
@@ -51,10 +51,12 @@ ARG PAPER_JAR_SHA256
# (internal/api/handlers_access.go) issues `lp user ...` over RCON, so a lobby built # (internal/api/handlers_access.go) issues `lp user ...` over RCON, so a lobby built
# without it answers every grant with "Unknown command" — a failure the operator only # without it answers every grant with "Unknown command" — a failure the operator only
# discovers in production, because the server itself starts and runs perfectly well. # discovers in production, because the server itself starts and runs perfectly well.
# Failing the build is the cheap place to notice. Resolved by URL rather than pinned # Failing the build is the cheap place to notice. Passed in rather than pinned here for
# here for the same reason PAPER_JAR_URL is: bootstrap.sh asks upstream for the current # the same reason PAPER_JAR_URL is: deploy/game-stack.lock names the build, so this file
# build, so this file does not go stale on every LuckPerms release. # does not change on every LuckPerms release. The digest is required like Paper's; the
# jar runs inside the lobby with the server's full permissions.
ARG LUCKPERMS_JAR_URL ARG LUCKPERMS_JAR_URL
ARG LUCKPERMS_JAR_SHA256
WORKDIR /paper WORKDIR /paper
RUN set -eu; \ RUN set -eu; \
if [ -z "${PAPER_JAR_URL:-}" ]; then \ if [ -z "${PAPER_JAR_URL:-}" ]; then \
@@ -66,11 +68,15 @@ RUN set -eu; \
if [ -z "${LUCKPERMS_JAR_URL:-}" ]; then \ if [ -z "${LUCKPERMS_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg LUCKPERMS_JAR_URL=<luckperms bukkit jar> is required" >&2; exit 1; \ echo "ERROR: --build-arg LUCKPERMS_JAR_URL=<luckperms bukkit jar> is required" >&2; exit 1; \
fi; \ fi; \
if [ -z "${LUCKPERMS_JAR_SHA256:-}" ]; then \
echo "ERROR: --build-arg LUCKPERMS_JAR_SHA256=<luckperms jar sha256> is required" >&2; exit 1; \
fi; \
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \ apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
mkdir -p /paper/plugins; \ mkdir -p /paper/plugins; \
curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \ curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \
echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \ echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \
curl -fSL "$LUCKPERMS_JAR_URL" -o /paper/plugins/LuckPerms.jar; \ curl -fSL "$LUCKPERMS_JAR_URL" -o /paper/plugins/LuckPerms.jar; \
echo "$LUCKPERMS_JAR_SHA256 /paper/plugins/LuckPerms.jar" | sha256sum -c; \
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \ apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \
echo "eula=true" > /paper/eula.txt echo "eula=true" > /paper/eula.txt
COPY --from=plugin /felis-paper.jar /paper/plugins/felis-paper.jar COPY --from=plugin /felis-paper.jar /paper/plugins/felis-paper.jar
@@ -82,6 +88,13 @@ COPY deploy/lobby/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
# The operator mounts the world PVC at /data. Runtime state lives there; /paper # The operator mounts the world PVC at /data. Runtime state lives there; /paper
# remains the immutable image seed copied into the volume by the entrypoint. # remains the immutable image seed copied into the volume by the entrypoint.
WORKDIR /data WORKDIR /data
# Run as the game uid (naming.GameUID in the Go tree). The operator pins the same uid in
# the pod securityContext whatever USER an image declares; declaring it here as well
# keeps a plain `docker run` of this image off root, and chowning the empty /data seed
# lets that run write its world. The jar seed above stays root-owned and read-only to
# the server.
RUN chown 1000:1000 /data
USER 1000:1000
# FELIS_GAME_PORT is the port the entrypoint pins Paper to; it MUST equal the operator's # FELIS_GAME_PORT is the port the entrypoint pins Paper to; it MUST equal the operator's
# GamePort (internal/operator/builders.go). Default 25565 — override only in lockstep # GamePort (internal/operator/builders.go). Default 25565 — override only in lockstep
+6 -2
View File
@@ -31,8 +31,12 @@ docker build -f deploy/lobby/Dockerfile \
--build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \ --build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \
--build-arg PAPER_JAR_SHA256=<sha256 of that jar> \ --build-arg PAPER_JAR_SHA256=<sha256 of that jar> \
-t felis-lobby:demo . -t felis-lobby:demo .
docker save felis-lobby:demo | sudo k3s ctr images import - # Publish into the cluster's registry (on the node; docker treats 127.0.0.1 as
# felis.toml → [velocity] lobby_image = "felis-lobby:demo" # insecure by default — or through a `kubectl -n felis port-forward svc/registry
# 5000:5000`, which is equivalent: only the path after the host matters).
docker tag felis-lobby:demo 127.0.0.1:5000/felis/lobby:demo
docker push 127.0.0.1:5000/felis/lobby:demo
# felis.toml → [velocity] lobby_image = "registry.felis.svc:5000/felis/lobby:demo"
sudo felis setup sudo felis setup
``` ```
+13 -6
View File
@@ -5,7 +5,7 @@
# internal/store/migrations/0019_recommended_paper.sql). It is NOT a system server: it # internal/store/migrations/0019_recommended_paper.sql). It is NOT a system server: it
# carries no felis-paper /menu plugin, no LuckPerms, and no forwarding-secret gate. # carries no felis-paper /menu plugin, no LuckPerms, and no forwarding-secret gate.
# #
# It writes NO Velocity forwarding config itself. The operator injects a root # It writes NO Velocity forwarding config itself. The operator injects a
# `felis init-forwarding` initContainer into every USER server (internal/operator/ # `felis init-forwarding` initContainer into every USER server (internal/operator/
# builders.go: buildStatefulSet) that writes config/paper-global.yml + server.properties # builders.go: buildStatefulSet) that writes config/paper-global.yml + server.properties
# online-mode=false onto the /data PVC before this container starts. That external step is # online-mode=false onto the /data PVC before this container starts. That external step is
@@ -14,11 +14,11 @@
# image a user brings is made joinable the same way. If the initContainer is absent (no # image a user brings is made joinable the same way. If the initContainer is absent (no
# FELIS_IMAGE configured) Paper boots as a standalone online server: degraded, not broken. # FELIS_IMAGE configured) Paper boots as a standalone online server: degraded, not broken.
# #
# Build (deploy/bootstrap.sh does this for you; PAPER_JAR_URL comes from PaperMC's Fill v3 # Build (deploy/bootstrap.sh does this for you, with the SAME Paper build the lobby uses —
# API — the SAME url the lobby build resolves, so this reuses it and adds no new dependency): # deploy/game-stack.lock names it — so this adds no new dependency):
# . <(grep '^PAPER_' deploy/game-stack.lock)
# docker build -f deploy/paper/Dockerfile \ # docker build -f deploy/paper/Dockerfile \
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-<ver>-<build>.jar \ # --build-arg PAPER_JAR_URL="$PAPER_JAR_URL" --build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
# --build-arg PAPER_JAR_SHA256=<that same sha — the objects/ path segment> \
# -t felis-paper:demo . # -t felis-paper:demo .
# docker save felis-paper:demo | sudo k3s ctr images import - # docker save felis-paper:demo | sudo k3s ctr images import -
# # felis.toml → recommended via 0019_recommended_paper.sql (no [velocity] key points here) # # felis.toml → recommended via 0019_recommended_paper.sql (no [velocity] key points here)
@@ -29,7 +29,7 @@
# 25-jre, not 21: Paper 26.2 declares java.version.minimum=25 (PaperMC Fill v3) and refuses # 25-jre, not 21: Paper 26.2 declares java.version.minimum=25 (PaperMC Fill v3) and refuses
# to boot on anything older. # to boot on anything older.
FROM eclipse-temurin:25-jre FROM eclipse-temurin:25-jre@sha256:bb036ed6cfdc57e3da7c22634d15f1b840d2caf76183861c80e81ca4b5104abb
ARG PAPER_JAR_URL ARG PAPER_JAR_URL
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces # Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
# that shape at build time. Checking the digest after the download turns a truncated or # that shape at build time. Checking the digest after the download turns a truncated or
@@ -54,6 +54,13 @@ COPY deploy/paper/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
# on the PVC. /paper stays the immutable image seed: the jar is never copied onto the # on the PVC. /paper stays the immutable image seed: the jar is never copied onto the
# volume, so the panel file editor (which sees only /data) cannot tamper with it. # volume, so the panel file editor (which sees only /data) cannot tamper with it.
WORKDIR /data WORKDIR /data
# Run as the game uid (naming.GameUID in the Go tree). The operator pins the same uid in
# the pod securityContext whatever USER an image declares; declaring it here as well
# keeps a plain `docker run` of this image off root, and chowning the empty /data seed
# lets that run write its world. The jar seed above stays root-owned and read-only to
# the server.
RUN chown 1000:1000 /data
USER 1000:1000
# FELIS_GAME_PORT is the port the entrypoint pins Paper to; it MUST equal the operator's # FELIS_GAME_PORT is the port the entrypoint pins Paper to; it MUST equal the operator's
# GamePort (internal/operator/builders.go). Default 25565 — override only in lockstep with # GamePort (internal/operator/builders.go). Default 25565 — override only in lockstep with
+1 -1
View File
@@ -3,7 +3,7 @@
# #
# A plain Paper backend for a user's OWN world — NOT a system server. Unlike deploy/limbo # A plain Paper backend for a user's OWN world — NOT a system server. Unlike deploy/limbo
# and deploy/lobby it writes no Velocity forwarding config and has no secret gate: the # and deploy/lobby it writes no Velocity forwarding config and has no secret gate: the
# operator injects a root `felis init-forwarding` initContainer that writes # operator injects a `felis init-forwarding` initContainer that writes
# config/paper-global.yml + server.properties online-mode=false onto /data BEFORE this # config/paper-global.yml + server.properties online-mode=false onto /data BEFORE this
# container starts, so forwarding is configured externally and this stays a drop-in Paper # container starts, so forwarding is configured externally and this stays a drop-in Paper
# image. With no initContainer (no FELIS_IMAGE) Paper just boots standalone-online — # image. With no initContainer (no FELIS_IMAGE) Paper just boots standalone-online —
+73
View File
@@ -0,0 +1,73 @@
#!/usr/bin/env bash
# Refreshes deploy/game-stack.lock to upstream's newest builds, hashed.
#
# bash deploy/update-game-stack-lock.sh # rewrite the lock in place
# bash deploy/update-game-stack-lock.sh --check # exit 1 if upstream moved on
#
# It runs the same resolver bootstrap.sh uses for FELIS_GAME_STACK=latest (the functions are
# lifted out of bootstrap.sh, so the two cannot drift), then pins Velocity's newest build of
# VELOCITY_LATEST_MINOR. Limbo and LuckPerms publish no digest, so their jars are downloaded
# and hashed here; Paper and Velocity come from Fill's content-addressed URLs.
#
# Review the diff before committing: MC_VERSION moves the login gate's protocol, and the
# lobby, plain-Paper image and every client follow it.
set -Eeuo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
BS="${here}/bootstrap.sh"
LOCK="${here}/game-stack.lock"
check=0
case "${1:-}" in
--check) check=1 ;;
"") ;;
*) printf 'usage: %s [--check]\n' "$0" >&2; exit 2 ;;
esac
log() { printf '[lock] %s\n' "$*" >&2; }
ok() { printf '[ ok ] %s\n' "$*" >&2; }
warn() { :; }
die() { printf '[fail] %s\n' "$*" >&2; exit 1; }
lift() { # function-name
local body
body="$(awk -v f="$1" '$0 ~ "^" f "\\(\\) \\{" {on=1} on {print} on && /^}/ {exit}' "$BS")"
[ -n "$body" ] || die "bootstrap.sh no longer defines $1"
eval "$body"
}
for fn in meta_get papermc_latest_jar luckperms_latest_jar url_sha256 resolve_latest_game_jars; do
lift "$fn"
done
eval "$(grep '^VELOCITY_LATEST_MINOR=' "$BS")"
[ -n "${VELOCITY_LATEST_MINOR:-}" ] || die "bootstrap.sh no longer sets VELOCITY_LATEST_MINOR"
resolve_latest_game_jars
log "resolving the newest Velocity ${VELOCITY_LATEST_MINOR} build"
velocity="$(papermc_latest_jar velocity "$VELOCITY_LATEST_MINOR")" \
|| die "no Velocity build for ${VELOCITY_LATEST_MINOR}"
# shellcheck disable=SC2034 # read back through ${!key} below
VELOCITY_VERSION="$VELOCITY_LATEST_MINOR"
# shellcheck disable=SC2034
VELOCITY_JAR_URL="${velocity% *}"
# shellcheck disable=SC2034
VELOCITY_JAR_SHA256="${velocity##* }"
tmp="$(mktemp)"
trap 'rm -f "$tmp"' EXIT
# The comment header is kept as it is; only the KEY=value lines are regenerated.
sed -n '/^#/p;/^#/!q' "$LOCK" > "$tmp"
for key in MC_VERSION LIMBO_VERSION LIMBO_JAR_URL LIMBO_JAR_SHA256 LIMBO_SCHEM_URL LIMBO_SCHEM_SHA256 \
PAPER_JAR_URL PAPER_JAR_SHA256 LUCKPERMS_JAR_URL LUCKPERMS_JAR_SHA256 \
VELOCITY_VERSION VELOCITY_JAR_URL VELOCITY_JAR_SHA256; do
printf '%s=%s\n' "$key" "${!key}" >> "$tmp"
done
if cmp -s "$tmp" "$LOCK"; then
ok "game-stack.lock already pins upstream's newest builds"
exit 0
fi
diff -u "$LOCK" "$tmp" >&2 || true
if [ "$check" = 1 ]; then
die "upstream has newer builds than game-stack.lock"
fi
cp "$tmp" "$LOCK"
ok "game-stack.lock updated; run go test . and the bootstrap tests, then commit"
+42 -14
View File
@@ -42,9 +42,20 @@ A grep across `*.md` and `*.go` returns both sets; only the Go ones are seams.
does not exist. `felis update` runs with a zero window, under which every does not exist. `felis update` runs with a zero window, under which every
`Scheduled` component degrades to a notify, so no path can currently claim an `Scheduled` component degrades to a notify, so no path can currently claim an
apply is under way. apply is under way.
- `internal/submit/blobstore.go:40` — the uploads PVC is mounted into felis-api but - `internal/submit/blobstore.go` — CLOSED 2026-09-22. The uploads PVC still cannot
not into the Kaniko build Pod, so a submitted context is durable at the derived cross namespaces, so the transport went through the API instead of a mount: the
location without yet being readable by the build that consumes it. derived context ref is now the internal-face URL
(`/api/v1/internal/submissions/{id}/context`, service-token gated), the build
Job's `context-fetch` initContainer streams it with `felis fetch-context` and
extracts under a zip-slip guard into a size-limited emptyDir, and Kaniko builds
`--context=/context`. The token reaches the build namespace through the same
Secret-replica mechanism the login gate uses (bootstrap + `felis setup`), and the
build egress lock allows exactly the control namespace on the internal port.
Uniform for local and s3:// stores — neither hands the sandboxed build Pod a
filesystem view or object-store credentials. Kaniko, Trivy and Trivy's two DBs
come from the registry's `mirror/` copies, which the installer and
felis-build-tools.timer keep current (`felis mirror-build-tools`,
docs/troubleshooting.md §8e); the `[registry]` keys override them.
## Built; only its I/O is unverifiable from this repo ## Built; only its I/O is unverifiable from this repo
@@ -58,21 +69,38 @@ or a real upstream account to run it against — not an implementation.
drives the whole flow through a fake. drives the whole flow through a fake.
- `internal/api/console.go:39`, `internal/api/logstream.go:236`, - `internal/api/console.go:39`, `internal/api/logstream.go:236`,
`internal/api/logstream.go:306`, `internal/fileedit/k8sjobs.go:45` — each needs a `internal/api/logstream.go:306`, `internal/fileedit/k8sjobs.go:45` — each needs a
live cluster (RCON, `pods/log` follow, a Job). live cluster (RCON, `pods/log` follow, a Job). **Verified live 2026-09-22/23
(auditfix7–25):** the RCON command spine (wake → probe → `command`/access
mutations/stop), the log SSE stream, and the fileedit Job have each run
end-to-end on the drill cluster.
- `internal/api/handlers_access.go:170,490` — parsing real vanilla and LuckPerms - `internal/api/handlers_access.go:170,490` — parsing real vanilla and LuckPerms
command output. command output. **Verified live 2026-09-23:** players / whitelist / banlist
parses matched a live Paper server's replies (LuckPerms not installed → the raw
reply falls through as documented; the input guards held on four negative cases).
## Deliberately accepted, not scheduled to close ## Deliberately accepted, not scheduled to close
These are decisions, not backlog. Each names the condition under which it would be These are decisions, not backlog. Each names the condition under which it would be
worth revisiting. worth revisiting.
- `internal/api/pgrepo.go:281` — the quota check and `ClaimServer` are two statements - ~~`internal/api/pgrepo.go:281` — the quota check and `ClaimServer` are two statements
(audit #4 TOCTOU). Closeable only against a real Postgres. (audit #4 TOCTOU). Closeable only against a real Postgres.~~ **Closed** — the gate
- `internal/api/api.go:671` — `cooldownLimiter` is process-local, so across N api moved inside `ClaimServer` (advisory lock + re-check + UPDATE in one transaction),
red-then-green in the pgint suite, which is exactly the real-Postgres harness this
line was waiting for.
- `internal/api/api.go:773` — `cooldownLimiter` is process-local, so across N api
replicas a caller could draw up to N OTP codes per window. The intra-replica burst replicas a caller could draw up to N OTP codes per window. The intra-replica burst
is closed; cross-replica bounding needs a shared store, out of scope for a is closed; cross-replica bounding needs a shared store, out of scope for a
single-replica install. single-replica install. Revisit before the api Deployment runs more than one
replica.
- `internal/submit/submit.go:524` — the per-user upload storage budget reads the
stored bytes, then writes. On one replica the API's per-user upload reservation
serializes it; across replicas a burst can overshoot by one blob per interleaved
upload, each still under the single-blob cap. The pending-submission cap no
longer has this shape: `CreateSubmission` counts and inserts under a
per-submitter advisory lock (pgint `TestSubmitPendingCapHoldsUnderConcurrency`).
Revisit with the cooldown above, before scaling api replicas: a reservation row
per upload in the same kind of transaction closes it.
- `internal/submit/submit.go:436` and `internal/submit/submit_test.go:351` — a - `internal/submit/submit.go:436` and `internal/submit/submit_test.go:351` — a
post-CAS `Approve` post-CAS `Approve`
failure leaves a row indistinguishable from the benign case, so `Approve` returns a failure leaves a row indistinguishable from the benign case, so `Approve` returns a
@@ -106,8 +134,8 @@ worth revisiting.
## Recorded outside the code ## Recorded outside the code
- `deploy/limbo/README.md:139` — no NetworkPolicy locks the minecraft-namespace - The minecraft-namespace egress is locked (`felis-server-egress`, DNS plus the
egress or the control-namespace ingress today, which is why the login pod reaches public internet with every private range and the node's own global addresses
`felis-api-internal:8081`. This is a conditional obligation rather than a seam: if excluded) and `felis-login-to-internal-api` opens the one platform path a game pod
a future deployment adds either lock, it must also open that path. Spec v4.1 §21 needs — login → felis-api:8081. Any new in-cluster service a game server must call
asks for those policies; `cmd/felis/manifests.go` renders the game-port one. needs its own allow policy next to that one (`internal/platform/netpol.go`).
+476 -43
View File
@@ -147,6 +147,19 @@ components:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
RateLimited:
description: >
This client address called the public sign-in doors faster than the per-address
limit allows (code rate_limited); Retry-After gives the seconds until the next
call is admitted. The address is the visitor header the install's edge writes
([auth] client_ip_header: CF-Connecting-IP behind the Cloudflare tunnel), else
the TCP peer; IPv6 clients share one limit per /64.
headers:
Retry-After:
schema: { type: integer }
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
AccessResult: AccessResult:
description: The structured access mutation succeeded; the raw RCON reply is in output. description: The structured access mutation succeeded; the raw RCON reply is in output.
content: content:
@@ -201,6 +214,49 @@ components:
nullable: true nullable: true
description: Window end (RFC3339, exclusive), or null when unset. description: Window end (RFC3339, exclusive), or null when unset.
DBBackupStatus:
type: object
description: >
The newest control-plane database backup the host recorded
(internal/api/handlers_dbbackup.go dbBackupView; the record itself is
internal/dbbackup Status, written by `felis db backup`).
required: [last, stale, max_age_seconds]
properties:
last:
type: object
nullable: true
description: Null until the first backup has been recorded.
required: [at, name, label, size_bytes, dir]
properties:
at:
type: string
format: date-time
description: When the bundle was written.
name:
type: string
description: Bundle file name, felis-db-<UTC stamp>-<label>.tar.
label:
type: string
enum: [daily, pre-migrate, pre-restore, manual]
size_bytes:
type: integer
format: int64
felis_version:
type: string
schema_version:
type: integer
description: Newest applied migration at backup time.
dir:
type: string
description: Backup directory on the host.
stale:
type: boolean
description: True when there is no record or it is older than max_age_seconds.
max_age_seconds:
type: integer
format: int64
description: The freshness limit (26h), shared with `felis db check` and FelisDBBackupStale.
PasskeyCredential: PasskeyCredential:
type: object type: object
description: > description: >
@@ -242,6 +298,18 @@ components:
endpointAddress: { type: string } endpointAddress: { type: string }
playersOnline: { type: integer, format: int32 } playersOnline: { type: integer, format: int32 }
playersMax: { type: integer, format: int32 } playersMax: { type: integer, format: int32 }
displayName: { type: string }
image: { type: string }
javaMemory: { type: string }
storageSize: { type: string }
cpu: { type: string }
idleStopSeconds:
type: integer
format: int32
description: Seconds the server may sit empty before idle auto-stop scales it down; 0 when it never idles out (off, RCON disabled, or a system server).
playerCountUnknown:
type: boolean
description: Present and true while the operator cannot read the player count over RCON; idle auto-stop waits until it can.
MyServerView: MyServerView:
type: object type: object
@@ -337,6 +405,19 @@ components:
build_id: build_id:
type: string type: string
description: image_builds.id, set only after the build hand-off succeeds. description: image_builds.id, set only after the build hand-off succeeds.
build_status:
type: string
enum: [pending, building, succeeded, failed, cancelled]
description: >-
The linked build's outcome, attached by the LIST routes
(/me/submissions, /submissions) — for a submitter this is the only
visible outlet for a failed build. Omitted until a build is linked
and its row is readable.
build_error:
type: string
description: >-
The build's recorded failure text (e.g. a CRITICAL CVE scan
failure), attached alongside build_status.
reviewed_by: { type: string } reviewed_by: { type: string }
reject_reason: { type: string } reject_reason: { type: string }
created_at: { type: string, format: date-time } created_at: { type: string, format: date-time }
@@ -454,6 +535,26 @@ paths:
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
/metrics:
get:
tags: [metrics]
operationId: metrics
summary: Prometheus metrics (felis_* collectors) on the internal face.
description: >-
Scrape-only infrastructure route, not a product API: the internal listener is
ClusterIP-only and a Prometheus scrape carries no token, the same stance as the
probes. Serves the felis_* exposition documented in troubleshooting §14; the
external face never serves it.
x-felis-face: [internal]
x-felis-tier: public
security: []
responses:
'200':
description: Prometheus text exposition format.
content:
text/plain:
schema: { type: string }
/session/minecraft/hasJoined: /session/minecraft/hasJoined:
get: get:
tags: [nano] tags: [nano]
@@ -540,7 +641,12 @@ paths:
tags: [admin-servers] tags: [admin-servers]
operationId: createServer operationId: createServer
summary: Create a server (admin). summary: Create a server (admin).
description: Requires the admin Access path; the image must be whitelisted. description: >-
Requires the admin Access path; the image must be whitelisted. An image in the
platform registry is stored pinned to the digest its tag names at creation
(name:tag@sha256:…), so a later push over the tag never moves the server;
400 image_not_in_registry when the registry lacks the tag, 503
registry_unavailable when it cannot be asked.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: admin x-felis-tier: admin
security: [{ accessJWT: [] }] security: [{ accessJWT: [] }]
@@ -607,6 +713,34 @@ paths:
'404': '404':
$ref: '#/components/responses/NotFound' $ref: '#/components/responses/NotFound'
/api/v1/internal/submissions/{id}/context:
get:
tags: [submissions-internal]
operationId: internalSubmissionContext
summary: Stream a submission's stored build-context tarball to the build Pod.
description: >-
The build Job's fetch initContainer cannot mount the control-plane uploads
PVC (a PVC does not cross namespaces) and holds no object-store
credentials, so the API that stored the blob streams it here. Served on
the internal face (service token, no Zero Trust).
x-felis-face: [internal]
x-felis-tier: service
security: [{ serviceToken: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The stored gzip tarball, verbatim.
content:
application/gzip:
schema: { type: string, format: binary }
'401':
$ref: '#/components/responses/Unauthorized'
'404':
$ref: '#/components/responses/NotFound'
'503':
$ref: '#/components/responses/ServiceUnavailable'
/api/v1/internal/servers/{name}/join-event: /api/v1/internal/servers/{name}/join-event:
post: post:
tags: [servers-internal] tags: [servers-internal]
@@ -680,6 +814,11 @@ paths:
$ref: '#/components/responses/Forbidden' $ref: '#/components/responses/Forbidden'
'404': '404':
$ref: '#/components/responses/NotFound' $ref: '#/components/responses/NotFound'
'409':
description: A restore, backup or file write holds the server's world volume (maintenance_in_progress); nothing was started.
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: Wake cooldown is still active for this server. description: Wake cooldown is still active for this server.
content: content:
@@ -1132,7 +1271,7 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'409': '409':
description: Server is not stopped (its world PVC is still mounted). description: Server is not stopped (not_stopped), or a restore, backup or file write already holds its world volume (maintenance_in_progress).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -1167,6 +1306,11 @@ paths:
$ref: '#/components/responses/Forbidden' $ref: '#/components/responses/Forbidden'
'404': '404':
$ref: '#/components/responses/NotFound' $ref: '#/components/responses/NotFound'
'409':
description: A restore, backup or file write holds the server's world volume (maintenance_in_progress); nothing was started.
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: Wake cooldown is still active. description: Wake cooldown is still active.
content: content:
@@ -1761,8 +1905,8 @@ paths:
array. It never reveals staffness: methods are computed identically for every array. It never reveals staffness: methods are computed identically for every
resolved account (no role branch), so a staff and a player address in the same resolved account (no role branch), so a staff and a player address in the same
credential state return byte-identical bodies. passkey is offered only when a credential state return byte-identical bodies. passkey is offered only when a
verifier is wired. Sends no mail and mutates nothing; not rate-limited at the app verifier is wired. Sends no mail and mutates nothing; bounded by the per-address
layer (volumetric abuse is bounded at the edge). Gated on local_auth_enabled. sign-in rate limit (429 rate_limited). Gated on local_auth_enabled.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: public x-felis-tier: public
security: [] security: []
@@ -1804,6 +1948,8 @@ paths:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429':
$ref: '#/components/responses/RateLimited'
/api/v1/auth/passkey/login/begin: /api/v1/auth/passkey/login/begin:
post: post:
@@ -1859,7 +2005,9 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: A passkey login for this recipient was started too recently (otp_resend_cooldown). description: >-
A passkey login for this recipient was started too recently (otp_resend_cooldown);
or this client address called the sign-in doors too often (rate_limited, with Retry-After).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -1931,6 +2079,8 @@ paths:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429':
$ref: '#/components/responses/RateLimited'
'503': '503':
description: No passkey verifier is wired on this deployment (passkey_unavailable). description: No passkey verifier is wired on this deployment (passkey_unavailable).
content: content:
@@ -1952,9 +2102,9 @@ paths:
userHandle inside the signed assertion at finish. The challenge cannot be userHandle inside the signed assertion at finish. The challenge cannot be
user-keyed, so it is stashed under login_id in a non-user-keyed store and echoed user-keyed, so it is stashed under login_id in a non-user-keyed store and echoed
back at finish. Mounted Public and gated on local_auth_enabled. There is no back at finish. Mounted Public and gated on local_auth_enabled. There is no
recipient or principal to key a per-caller cooldown on (that volumetric limiting recipient or principal to key a per-caller cooldown on, so one client is bounded
is delegated to the edge), so the server-side brake is a hard global cap on live by the per-address sign-in rate limit (429 rate_limited) and the table by a hard
challenges (429 too_many_challenges). Inert for a credential until its owner global cap on live challenges (429 too_many_challenges). Inert for a credential until its owner
enrolls a resident passkey; email-OTP and username-first passkey remain the enrolls a resident passkey; email-OTP and username-first passkey remain the
fallbacks, so no authenticator is ever locked out. fallbacks, so no authenticator is ever locked out.
x-felis-face: [external] x-felis-face: [external]
@@ -2000,8 +2150,9 @@ paths:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: >- description: >-
Too many discoverable logins are in flight server-wide; the global cap is hit Too many discoverable logins are in flight server-wide (too_many_challenges;
(too_many_challenges). No per-recipient signal is leaked — the cap is global. the cap is global, so no per-recipient signal leaks); or this client address
called the sign-in doors too often (rate_limited, with Retry-After).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -2078,6 +2229,8 @@ paths:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429':
$ref: '#/components/responses/RateLimited'
'503': '503':
description: No passkey verifier is wired on this deployment (passkey_unavailable). description: No passkey verifier is wired on this deployment (passkey_unavailable).
content: content:
@@ -2095,7 +2248,9 @@ paths:
purpose. An address with no account returns the SAME 202 with no code minted, purpose. An address with no account returns the SAME 202 with no code minted,
and the per-recipient cooldown is kept on that path too, so probing reveals and the per-recipient cooldown is kept on that path too, so probing reveals
nothing (existence is learnt only at the sanctioned /auth/options oracle). nothing (existence is learnt only at the sanctioned /auth/options oracle).
Gated on local_auth_enabled. An account that spent its daily wrong-code budget (10 per 24h, across every
code) also gets the same 202 and no mail until the window ends. Gated on
local_auth_enabled.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: public x-felis-tier: public
security: [] security: []
@@ -2137,7 +2292,10 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: A code for this recipient was requested too recently (otp_resend_cooldown). description: >-
A code for this recipient was requested too recently (otp_resend_cooldown);
or this client address called the sign-in doors too often (rate_limited, with Retry-After);
or the install-wide mail budget is spent (mail_rate_limited, with Retry-After).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -2154,7 +2312,10 @@ paths:
under the login purpose, and on success mints a host-only felis_session. An under the login purpose, and on success mints a host-only felis_session. An
unknown address, a wrong or expired code, and an attempt-exhausted code all unknown address, a wrong or expired code, and an attempt-exhausted code all
return the IDENTICAL 400 invalid_code, so the door is not an existence or return the IDENTICAL 400 invalid_code, so the door is not an existence or
lockout oracle. Staff are refused (403) — but only AFTER a valid code is lockout oracle. The 10th wrong code in 24h locks the door for that account
until the window ends (the right code then also reads as invalid_code); the
owner is told by mail once, and the lock is audited as auth.otp.locked.
Staff are refused (403) — but only AFTER a valid code is
redeemed, so only the account owner can ever reach that refusal. redeemed, so only the account owner can ever reach that refusal.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: public x-felis-tier: public
@@ -2199,6 +2360,8 @@ paths:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429':
$ref: '#/components/responses/RateLimited'
/api/v1/auth/op-login/start: /api/v1/auth/op-login/start:
post: post:
@@ -2210,7 +2373,8 @@ paths:
staff address, opens an op_login request, and mails a one-time code under the staff address, opens an op_login request, and mails a one-time code under the
op_login purpose, returning the request handle the browser polls. A non-staff op_login purpose, returning the request handle the browser polls. A non-staff
or unknown address gets the SAME 202 with a random, non-persisted handle and no or unknown address gets the SAME 202 with a random, non-persisted handle and no
mail, so this never becomes a staff-enumeration oracle. Gated on mail, so this never becomes a staff-enumeration oracle. A staff account that
spent its daily wrong-code budget gets the same neutral 202. Gated on
local_auth_enabled. local_auth_enabled.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: public x-felis-tier: public
@@ -2253,7 +2417,10 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: A code for this recipient was requested too recently (otp_resend_cooldown). description: >-
A code for this recipient was requested too recently (otp_resend_cooldown);
or this client address called the sign-in doors too often (rate_limited, with Retry-After);
or the install-wide mail budget is spent (mail_rate_limited, with Retry-After).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -2301,7 +2468,7 @@ paths:
Public, pre-session final leg: mints a host-only staff session only when BOTH Public, pre-session final leg: mints a host-only staff session only when BOTH
factors have landed — the request is approved-and-live AND the mailed code factors have landed — the request is approved-and-live AND the mailed code
verifies. Every failure (unknown handle, not-yet-approved, wrong or locked code, verifies. Every failure (unknown handle, not-yet-approved, wrong or locked code,
lost race) collapses into one uniform 400 op_login_invalid, so a code-less an account past its daily wrong-code budget, lost race) collapses into one uniform 400 op_login_invalid, so a code-less
caller learns nothing. Admin is re-asserted before the session is issued. caller learns nothing. Admin is re-asserted before the session is issued.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: public x-felis-tier: public
@@ -2347,6 +2514,8 @@ paths:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429':
$ref: '#/components/responses/RateLimited'
/api/v1/auth/setup/redeem: /api/v1/auth/setup/redeem:
post: post:
@@ -2404,6 +2573,8 @@ paths:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429':
$ref: '#/components/responses/RateLimited'
/api/v1/auth/setup/status: /api/v1/auth/setup/status:
get: get:
@@ -2520,6 +2691,8 @@ paths:
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429':
$ref: '#/components/responses/RateLimited'
/api/v1/me: /api/v1/me:
get: get:
@@ -2653,6 +2826,31 @@ paths:
'403': '403':
$ref: '#/components/responses/Forbidden' $ref: '#/components/responses/Forbidden'
/api/v1/platform/db-backup:
get:
tags: [admin-updates]
operationId: getDBBackup
summary: Freshness of the newest control-plane database backup (admin).
description: >-
What the host's felis-db-backup.timer (or a manual `felis db backup`)
last recorded in platform_settings. last is null before the first
backup; stale is true then, and whenever the newest backup is older than
max_age_seconds. Read-only: backups run on the host, never through the API.
x-felis-face: [external]
x-felis-tier: admin
security: [{ accessJWT: [] }]
responses:
'200':
description: The newest recorded backup and whether it is stale.
content:
application/json:
schema:
$ref: '#/components/schemas/DBBackupStatus'
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/Forbidden'
/api/v1/fleet: /api/v1/fleet:
get: get:
tags: [admin-servers] tags: [admin-servers]
@@ -2690,6 +2888,13 @@ paths:
The owner's display identity (email, or username when The owner's display identity (email, or username when
the address is absent). Absent for an unclaimed server the address is absent). Absent for an unclaimed server
or when the best-effort owner lookup failed. or when the best-effort owner lookup failed.
system:
type: boolean
description: >-
True for a platform-provisioned system service (the login
gate, the lobby). Their reserved names are rejected by
every per-server route, so the cockpit renders them
read-only instead of offering actions that would 400.
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'403': '403':
@@ -2760,7 +2965,7 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'409': '409':
description: Submission has already been reviewed. description: Server is not stopped (not_stopped), or a restore, backup or file write already holds its world volume (maintenance_in_progress).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -2771,11 +2976,12 @@ paths:
post: post:
tags: [backups] tags: [backups]
operationId: backupNow operationId: backupNow
summary: Back up a server's world on demand (owner-or-admin; server must be stopped). summary: Back up a server's data volume on demand (owner-or-admin; server must be stopped).
description: >- description: >-
Snapshots the server's world into the archive store as a first-class Snapshots the server's whole data volume (worlds, config, plugins/mods,
world_backups row (reason "manual"), restorable later like an inactivity jars, libraries — not just world folders) into the archive store as a
backup. The world PVC is RWO and held by a running server, so the server must first-class world_backups row (reason "manual"), restorable later like an
inactivity backup. A restore replaces the volume with the archive. The world PVC is RWO and held by a running server, so the server must
be fully stopped first (409 not_stopped otherwise). The backup runs be fully stopped first (409 not_stopped otherwise). The backup runs
asynchronously as a Job, so success is 202 (backing_up). asynchronously as a Job, so success is 202 (backing_up).
x-felis-face: [external] x-felis-face: [external]
@@ -2804,7 +3010,57 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'409': '409':
description: Server is not stopped (its world PVC is still mounted). description: Server is not stopped (not_stopped), or a restore, backup or file write already holds its world volume (maintenance_in_progress).
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'503':
$ref: '#/components/responses/ServiceUnavailable'
# -------------------------------------------------- async job status (app) ---
/api/v1/servers/{name}/jobs:
get:
tags: [backups]
operationId: listServerJobs
summary: Latest async world operations (backup/restore) for a server (owner-or-admin).
description: >-
Backup and restore run as cluster Jobs, so a 202 that later failed left
its only trace in the Job object. This route projects the newest such
Jobs, newest first, so failures are observable without kubectl. State is
"running" | "succeeded" | "failed".
x-felis-face: [external]
x-felis-tier: app
security: [{ accessJWT: [] }]
parameters:
- { name: name, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The server's newest backup/restore jobs.
content:
application/json:
schema:
type: object
required: [server, jobs]
properties:
server: { type: string }
jobs:
type: array
items:
type: object
required: [name, kind, state]
properties:
name: { type: string }
kind: { type: string, enum: [backup, restore] }
state: { type: string, enum: [running, succeeded, failed] }
message: { type: string }
started_at: { type: string, format: date-time }
finished_at: { type: string, format: date-time }
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/Forbidden'
'404':
description: Unknown server.
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -3003,7 +3259,7 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'409': '409':
description: Server is not stopped (its world PVC is still mounted). description: Server is not stopped (not_stopped), or a restore, backup or file write already holds its world volume (maintenance_in_progress).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -3571,6 +3827,15 @@ paths:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'429':
description: >-
Resend requested before the cooldown elapsed (otp_resend_cooldown); or the
account spent its daily wrong-code budget (otp_account_locked, with
Retry-After); or the install-wide mail budget is spent
(mail_rate_limited, with Retry-After).
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'502': '502':
$ref: '#/components/responses/MailUndeliverable' $ref: '#/components/responses/MailUndeliverable'
@@ -3582,8 +3847,10 @@ paths:
description: > description: >
Consumes a previously delivered code for the authenticated principal. On Consumes a previously delivered code for the authenticated principal. On
success the user's email is written and email_verified is set true. Too many success the user's email is written and email_verified is set true. Too many
incorrect attempts lock the code (429); an unknown, expired, consumed, or incorrect attempts lock the code (429 otp_locked); 10 wrong codes in 24h,
mismatched code is a 400. counted across every code, lock the account's email-code door until the
window ends (429 otp_account_locked with Retry-After). An unknown, expired,
consumed, or mismatched code is a 400.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: app x-felis-tier: app
security: [{ accessJWT: [] }] security: [{ accessJWT: [] }]
@@ -3615,7 +3882,9 @@ paths:
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'429': '429':
description: Too many incorrect attempts; the code is locked. description: >-
Too many incorrect attempts on this code (otp_locked), or the account's
daily wrong-code budget is spent (otp_account_locked, with Retry-After).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -3873,7 +4142,11 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: Resend requested before the cooldown elapsed. description: >-
Resend requested before the cooldown elapsed (otp_resend_cooldown), or the
account's daily wrong-code budget is spent (otp_account_locked, with
Retry-After); or the install-wide mail budget is spent
(mail_rate_limited, with Retry-After).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -3888,8 +4161,9 @@ paths:
description: > description: >
Consumes the fresh migrate-purpose email code for the caller's initiated Consumes the fresh migrate-purpose email code for the caller's initiated
migration and advances it to confirmed with confirm_factor email_otp. Too many migration and advances it to confirmed with confirm_factor email_otp. Too many
wrong attempts lock the code (429 otp_locked); an unknown, expired, consumed, or wrong attempts lock the code (429 otp_locked), and 10 wrong codes in 24h lock
mismatched code is a 400 invalid_code. the account's email-code door (429 otp_account_locked with Retry-After); an
unknown, expired, consumed, or mismatched code is a 400 invalid_code.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: app x-felis-tier: app
security: [{ accessJWT: [] }] security: [{ accessJWT: [] }]
@@ -3930,7 +4204,10 @@ paths:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'429': '429':
description: The code is locked after too many wrong attempts (otp_locked). description: >-
The code is locked after too many wrong attempts (otp_locked), or the
account's daily wrong-code budget is spent (otp_account_locked, with
Retry-After).
content: content:
application/json: application/json:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
@@ -4147,18 +4424,30 @@ paths:
$ref: '#/components/responses/BadRequest' $ref: '#/components/responses/BadRequest'
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'403':
description: >-
The per-user submission allowance is spent — too many of the
caller's submissions are awaiting review, or their stored-upload
budget is full (submission_quota_exceeded).
'429':
description: >-
A submission was created within the per-user cooldown window
(submission_cooldown).
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
get: get:
tags: [submissions] tags: [submissions]
operationId: mySubmissions operationId: mySubmissions
summary: List the caller's own modpack submissions (user-directed lane over §16). summary: List the caller's own modpack submissions with each linked build's outcome (user-directed lane over §16).
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: app x-felis-tier: app
security: [{ accessJWT: [] }] security: [{ accessJWT: [] }]
responses: responses:
'200': '200':
description: The caller's submissions, newest first. description: >-
The caller's submissions, newest first; rows with a linked build
additionally carry build_status/build_error so the submitter can see
whether their build succeeded or failed (and why).
content: content:
application/json: application/json:
schema: schema:
@@ -4185,8 +4474,10 @@ paths:
principal; a submission the caller does not own is reported as 404, so principal; a submission the caller does not own is reported as 404, so
this endpoint cannot upload to or probe another user's submission. Only a this endpoint cannot upload to or probe another user's submission. Only a
pending_review submission accepts a context (409 otherwise); a wrong-format pending_review submission accepts a context (409 otherwise); a wrong-format
or oversize body is rejected with 400. Returns 503 when the deployment's or oversize body is rejected with 400, and an upload that would push the
context store has no implemented upload transport. caller past their per-user stored-context budget is refused with 403
before the excess is persisted. Returns 503 when the deployment's context
store has no implemented upload transport.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: app x-felis-tier: app
security: [{ accessJWT: [] }] security: [{ accessJWT: [] }]
@@ -4207,10 +4498,53 @@ paths:
$ref: '#/components/responses/BadRequest' $ref: '#/components/responses/BadRequest'
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'403':
description: >-
The upload would exceed the caller's per-user stored-context budget
(submission_quota_exceeded).
'404': '404':
$ref: '#/components/responses/NotFound' $ref: '#/components/responses/NotFound'
'409': '409':
$ref: '#/components/responses/Conflict' $ref: '#/components/responses/Conflict'
'429':
description: >-
An upload was accepted within the per-user cooldown window
(submission_cooldown).
'503':
$ref: '#/components/responses/ServiceUnavailable'
/api/v1/me/submissions/{id}:
delete:
tags: [submissions]
operationId: withdrawSubmission
summary: Withdraw your own pending submission (user side; user-directed lane over §16).
description: >-
Retracts the caller's own submission while it is still pending review:
the row and its uploaded build context are deleted, freeing the pending
slot and the per-user storage budget for a fresh submission. A reviewed
submission is frozen (409 — its build may already be consuming the
context), and a submission the caller does not own reads back as 404, so
this endpoint cannot probe or clear another user's uploads.
x-felis-face: [external]
x-felis-tier: app
security: [{ accessJWT: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The withdrawn submission, as it was before the deletion.
content:
application/json:
schema: { $ref: '#/components/schemas/Submission' }
'401':
$ref: '#/components/responses/Unauthorized'
'404':
$ref: '#/components/responses/NotFound'
'409':
description: Submission has already been reviewed and cannot be withdrawn.
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
@@ -4233,9 +4567,21 @@ paths:
type: object type: object
description: Only the supplied fields are patched; an empty patch is rejected. description: Only the supplied fields are patched; an empty patch is rejected.
properties: properties:
display_name: { type: string } displayName: { type: string }
autostart_policy: { type: string } autostartPolicy: { type: string }
image: { type: string } image:
type: string
description: >-
Re-admitted against the whitelist (a pinned name:tag@sha256:… ref is
admitted by its name:tag) and pinned like create does. A pin equal to
the current image is no change; any other needs confirmImageChange.
confirmImageChange:
type: boolean
description: >-
Acknowledges that the new image opens the world with its Minecraft
version, whose chunk upgrades the old one cannot read. Without it an
image that would move the server is refused with 409
image_change_unconfirmed. The audit row records image_from/image_to.
memory: { type: string } memory: { type: string }
storage: storage:
type: string type: string
@@ -4244,9 +4590,13 @@ paths:
type: object type: object
properties: properties:
cpu: { type: string } cpu: { type: string }
cpu_request: { type: string } cpuRequest: { type: string }
memory: { type: string } memory: { type: string }
memory_request: { type: string } memoryRequest: { type: string }
idleStopSeconds:
type: integer
format: int32
description: Idle auto-stop. 0 turns it off; otherwise the server stops after this many seconds with nobody online (60–86400, else 400 bad_idle_stop).
responses: responses:
'200': '200':
description: Patched. description: Patched.
@@ -4268,6 +4618,10 @@ paths:
$ref: '#/components/responses/Forbidden' $ref: '#/components/responses/Forbidden'
'404': '404':
$ref: '#/components/responses/NotFound' $ref: '#/components/responses/NotFound'
'409':
$ref: '#/components/responses/Conflict'
'503':
$ref: '#/components/responses/ServiceUnavailable'
/api/v1/images/build: /api/v1/images/build:
post: post:
@@ -4285,10 +4639,24 @@ paths:
type: object type: object
required: [image_ref, dockerfile, context_ref] required: [image_ref, dockerfile, context_ref]
properties: properties:
image_ref: { type: string } image_ref:
dockerfile: { type: string } type: string
context_ref: { type: string } description: Push target under the internal registry (e.g. registry.felis.svc:5000/foo:1.0).
base_image: { type: string } dockerfile:
type: string
description: >-
Audit archive of the recipe, recorded on the build row and shown in the
panel — the executed Dockerfile is the file named `Dockerfile` at the
root of the context tarball (Kaniko runs --dockerfile=Dockerfile), so
this field is never executed.
context_ref:
type: string
description: >-
Location of the uploaded gzip build context; its root must contain the
Dockerfile that gets executed.
base_image:
type: string
description: Resolved FROM, recorded for audit only — not a build gate.
responses: responses:
'202': '202':
description: Build accepted. description: Build accepted.
@@ -4522,6 +4890,39 @@ paths:
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
/api/v1/submissions/{id}/context:
get:
tags: [submissions]
operationId: downloadSubmissionContext
summary: Download a submission's uploaded build context (admin; user-directed lane over §16).
description: >-
The reviewer's read path to the artifact they are about to approve: the
executed Dockerfile lives inside this tarball (Kaniko runs the context's
root `Dockerfile`), so without it the human gate would be blind. Streams
the stored context.tar.gz verbatim with an attachment disposition — the
same bytes the build Pod fetches over the internal face. 404 when the
submission is unknown or has no uploaded context; 503 when the
deployment's context store has no implemented transport.
x-felis-face: [external]
x-felis-tier: admin
security: [{ accessJWT: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The stored build context (gzip tarball), served as an attachment.
content:
application/gzip:
schema: { type: string, format: binary }
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/Forbidden'
'404':
$ref: '#/components/responses/NotFound'
'503':
$ref: '#/components/responses/ServiceUnavailable'
/api/v1/submissions/{id}/reject: /api/v1/submissions/{id}/reject:
post: post:
tags: [submissions] tags: [submissions]
@@ -4562,3 +4963,35 @@ paths:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
/api/v1/submissions/{id}:
delete:
tags: [submissions]
operationId: deleteSubmission
summary: Retire a submission outright — row and uploaded context (admin; user-directed lane over §16).
description: >-
Removes the submission and its uploaded build context, any status — the
lane's only lifecycle valve, and the path that reclaims a rejected or
consumed upload from the uploads PVC. The reviewer identity is recorded
in the audit event, not on the (now deleted) row. Deleting an approved
submission whose build is still running fails that build's context
fetch; the admin has explicitly chosen to retire the artifact.
x-felis-face: [external]
x-felis-tier: admin
security: [{ accessJWT: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The deleted submission, as it was before the deletion.
content:
application/json:
schema: { $ref: '#/components/schemas/Submission' }
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/Forbidden'
'404':
$ref: '#/components/responses/NotFound'
'503':
$ref: '#/components/responses/ServiceUnavailable'
+4 -4
View File
@@ -87,7 +87,7 @@ sequenceDiagram
else quota available else quota available
Repo-->>API: true Repo-->>API: true
API->>Repo: ClaimServer(name, user_id) API->>Repo: ClaimServer(name, user_id)
Note over Repo: SELECT EXISTS(server); then atomic UPDATE servers SET owner_id=$2, claimed_at=now() WHERE name=$1 AND owner_id IS NULL AND deleted_at IS NULL Note over Repo: ONE transaction: pg_advisory_xact_lock(user_id) serializes this user's claim lane; SELECT FROM servers WHERE name=$1 AND deleted_at IS NULL FOR UPDATE; re-run the four-dimension quota gate (authoritative — the pre-check above is a fast path); then UPDATE servers SET owner_id=$2, claimed_at=now() WHERE name=$1 AND owner_id IS NULL AND deleted_at IS NULL
alt server missing alt server missing
Repo-->>API: ErrNotFound Repo-->>API: ErrNotFound
API-->>Panel: 404 not_found API-->>Panel: 404 not_found
@@ -137,14 +137,14 @@ sequenceDiagram
Panel->>APIExternal: POST /api/v1/account/link/verify {code} Panel->>APIExternal: POST /api/v1/account/link/verify {code}
APIExternal->>APIExternal: trim and uppercase code APIExternal->>APIExternal: trim and uppercase code
APIExternal->>Repo: VerifyLinkCode(user_id, code, now) APIExternal->>Repo: VerifyLinkCode(user_id, code, now)
Repo->>Repo: SELECT non-expired code Repo->>Repo: SELECT mc_uuid, auth_source FROM account_link_codes WHERE code=$1 AND expires_at>$2 FOR UPDATE
alt missing or expired code alt missing or expired code
Repo-->>APIExternal: ErrLinkCodeInvalid Repo-->>APIExternal: ErrLinkCodeInvalid
APIExternal-->>Panel: 400 invalid_code APIExternal-->>Panel: 400 invalid_code
else UUID linked to another user else UUID linked to a different, live user
Repo-->>APIExternal: ErrConflict Repo-->>APIExternal: ErrConflict
APIExternal-->>Panel: 409 already_linked APIExternal-->>Panel: 409 already_linked
else valid code else valid code (re-verify by the same user is idempotent; a retired/soft-deleted owner's link is taken over)
Repo->>Repo: INSERT account_links(user_id, mc_uuid, auth_source) ON CONFLICT (user_id, mc_uuid) DO UPDATE auth_source Repo->>Repo: INSERT account_links(user_id, mc_uuid, auth_source) ON CONFLICT (user_id, mc_uuid) DO UPDATE auth_source
Repo->>Repo: DELETE account_link_codes WHERE code=$1 Repo->>Repo: DELETE account_link_codes WHERE code=$1
Repo-->>APIExternal: mc_uuid, auth_source Repo-->>APIExternal: mc_uuid, auth_source
+1269 -34
View File
File diff suppressed because it is too large. Load diff
+1 -1
View File
@@ -9,6 +9,7 @@ require (
github.com/charmbracelet/huh v1.0.0 github.com/charmbracelet/huh v1.0.0
github.com/charmbracelet/lipgloss v1.1.0 github.com/charmbracelet/lipgloss v1.1.0
github.com/descope/virtualwebauthn v1.0.5 github.com/descope/virtualwebauthn v1.0.5
github.com/go-logr/logr v1.4.2
github.com/go-webauthn/webauthn v0.17.4 github.com/go-webauthn/webauthn v0.17.4
github.com/golang-jwt/jwt/v5 v5.3.1 github.com/golang-jwt/jwt/v5 v5.3.1
github.com/google/uuid v1.6.0 github.com/google/uuid v1.6.0
@@ -42,7 +43,6 @@ require (
github.com/erikgeiser/coninput v0.0.0-20211004153227-1c3628e74d0f // indirect github.com/erikgeiser/coninput v0.0.0-20211004153227-1c3628e74d0f // indirect
github.com/evanphx/json-patch/v5 v5.9.0 // indirect github.com/evanphx/json-patch/v5 v5.9.0 // indirect
github.com/fxamacker/cbor/v2 v2.9.2 // indirect github.com/fxamacker/cbor/v2 v2.9.2 // indirect
github.com/go-logr/logr v1.4.2 // indirect
github.com/go-openapi/jsonpointer v0.19.6 // indirect github.com/go-openapi/jsonpointer v0.19.6 // indirect
github.com/go-openapi/jsonreference v0.20.2 // indirect github.com/go-openapi/jsonreference v0.20.2 // indirect
github.com/go-openapi/swag v0.22.4 // indirect github.com/go-openapi/swag v0.22.4 // indirect
+138 -19
View File
@@ -33,6 +33,11 @@ type API struct {
// still exercised even before the subsystem is wired in. // still exercised even before the subsystem is wired in.
Builder ImageBuilder Builder ImageBuilder
// Images pins a whitelisted image ref to the digest it names when a server is
// created or its image is changed (internal/imagepin), so a later push over
// the same tag never reaches an existing world. Nil stores refs as given.
Images ImagePinner
// Console is the synchronous RCON write channel (spec §8 写=RCON). It is // Console is the synchronous RCON write channel (spec §8 写=RCON). It is
// wired in production (cmd/felis); a nil Console makes the command route report // wired in production (cmd/felis); a nil Console makes the command route report
// 503 rather than panic, so the ownership boundary is still exercised in tests. // 503 rather than panic, so the ownership boundary is still exercised in tests.
@@ -63,6 +68,11 @@ type API struct {
// authorization boundary is exercised before the backup-Job executor is wired. // authorization boundary is exercised before the backup-Job executor is wired.
Backuper Backuper Backuper Backuper
// JobStatus reads the latest backup/restore Job outcomes for GET
// /servers/{name}/jobs — the status outlet for async failures (the enqueue
// endpoints only answer 202). Optional: nil → that route reports 503.
JobStatus JobStatusReader
// Files is the server file editor (list / read / write a file in a stopped // Files is the server file editor (list / read / write a file in a stopped
// server's world volume — the "one wrong line in server.properties" repair). // server's world volume — the "one wrong line in server.properties" repair).
// Like Restorer and Backuper it is optional: when nil the file routes report // Like Restorer and Backuper it is optional: when nil the file routes report
@@ -118,6 +128,23 @@ type API struct {
// on the wake lever). Zero disables throttling. // on the wake lever). Zero disables throttling.
WakeCooldown time.Duration WakeCooldown time.Duration
// BackupCooldown spaces out an owner's on-demand backups of one server, and
// BackupStoreCap refuses them once the present backups reach [archive]
// max_local_bytes (data-durability-9): each archive lands on the node disk
// the worlds and the database share. Admins and the break-glass console are
// exempt. Zero disables each lever.
BackupCooldown time.Duration
BackupStoreCap int64
// SubmitCreateCooldown / SubmitUploadCooldown throttle the user-modpack
// submission lane per user: create bounds how quickly review-queue rows can
// appear, upload bounds how often a user may stream a (up to 1 GiB) build
// context. The keys are separate, so the lane's normal shape — create, then
// upload — is never blocked by its own throttle. Zero disables each lever
// (the same idiom as WakeCooldown); cmd/felis wires positive values.
SubmitCreateCooldown time.Duration
SubmitUploadCooldown time.Duration
// MaxRunningServers caps how many servers may be desired-Running cluster-wide // MaxRunningServers caps how many servers may be desired-Running cluster-wide
// (spec §9.1: the concurrency-上限 lever hanging on the same wake chokepoint as // (spec §9.1: the concurrency-上限 lever hanging on the same wake chokepoint as
// cooldown and autostartPolicy). Zero — the default — disables it: §9.2 wires // cooldown and autostartPolicy). Zero — the default — disables it: §9.2 wires
@@ -144,6 +171,16 @@ type API struct {
// Consumed by handleHasJoined (handlers_hasjoined.go). // Consumed by handleHasJoined (handlers_hasjoined.go).
AuthSources []AuthSource AuthSources []AuthSource
// AuthDoorLimit bounds how often one client address may call the public
// pre-session auth doors (ratelimit.go). MailLimit bounds all mail the API
// sends, install-wide. Zero values disable them; cmd/felis wires both.
AuthDoorLimit RateLimit
MailLimit RateLimit
// ClientIPHeader names the header the install's edge writes the client
// address into (CF-Connecting-IP behind the Cloudflare tunnel,
// X-Forwarded-For behind an operator proxy). Empty means the TCP peer.
ClientIPHeader string
// Now is the clock, injectable for tests. Defaults to time.Now. // Now is the clock, injectable for tests. Defaults to time.Now.
Now func() time.Time Now func() time.Time
@@ -153,8 +190,16 @@ type API struct {
otpCooldownOnce sync.Once otpCooldownOnce sync.Once
otpCooldown *cooldownLimiter otpCooldown *cooldownLimiter
submitCooldownOnce sync.Once
submitCooldown *cooldownLimiter
streamCapOnce sync.Once streamCapOnce sync.Once
streamCap *streamLimiter streamCap *streamLimiter
authDoorOnce sync.Once
authDoorBuckets *bucketSet
mailOnce sync.Once
mailBuckets *bucketSet
} }
// panelURL returns the public player-console origin ("https://console.<root>"), // panelURL returns the public player-console origin ("https://console.<root>"),
@@ -199,6 +244,20 @@ func (a *API) otpLimiter() *cooldownLimiter {
return a.otpCooldown return a.otpCooldown
} }
// submitLimiter lazily builds a SEPARATE cooldown limiter for the user-modpack
// submission lane, so its throttles never share state with the wake or OTP
// keyspaces. One limiter backs both levers with prefixed keys (see the
// submissionCreateKey/UploadKey constants), so create and upload never contend
// with each other. Like the other cooldowns it is process-local; with multiple
// api replicas the effective spacing is per-replica, the same accepted
// KNOWN-LIMITATION the OTP resend throttle carries.
func (a *API) submitLimiter() *cooldownLimiter {
a.submitCooldownOnce.Do(func() {
a.submitCooldown = &cooldownLimiter{now: a.now, last: map[string]time.Time{}}
})
return a.submitCooldown
}
// streamGate lazily builds the per-principal SSE stream cap bound to // streamGate lazily builds the per-principal SSE stream cap bound to
// MaxStreamsPerPrincipal. A zero cap yields a disabled limiter that admits every // MaxStreamsPerPrincipal. A zero cap yields a disabled limiter that admits every
// stream, so a deployment (or test) that leaves it unset pays nothing. // stream, so a deployment (or test) that leaves it unset pays nothing.
@@ -255,17 +314,29 @@ type apiRoute struct {
// whose EmailVerified is false is restricted to these routes only. // whose EmailVerified is false is restricted to these routes only.
SetupAllowed bool SetupAllowed bool
// AuthDoor marks a public pre-session auth door: it is rate limited per
// client address (throttleAuthDoor). The op-login status poll is left off,
// since the browser calls it every few seconds while it waits.
AuthDoor bool
h http.HandlerFunc h http.HandlerFunc
} }
// internalAPIRoutes is the internal face's served route table (spec §7, §14): // internalAPIRoutes is the internal face's served route table (spec §7, §14):
// service-token auth, never Zero Trust. It carries both health probes. // service-token auth, never Zero Trust. It carries both health probes and the
// metrics scrape.
func (a *API) internalAPIRoutes() []apiRoute { func (a *API) internalAPIRoutes() []apiRoute {
return []apiRoute{ return []apiRoute{
{Method: "GET", Pattern: "/healthz", Public: true, h: a.handleHealthz}, {Method: "GET", Pattern: "/healthz", Public: true, h: a.handleHealthz},
{Method: "GET", Pattern: "/readyz", Public: true, h: a.handleReadyz}, {Method: "GET", Pattern: "/readyz", Public: true, h: a.handleReadyz},
// Prometheus scrape (felis_* collectors); public because a scrape carries
// no token, internal-only so it is never exposed off-cluster.
{Method: "GET", Pattern: "/metrics", Public: true, h: a.handleMetrics},
{Method: "GET", Pattern: "/api/v1/servers", h: a.handleListServers}, {Method: "GET", Pattern: "/api/v1/servers", h: a.handleListServers},
// The build Pod's context-fetch initContainer streams a submission's stored
// modpack through this route (build namespace cannot mount the uploads PVC).
{Method: "GET", Pattern: "/api/v1/internal/submissions/{id}/context", h: a.handleInternalSubmissionContext},
{Method: "POST", Pattern: "/api/v1/internal/servers/{name}/ready", h: a.handleReady}, {Method: "POST", Pattern: "/api/v1/internal/servers/{name}/ready", h: a.handleReady},
{Method: "POST", Pattern: "/api/v1/internal/servers/{name}/join-event", h: a.handleJoinEvent}, {Method: "POST", Pattern: "/api/v1/internal/servers/{name}/join-event", h: a.handleJoinEvent},
// Domain-autostart (spec §9.1, §14): velocity drives the wake lever and polls // Domain-autostart (spec §9.1, §14): velocity drives the wake lever and polls
@@ -343,21 +414,21 @@ func (a *API) externalAPIRoutes() []apiRoute {
// counter-slice to the anti-enumeration doors — the ONE sanctioned place existence // counter-slice to the anti-enumeration doors — the ONE sanctioned place existence
// is disclosed — but it never reveals staffness (methods computed with no role // is disclosed — but it never reveals staffness (methods computed with no role
// branch, so a staff and a player address in the same state are indistinguishable). // branch, so a staff and a player address in the same state are indistinguishable).
{Method: "POST", Pattern: "/api/v1/auth/options", Public: true, h: a.handleAuthOptions}, {Method: "POST", Pattern: "/api/v1/auth/options", Public: true, AuthDoor: true, h: a.handleAuthOptions},
{Method: "POST", Pattern: "/api/v1/auth/setup/redeem", Public: true, h: a.handleSetupRedeem}, {Method: "POST", Pattern: "/api/v1/auth/setup/redeem", Public: true, AuthDoor: true, h: a.handleSetupRedeem},
{Method: "GET", Pattern: "/api/v1/auth/setup/status", SetupAllowed: true, h: a.handleSetupStatus}, {Method: "GET", Pattern: "/api/v1/auth/setup/status", SetupAllowed: true, h: a.handleSetupStatus},
{Method: "POST", Pattern: "/api/v1/auth/passkey/login/begin", Public: true, h: a.handlePasskeyLoginBegin}, {Method: "POST", Pattern: "/api/v1/auth/passkey/login/begin", Public: true, AuthDoor: true, h: a.handlePasskeyLoginBegin},
{Method: "POST", Pattern: "/api/v1/auth/passkey/login/finish", Public: true, h: a.handlePasskeyLoginFinish}, {Method: "POST", Pattern: "/api/v1/auth/passkey/login/finish", Public: true, AuthDoor: true, h: a.handlePasskeyLoginFinish},
// Discoverable ("usernameless") passkey login (task #40): the from-zero sibling of the // Discoverable ("usernameless") passkey login (task #40): the from-zero sibling of the
// email-first pair above — no identifier typed, the account is resolved from the // email-first pair above — no identifier typed, the account is resolved from the
// userHandle inside the signed assertion (handlers_passkey_discoverable.go). // userHandle inside the signed assertion (handlers_passkey_discoverable.go).
{Method: "POST", Pattern: "/api/v1/auth/passkey/login/discoverable/begin", Public: true, h: a.handlePasskeyLoginDiscoverableBegin}, {Method: "POST", Pattern: "/api/v1/auth/passkey/login/discoverable/begin", Public: true, AuthDoor: true, h: a.handlePasskeyLoginDiscoverableBegin},
{Method: "POST", Pattern: "/api/v1/auth/passkey/login/discoverable/finish", Public: true, h: a.handlePasskeyLoginDiscoverableFinish}, {Method: "POST", Pattern: "/api/v1/auth/passkey/login/discoverable/finish", Public: true, AuthDoor: true, h: a.handlePasskeyLoginDiscoverableFinish},
{Method: "POST", Pattern: "/api/v1/auth/email/start", Public: true, h: a.handleLoginEmailStart}, {Method: "POST", Pattern: "/api/v1/auth/email/start", Public: true, AuthDoor: true, h: a.handleLoginEmailStart},
{Method: "POST", Pattern: "/api/v1/auth/email/verify", Public: true, h: a.handleLoginEmailVerify}, {Method: "POST", Pattern: "/api/v1/auth/email/verify", Public: true, AuthDoor: true, h: a.handleLoginEmailVerify},
{Method: "POST", Pattern: "/api/v1/auth/op-login/start", Public: true, h: a.handleOpLoginStart}, {Method: "POST", Pattern: "/api/v1/auth/op-login/start", Public: true, AuthDoor: true, h: a.handleOpLoginStart},
{Method: "GET", Pattern: "/api/v1/auth/op-login/status/{id}", Public: true, h: a.handleOpLoginStatus}, {Method: "GET", Pattern: "/api/v1/auth/op-login/status/{id}", Public: true, h: a.handleOpLoginStatus},
{Method: "POST", Pattern: "/api/v1/auth/op-login/finish", Public: true, h: a.handleOpLoginFinish}, {Method: "POST", Pattern: "/api/v1/auth/op-login/finish", Public: true, AuthDoor: true, h: a.handleOpLoginFinish},
// Player-console bootstrap (console-tier access model): the account-less // Player-console bootstrap (console-tier access model): the account-less
// player's door into console.<root_domain>. Public — like login there is no prior // player's door into console.<root_domain>. Public — like login there is no prior
// principal — and session-minting, but the artifact it consumes is a one-time // principal — and session-minting, but the artifact it consumes is a one-time
@@ -365,7 +436,7 @@ func (a *API) externalAPIRoutes() []apiRoute {
// possession already proves a Minecraft identity. A code whose UUID belongs to // possession already proves a Minecraft identity. A code whose UUID belongs to
// staff is refused (403) so this never yields an admin session; op.console stays // staff is refused (403) so this never yields an admin session; op.console stays
// behind Zero Trust (handlers_onboard.go). // behind Zero Trust (handlers_onboard.go).
{Method: "POST", Pattern: "/api/v1/auth/bind", Public: true, h: a.handleBindRedeem}, {Method: "POST", Pattern: "/api/v1/auth/bind", Public: true, AuthDoor: true, h: a.handleBindRedeem},
// App-auth tier: operations on your own servers (spec §14). // App-auth tier: operations on your own servers (spec §14).
{Method: "POST", Pattern: "/api/v1/servers/{name}/wake", h: a.handleWake}, {Method: "POST", Pattern: "/api/v1/servers/{name}/wake", h: a.handleWake},
@@ -407,6 +478,7 @@ func (a *API) externalAPIRoutes() []apiRoute {
// owned), and restore is gated by owner-or-admin PLUS a former-owner match, so // owned), and restore is gated by owner-or-admin PLUS a former-owner match, so
// neither sits behind adminOnly. // neither sits behind adminOnly.
{Method: "GET", Pattern: "/api/v1/backups", h: a.handleListBackups}, {Method: "GET", Pattern: "/api/v1/backups", h: a.handleListBackups},
{Method: "GET", Pattern: "/api/v1/servers/{name}/jobs", h: a.handleServerJobs},
{Method: "POST", Pattern: "/api/v1/servers/{name}/restore-backup", h: a.handleRestoreBackup}, {Method: "POST", Pattern: "/api/v1/servers/{name}/restore-backup", h: a.handleRestoreBackup},
{Method: "POST", Pattern: "/api/v1/servers/{name}/backup", h: a.handleBackupNow}, {Method: "POST", Pattern: "/api/v1/servers/{name}/backup", h: a.handleBackupNow},
// Server file editor: list / read / write a file in a STOPPED server's world // Server file editor: list / read / write a file in a STOPPED server's world
@@ -479,6 +551,11 @@ func (a *API) externalAPIRoutes() []apiRoute {
// App-tier and owner-scoped (the id must belong to the principal), exactly // App-tier and owner-scoped (the id must belong to the principal), exactly
// like the create/list routes above. // like the create/list routes above.
{Method: "POST", Pattern: "/api/v1/me/submissions/{id}/context", h: a.handleUploadSubmissionContext}, {Method: "POST", Pattern: "/api/v1/me/submissions/{id}/context", h: a.handleUploadSubmissionContext},
// Withdraw the caller's OWN pending submission: the row and its uploaded
// context are deleted, freeing the pending slot and storage budget. Same
// owner-scoping as the upload route — a reviewed submission is frozen (409)
// and another user's id is invisible (404).
{Method: "DELETE", Pattern: "/api/v1/me/submissions/{id}", h: a.handleWithdrawSubmission},
// Admin (Zero-Trust) tier: create / mutate spec / image admission. These gate // Admin (Zero-Trust) tier: create / mutate spec / image admission. These gate
// on Principal.IsAdmin() inside the handler via the adminOnly wrapper, so the // on Principal.IsAdmin() inside the handler via the adminOnly wrapper, so the
// boundary is exercised even where the body is a later-phase stub. // boundary is exercised even where the body is a later-phase stub.
@@ -508,12 +585,22 @@ func (a *API) externalAPIRoutes() []apiRoute {
{Method: "GET", Pattern: "/api/v1/submissions", Admin: true, h: a.handleListSubmissions}, {Method: "GET", Pattern: "/api/v1/submissions", Admin: true, h: a.handleListSubmissions},
{Method: "POST", Pattern: "/api/v1/submissions/{id}/approve", Admin: true, h: a.handleApproveSubmission}, {Method: "POST", Pattern: "/api/v1/submissions/{id}/approve", Admin: true, h: a.handleApproveSubmission},
{Method: "POST", Pattern: "/api/v1/submissions/{id}/reject", Admin: true, h: a.handleRejectSubmission}, {Method: "POST", Pattern: "/api/v1/submissions/{id}/reject", Admin: true, h: a.handleRejectSubmission},
// Retire a submission outright (row + uploaded context), any status. The
// lane's lifecycle valve: without it, rejected/consumed uploads accumulated
// on the uploads PVC forever — there is no other delete path.
{Method: "DELETE", Pattern: "/api/v1/submissions/{id}", Admin: true, h: a.handleDeleteSubmission},
// The reviewer's read path to the uploaded blob: the executed Dockerfile
// lives inside it, so approval would otherwise be blind.
{Method: "GET", Pattern: "/api/v1/submissions/{id}/context", Admin: true, h: a.handleAdminSubmissionContext},
// Auto-update maintenance window (spec §B; decision core internal/updates). // Auto-update maintenance window (spec §B; decision core internal/updates).
// Admin-tier: it governs whether Felis may apply an update to itself, so setting // Admin-tier: it governs whether Felis may apply an update to itself, so setting
// it requires the admin Zero-Trust path, not a mere session. API+persistence // it requires the admin Zero-Trust path, not a mere session. API+persistence
// only — the runner/executors that consume the window are still INTEGRATION-ONLY. // only — the runner/executors that consume the window are still INTEGRATION-ONLY.
{Method: "GET", Pattern: "/api/v1/updates/window", Admin: true, h: a.handleGetUpdateWindow}, {Method: "GET", Pattern: "/api/v1/updates/window", Admin: true, h: a.handleGetUpdateWindow},
{Method: "PUT", Pattern: "/api/v1/updates/window", Admin: true, h: a.handleSetUpdateWindow}, {Method: "PUT", Pattern: "/api/v1/updates/window", Admin: true, h: a.handleSetUpdateWindow},
// Control-plane database backup freshness, as the host's felis-db-backup.timer
// last recorded it. Admin-tier: it names the host backup directory.
{Method: "GET", Pattern: "/api/v1/platform/db-backup", Admin: true, h: a.handleGetDBBackup},
// User admin (spec §7, owner-only). Every route gates on the admin Zero-Trust // User admin (spec §7, owner-only). Every route gates on the admin Zero-Trust
// path AND the owner role: listing, mutating, disabling, or deleting users is // path AND the owner role: listing, mutating, disabling, or deleting users is
@@ -560,7 +647,11 @@ func (a *API) buildFace(routes []apiRoute, guard func(http.Handler) http.Handler
for _, rt := range routes { for _, rt := range routes {
pattern := rt.Method + " " + rt.Pattern pattern := rt.Method + " " + rt.Pattern
if rt.Public { if rt.Public {
mux.HandleFunc(pattern, rt.h) h := rt.h
if rt.AuthDoor {
h = a.throttleAuthDoor(h)
}
mux.HandleFunc(pattern, h)
continue continue
} }
h := rt.h h := rt.h
@@ -673,10 +764,35 @@ func principalFromContext(ctx context.Context) *Principal {
// reserve/release pair closes the intra-replica concurrent burst (the bug fixed in // reserve/release pair closes the intra-replica concurrent burst (the bug fixed in
// #35); cross-replica bounding would need a shared store (out of scope for the // #35); cross-replica bounding would need a shared store (out of scope for the
// single-replica demo). // single-replica demo).
//
// Entries older than the longest window the limiter has been asked about can
// no longer block anything, so checks sweep them out (at most once per
// bucketSweepEvery). Without that, every distinct address typed into a public
// door, whose neutral branch keeps its reservation, stayed in the map for the
// life of the process.
type cooldownLimiter struct { type cooldownLimiter struct {
mu sync.Mutex mu sync.Mutex
now func() time.Time now func() time.Time
last map[string]time.Time last map[string]time.Time
maxWindow time.Duration
swept time.Time
}
// noteWindow widens the retention to window and sweeps stale entries when due.
// The caller holds mu.
func (c *cooldownLimiter) noteWindow(window time.Duration, now time.Time) {
if window > c.maxWindow {
c.maxWindow = window
}
if c.maxWindow <= 0 || now.Sub(c.swept) < bucketSweepEvery {
return
}
c.swept = now
for k, t := range c.last {
if now.Sub(t) >= c.maxWindow {
delete(c.last, k)
}
}
} }
// allowed reports whether name may wake now WITHOUT recording the attempt. A // allowed reports whether name may wake now WITHOUT recording the attempt. A
@@ -691,7 +807,9 @@ func (c *cooldownLimiter) allowed(name string, window time.Duration) bool {
} }
c.mu.Lock() c.mu.Lock()
defer c.mu.Unlock() defer c.mu.Unlock()
if last, ok := c.last[name]; ok && c.now().Sub(last) < window { now := c.now()
c.noteWindow(window, now)
if last, ok := c.last[name]; ok && now.Sub(last) < window {
return false return false
} }
return true return true
@@ -724,10 +842,11 @@ func (c *cooldownLimiter) reserve(name string, window time.Duration) (time.Time,
} }
c.mu.Lock() c.mu.Lock()
defer c.mu.Unlock() defer c.mu.Unlock()
if last, ok := c.last[name]; ok && c.now().Sub(last) < window { t := c.now()
c.noteWindow(window, t)
if last, ok := c.last[name]; ok && t.Sub(last) < window {
return time.Time{}, false return time.Time{}, false
} }
t := c.now()
c.last[name] = t c.last[name] = t
return t, true return t, true
} }
+346 -19
View File
@@ -38,8 +38,20 @@ type fakeRepo struct {
owners map[string]string owners map[string]string
ownersErr error ownersErr error
claimOK map[string]bool // name -> claim succeeds; absent name -> ErrNotFound claimOK map[string]bool // name -> claim succeeds; absent name -> ErrNotFound
audits []AuditEntry // claimQuotaRefuse simulates ClaimServer's atomic quota gate (audit #4)
joins []string // refusing a name whose advisory pre-check already passed.
claimQuotaRefuse map[string]bool
// serverResources / resourceUpdates mirror the cached resource columns:
// ServerResources is what the resize path reads (to preserve storage), and
// UpdateServerResources records the write for assertions.
serverResources map[string]ResourceSpec
resourceUpdates map[string]ResourceSpec
audits []AuditEntry
failAudit error // Audit fails with it (a store outage)
// backupRequested mirrors the newest backup.create audit row per server,
// stamped by Audit with the wall clock (LastBackupRequest).
backupRequested map[string]time.Time
joins []string
// create-server seeding (spec §15) // create-server seeding (spec §15)
seeded map[string]bool // name -> servers row exists seeded map[string]bool // name -> servers row exists
aliases map[string]string // subdomain -> bound server name aliases map[string]string // subdomain -> bound server name
@@ -59,9 +71,15 @@ type fakeRepo struct {
staff map[string]*StaffUser // username -> staff login row staff map[string]*StaffUser // username -> staff login row
sessions map[string]*fakeSession // token_hash -> session sessions map[string]*fakeSession // token_hash -> session
settings map[string][]byte // key -> jsonb value settings map[string][]byte // key -> jsonb value
// failSessionUser / failGetSetting force those reads to fail with a generic
// (non-ErrNotFound) error, simulating a store outage for the 503 auth path.
failSessionUser error
failGetSetting error
// player email OTPs (spec §B2). Keyed by row id; the verify path scans for the // player email OTPs (spec §B2). Keyed by row id; the verify path scans for the
// newest live (user, purpose) just as the PG query does. // newest live (user, purpose) just as the PG query does.
otps map[string]*fakeEmailOTP otps map[string]*fakeEmailOTP
// otpBudget mirrors otp_failure_windows, keyed user|purpose.
otpBudget map[string]*fakeOTPBudget
// op-login requests (spec §B op-login). opLogins mirrors op_login_requests keyed // op-login requests (spec §B op-login). opLogins mirrors op_login_requests keyed
// by id; the in-game approve/finish paths mutate status/consumed in place, and // by id; the in-game approve/finish paths mutate status/consumed in place, and
// tests plant rows directly to drive the status/finish/pending-list paths. // tests plant rows directly to drive the status/finish/pending-list paths.
@@ -91,6 +109,10 @@ type fakeRepo struct {
// user admin fakes // user admin fakes
seededUsers []seededUser seededUsers []seededUser
fakeQuotas map[string]*QuotaView fakeQuotas map[string]*QuotaView
// deletedIDs remembers soft-deleted user ids: DeleteUser drops the row from
// seededUsers (so listings hide it, mirroring the WHERE deleted_at IS NULL
// query), and this set keeps the account dead for the liveness guards.
deletedIDs map[string]bool
// pingErr, when non-nil, is returned by Ping to simulate DB liveness check // pingErr, when non-nil, is returned by Ping to simulate DB liveness check
// failures in /readyz tests. // failures in /readyz tests.
pingErr error pingErr error
@@ -135,6 +157,37 @@ type fakeDataHold struct {
expiresAt time.Time expiresAt time.Time
} }
// fakeOTPBudget mirrors an otp_failure_windows row.
type fakeOTPBudget struct {
windowStart time.Time
failures int
}
// OTPLockedUntil mirrors PGRepo.OTPLockedUntil through the shared otpLockEnd rule.
func (f *fakeRepo) OTPLockedUntil(_ context.Context, userID, purpose string, now time.Time) (time.Time, error) {
b := f.otpBudget[userID+"|"+purpose]
if b == nil {
return time.Time{}, nil
}
return otpLockEnd(b.windowStart, b.failures, now), nil
}
// chargeOTP mirrors chargeOTPMismatch: one wrong guess on the code and the budget.
func (f *fakeRepo) chargeOTP(live *fakeEmailOTP, now time.Time) error {
live.attempts++
key := live.userID + "|" + live.purpose
b := f.otpBudget[key]
if b == nil || !b.windowStart.Add(otpFailureWindow).After(now) {
b = &fakeOTPBudget{windowStart: now}
f.otpBudget[key] = b
}
b.failures++
if b.failures == otpFailureBudget {
return &OTPAccountLockedError{Until: b.windowStart.Add(otpFailureWindow), JustLocked: true}
}
return ErrOTPInvalid
}
// fakeEmailOTP mirrors an email_otps row: only the code hash is held (never the // fakeEmailOTP mirrors an email_otps row: only the code hash is held (never the
// digits), attempts caps brute force, consumed marks single-use, and createdAt // digits), attempts caps brute force, consumed marks single-use, and createdAt
// orders the newest-live lookup. // orders the newest-live lookup.
@@ -202,14 +255,16 @@ func newFakeRepo() *fakeRepo {
allowlist: map[string]map[string]bool{}, allowUUID: map[string]map[string]bool{}, allowlist: map[string]map[string]bool{}, allowUUID: map[string]map[string]bool{},
mine: map[string][]MyServerView{}, mine: map[string][]MyServerView{},
owners: map[string]string{}, owners: map[string]string{},
claimOK: map[string]bool{}, claimOK: map[string]bool{}, claimQuotaRefuse: map[string]bool{},
seeded: map[string]bool{}, aliases: map[string]string{}, serverResources: map[string]ResourceSpec{}, resourceUpdates: map[string]ResourceSpec{},
seeded: map[string]bool{}, aliases: map[string]string{},
linkCodes: map[string]fakeLinkCode{}, links: map[string]string{}, linkCodes: map[string]fakeLinkCode{}, links: map[string]string{},
linkAuthSource: map[string]string{}, linkAuthSource: map[string]string{},
staff: map[string]*StaffUser{}, staff: map[string]*StaffUser{},
sessions: map[string]*fakeSession{}, sessions: map[string]*fakeSession{},
settings: map[string][]byte{}, settings: map[string][]byte{},
otps: map[string]*fakeEmailOTP{}, otps: map[string]*fakeEmailOTP{},
otpBudget: map[string]*fakeOTPBudget{},
opLogins: map[string]*fakeOpLogin{}, opLogins: map[string]*fakeOpLogin{},
setupTokens: map[string]fakeSetupToken{}, setupTokens: map[string]fakeSetupToken{},
blacklist: map[string]bool{}, blacklist: map[string]bool{},
@@ -220,6 +275,7 @@ func newFakeRepo() *fakeRepo {
discoverableChallenges: map[string]*fakeDiscoverableChallenge{}, discoverableChallenges: map[string]*fakeDiscoverableChallenge{},
fakeQuotas: map[string]*QuotaView{}, fakeQuotas: map[string]*QuotaView{},
migrations: map[string]*fakeMigration{}, migrations: map[string]*fakeMigration{},
deletedIDs: map[string]bool{},
} }
} }
@@ -245,10 +301,13 @@ func (f *fakeRepo) QuotaCheck(_ context.Context, userID string, _ string, _ Reso
return f.QuotaAvailable(context.TODO(), userID) return f.QuotaAvailable(context.TODO(), userID)
} }
func (f *fakeRepo) UpdateServerResources(_ context.Context, _ string, _, _, _ int) error { return nil } func (f *fakeRepo) UpdateServerResources(_ context.Context, name string, cpu, mem, stor int) error {
f.resourceUpdates[name] = ResourceSpec{CPUMilli: cpu, MemoryMB: mem, StorageMB: stor}
return nil
}
func (f *fakeRepo) ServerResources(_ context.Context, _ string) (ResourceSpec, error) { func (f *fakeRepo) ServerResources(_ context.Context, name string) (ResourceSpec, error) {
return ResourceSpec{}, nil return f.serverResources[name], nil
} }
func (f *fakeRepo) CreateLinkCode(_ context.Context, code, mcUUID, authSource string, expiresAt time.Time) error { func (f *fakeRepo) CreateLinkCode(_ context.Context, code, mcUUID, authSource string, expiresAt time.Time) error {
f.linkCodes[code] = fakeLinkCode{mcUUID: mcUUID, authSource: authSource, expiresAt: expiresAt} f.linkCodes[code] = fakeLinkCode{mcUUID: mcUUID, authSource: authSource, expiresAt: expiresAt}
@@ -266,7 +325,13 @@ func (f *fakeRepo) VerifyLinkCode(_ context.Context, userID, code string, now ti
return "", "", ErrLinkCodeInvalid return "", "", ErrLinkCodeInvalid
} }
if existing, ok := f.links[rec.mcUUID]; ok && existing != userID { if existing, ok := f.links[rec.mcUUID]; ok && existing != userID {
return "", "", ErrConflict // do not consume another user's pending code // A soft-deleted link's identity is unclaimed: the fresh in-game code lets a
// live caller take it over (mirrors PGRepo). Disabled-but-not-deleted stays a
// conflict — takeover there would bypass the lockout. Neither arm consumes
// the code.
if !f.seededDeleted(existing) {
return "", "", ErrConflict
}
} }
f.links[rec.mcUUID] = userID f.links[rec.mcUUID] = userID
f.linkAuthSource[rec.mcUUID] = rec.authSource // copy/refresh, mirrors DO UPDATE f.linkAuthSource[rec.mcUUID] = rec.authSource // copy/refresh, mirrors DO UPDATE
@@ -299,6 +364,9 @@ func (f *fakeRepo) RedeemPlayerBindCode(_ context.Context, newUserID, code strin
return "", "", "", ErrPlayerBindForbidden // staff must use op.console; do not consume return "", "", "", ErrPlayerBindForbidden // staff must use op.console; do not consume
} }
} }
if f.seededDead(existing) {
return "", "", "", ErrPlayerAccountRetired // dead account; do not consume
}
delete(f.linkCodes, code) delete(f.linkCodes, code)
return existing, rec.mcUUID, rec.authSource, nil return existing, rec.mcUUID, rec.authSource, nil
} }
@@ -343,12 +411,21 @@ func (f *fakeRepo) VerifyEmailOTP(_ context.Context, userID, purpose, codeHash s
if !live.expiresAt.After(now) { if !live.expiresAt.After(now) {
return "", ErrOTPInvalid return "", ErrOTPInvalid
} }
if until, _ := f.OTPLockedUntil(context.Background(), userID, purpose, now); !until.IsZero() {
return "", &OTPAccountLockedError{Until: until}
}
if live.attempts >= otpMaxAttempts { if live.attempts >= otpMaxAttempts {
return "", ErrOTPLocked return "", ErrOTPLocked
} }
if live.codeHash != codeHash { if live.codeHash != codeHash {
live.attempts++ // a typo costs an attempt but does not consume the code return "", f.chargeOTP(live, now) // a typo costs an attempt but does not consume the code
return "", ErrOTPInvalid }
// A DIFFERENT verified holder of the same address → ErrEmailTaken, code left
// live — mirrors PGRepo's guard + the users_verified_email_unique index.
for _, u := range f.staff {
if u.ID != userID && u.EmailVerified && strings.EqualFold(u.Email, live.email) {
return "", ErrEmailTaken
}
} }
live.consumed = true live.consumed = true
for _, u := range f.staff { // flip the user row verified (UPDATE users ...) for _, u := range f.staff { // flip the user row verified (UPDATE users ...)
@@ -595,7 +672,7 @@ func (f *fakeRepo) UUIDInAllowlist(_ context.Context, n, uuid string) (bool, err
return f.allowUUID[n][uuid], nil return f.allowUUID[n][uuid], nil
} }
func (f *fakeRepo) UserByMCUUID(_ context.Context, uuid string) (string, error) { func (f *fakeRepo) UserByMCUUID(_ context.Context, uuid string) (string, error) {
if u, ok := f.links[uuid]; ok { if u, ok := f.links[uuid]; ok && !f.seededDead(u) {
return u, nil return u, nil
} }
return "", ErrNotFound return "", ErrNotFound
@@ -619,15 +696,15 @@ func (f *fakeRepo) IsUsernameBlacklisted(_ context.Context, mcUUID string) (bool
} }
// IsProtectedAdminLink mirrors PGRepo's JOIN of account_links to users: linked, // IsProtectedAdminLink mirrors PGRepo's JOIN of account_links to users: linked,
// auth_source 'thirdparty', and the linked user an admin — no password-hash test, so // auth_source 'thirdparty', and the linked user staff (admin OR owner) — no
// an SSO Operator (role='admin', with no password) is protected like any other. // password-hash test, so an SSO Operator or the Owner is protected like any other.
func (f *fakeRepo) IsProtectedAdminLink(_ context.Context, mcUUID string) (bool, error) { func (f *fakeRepo) IsProtectedAdminLink(_ context.Context, mcUUID string) (bool, error) {
userID, ok := f.links[mcUUID] userID, ok := f.links[mcUUID]
if !ok || f.linkAuthSource[mcUUID] != authSourceThirdParty { if !ok || f.linkAuthSource[mcUUID] != authSourceThirdParty {
return false, nil return false, nil
} }
for _, u := range f.staff { for _, u := range f.staff {
if u.ID == userID && u.Role == "admin" { if u.ID == userID && staffRole(u.Role) {
return true, nil return true, nil
} }
} }
@@ -638,6 +715,9 @@ func (f *fakeRepo) ClaimServer(_ context.Context, n, u string) (bool, error) {
if !present { if !present {
return false, ErrNotFound return false, ErrNotFound
} }
if ok && f.claimQuotaRefuse[n] {
return false, ErrQuotaExceeded // mirrors the atomic gate losing the race
}
return ok, nil return ok, nil
} }
func (f *fakeRepo) RecordJoin(_ context.Context, n, uuid string) error { func (f *fakeRepo) RecordJoin(_ context.Context, n, uuid string) error {
@@ -669,10 +749,36 @@ func (f *fakeRepo) SeedServer(_ context.Context, name, subdomain string, _, _, _
} }
func (f *fakeRepo) Ping(_ context.Context) error { return f.pingErr } func (f *fakeRepo) Ping(_ context.Context) error { return f.pingErr }
func (f *fakeRepo) Audit(_ context.Context, e AuditEntry) error { func (f *fakeRepo) Audit(_ context.Context, e AuditEntry) error {
if f.failAudit != nil {
return f.failAudit
}
f.audits = append(f.audits, e) f.audits = append(f.audits, e)
if e.Action == "backup.create" {
if f.backupRequested == nil {
f.backupRequested = map[string]time.Time{}
}
f.backupRequested[e.ServerName] = time.Now()
}
return nil return nil
} }
func (f *fakeRepo) LastBackupRequest(_ context.Context, serverName string, since time.Time) (time.Time, error) {
if at, ok := f.backupRequested[serverName]; ok && !at.Before(since) {
return at, nil
}
return time.Time{}, nil
}
func (f *fakeRepo) BackupStoreBytes(context.Context) (int64, error) {
var n int64
for _, b := range f.backups {
if b.view.Status == "present" {
n += b.view.SizeBytes
}
}
return n, nil
}
// AllBackups / BackupsForUser / LatestBackup mirror the PG queries' contract so // AllBackups / BackupsForUser / LatestBackup mirror the PG queries' contract so
// the hermetic tests can't pass against a too-lenient fake: only status='present' // the hermetic tests can't pass against a too-lenient fake: only status='present'
// rows are visible, the user scope is the former_owner column, and LatestBackup // rows are visible, the user scope is the former_owner column, and LatestBackup
@@ -768,14 +874,21 @@ func (f *fakeRepo) CreateSession(_ context.Context, tokenHash, userID string, ex
return nil return nil
} }
func (f *fakeRepo) SessionUser(_ context.Context, tokenHash string, now time.Time) (*SessionedUser, error) { func (f *fakeRepo) SessionUser(_ context.Context, tokenHash string, now time.Time) (*SessionedUser, error) {
if f.failSessionUser != nil {
return nil, f.failSessionUser
}
s, ok := f.sessions[tokenHash] s, ok := f.sessions[tokenHash]
if !ok || s.revoked || !s.expiresAt.After(now) { if !ok || s.revoked || !s.expiresAt.After(now) {
return nil, ErrNotFound return nil, ErrNotFound
} }
for _, u := range f.staff { for _, u := range f.staff {
if u.ID == s.userID { if u.ID == s.userID {
if f.seededDead(u.ID) {
return nil, ErrNotFound
}
return &SessionedUser{ return &SessionedUser{
ID: u.ID, Email: u.Email, Role: u.Role, ID: u.ID, Username: u.Username, Email: u.Email, Role: u.Role,
EmailVerified: u.EmailVerified,
}, nil }, nil
} }
} }
@@ -788,6 +901,9 @@ func (f *fakeRepo) RevokeSession(_ context.Context, tokenHash string) error {
return nil return nil
} }
func (f *fakeRepo) GetSetting(_ context.Context, key string) ([]byte, error) { func (f *fakeRepo) GetSetting(_ context.Context, key string) ([]byte, error) {
if f.failGetSetting != nil {
return nil, f.failGetSetting
}
if v, ok := f.settings[key]; ok { if v, ok := f.settings[key]; ok {
return v, nil return v, nil
} }
@@ -863,11 +979,29 @@ func (f *fakeRepo) ListUsers(_ context.Context, opts ListUsersOpts) ([]UserView,
} }
func (f *fakeRepo) UserDetail(_ context.Context, userID string) (*UserDetail, error) { func (f *fakeRepo) UserDetail(_ context.Context, userID string) (*UserDetail, error) {
deletedAt := time.Unix(1_700_000_000, 0)
for _, su := range f.seededUsers { for _, su := range f.seededUsers {
if su.view.ID == userID { if su.view.ID == userID {
return &su.detail, nil return &su.detail, nil
} }
} }
// Legacy fixtures seeded only into f.staff are live accounts (nothing marked
// them disabled or deleted), so detail reads must resolve them too — the
// liveness guards (discoverable login, owner protection) treat "unknown" as a
// fault, and these fixtures are known.
for _, u := range f.staff {
if u.ID == userID {
if f.deletedIDs[userID] {
// A soft-deleted account still HAS a detail row; it is flagged, not gone.
return &UserDetail{UserView: UserView{
ID: u.ID, Username: u.Username, Role: u.Role, Disabled: true,
}, DeletedAt: &deletedAt}, nil
}
return &UserDetail{UserView: UserView{
ID: u.ID, Username: u.Username, Email: u.Email, Role: u.Role,
}}, nil
}
}
return nil, ErrNotFound return nil, ErrNotFound
} }
@@ -904,6 +1038,12 @@ func (f *fakeRepo) UpdateUser(_ context.Context, userID string, patch UpdateUser
f.seededUsers[i].detail.Username = *patch.Username f.seededUsers[i].detail.Username = *patch.Username
} }
if patch.Email != nil { if patch.Email != nil {
// Changing the address voids the proof of it, exactly like PGRepo:
// only VerifyEmailOTP may assert a verified address.
if *patch.Email != f.seededUsers[i].view.Email {
f.seededUsers[i].view.EmailVerified = false
f.seededUsers[i].detail.EmailVerified = false
}
f.seededUsers[i].view.Email = *patch.Email f.seededUsers[i].view.Email = *patch.Email
f.seededUsers[i].detail.Email = *patch.Email f.seededUsers[i].detail.Email = *patch.Email
} }
@@ -921,6 +1061,20 @@ func (f *fakeRepo) DeleteUser(_ context.Context, userID, _ string) error {
for i, su := range f.seededUsers { for i, su := range f.seededUsers {
if su.view.ID == userID { if su.view.ID == userID {
f.seededUsers = append(f.seededUsers[:i], f.seededUsers[i+1:]...) f.seededUsers = append(f.seededUsers[:i], f.seededUsers[i+1:]...)
f.deletedIDs[userID] = true
// Mirror PGRepo: deletion severs the account's identity assets so the
// closed account keeps neither a login credential nor a MC-UUID claim.
for uuid, uid := range f.links {
if uid == userID {
delete(f.links, uuid)
delete(f.linkAuthSource, uuid)
}
}
for cid, cred := range f.passkeyCreds {
if cred.UserID == userID {
delete(f.passkeyCreds, cid)
}
}
return nil return nil
} }
} }
@@ -1068,7 +1222,52 @@ func (f *fakeRepo) RedeemMigration(_ context.Context, targetUserID, codeHash str
// ---- quota admin fakes ---- // ---- quota admin fakes ----
// liveUserExists mirrors PGRepo.requireLiveUser: the admin quota/link fakes
// only act on a live seeded row, so a unit test can drive the unknown-user 404
// the real FK would otherwise turn into a 500.
func (f *fakeRepo) liveUserExists(id string) bool {
for _, su := range f.seededUsers {
if su.view.ID == id && su.detail.DeletedAt == nil {
return true
}
}
return false
}
// seededDead mirrors PGRepo's liveness filters (audit #33): a seeded user that was
// disabled or soft-deleted is dead for the login doors and session validation. A
// fixture that was never seeded (legacy tests put it straight into f.staff) is
// treated as live, matching the fakes' pre-existing behavior.
func (f *fakeRepo) seededDead(id string) bool {
if f.deletedIDs[id] {
return true
}
for _, su := range f.seededUsers {
if su.view.ID == id {
return su.view.Disabled || su.detail.DeletedAt != nil
}
}
return false
}
// seededDeleted is the narrower liveness query: soft-deleted only (a disabled
// account still holds its identity, mirroring VerifyLinkCode's takeover rule).
func (f *fakeRepo) seededDeleted(id string) bool {
if f.deletedIDs[id] {
return true
}
for _, su := range f.seededUsers {
if su.view.ID == id {
return su.detail.DeletedAt != nil
}
}
return false
}
func (f *fakeRepo) GetQuotas(_ context.Context, userID string) (*QuotaView, error) { func (f *fakeRepo) GetQuotas(_ context.Context, userID string) (*QuotaView, error) {
if !f.liveUserExists(userID) {
return nil, ErrNotFound
}
v := &QuotaView{UserID: userID} v := &QuotaView{UserID: userID}
if f.fakeQuotas == nil { if f.fakeQuotas == nil {
return v, nil return v, nil
@@ -1083,6 +1282,9 @@ func (f *fakeRepo) GetQuotas(_ context.Context, userID string) (*QuotaView, erro
} }
func (f *fakeRepo) SetQuotas(_ context.Context, userID string, qi QuotaInput, _ string) (*QuotaView, error) { func (f *fakeRepo) SetQuotas(_ context.Context, userID string, qi QuotaInput, _ string) (*QuotaView, error) {
if !f.liveUserExists(userID) {
return nil, ErrNotFound
}
if f.fakeQuotas == nil { if f.fakeQuotas == nil {
f.fakeQuotas = map[string]*QuotaView{} f.fakeQuotas = map[string]*QuotaView{}
} }
@@ -1135,6 +1337,9 @@ func (f *fakeRepo) UnlinkAccount(_ context.Context, userID, mcUUID string) error
} }
func (f *fakeRepo) LinkAccount(_ context.Context, userID, mcUUID, authSource string) error { func (f *fakeRepo) LinkAccount(_ context.Context, userID, mcUUID, authSource string) error {
if !f.liveUserExists(userID) {
return ErrNotFound
}
if existing, ok := f.links[mcUUID]; ok && existing != userID { if existing, ok := f.links[mcUUID]; ok && existing != userID {
return ErrConflict return ErrConflict
} }
@@ -1153,7 +1358,7 @@ func (f *fakeRepo) LinkAccount(_ context.Context, userID, mcUUID, authSource str
// is indistinguishable from no account: both yield ErrNotFound. // is indistinguishable from no account: both yield ErrNotFound.
func (f *fakeRepo) UserByEmail(_ context.Context, email string) (*StaffUser, error) { func (f *fakeRepo) UserByEmail(_ context.Context, email string) (*StaffUser, error) {
for _, u := range f.staff { for _, u := range f.staff {
if u.EmailVerified && strings.EqualFold(u.Email, email) { if u.EmailVerified && strings.EqualFold(u.Email, email) && !f.seededDead(u.ID) {
su := *u su := *u
return &su, nil return &su, nil
} }
@@ -1181,12 +1386,14 @@ func (f *fakeRepo) ConsumeLoginEmailOTP(_ context.Context, userID, purpose, code
if live == nil || !live.expiresAt.After(now) { if live == nil || !live.expiresAt.After(now) {
return ErrOTPInvalid return ErrOTPInvalid
} }
if until, _ := f.OTPLockedUntil(context.Background(), userID, purpose, now); !until.IsZero() {
return &OTPAccountLockedError{Until: until}
}
if live.attempts >= otpMaxAttempts { if live.attempts >= otpMaxAttempts {
return ErrOTPLocked return ErrOTPLocked
} }
if live.codeHash != codeHash { if live.codeHash != codeHash {
live.attempts++ // a typo costs an attempt but does not consume the code return f.chargeOTP(live, now) // a typo costs an attempt but does not consume the code
return ErrOTPInvalid
} }
live.consumed = true live.consumed = true
return nil return nil
@@ -1313,14 +1520,22 @@ type fakeCluster struct {
desired map[string]v1alpha1.DesiredState desired map[string]v1alpha1.DesiredState
created map[string]CreateServerInput // name -> the validated input it was created from created map[string]CreateServerInput // name -> the validated input it was created from
patched map[string]ServerSpecPatch // name -> the validated spec patch it received patched map[string]ServerSpecPatch // name -> the validated spec patch it received
noWorld map[string]bool // server names modeled WITHOUT a world volume (never started / reaped)
createErr error createErr error
pingErr error pingErr error
// maintErr / wakeErr: what AcquireMaintenance / SetDesiredState(Running)
// return for a server (the world-volume lock, internal/maintenance).
maintErr map[string]error
wakeErr map[string]error
acquired []string // "name:kind" per admitted AcquireMaintenance
released []string // names per ReleaseMaintenance
} }
func newFakeCluster() *fakeCluster { func newFakeCluster() *fakeCluster {
return &fakeCluster{byName: map[string]*ServerInfo{}, bySub: map[string]*ServerInfo{}, return &fakeCluster{byName: map[string]*ServerInfo{}, bySub: map[string]*ServerInfo{},
desired: map[string]v1alpha1.DesiredState{}, created: map[string]CreateServerInput{}, desired: map[string]v1alpha1.DesiredState{}, created: map[string]CreateServerInput{},
patched: map[string]ServerSpecPatch{}} patched: map[string]ServerSpecPatch{}, noWorld: map[string]bool{},
maintErr: map[string]error{}, wakeErr: map[string]error{}}
} }
func (c *fakeCluster) GetServer(_ context.Context, n string) (*ServerInfo, error) { func (c *fakeCluster) GetServer(_ context.Context, n string) (*ServerInfo, error) {
if s, ok := c.byName[n]; ok { if s, ok := c.byName[n]; ok {
@@ -1336,10 +1551,31 @@ func (c *fakeCluster) GetBySubdomain(_ context.Context, s string) (*ServerInfo,
} }
func (c *fakeCluster) ListServers(_ context.Context) ([]ServerInfo, error) { return c.list, nil } func (c *fakeCluster) ListServers(_ context.Context) ([]ServerInfo, error) { return c.list, nil }
func (c *fakeCluster) Ping(_ context.Context) error { return c.pingErr } func (c *fakeCluster) Ping(_ context.Context) error { return c.pingErr }
// WorldVolumeExists models the world PVC: present unless the test named the
// server in noWorld (never started / already reaped).
func (c *fakeCluster) WorldVolumeExists(_ context.Context, n string) (bool, error) {
return !c.noWorld[n], nil
}
func (c *fakeCluster) SetDesiredState(_ context.Context, n string, s v1alpha1.DesiredState) error { func (c *fakeCluster) SetDesiredState(_ context.Context, n string, s v1alpha1.DesiredState) error {
if err := c.wakeErr[n]; err != nil && s == v1alpha1.DesiredRunning {
return err
}
c.desired[n] = s c.desired[n] = s
return nil return nil
} }
func (c *fakeCluster) AcquireMaintenance(_ context.Context, n, kind string) error {
if err := c.maintErr[n]; err != nil {
return err
}
c.acquired = append(c.acquired, n+":"+kind)
return nil
}
func (c *fakeCluster) ReleaseMaintenance(_ context.Context, n string) error {
c.released = append(c.released, n)
return nil
}
func (c *fakeCluster) CreateServer(_ context.Context, in CreateServerInput) error { func (c *fakeCluster) CreateServer(_ context.Context, in CreateServerInput) error {
if c.createErr != nil { if c.createErr != nil {
return c.createErr return c.createErr
@@ -1366,6 +1602,9 @@ func (c *fakeCluster) PatchServerSpec(_ context.Context, n string, p ServerSpecP
if p.AutostartPolicy != nil { if p.AutostartPolicy != nil {
info.AutostartPolicy = string(*p.AutostartPolicy) info.AutostartPolicy = string(*p.AutostartPolicy)
} }
if p.IdleStopSeconds != nil {
info.IdleStopSeconds = *p.IdleStopSeconds
}
return nil return nil
} }
@@ -1671,6 +1910,44 @@ func TestFleetAdminRead(t *testing.T) {
} }
}) })
t.Run("system services are marked read-only", func(t *testing.T) {
// The login gate and the lobby carry reserved names, so every per-server
// route rejects them; the fleet row must say "system" so the cockpit
// renders them without actions that would 400.
sysCl := newFakeCluster()
sysCl.list = []ServerInfo{
{Name: "login", Phase: "Running", Ready: true},
{Name: "lobby", Phase: "Running", Ready: true},
{Name: "survival", Phase: "Stopped"},
}
api := newTestAPI(newFakeRepo(), sysCl)
api.External = staticExternal{p: &Principal{UserID: "a1", Email: "[email protected]",
Role: "admin", ViaAdminAccess: true}}
w := do(api.ExternalHandler(), "GET", "/api/v1/fleet", "", nil)
if w.Code != http.StatusOK {
t.Fatalf("code = %d, want 200 (%s)", w.Code, w.Body.String())
}
var got struct {
Servers []struct {
Name string `json:"name"`
System bool `json:"system"`
} `json:"servers"`
}
if err := json.Unmarshal(w.Body.Bytes(), &got); err != nil {
t.Fatalf("body not JSON: %v", err)
}
byName := map[string]bool{}
for _, r := range got.Servers {
byName[r.Name] = r.System
}
if !byName["login"] || !byName["lobby"] {
t.Errorf("system flags = %+v, want login+lobby marked", byName)
}
if byName["survival"] {
t.Errorf("survival marked system; only platform services are")
}
})
t.Run("owner merges for claimed, absent for unclaimed", func(t *testing.T) { t.Run("owner merges for claimed, absent for unclaimed", func(t *testing.T) {
repo := newFakeRepo() repo := newFakeRepo()
// Only "survival" is claimed; "creative"/"skyblock" stay unowned. // Only "survival" is claimed; "creative"/"skyblock" stay unowned.
@@ -1760,6 +2037,21 @@ func TestClaimStateMachine(t *testing.T) {
t.Fatalf("code = %d body %s", w.Code, w.Body.String()) t.Fatalf("code = %d body %s", w.Code, w.Body.String())
} }
}) })
t.Run("atomic gate refusal -> 403 quota_exceeded", func(t *testing.T) {
// The advisory pre-check passed, but ClaimServer's serialized re-check
// (audit #4) refuses: the caller must see the same 403, not a 500.
repo := newFakeRepo()
repo.linked["u1"] = true
repo.quota["u1"] = true
repo.claimOK["survival"] = true
repo.claimQuotaRefuse["survival"] = true
api := newTestAPI(repo, newFakeCluster())
api.External = staticExternal{p: user}
w := do(api.ExternalHandler(), "POST", "/api/v1/servers/survival/claim", "", nil)
if w.Code != http.StatusForbidden || decodeErr(t, w) != "quota_exceeded" {
t.Fatalf("code = %d body %s, want 403 quota_exceeded", w.Code, w.Body.String())
}
})
t.Run("already claimed -> 409", func(t *testing.T) { t.Run("already claimed -> 409", func(t *testing.T) {
repo := newFakeRepo() repo := newFakeRepo()
repo.linked["u1"] = true repo.linked["u1"] = true
@@ -2195,3 +2487,38 @@ func TestAccessVerifier(t *testing.T) {
} }
}) })
} }
// TestSessionAuthOutageIs503Not401: a session-store outage must surface as 503
// auth_unavailable, not a 401 that reads as "please log in again". Both failure
// points are covered — the local_auth_enabled read and the session row read —
// plus the regression that a genuinely missing session still answers 401.
func TestSessionAuthOutageIs503Not401(t *testing.T) {
apiWith := func(repo *fakeRepo) *API {
a := newTestAPI(repo, newFakeCluster())
a.External = SessionAuth{Repo: repo, RootDomain: testRoot, AdminHostname: "op.console." + testRoot}
return a
}
cookie := map[string]string{"Cookie": sessionCookieName + "=any"}
outage := errors.New("dial tcp 10.0.0.5:5432: connect: connection refused")
repo := newFakeRepo()
repo.settings[LocalAuthEnabledKey] = []byte("true")
repo.failGetSetting = outage
if w := do(apiWith(repo).ExternalHandler(), "GET", "/api/v1/me", "", cookie); w.Code != http.StatusServiceUnavailable || decodeErr(t, w) != "auth_unavailable" {
t.Fatalf("settings read outage = %d body %s, want 503 auth_unavailable", w.Code, w.Body.String())
}
repo = newFakeRepo()
repo.settings[LocalAuthEnabledKey] = []byte("true")
repo.failSessionUser = outage
if w := do(apiWith(repo).ExternalHandler(), "GET", "/api/v1/me", "", cookie); w.Code != http.StatusServiceUnavailable || decodeErr(t, w) != "auth_unavailable" {
t.Fatalf("session read outage = %d body %s, want 503 auth_unavailable", w.Code, w.Body.String())
}
// Regression: fail-closed auth (missing/invalid session) stays a 401.
repo = newFakeRepo()
repo.settings[LocalAuthEnabledKey] = []byte("true")
if w := do(apiWith(repo).ExternalHandler(), "GET", "/api/v1/me", "", cookie); w.Code != http.StatusUnauthorized || decodeErr(t, w) != "unauthorized" {
t.Fatalf("missing session = %d body %s, want 401 unauthorized", w.Code, w.Body.String())
}
}
+185
View File
@@ -0,0 +1,185 @@
package api
import (
"cmp"
"context"
"encoding/json"
"errors"
"log"
"net/http"
"strings"
"time"
"unicode/utf8"
"felis.lolicon.best/internal/metrics"
)
// Audit rows (audit_logs, spec §6).
//
// Every row an HTTP request writes carries the acting account's id
// (actor_user_id), the caller's address (the one the sign-in limit keys on) and
// user agent; actor is display text. A failed write never fails the operation
// it records, which already happened, but it is logged and counted
// (felis_audit_write_failures_total, FelisAuditWriteFailing): a silent drop is
// how a database blip erases the trail.
const (
// auditWriteTimeout bounds one audit insert. The write outlives the caller's
// request context, so a client that hangs up right after the action cannot
// cancel its own audit row.
auditWriteTimeout = 5 * time.Second
// auditUserAgentMax bounds the stored user agent; the header is the caller's
// to write.
auditUserAgentMax = 256
// anonymousActor names a caller no account was resolved for.
anonymousActor = "anonymous"
)
// auditActor is the display name for a principal: an email only when something
// vouches for it (an Access JWT, or a session whose address was verified), else
// the username. A player can set their address to anyone's before verifying it,
// so an unverified email would let them sign rows as that person.
func auditActor(p *Principal) string {
switch {
case p == nil:
return anonymousActor
case p.Email != "" && (p.EmailVerified || !p.ViaSession):
return p.Email
case p.Username != "":
return p.Username
case p.UserID != "":
return p.UserID
}
return anonymousActor
}
// audit records an action by the signed-in caller. target is the object acted
// on (a server name, a user or credential id) and lands in server_name.
func (a *API) audit(r *http.Request, action, target string) {
p := principalFromContext(r.Context())
e := AuditEntry{Actor: auditActor(p), Action: action, ServerName: target}
if p != nil {
e.ActorUserID = p.UserID
}
a.auditEntry(r, e)
}
// auditImageChange records a confirmed image change as server.patch with the
// image it replaced and the one it set, so the audit log alone can say which
// build a world ran before it was moved.
func (a *API) auditImageChange(r *http.Request, server, from, to string) {
p := principalFromContext(r.Context())
e := AuditEntry{Actor: auditActor(p), Action: "server.patch", ServerName: server}
if p != nil {
e.ActorUserID = p.UserID
}
e.Payload = auditPayload(map[string]any{"image_from": from, "image_to": to})
a.auditEntry(r, e)
}
// auditAccount records an action a pre-session door took for the account it
// resolved (u nil: none was). The username is the actor: the door has not yet
// proven anything about the address.
func (a *API) auditAccount(r *http.Request, u *StaffUser, action, target string) {
e := AuditEntry{Actor: anonymousActor, Action: action, ServerName: target}
if u != nil {
e.Actor, e.ActorUserID = u.Username, u.ID
}
a.auditEntry(r, e)
}
// auditEntry fills the request detail into e and writes it. Source defaults to
// external; internal callers set it and the component actor themselves.
func (a *API) auditEntry(r *http.Request, e AuditEntry) {
if e.Source == "" {
e.Source = "external"
}
e.RequestID = requestIDFromContext(r.Context())
if e.Source == "external" {
if ip := a.clientIP(r); ip.IsValid() {
e.ClientIP = ip.String()
}
e.UserAgent = truncateUTF8(r.UserAgent(), auditUserAgentMax)
}
a.writeAudit(r.Context(), e)
}
// writeAudit inserts e, logging and counting a failure.
func (a *API) writeAudit(ctx context.Context, e AuditEntry) {
ctx, cancel := context.WithTimeout(context.WithoutCancel(ctx), auditWriteTimeout)
defer cancel()
if err := a.Repo.Audit(ctx, e); err != nil {
metrics.AuditWriteFailuresTotal.Inc()
log.Printf("audit: lost %s by %s (user %q, request_id=%s): %v",
e.Action, e.Actor, e.ActorUserID, e.RequestID, err)
}
}
// authFailure records one refused sign-in attempt: felis_auth_failures_total
// by door and reason, and an auth.<door>.failed row naming the account when
// the door resolved one (u nil: the signed-in caller if any, else anonymous).
// The doors keep their answers uniform so a prober learns nothing; the reason
// is for the operator.
func (a *API) authFailure(r *http.Request, door, reason string, u *StaffUser) {
metrics.AuthFailuresTotal.WithLabelValues(door, reason).Inc()
e := AuditEntry{Action: "auth." + door + ".failed", Payload: auditPayload(map[string]any{"reason": reason})}
switch p := principalFromContext(r.Context()); {
case u != nil:
e.Actor, e.ActorUserID = cmp.Or(u.Username, u.ID), u.ID
case p != nil:
e.Actor, e.ActorUserID = auditActor(p), p.UserID
default:
e.Actor = anonymousActor
}
a.auditEntry(r, e)
}
// passkeyCloneRejected records an assertion refused for a regressed signature
// counter. It keeps its own action so a cloned authenticator stands out from
// ordinary failures, and counts as a failure of its door. u nil: a signed-in
// step-up, attributed to the caller.
func (a *API) passkeyCloneRejected(r *http.Request, door string, u *StaffUser, credentialID string) {
metrics.AuthFailuresTotal.WithLabelValues(door, "clone_rejected").Inc()
if u == nil {
a.audit(r, "auth.passkey_clone_rejected", credentialID)
return
}
a.auditAccount(r, u, "auth.passkey_clone_rejected", credentialID)
}
// isOTPRefusal reports whether err is a refused code (wrong, spent, or the
// account's budget locked), as opposed to a fault.
func isOTPRefusal(err error) bool {
return errors.Is(err, ErrOTPInvalid) || errors.Is(err, ErrOTPLocked) || errors.Is(err, ErrOTPAccountLocked)
}
// otpFailureReason names a refused code for authFailure.
func otpFailureReason(err error) string {
switch {
case errors.Is(err, ErrOTPAccountLocked):
return "account_locked"
case errors.Is(err, ErrOTPLocked):
return "code_locked"
}
return "bad_code"
}
// auditPayload marshals a small detail map for AuditEntry.Payload.
func auditPayload(v map[string]any) []byte {
b, _ := json.Marshal(v)
return b
}
// truncateUTF8 makes s valid UTF-8 (a header may carry any byte, a text
// column refuses invalid sequences) and cuts it to at most n bytes on a rune
// boundary.
func truncateUTF8(s string, n int) string {
s = strings.ToValidUTF8(s, "\uFFFD")
if len(s) <= n {
return s
}
for n > 0 && !utf8.RuneStart(s[n]) {
n--
}
return s[:n]
}
+177
View File
@@ -0,0 +1,177 @@
package api
import (
"errors"
"net/http"
"strings"
"testing"
"time"
"github.com/prometheus/client_golang/prometheus/testutil"
"felis.lolicon.best/internal/metrics"
)
// Audit attribution: a row names the acting account by id and never by an
// address the caller merely asserted; refused sign-ins, throttling and logouts
// leave rows; a failed write is counted instead of vanishing.
func TestAuditActorIgnoresUnverifiedEmail(t *testing.T) {
cases := []struct {
name string
p *Principal
want string
}{
{"verified session email", &Principal{UserID: "u1", Username: "alice", Email: "[email protected]", EmailVerified: true, ViaSession: true}, "[email protected]"},
{"unverified session email", &Principal{UserID: "u2", Username: "mallory", Email: "[email protected]", ViaSession: true}, "mallory"},
{"access jwt email", &Principal{UserID: "sub", Email: "[email protected]"}, "[email protected]"},
{"no email", &Principal{UserID: "u3", Username: "bob", ViaSession: true}, "bob"},
{"id only", &Principal{UserID: "u4", ViaSession: true}, "u4"},
{"nobody", nil, anonymousActor},
}
for _, c := range cases {
if got := auditActor(c.p); got != c.want {
t.Errorf("%s: auditActor = %q, want %q", c.name, got, c.want)
}
}
}
// A player who sets their address to the owner's still signs every row as
// themselves, by username and by id.
func TestAuditCannotBeSignedWithAnotherPersonsEmail(t *testing.T) {
repo := newFakeRepo()
repo.settings[LocalAuthEnabledKey] = []byte("true")
repo.staff["owner"] = &StaffUser{ID: "u1", Username: "owner", Email: "[email protected]", Role: "owner", EmailVerified: true}
repo.staff["mallory"] = &StaffUser{ID: "u2", Username: "mallory", Email: "[email protected]", Role: "user", EmailVerified: true}
repo.sessions[hashCookie("tok")] = &fakeSession{userID: "u2", expiresAt: time.Unix(1_700_000_000, 0).Add(time.Hour)}
api := newTestAPI(repo, newFakeCluster())
api.External = SessionAuth{Repo: repo, RootDomain: testRoot, Now: api.now}
api.ClientIPHeader = "CF-Connecting-IP"
eh := api.ExternalHandler()
hdr := map[string]string{
"Content-Type": "application/json", "Cookie": sessionCookieName + "=tok",
"CF-Connecting-IP": "203.0.113.5", "User-Agent": "probe/1.0",
}
for _, email := range []string{"[email protected]", "[email protected]"} {
if w := do(eh, "POST", "/api/v1/account/email", `{"email":"`+email+`"}`, hdr); w.Code != http.StatusOK {
t.Fatalf("set email = %d (%s)", w.Code, w.Body.String())
}
}
// The first write ran while the address was still verified; the second
// ran with the owner's address set and unverified.
last := repo.audits[len(repo.audits)-1]
if last.Actor != "mallory" || last.ActorUserID != "u2" {
t.Fatalf("audit after spoofing = actor %q user %q, want mallory/u2", last.Actor, last.ActorUserID)
}
if last.ClientIP != "203.0.113.5" || last.UserAgent != "probe/1.0" || last.RequestID == "" {
t.Fatalf("audit request detail = %+v", last)
}
}
func TestSignInFailuresAreAuditedAndCounted(t *testing.T) {
api, repo, mailer := seedLoginEmailAPI(t)
eh := api.ExternalHandler()
noAccount := metrics.AuthFailuresTotal.WithLabelValues("login_email", "no_account")
badCode := metrics.AuthFailuresTotal.WithLabelValues("login_email", "bad_code")
n0, b0 := testutil.ToFloat64(noAccount), testutil.ToFloat64(badCode)
do(eh, "POST", "/api/v1/auth/email/verify", `{"email":"[email protected]","code":"123456"}`, jsonHeader)
if w := do(eh, "POST", "/api/v1/auth/email/start", `{"email":"[email protected]"}`, jsonHeader); w.Code != http.StatusAccepted {
t.Fatalf("start = %d", w.Code)
}
wrong := "000000"
if mailer.code == wrong {
wrong = "111111"
}
do(eh, "POST", "/api/v1/auth/email/verify", `{"email":"[email protected]","code":"`+wrong+`"}`, jsonHeader)
if got := testutil.ToFloat64(noAccount) - n0; got != 1 {
t.Errorf("no_account failures counted %v, want 1", got)
}
if got := testutil.ToFloat64(badCode) - b0; got != 1 {
t.Errorf("bad_code failures counted %v, want 1", got)
}
var failed []AuditEntry
for _, e := range repo.audits {
if e.Action == "auth.login_email.failed" {
failed = append(failed, e)
}
}
if len(failed) != 2 {
t.Fatalf("failure audits = %+v, want 2", failed)
}
if failed[0].Actor != anonymousActor || failed[0].ActorUserID != "" || !strings.Contains(string(failed[0].Payload), "no_account") {
t.Errorf("unknown-address failure = %+v", failed[0])
}
if failed[1].Actor != "player" || failed[1].ActorUserID != "u1" || !strings.Contains(string(failed[1].Payload), "bad_code") {
t.Errorf("wrong-code failure = %+v", failed[1])
}
}
func TestThrottleAuditsOncePerEpisode(t *testing.T) {
api, repo, _ := seedLoginEmailAPI(t)
api.AuthDoorLimit = RateLimit{Burst: 1, PerMinute: 1}
clock := time.Unix(1_700_000_000, 0)
api.Now = func() time.Time { return clock }
eh := api.ExternalHandler()
options := func() int {
return do(eh, "POST", "/api/v1/auth/options", `{"email":"[email protected]"}`, jsonHeader).Code
}
throttled := func() int {
n := 0
for _, e := range repo.audits {
if e.Action == "auth.rate_limited" {
n++
}
}
return n
}
options()
for i := 0; i < 5; i++ {
if c := options(); c != http.StatusTooManyRequests {
t.Fatalf("call %d = %d, want 429", i+2, c)
}
}
if n := throttled(); n != 1 {
t.Fatalf("5 refusals left %d audit rows, want 1", n)
}
clock = clock.Add(time.Minute)
options()
options()
if n := throttled(); n != 2 {
t.Fatalf("a second episode left %d rows in total, want 2", n)
}
}
func TestAuditWriteFailureIsCountedNotFatal(t *testing.T) {
api, repo, _ := seedLoginEmailAPI(t)
repo.failAudit = errors.New("db down")
before := testutil.ToFloat64(metrics.AuditWriteFailuresTotal)
if w := do(api.ExternalHandler(), "POST", "/api/v1/auth/email/start", `{"email":"[email protected]"}`, jsonHeader); w.Code != http.StatusAccepted {
t.Fatalf("start with the audit store down = %d, want 202", w.Code)
}
if got := testutil.ToFloat64(metrics.AuditWriteFailuresTotal) - before; got != 1 {
t.Fatalf("audit write failures counted %v, want 1", got)
}
}
func TestLogoutAuditsTheLiveSession(t *testing.T) {
api, repo, _ := seedLoginEmailAPI(t)
repo.sessions[hashCookie("tok")] = &fakeSession{userID: "u1", expiresAt: time.Unix(1_700_000_000, 0).Add(time.Hour)}
eh := api.ExternalHandler()
cookie := map[string]string{"Content-Type": "application/json", "Cookie": sessionCookieName + "=tok"}
do(eh, "POST", "/api/v1/auth/logout", "", cookie)
do(eh, "POST", "/api/v1/auth/logout", "", cookie) // already revoked: no second row
if len(repo.audits) != 1 || repo.audits[0].Action != "auth.logout" || repo.audits[0].ActorUserID != "u1" || repo.audits[0].Actor != "player" {
t.Fatalf("logout audits = %+v, want one auth.logout by player/u1", repo.audits)
}
}
func TestTruncateUTF8KeepsRunesWhole(t *testing.T) {
if got := truncateUTF8("ab\xffc", 10); got != "ab�c" {
t.Errorf("invalid byte = %q", got)
}
if got := truncateUTF8("猫猫", 4); got != "猫" {
t.Errorf("cut mid-rune = %q, want 猫", got)
}
}
+7 -4
View File
@@ -15,15 +15,18 @@ import (
type Principal struct { type Principal struct {
// UserID is the stable web identity (SSO subject → users.id). // UserID is the stable web identity (SSO subject → users.id).
UserID string UserID string
// Email is the audited actor identity (spec §14: audit actor = Access email). // Username is the account's login name; empty for an Access-JWT caller.
Username string
// Email is the account's address. Only an Access JWT or EmailVerified vouches
// for it: a player can set any address before verifying it (auditActor).
Email string Email string
// Role is "admin" or "user" (mirrors users.role). // Role is "owner", "admin", or "user" (mirrors users.role).
Role string Role string
// ViaAdminAccess is true only when the request arrived through an admin-graded // ViaAdminAccess is true only when the request arrived through an admin-graded
// path: the admin.* Zero-Trust hostname (Cloudflare Access, the remote face) OR // path: the admin.* Zero-Trust hostname (Cloudflare Access, the remote face) OR
// a local session presented on the op.console host (SessionAuth, the // a local session presented on the op.console host (SessionAuth, the
// passwordless face). Admin-tier operations require it in addition to // passwordless face). Admin-tier operations require it in addition to
// Role=="admin" (spec §14: ZT is graded by operation). A role=admin session // a staff role (spec §14: ZT is graded by operation). A staff session
// arriving on the player console (console.*) never sets it. // arriving on the player console (console.*) never sets it.
ViaAdminAccess bool ViaAdminAccess bool
// EmailVerified mirrors users.email_verified. The lockdown middleware gates // EmailVerified mirrors users.email_verified. The lockdown middleware gates
@@ -49,7 +52,7 @@ func staffRole(role string) bool {
} }
// IsAdmin reports whether the principal may perform admin-tier operations. // IsAdmin reports whether the principal may perform admin-tier operations.
// Both the role claim and the admin Access path are required: a role=admin // Both the role claim and the admin Access path are required: a staff
// session arriving on panel.* must not bypass the Zero-Trust boundary. // session arriving on panel.* must not bypass the Zero-Trust boundary.
// An owner implicitly passes this check (the owner role is a superset of admin). // An owner implicitly passes this check (the owner role is a superset of admin).
func (p *Principal) IsAdmin() bool { func (p *Principal) IsAdmin() bool {
+26 -1
View File
@@ -28,6 +28,12 @@ type ServerInfo struct {
JavaMemory string `json:"javaMemory,omitempty"` JavaMemory string `json:"javaMemory,omitempty"`
StorageSize string `json:"storageSize,omitempty"` StorageSize string `json:"storageSize,omitempty"`
CPU string `json:"cpu,omitempty"` CPU string `json:"cpu,omitempty"`
// IdleStopSeconds is how long the server may sit empty before idle
// auto-stop scales it down; 0 means it never idles out.
IdleStopSeconds int32 `json:"idleStopSeconds"`
// PlayerCountUnknown is true while the operator cannot read the player
// count over RCON; idle auto-stop waits until it can.
PlayerCountUnknown bool `json:"playerCountUnknown,omitempty"`
} }
// CreateServerInput is the validated, structured create-server form (spec §15). // CreateServerInput is the validated, structured create-server form (spec §15).
@@ -67,6 +73,9 @@ type ServerSpecPatch struct {
// (felis-api resolves both from the same form) or both stay nil. // (felis-api resolves both from the same form) or both stay nil.
JavaMemory *string JavaMemory *string
Resources *corev1.ResourceRequirements Resources *corev1.ResourceRequirements
// IdleStopSeconds sets idle auto-stop: 0 turns it off, anything else is the
// empty duration before the stop (already range-checked).
IdleStopSeconds *int32
} }
// Cluster is the lifecycle-layer access the API depends on: reads of the // Cluster is the lifecycle-layer access the API depends on: reads of the
@@ -81,6 +90,12 @@ type Cluster interface {
// GetServer reads one MinecraftServer's lifecycle view, or ErrNotFound. // GetServer reads one MinecraftServer's lifecycle view, or ErrNotFound.
GetServer(ctx context.Context, name string) (*ServerInfo, error) GetServer(ctx context.Context, name string) (*ServerInfo, error)
// WorldVolumeExists reports whether the server's world PVC exists in the
// server namespace. A server that never started — or whose world the
// retention reaper already archived and deleted — has no claim, and a
// backup/restore Job would hang Pending on the missing volume with nothing
// ever recorded, so both handlers refuse those up front.
WorldVolumeExists(ctx context.Context, name string) (bool, error)
// GetBySubdomain finds the MinecraftServer whose spec.subdomain matches, or // GetBySubdomain finds the MinecraftServer whose spec.subdomain matches, or
// ErrNotFound. // ErrNotFound.
GetBySubdomain(ctx context.Context, subdomain string) (*ServerInfo, error) GetBySubdomain(ctx context.Context, subdomain string) (*ServerInfo, error)
@@ -88,8 +103,18 @@ type Cluster interface {
// velocity registration pull (spec §7 GET /servers). // velocity registration pull (spec §7 GET /servers).
ListServers(ctx context.Context) ([]ServerInfo, error) ListServers(ctx context.Context) ([]ServerInfo, error)
// SetDesiredState flips spec.desiredState — the only write the API performs // SetDesiredState flips spec.desiredState — the only write the API performs
// against the CRD (spec §9.1). It is idempotent. // against the CRD (spec §9.1). It is idempotent. Flipping to Running returns a
// *MaintenanceBusyError (errors.Is ErrMaintenanceInProgress) while a restore,
// backup or file write holds the world volume.
SetDesiredState(ctx context.Context, name string, state v1alpha1.DesiredState) error SetDesiredState(ctx context.Context, name string, state v1alpha1.DesiredState) error
// AcquireMaintenance admits one world-volume operation (internal/maintenance
// kind): ErrNotStopped unless the server is fully stopped, a
// *MaintenanceBusyError while another operation holds the volume. The check
// and the lock are one atomic write against a concurrent wake.
AcquireMaintenance(ctx context.Context, name, kind string) error
// ReleaseMaintenance drops the admission lock once the operation's Job exists
// (or could not be created). It is idempotent.
ReleaseMaintenance(ctx context.Context, name string) error
// CreateServer creates a MinecraftServer CRD from the validated form (spec // CreateServer creates a MinecraftServer CRD from the validated form (spec
// §15). It returns ErrConflict if a server of that name already exists. // §15). It returns ErrConflict if a server of that name already exists.
CreateServer(ctx context.Context, in CreateServerInput) error CreateServer(ctx context.Context, in CreateServerInput) error
+79 -8
View File
@@ -6,6 +6,8 @@ import (
"fmt" "fmt"
"log" "log"
"net/http" "net/http"
"strconv"
"time"
) )
// Sentinel errors the repository and cluster layers return so handlers can map // Sentinel errors the repository and cluster layers return so handlers can map
@@ -15,6 +17,13 @@ var (
ErrNotFound = errors.New("not found") ErrNotFound = errors.New("not found")
// ErrConflict means an atomic precondition failed (e.g. claim lost the race). // ErrConflict means an atomic precondition failed (e.g. claim lost the race).
ErrConflict = errors.New("conflict") ErrConflict = errors.New("conflict")
// ErrQuotaExceeded means an ownership write would push the user over a quota
// cap (spec §9.3). ClaimServer — the atomic gate — returns it when a claim
// passes the handler's advisory pre-check but loses the serialized re-check
// (two concurrent claims by one user); handlers map it to a 403
// quota_exceeded, the same answer the pre-check gives, so the CONCURRENT case
// and the SEQUENTIAL case are indistinguishable to the caller.
ErrQuotaExceeded = errors.New("server quota exhausted")
// ErrLinkCodeInvalid means an account-link code is unknown or expired (spec // ErrLinkCodeInvalid means an account-link code is unknown or expired (spec
// §10). It is a client error (the verify endpoint exists; the code is bad), so // §10). It is a client error (the verify endpoint exists; the code is bad), so
// handlers map it to 400, not 404. // handlers map it to 400, not 404.
@@ -37,6 +46,11 @@ var (
// can answer 429 (back off / request a new code) rather than inviting another // can answer 429 (back off / request a new code) rather than inviting another
// guess against a code that will never accept one. // guess against a code that will never accept one.
ErrOTPLocked = errors.New("email code locked: too many attempts") ErrOTPLocked = errors.New("email code locked: too many attempts")
// ErrOTPAccountLocked means the (user, purpose) has spent its wrong-code budget
// for the current window (otpFailureBudget): every code for that door is refused,
// the right one included, until the window ends. The repo returns it as an
// *OTPAccountLockedError carrying the end of the lock.
ErrOTPAccountLocked = errors.New("email codes locked for this account: too many wrong codes")
// ErrPasskeyChallengeInvalid means a passkey enrollment ceremony cannot be // ErrPasskeyChallengeInvalid means a passkey enrollment ceremony cannot be
// finished: there is no live (unconsumed, unexpired) challenge for the caller and // finished: there is no live (unconsumed, unexpired) challenge for the caller and
// purpose (Phase 6 WebAuthn bind). Like ErrOTPInvalid it is a client error — the // purpose (Phase 6 WebAuthn bind). Like ErrOTPInvalid it is a client error — the
@@ -44,14 +58,22 @@ var (
// consumed, or expired) — so handlers map it to 400, not 404. // consumed, or expired) — so handlers map it to 400, not 404.
ErrPasskeyChallengeInvalid = errors.New("passkey challenge invalid or expired") ErrPasskeyChallengeInvalid = errors.New("passkey challenge invalid or expired")
// ErrPlayerBindForbidden means a public Bind-Code redemption resolved to a STAFF // ErrPlayerBindForbidden means a public Bind-Code redemption resolved to a STAFF
// account (role=admin), which the player-console bootstrap refuses (console-tier // account (admin or owner), which the player-console bootstrap refuses
// access model). Operators authenticate at op.console behind Zero Trust, never via // (console-tier access model). Staff authenticate at op.console behind Zero Trust,
// the account-less console.<root_domain> door, so the public bootstrap provably // never via the account-less console.<root_domain> door, so the public bootstrap
// never mints a session for an admin identity. It is distinct from ErrConflict so // provably never mints a session for a staff identity. It is distinct from
// the handler answers 403 (wrong door) rather than 409 (already linked). // ErrConflict so the handler answers 403 (wrong door) rather than 409.
ErrPlayerBindForbidden = errors.New("bind code belongs to a staff account") ErrPlayerBindForbidden = errors.New("bind code belongs to a staff account")
// ErrPlayerAccountRetired means a Bind-Code redemption resolved to an account the
// platform has closed: an owner soft-deleted it, or it is disabled (locked out).
// Reusing the row would mint a fresh session for a dead account — the same
// resurrection the login doors refuse by resolving only live accounts — so the
// redeemer gets an explicit 403 instead. The code is NOT consumed, so re-enabling
// the account and retrying still works within the code's TTL.
ErrPlayerAccountRetired = errors.New("player account is retired or disabled")
// ErrEmailTaken means a verified email would collide with another account's // ErrEmailTaken means a verified email would collide with another account's
// already-verified address (spec §B email-first login foundation, migration 0010). // already-verified address (spec §B email-first login foundation; the
// users_verified_email_unique index ships in migration 0020).
// VerifyEmailOTP returns it — WITHOUT consuming the code, since the address, not // VerifyEmailOTP returns it — WITHOUT consuming the code, since the address, not
// the code, is the problem — when a DIFFERENT user has already proven the same // the code, is the problem — when a DIFFERENT user has already proven the same
// address case-insensitively. It is the clean, application-level counterpart of // address case-insensitively. It is the clean, application-level counterpart of
@@ -67,8 +89,40 @@ var (
// sentinels so the handler answers 429 (a transient "too busy, retry" — the cap self-clears // sentinels so the handler answers 429 (a transient "too busy, retry" — the cap self-clears
// as challenges expire), never a 400 that invites an immediate retry. // as challenges expire), never a 400 that invites an immediate retry.
ErrTooManyDiscoverableChallenges = errors.New("too many discoverable login challenges in flight") ErrTooManyDiscoverableChallenges = errors.New("too many discoverable login challenges in flight")
// ErrNotStopped means a world-volume operation was refused because the server is
// not fully stopped: desiredState is not Stopped, or its pod is still shutting
// down (phase Stopping) and holds the volume while it saves.
ErrNotStopped = errors.New("server is not stopped")
// ErrMaintenanceInProgress means another operation holds the server's world
// volume (internal/maintenance). Cluster methods return it wrapped in a
// *MaintenanceBusyError that names the holder.
ErrMaintenanceInProgress = errors.New("world maintenance in progress")
) )
// MaintenanceBusyError names what holds a server's world volume. errors.Is
// matches it against ErrMaintenanceInProgress.
type MaintenanceBusyError struct{ Kind string }
func (e *MaintenanceBusyError) Error() string {
return "world maintenance in progress: " + e.Kind
}
func (e *MaintenanceBusyError) Is(target error) bool { return target == ErrMaintenanceInProgress }
// OTPAccountLockedError is ErrOTPAccountLocked with its detail. JustLocked is set
// only on the wrong guess that spent the budget, so the handler notifies and
// audits the lock exactly once.
type OTPAccountLockedError struct {
Until time.Time
JustLocked bool
}
func (e *OTPAccountLockedError) Error() string {
return ErrOTPAccountLocked.Error() + " until " + e.Until.UTC().Format(time.RFC3339)
}
func (e *OTPAccountLockedError) Is(target error) bool { return target == ErrOTPAccountLocked }
// apiError is a handler-level error carrying an HTTP status and a stable, // apiError is a handler-level error carrying an HTTP status and a stable,
// machine-readable code. The error envelope matches the platform convention: // machine-readable code. The error envelope matches the platform convention:
// //
@@ -77,10 +131,19 @@ type apiError struct {
status int status int
code string code string
msg string msg string
// wait, when positive, is sent as Retry-After (whole seconds, rounded up).
wait time.Duration
} }
func (e *apiError) Error() string { return e.msg } func (e *apiError) Error() string { return e.msg }
// retryAfter returns a copy of e that tells the client when to retry.
func (e *apiError) retryAfter(d time.Duration) *apiError {
c := *e
c.wait = d
return &c
}
// newError builds an apiError with a formatted message. // newError builds an apiError with a formatted message.
func newError(status int, code, format string, a ...any) *apiError { func newError(status int, code, format string, a ...any) *apiError {
return &apiError{status: status, code: code, msg: fmt.Sprintf(format, a...)} return &apiError{status: status, code: code, msg: fmt.Sprintf(format, a...)}
@@ -89,8 +152,13 @@ func newError(status int, code, format string, a ...any) *apiError {
// Common errors reused across handlers. // Common errors reused across handlers.
var ( var (
errUnauthorized = newError(http.StatusUnauthorized, "unauthorized", "authentication required") errUnauthorized = newError(http.StatusUnauthorized, "unauthorized", "authentication required")
errForbidden = newError(http.StatusForbidden, "forbidden", "not permitted") // errAuthUnavailable answers when the session store itself is unreachable
errBadRequest = newError(http.StatusBadRequest, "bad_request", "invalid request") // (Postgres down): an outage is not a credential verdict, so the caller gets
// 503 "retry" instead of a 401 that reads as "log in again".
errAuthUnavailable = newError(http.StatusServiceUnavailable, "auth_unavailable",
"authentication is temporarily unavailable; retry shortly")
errForbidden = newError(http.StatusForbidden, "forbidden", "not permitted")
errBadRequest = newError(http.StatusBadRequest, "bad_request", "invalid request")
) )
// writeJSON writes v as an indented JSON body with the given status. // writeJSON writes v as an indented JSON body with the given status.
@@ -118,6 +186,9 @@ func writeError(w http.ResponseWriter, r *http.Request, err error) {
r.Method, r.URL.Path, requestIDFromContext(r.Context()), err) r.Method, r.URL.Path, requestIDFromContext(r.Context()), err)
ae = newError(http.StatusInternalServerError, "internal", "internal error") ae = newError(http.StatusInternalServerError, "internal", "internal error")
} }
if ae.wait > 0 {
w.Header().Set("Retry-After", strconv.FormatInt(int64((ae.wait+time.Second-1)/time.Second), 10))
}
body := map[string]any{ body := map[string]any{
"error": map[string]string{ "error": map[string]string{
"code": ae.code, "code": ae.code,
+5 -5
View File
@@ -159,7 +159,7 @@ func (a *API) handleAccessWhitelist(w http.ResponseWriter, r *http.Request) {
if !ok { if !ok {
return return
} }
a.audit(r, principalFromContext(r.Context()).Email, "access.whitelist."+body.Action, name) a.audit(r, "access.whitelist."+body.Action, name)
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"name": name, "action": body.Action, "player": body.Player, "output": out, "name": name, "action": body.Action, "player": body.Player, "output": out,
}) })
@@ -249,7 +249,7 @@ func (a *API) handleAccessBan(w http.ResponseWriter, r *http.Request) {
if !ok { if !ok {
return return
} }
a.audit(r, principalFromContext(r.Context()).Email, "access.ban."+body.Action, name) a.audit(r, "access.ban."+body.Action, name)
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"name": name, "action": body.Action, "player": body.Player, "output": out, "name": name, "action": body.Action, "player": body.Player, "output": out,
}) })
@@ -282,7 +282,7 @@ func (a *API) handleAccessKick(w http.ResponseWriter, r *http.Request) {
if !ok { if !ok {
return return
} }
a.audit(r, principalFromContext(r.Context()).Email, "access.kick", name) a.audit(r, "access.kick", name)
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"name": name, "player": body.Player, "output": out, "name": name, "player": body.Player, "output": out,
}) })
@@ -354,7 +354,7 @@ func (a *API) handleAccessPermission(w http.ResponseWriter, r *http.Request) {
if !ok { if !ok {
return return
} }
a.audit(r, principalFromContext(r.Context()).Email, "access.permission."+body.Action, name) a.audit(r, "access.permission."+body.Action, name)
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"name": name, "action": body.Action, "player": body.Player, "name": name, "action": body.Action, "player": body.Player,
"node": body.Node, "output": out, "node": body.Node, "output": out,
@@ -397,7 +397,7 @@ func (a *API) handleAccessGroup(w http.ResponseWriter, r *http.Request) {
if !ok { if !ok {
return return
} }
a.audit(r, principalFromContext(r.Context()).Email, "access.group."+body.Action, name) a.audit(r, "access.group."+body.Action, name)
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"name": name, "action": body.Action, "player": body.Player, "group": body.Group, "output": out, "name": name, "action": body.Action, "player": body.Player, "group": body.Group, "output": out,
}) })
+1 -1
View File
@@ -229,7 +229,7 @@ func (a *API) handleLinkVerify(w http.ResponseWriter, r *http.Request) {
writeError(w, r, err) writeError(w, r, err)
return return
} }
a.audit(r, p.Email, "account.link", "") a.audit(r, "account.link", "")
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"linked": true, "mc_uuid": mcUUID, "auth_source": authSource, "linked": true, "mc_uuid": mcUUID, "auth_source": authSource,
}) })
+29 -12
View File
@@ -114,11 +114,10 @@ func (a *API) handleMigrateStart(w http.ResponseWriter, r *http.Request) {
return return
} }
// Internal-face event: attribute to the in-game initiator, Source 'internal'. // Internal-face event: attribute to the in-game initiator, Source 'internal'.
_ = a.Repo.Audit(r.Context(), AuditEntry{ a.auditEntry(r, AuditEntry{
Actor: "mc:" + mcUUID, Actor: "mc:" + mcUUID,
Source: "internal", Source: "internal",
Action: "account.migrate.start", Action: "account.migrate.start",
RequestID: requestIDFromContext(r.Context()),
}) })
writeJSON(w, http.StatusCreated, map[string]any{"started": true, "state": "initiated"}) writeJSON(w, http.StatusCreated, map[string]any{"started": true, "state": "initiated"})
} }
@@ -206,6 +205,13 @@ func (a *API) handleMigrateConfirmOTPStart(w http.ResponseWriter, r *http.Reques
} }
// Per-recipient cooldown, namespaced apart from the other OTP doors so they never // Per-recipient cooldown, namespaced apart from the other OTP doors so they never
// perturb each other's throttle. // perturb each other's throttle.
if until, err := a.Repo.OTPLockedUntil(r.Context(), p.UserID, otpPurposeMigrate, a.now()); err != nil {
writeError(w, r, err)
return
} else if !until.IsZero() {
writeOTPAccountLocked(w, r, until, a.now())
return
}
emailKey := "migrate:confirm:" + strings.ToLower(p.Email) emailKey := "migrate:confirm:" + strings.ToLower(p.Email)
lim := a.otpLimiter() lim := a.otpLimiter()
emailAt, ok := lim.reserve(emailKey, otpResendCooldown) emailAt, ok := lim.reserve(emailKey, otpResendCooldown)
@@ -240,7 +246,7 @@ func (a *API) handleMigrateConfirmOTPStart(w http.ResponseWriter, r *http.Reques
return return
} }
committed = true committed = true
a.audit(r, auditActor(p), "account.migrate.confirm_otp_sent", "") a.audit(r, "account.migrate.confirm_otp_sent", "")
writeJSON(w, http.StatusAccepted, map[string]any{"sent": true, "expires_at": expiresAt.UTC()}) writeJSON(w, http.StatusAccepted, map[string]any{"sent": true, "expires_at": expiresAt.UTC()})
} }
@@ -267,7 +273,16 @@ func (a *API) handleMigrateConfirmOTPVerify(w http.ResponseWriter, r *http.Reque
if _, ok := a.requireInitiatedMigration(w, r, p.UserID); !ok { if _, ok := a.requireInitiatedMigration(w, r, p.UserID); !ok {
return return
} }
switch err := a.Repo.ConsumeLoginEmailOTP(r.Context(), p.UserID, otpPurposeMigrate, otpCodeHash(code), a.now()); { var lock *OTPAccountLockedError
err := a.Repo.ConsumeLoginEmailOTP(r.Context(), p.UserID, otpPurposeMigrate, otpCodeHash(code), a.now())
if isOTPRefusal(err) {
a.authFailure(r, "migrate_confirm", otpFailureReason(err), nil)
}
switch {
case errors.As(err, &lock):
a.noteOTPLock(r, err, p.UserID, otpPurposeMigrate)
writeOTPAccountLocked(w, r, lock.Until, a.now())
return
case errors.Is(err, ErrOTPLocked): case errors.Is(err, ErrOTPLocked):
writeError(w, r, newError(http.StatusTooManyRequests, "otp_locked", writeError(w, r, newError(http.StatusTooManyRequests, "otp_locked",
"too many incorrect attempts; request a new code")) "too many incorrect attempts; request a new code"))
@@ -288,7 +303,7 @@ func (a *API) handleMigrateConfirmOTPVerify(w http.ResponseWriter, r *http.Reque
writeError(w, r, err) writeError(w, r, err)
return return
} }
a.audit(r, auditActor(p), "account.migrate.confirmed", "") a.audit(r, "account.migrate.confirmed", "")
writeJSON(w, http.StatusOK, map[string]any{"confirmed": true}) writeJSON(w, http.StatusOK, map[string]any{"confirmed": true})
} }
@@ -375,6 +390,7 @@ func (a *API) handleMigrateConfirmPasskeyFinish(w http.ResponseWriter, r *http.R
sessionData, err := a.Repo.ConsumePasskeyChallengeByUser(r.Context(), p.UserID, passkeyPurposeMigrate, a.now()) sessionData, err := a.Repo.ConsumePasskeyChallengeByUser(r.Context(), p.UserID, passkeyPurposeMigrate, a.now())
if err != nil { if err != nil {
if errors.Is(err, ErrPasskeyChallengeInvalid) { if errors.Is(err, ErrPasskeyChallengeInvalid) {
a.authFailure(r, "migrate_passkey", "challenge_invalid", nil)
writeError(w, r, newError(http.StatusBadRequest, "passkey_login_invalid", writeError(w, r, newError(http.StatusBadRequest, "passkey_login_invalid",
"passkey confirmation could not be completed; begin again")) "passkey confirmation could not be completed; begin again"))
return return
@@ -389,6 +405,7 @@ func (a *API) handleMigrateConfirmPasskeyFinish(w http.ResponseWriter, r *http.R
} }
va, err := a.Passkey.FinishLogin(migratePasskeyUser(p, creds), sessionData, bytes.NewReader(req.Assertion)) va, err := a.Passkey.FinishLogin(migratePasskeyUser(p, creds), sessionData, bytes.NewReader(req.Assertion))
if err != nil { if err != nil {
a.authFailure(r, "migrate_passkey", "bad_assertion", nil)
writeError(w, r, newError(http.StatusBadRequest, "passkey_login_invalid", writeError(w, r, newError(http.StatusBadRequest, "passkey_login_invalid",
"passkey confirmation could not be completed; begin again")) "passkey confirmation could not be completed; begin again"))
return return
@@ -400,7 +417,7 @@ func (a *API) handleMigrateConfirmPasskeyFinish(w http.ResponseWriter, r *http.R
// next login. // next login.
if err := a.applyAssertionCounter(r.Context(), va); err != nil { if err := a.applyAssertionCounter(r.Context(), va); err != nil {
if errors.Is(err, errPasskeyClonedAuthenticator) { if errors.Is(err, errPasskeyClonedAuthenticator) {
a.audit(r, auditActor(p), "auth.passkey_clone_rejected", va.CredentialID) a.passkeyCloneRejected(r, "migrate_passkey", nil, va.CredentialID)
writeError(w, r, newError(http.StatusBadRequest, "passkey_login_invalid", writeError(w, r, newError(http.StatusBadRequest, "passkey_login_invalid",
"passkey confirmation could not be completed; begin again")) "passkey confirmation could not be completed; begin again"))
return return
@@ -417,7 +434,7 @@ func (a *API) handleMigrateConfirmPasskeyFinish(w http.ResponseWriter, r *http.R
writeError(w, r, err) writeError(w, r, err)
return return
} }
a.audit(r, auditActor(p), "account.migrate.confirmed", "") a.audit(r, "account.migrate.confirmed", "")
writeJSON(w, http.StatusOK, map[string]any{"confirmed": true}) writeJSON(w, http.StatusOK, map[string]any{"confirmed": true})
} }
@@ -494,7 +511,7 @@ func (a *API) handleMigrateIssueCode(w http.ResponseWriter, r *http.Request) {
writeError(w, r, err) writeError(w, r, err)
return return
} }
a.audit(r, auditActor(p), "account.migrate.code_issued", targetID) a.audit(r, "account.migrate.code_issued", targetID)
writeJSON(w, http.StatusCreated, map[string]any{"code": code, "expires_at": expiresAt.UTC()}) writeJSON(w, http.StatusCreated, map[string]any{"code": code, "expires_at": expiresAt.UTC()})
} }
@@ -532,7 +549,7 @@ func (a *API) handleMigrateRedeem(w http.ResponseWriter, r *http.Request) {
writeError(w, r, err) writeError(w, r, err)
return return
} }
a.audit(r, auditActor(p), "account.migrate.redeemed", sourceUserID) a.audit(r, "account.migrate.redeemed", sourceUserID)
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"migrated": true, "migrated": true,
"servers_moved": len(moved), "servers_moved": len(moved),
+54
View File
@@ -1,7 +1,9 @@
package api package api
import ( import (
"context"
"encoding/json" "encoding/json"
"errors"
"net/http" "net/http"
"net/http/httptest" "net/http/httptest"
"strings" "strings"
@@ -252,6 +254,58 @@ func TestLinkVerifyIdempotent(t *testing.T) {
} }
} }
// A link whose account was SOFT-DELETED is unclaimed: a fresh in-game code lets a
// live account take it over (the migrated-source path — retire keeps the link but
// kills the account), while a merely disabled holder keeps its identity so the
// lockout cannot be re-linked away, and neither dead link has in-game standing
// (UserByMCUUID reads it exactly like an unlinked UUID). Audit #33's in-game half.
func TestLinkVerifyTakesOverDeletedLinkOnly(t *testing.T) {
ctx := context.Background()
const mcGone = "55555555-5555-5555-5555-555555555555"
const mcLocked = "66666666-6666-6666-6666-666666666666"
user := &Principal{UserID: "u-take", Email: "[email protected]", Role: "user"}
repo := newFakeRepo()
repo.seedUser(UserView{ID: "u-gone", Username: "gone", Email: "[email protected]", Role: "user"})
repo.seedUser(UserView{ID: "u-locked", Username: "locked", Email: "[email protected]", Role: "user"})
if err := repo.DeleteUser(ctx, "u-gone", "test"); err != nil {
t.Fatalf("DeleteUser: %v", err)
}
repo.links[mcGone] = "u-gone" // a retired source's link outlives the account
if _, err := repo.UserByMCUUID(ctx, mcGone); !errors.Is(err, ErrNotFound) {
t.Fatalf("UserByMCUUID(deleted link) = %v, want ErrNotFound (no in-game standing)", err)
}
repo.linkCodes["TAKEOVER"] = fakeLinkCode{mcUUID: mcGone, expiresAt: time.Unix(1_700_000_600, 0)}
api := newTestAPI(repo, newFakeCluster())
api.External = staticExternal{p: user}
if w := do(api.ExternalHandler(), "POST", "/api/v1/account/link/verify", `{"code":"TAKEOVER"}`, nil); w.Code != http.StatusOK {
t.Fatalf("takeover verify: code = %d, want 200 (%s)", w.Code, w.Body.String())
}
if repo.links[mcGone] != "u-take" {
t.Errorf("links[%s] = %q after takeover, want u-take", mcGone, repo.links[mcGone])
}
// A disabled (not deleted) holder keeps the identity: 409, code survives, link unmoved.
if err := repo.SetUserDisabled(ctx, "u-locked", true); err != nil {
t.Fatalf("disable: %v", err)
}
repo.links[mcLocked] = "u-locked"
repo.linkCodes["LOCKED12"] = fakeLinkCode{mcUUID: mcLocked, expiresAt: time.Unix(1_700_000_600, 0)}
if _, err := repo.UserByMCUUID(ctx, mcLocked); !errors.Is(err, ErrNotFound) {
t.Fatalf("UserByMCUUID(disabled link) = %v, want ErrNotFound", err)
}
if w := do(api.ExternalHandler(), "POST", "/api/v1/account/link/verify", `{"code":"LOCKED12"}`, nil); w.Code != http.StatusConflict {
t.Fatalf("takeover of a disabled holder: code = %d, want 409 (%s)", w.Code, w.Body.String())
}
if repo.links[mcLocked] != "u-locked" {
t.Errorf("disabled holder's link moved to %q", repo.links[mcLocked])
}
if _, ok := repo.linkCodes["LOCKED12"]; !ok {
t.Error("refused verify consumed the code")
}
}
// TestLinkAuthSourcePropagates proves auth_source survives the whole §10 flow: a // TestLinkAuthSourcePropagates proves auth_source survives the whole §10 flow: a
// thirdparty source captured in-game at mint reaches the durable link and the // thirdparty source captured in-game at mint reaches the durable link and the
// verify response — the value the web side can never originate itself. // verify response — the value the web side can never originate itself.
+13 -3
View File
@@ -1,6 +1,10 @@
package api package api
import "net/http" import (
"net/http"
"felis.lolicon.best/internal/metrics"
)
// Passwordless auth handlers (spec §B). Staff (Owner/Operator) authenticate via // Passwordless auth handlers (spec §B). Staff (Owner/Operator) authenticate via
// email-OTP / passkey + in-game approve on op.console; players via bind code or // email-OTP / passkey + in-game approve on op.console; players via bind code or
@@ -10,10 +14,16 @@ import "net/http"
// handleLogout revokes the presented session and clears the cookie (spec §B). It // handleLogout revokes the presented session and clears the cookie (spec §B). It
// is mounted Public and idempotent: it reads the cookie directly, so it works even // is mounted Public and idempotent: it reads the cookie directly, so it works even
// when the session has already expired and never errors on a missing one. // when the session has already expired and never errors on a missing one. Ending
// a live session is audited under its account; a dead cookie leaves no row.
func (a *API) handleLogout(w http.ResponseWriter, r *http.Request) { func (a *API) handleLogout(w http.ResponseWriter, r *http.Request) {
if c, err := r.Cookie(sessionCookieName); err == nil && c.Value != "" { if c, err := r.Cookie(sessionCookieName); err == nil && c.Value != "" {
_ = a.Repo.RevokeSession(r.Context(), hashCookie(c.Value)) hash := hashCookie(c.Value)
u, uerr := a.Repo.SessionUser(r.Context(), hash, a.now())
if err := a.Repo.RevokeSession(r.Context(), hash); err == nil && uerr == nil {
metrics.SessionsRevokedTotal.WithLabelValues("logout").Inc()
a.auditEntry(r, AuditEntry{Actor: u.Username, ActorUserID: u.ID, Action: "auth.logout"})
}
} }
clearSessionCookie(w) clearSessionCookie(w)
writeJSON(w, http.StatusOK, map[string]any{"ok": true}) writeJSON(w, http.StatusOK, map[string]any{"ok": true})
+34 -10
View File
@@ -24,12 +24,9 @@ import (
// this separator). // this separator).
// - No principal. The throttle cannot key off a user id (there is none yet); it // - No principal. The throttle cannot key off a user id (there is none yet); it
// keys off the typed recipient address, the same anti-bomb dimension the onboard // keys off the typed recipient address, the same anti-bomb dimension the onboard
// start uses. Per-source (client-IP) aggregate limiting is deliberately NOT done // start uses. Volume from one client is bounded separately by the per-address
// here: cooldownLimiter is a one-per-window primitive, so keying it on client IP // token bucket every public auth door sits behind (throttleAuthDoor), and total
// would false-positive on shared egress (CGNAT / office NAT), and behind // mail by the install-wide mail budget (ratelimit.go).
// Cloudflare RemoteAddr is the proxy anyway. The only real harm — bombing one
// mailbox — is already bounded per recipient; volumetric per-source limiting
// belongs at the edge.
// - Refuse staff. Like handleBindRedeem this public door provably never mints a // - Refuse staff. Like handleBindRedeem this public door provably never mints a
// session for an admin identity: op.console stays behind Zero Trust (and its own // session for an admin identity: op.console stays behind Zero Trust (and its own
// in-game approval gate). The refusal happens only AFTER a valid code is // in-game approval gate). The refusal happens only AFTER a valid code is
@@ -84,6 +81,13 @@ func (a *API) handleLoginEmailStart(w http.ResponseWriter, r *http.Request) {
return return
} }
// The install-wide mail budget is checked before the address is resolved,
// so while it is spent every address gets the same 429.
if err := a.checkMailBudget(); err != nil {
writeError(w, r, err)
return
}
// Atomically reserve the per-recipient cooldown BEFORE any work, so a burst of // Atomically reserve the per-recipient cooldown BEFORE any work, so a burst of
// truly concurrent starts yields exactly one winner and each admitted send is one // truly concurrent starts yields exactly one winner and each admitted send is one
// real, non-idempotent email. The key is namespaced apart from the onboard door's // real, non-idempotent email. The key is namespaced apart from the onboard door's
@@ -126,6 +130,19 @@ func (a *API) handleLoginEmailStart(w http.ResponseWriter, r *http.Request) {
return return
} }
// A locked door (wrong-code budget spent) gets the same neutral 202 and no
// mail: the owner was told by the lock notice, and a distinct answer here
// would tell a prober the address has an account.
switch until, err := a.Repo.OTPLockedUntil(r.Context(), u.ID, otpPurposeLogin, a.now()); {
case err != nil:
writeError(w, r, err)
return
case !until.IsZero():
committed = true
writeJSON(w, http.StatusAccepted, map[string]any{"sent": true, "expires_at": expiresAt.UTC()})
return
}
code, err := newEmailOTP() code, err := newEmailOTP()
if err != nil { if err != nil {
writeError(w, r, err) writeError(w, r, err)
@@ -150,7 +167,7 @@ func (a *API) handleLoginEmailStart(w http.ResponseWriter, r *http.Request) {
return return
} }
committed = true committed = true
a.audit(r, u.Username, "auth.login_email.otp_sent", "") a.auditAccount(r, u, "auth.login_email.otp_sent", "")
writeJSON(w, http.StatusAccepted, map[string]any{"sent": true, "expires_at": expiresAt.UTC()}) writeJSON(w, http.StatusAccepted, map[string]any{"sent": true, "expires_at": expiresAt.UTC()})
} }
@@ -202,6 +219,7 @@ func (a *API) handleLoginEmailVerify(w http.ResponseWriter, r *http.Request) {
// Uniform with a wrong code: a caller probing whether an address has an account // Uniform with a wrong code: a caller probing whether an address has an account
// gets the same invalid_code either way. (The /auth/options oracle is the // gets the same invalid_code either way. (The /auth/options oracle is the
// sanctioned place to learn existence; this door does not double as one.) // sanctioned place to learn existence; this door does not double as one.)
a.authFailure(r, "login_email", "no_account", nil)
writeError(w, r, newError(http.StatusBadRequest, "invalid_code", "email code is invalid or expired")) writeError(w, r, newError(http.StatusBadRequest, "invalid_code", "email code is invalid or expired"))
return return
case err != nil: case err != nil:
@@ -220,7 +238,10 @@ func (a *API) handleLoginEmailVerify(w http.ResponseWriter, r *http.Request) {
// verified; touching the row here would let a stale OTP-snapshot address overwrite // verified; touching the row here would let a stale OTP-snapshot address overwrite
// the live one and could 500 a correct code on a spurious collision. // the live one and could 500 a correct code on a spurious collision.
switch err := a.Repo.ConsumeLoginEmailOTP(r.Context(), u.ID, otpPurposeLogin, otpCodeHash(code), a.now()); { switch err := a.Repo.ConsumeLoginEmailOTP(r.Context(), u.ID, otpPurposeLogin, otpCodeHash(code), a.now()); {
case errors.Is(err, ErrOTPInvalid), errors.Is(err, ErrOTPLocked): case errors.Is(err, ErrOTPInvalid), errors.Is(err, ErrOTPLocked), errors.Is(err, ErrOTPAccountLocked):
// The account lock answers the same way; its owner hears about it by mail.
a.noteOTPLock(r, err, u.ID, otpPurposeLogin)
a.authFailure(r, "login_email", otpFailureReason(err), u)
// Both a wrong/expired code and an attempt-exhausted one return the SAME 400 // Both a wrong/expired code and an attempt-exhausted one return the SAME 400
// invalid_code, byte-identical to the unknown-account branch above. Surfacing // invalid_code, byte-identical to the unknown-account branch above. Surfacing
// otp_locked as a distinct 429 (as the authenticated onboarding door does) would // otp_locked as a distinct 429 (as the authenticated onboarding door does) would
@@ -241,7 +262,10 @@ func (a *API) handleLoginEmailVerify(w http.ResponseWriter, r *http.Request) {
// Code redeemed. Refuse staff here — never before the verify — so op.console keeps // Code redeemed. Refuse staff here — never before the verify — so op.console keeps
// its Zero-Trust + in-game-approval gates and this public door provably yields only // its Zero-Trust + in-game-approval gates and this public door provably yields only
// a role=user player session (mirrors handleBindRedeem's refuse-staff contract). // a role=user player session (mirrors handleBindRedeem's refuse-staff contract).
if u.Role == "admin" { // Staff means anything above role=user: an admin OR the role=owner identity. The
// player door must yield only player sessions.
if u.Role != "user" {
a.authFailure(r, "login_email", "staff_account", u)
writeError(w, r, newError(http.StatusForbidden, "staff_account", writeError(w, r, newError(http.StatusForbidden, "staff_account",
"that account is staff; sign in at the operator console")) "that account is staff; sign in at the operator console"))
return return
@@ -258,7 +282,7 @@ func (a *API) handleLoginEmailVerify(w http.ResponseWriter, r *http.Request) {
return return
} }
setSessionCookie(w, token, expires) setSessionCookie(w, token, expires)
a.audit(r, u.Username, "auth.login_email", "") a.auditAccount(r, u, "auth.login_email", "")
writeJSON(w, http.StatusOK, map[string]any{ writeJSON(w, http.StatusOK, map[string]any{
"user_id": u.ID, "user_id": u.ID,
"role": u.Role, "role": u.Role,
Loaded 100 of 347 files, more files were not shown because too many files have changed in this diff. Show more