Commit Graph
323 Commits
Author SHA1 Message Date
Lemon-miaow e75448a118 feat(platform): registry/uploads/backup PVC 容量可配置 2026-09-24 23:10:40 +08:00
Lemon-miaow 720ab81bf8 feat(submit): 上传存储全局上限、磁盘余量检查,被拒上下文 7 天后回收 2026-09-24 23:10:40 +08:00
Lemon-miaow 9baa0a10af fix(registrygate): 启动后先等写入静默再开放 GC 窗口 2026-09-24 23:10:28 +08:00
Lemon-miaow 151c9d2e30 feat(registry): api 定期删除无引用 manifest,gate 提供 manifest 索引 2026-09-24 22:58:07 +08:00
Lemon-miaow 05e8c64e47 fix(imagepin): 带 digest 的镜像引用也向 registry 确认仍存在 2026-09-24 22:58:07 +08:00
Lemon-miaow 7f772bccbb feat(registry): 开启 manifest 删除,GC sidecar 在 gate 只读窗口内回收 blob 2026-09-24 22:46:37 +08:00
Lemon-miaow 5cadbd40a9 fix(submit): 待审上限在事务内加咨询锁原子检查,跨副本不超额 2026-09-24 22:26:30 +08:00
Lemon-miaow a27d76ae1d docs(build): 修正上下文存储与构建镜像拉取的过时描述 2026-09-24 22:26:29 +08:00
Lemon-miaow e836a73a8a fix(build): 构建并发上限与排队,构建命名空间加 ResourceQuota,SyncAll 逐个容错 2026-09-24 22:26:21 +08:00
Lemon-miaow e521cf5976 fix(submit): 审阅绑定上下文 sha256,批准须带摘要,构建只从内部 API 取上下文并校验字节 2026-09-24 22:24:50 +08:00
Lemon-miaow 13b65e19ec fix(build): 构建 pod 等出口策略生效再运行,加 seccomp、可选 user namespace 与磁盘上限,上下文解包限总字节与条目数 2026-09-24 20:42:14 +08:00
Lemon-miaow 9bda3a52fa fix(backup): 手动备份按服冷却、每服保留上限与独立保留期,容量驱逐不再删除回收世界的唯一副本 2026-09-24 19:58:59 +08:00
Lemon-miaow f69cdec9d4 fix(reaper): 有服务器处理失败或过期备份删不掉时以非零码退出,让 Job 失败告警触达运维 2026-09-24 19:34:09 +08:00
Lemon-miaow 22becd6859 fix(reaper): 认领时重置活跃时钟与预警,回收只复用本次认领后的 reaper 归档 2026-09-24 19:31:23 +08:00
Lemon-miaow 47890ca913 feat(offsite): 世界归档与数据库备份加密同步到异地 S3,reaper 确认异地副本后才删除世界 2026-09-24 19:25:16 +08:00
Lemon-miaow d17524cd67 feat(watchdog): 主机侧巡检定时器按异常邮件通知平台所有者,operator 增加 phase 与 build_info 指标、卡死存活探针与告警规则 2026-09-24 18:06:19 +08:00
Lemon-miaow 0c29ab3a96 fix(api): 未改动任何字段的 PATCH 返回空 patched 数组而非 null 2026-09-24 17:04:02 +08:00
Lemon-miaow 215bfd78d7 fix(images): 服务器镜像在创建时固定到仓库 digest,更换镜像需确认备份,安装器重建前先固定旧服并推送不可变版本标签 2026-09-24 16:57:38 +08:00
Lemon-miaow 857b6da886 fix(operator): 删除依赖 rcon-cli 的 preStop 死代码,缩容前由 operator 经 RCON 执行 save-all flush 2026-09-24 16:39:22 +08:00
Lemon-miaow 5521e498a9 fix(platform): minecraft 命名空间强制 PodSecurity baseline,reaper 世界根目录改走静态 hostPath PV 2026-09-24 16:28:12 +08:00
Lemon-miaow 346a93921e fix(operator): 游戏 Pod 改以 UID 1000 运行并丢弃全部能力,prepare-data 初始化容器修正旧存档属主 2026-09-24 16:23:49 +08:00
Lemon-miaow c1796bea17 feat(servers): 新建服务器默认空闲 10 分钟自动停服,面板可调,converge 回填旧服,系统服不休眠 2026-09-24 16:12:53 +08:00
Lemon-miaow bd909d5858 fix(operator): 读不出玩家数时暂停空闲自动停服,支持 1.12/Bukkit/EssentialsX 与颜色码 2026-09-24 16:06:55 +08:00
Lemon-miaow 857c73a66a fix(audit): 审计按账号 id 归属并记录来源 IP/UA,登录失败与限速入审计和指标,写入失败计数告警 2026-09-24 16:02:44 +08:00
Lemon-miaow c4e4953f3d fix(auth): 公开登录门按来源限速并设全站发信上限,冷却表定期清理 2026-09-24 15:51:42 +08:00
Lemon-miaow 15f729ffea fix(auth): 邮箱验证码按账号累计错误次数封顶并通知 2026-09-24 15:37:01 +08:00
Lemon-miaow c7db7d4126 feat(db): 控制面 PG 定时备份、迁移前快照与原子恢复 2026-09-24 15:19:42 +08:00
Lemon-miaow abfe60d62d fix(api): wake 与回档/备份/改文件按服务器互斥 2026-09-24 14:39:02 +08:00
Lemon-miaow 7819e5de50 feat(netpol): 锁定游戏服出站并为 registry 加入站围栏 2026-09-24 14:25:17 +08:00
Lemon-miaow 3424852a39 feat(registry): 写入改走鉴权网关,构建先扫描再推送 2026-09-24 14:25:17 +08:00
Lemon-miaow 854320ac3f feat(submit): give the upload lane a lifecycle — withdraw + admin delete (#76)
Nothing ever removed a submission: users could not retract a pending row, no
route deleted blobs or rows, and the reaper never touches uploads — so every
upload accumulated on the 5 GiB PVC forever and the only cleanup was SQL or
kubectl against the store.

- Blobs.Delete on both transports (local: RemoveAll of the id-namespaced dir,
  id re-validated at the boundary; S3: idempotent object DELETE).
- Store: DeleteSubmission (admin, any status) and DeletePendingSubmission
  (owner+pending CAS — a reviewed row can never be withdrawn out from under
  its build).
- Manager.Delete / Manager.Withdraw delete the ROW first (under the CAS for
  withdraw) and the blob after, so a live row can never point at a reaped
  blob; a cleanup failure names the orphan explicitly instead of failing mute.
- API: DELETE /me/submissions/{id} (withdraw, app tier) and
  DELETE /api/v1/submissions/{id} (admin) both return the row as it was;
  audit events submission.withdraw / submission.delete; openapi documents both
  paths; admin route pinned in the admin-only table.
- Panel: two-step withdraw on a pending row (frees the pending slot and the
  storage budget); two-step delete on every admin row; zh/en copy; wire tests.

Unit: submit (withdraw happy path / wrong owner / reviewed row / no transport /
blob-cleanup failure), local+S3 delete idempotence, api handlers (200/404/409/
503 + route tier); pgint: withdraw CAS + admin delete exactly-once.
go vet/go test/gofmt clean; panel vitest 120 + typecheck green.
2026-09-24 10:24:23 +08:00
Lemon-miaow ad4d256d8f fix(submit): bound the untrusted upload lane — per-user caps + throttles (#75)
A logged-in user could file submissions without bound and stream a 1 GiB
context per submission. The only limits were the single-blob size cap and the
5 GiB uploads PVC (platform/workloads.go); nothing counted a user's rows or
bytes, so one account could fill the volume and every other user's upload
would start failing.

- Create: per-user pending_review cap (default 5) — the review queue cannot
  be parked full of one account's rows. Check-then-insert, documented soft.
- UploadContext: per-user stored-context budget (default 2 GiB) charged
  against the blob store's REAL sizes (new Blobs.Size on local/S3 stores), so
  the sum cannot drift from the volume; the write is capped at the remaining
  budget, so the excess is refused before it is persisted, and a re-upload is
  charged only for its new bytes.
- API: per-user create/upload throttles (30s/15s, cmd/felis-wired) on a
  dedicated cooldown keyspace, reserve→release so a failed attempt never
  burns the window and a burst collapses to one winner; ErrQuotaExceeded →
  403 submission_quota_exceeded (distinct from the 400 an oversize blob
  gets), 429 submission_cooldown for the throttles.
- Panel: zh/en copy for both codes; openapi documents 403/429 on the two
  user routes; pgint covers the pending-queue count.

Unit tests: submit package (cap, budget boundary/exact-fit/replacement,
oversize-vs-quota split) and api handlers (quota 403 both paths, throttle
429 + recovery + failure-release). go vet/go test/gofmt clean; panel
vitest 118 + typecheck green.
2026-09-24 10:14:42 +08:00
Lemon-miaow 96aa8176cd test(api): make the console-disconnect teardown case wait for the write, not just the read (#74) 2026-09-24 03:16:40 +08:00
Lemon-miaow b14bacbfc6 fix(build): mirror the Trivy Java DB — jar-bearing builds failed closed at the scan gate (#72) 2026-09-24 03:11:01 +08:00
Lemon-miaow 6e47730501 fix(build): allow kaniko to unpack base-image layers — drop-ALL removed the caps the tar apply needs (#71) 2026-09-24 03:05:07 +08:00
Lemon-miaow ac403b9cd3 fix(build): let kaniko pull base images from the plain-HTTP registry — push-only insecure flags broke every FROM registry.felis.svc build (#70) 2026-09-24 02:56:56 +08:00
Lemon-miaow 2755e41ff3 fix(build): reap finished build Jobs — they accumulated forever
Every other Job family carries a TTLSecondsAfterFinished (fileedit 2m,
backup/restore 10m) but the build lane never set one: each build left a
completed Job + Pod in felis-build indefinitely (5 already on the drill
cluster, oldest 26h), growing etcd and — since completed pods count
against the node's pod budget (110 on stock k3s) — eventually blocking
new builds.  The code even anticipated GC it never got ("JobUnknown
means the Job was not found (e.g. GC'd)").

Set a deliberately long TTL (7 days): the kaniko log is the admin
failure-triage surface (GET /images/build/{id}/logs), so the window
keeps a week of logs while bounding steady-state pods; terminal builds
are idempotent under Sync, so a late log-404 is the only cost.
2026-09-24 02:03:20 +08:00
Lemon-miaow 4d4cdd6ea7 fix(api,panel): refuse file operations on a server without a world volume (#58)
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
2026-09-23 22:17:08 +08:00
Lemon-miaow 9dad61f508 fix(nano): screen unusable profiles per source instead of stopping the ladder (#55)
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.

Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.

Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
2026-09-23 21:21:21 +08:00
Lemon-miaow 2c6739ad76 fix(updater,install,docs): follow the move to FelisMC/Felis (#53)
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.

Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
2026-09-23 20:58:34 +08:00
Lemon-miaow 13d64e0000 feat(bootstrap): host every built image in the internal registry — GC-durable pulls
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:

- Every image the installer builds is now a registry ref
  (registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
  containerd under that exact name (first boot needs no registry round-trip)
  and mirrored into the registry after deploy_bundle (push_image_to_registry:
  push endpoint 127.0.0.1:5000, and only the path after the host matters to the
  registry — a push there lands where kubelet's mirrored pull looks). A ref
  outside the registry is warned about, not silently unmirrored.

- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
  registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
  registry Deployment binds (node containerd cannot dial the Service VIP — live
  drill: "Empty reply"). k3s regenerates containerd config only at agent start,
  so a CONTENT change restarts k3s and an identical file (every re-run)
  restarts nothing.

- import_registry_image caches registry:2 into containerd so the registry
  Deployment can start on a box that cannot reach Docker Hub.

- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
  'felis-paper:demo') at the registry refs — a user server created from those
  rows must not strand when GC collects the bare tag. Only recommended rows
  still holding the old seed are touched; enabled is preserved; a pre-existing
  target row wins over a duplicate.

bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
2026-09-23 19:02:58 +08:00
Lemon-miaow a9b275abbb fix(platform): registry OOM (audit #46) + loopback hostPort — the node-side pull path
Two changes to the registry Deployment, both prerequisite to GC-durable images:

- Dedicated resource template: the control plane's 256Mi memory limit was a
  live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
  mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
  same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.

- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
  reach the Service VIP (live stack: "Empty reply"), so the node-side pull
  path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
  http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
  the plain-HTTP registry off every other interface.

Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
2026-09-23 18:55:23 +08:00
Lemon-miaow 168a37542b feat(api,panel): reviewer context download for submissions; dockerfile field documented as audit-only (audit #45) 2026-09-23 16:41:55 +08:00
Lemon-miaow 94f71eea19 feat(api): serve felis_* metrics on the internal face (build-failure counter's only scrape path) 2026-09-23 16:17:50 +08:00
Lemon-miaow 508a1c02da fix(api): refuse backup/restore before a missing world volume
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
2026-09-23 07:49:00 +08:00
Lemon-miaow 55d515d41f fix(provisioning): keep the Owner seat single; clash on the operator name stays retryable
Two defects live-drilled in the break-glass staff provisioning:

- An Owner reset that typed any username other than the occupied seat took
  UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
  existing seat — possibly the compromised account the reset was meant to
  replace — live; every owner row is undeletable through the panel, so the tier
  could never converge back to one. provisionOwner now refuses with
  ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
  form); bootstrap still mints, and the seat's own username still resets in
  place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
  console keys its rename prompt off api.ErrConflict — the "choose another
  name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
  the contract; PGRepo had drifted). Map the unique violation to ErrConflict
  and pin it in pgint.

Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
2026-09-23 07:48:54 +08:00
Lemon-miaow daf760220b fix(cli): pin the reaper to its storage node; drop the stale uid-1000 note
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
2026-09-23 06:58:29 +08:00
Lemon-miaow 2f90851c03 fix(panel): mark platform system services read-only in the fleet table
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
2026-09-23 06:51:45 +08:00
Lemon-miaow 72c4aa3895 fix(submissions): surface each linked build's outcome to the submitter
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
2026-09-23 06:31:13 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00