Commit Graph
144 Commits
Author SHA1 Message Date
Lemon-miaow e836a73a8a fix(build): 构建并发上限与排队,构建命名空间加 ResourceQuota,SyncAll 逐个容错 2026-09-24 22:26:21 +08:00
Lemon-miaow e521cf5976 fix(submit): 审阅绑定上下文 sha256,批准须带摘要,构建只从内部 API 取上下文并校验字节 2026-09-24 22:24:50 +08:00
Lemon-miaow 13b65e19ec fix(build): 构建 pod 等出口策略生效再运行,加 seccomp、可选 user namespace 与磁盘上限,上下文解包限总字节与条目数 2026-09-24 20:42:14 +08:00
Lemon-miaow 9bda3a52fa fix(backup): 手动备份按服冷却、每服保留上限与独立保留期,容量驱逐不再删除回收世界的唯一副本 2026-09-24 19:58:59 +08:00
Lemon-miaow f69cdec9d4 fix(reaper): 有服务器处理失败或过期备份删不掉时以非零码退出,让 Job 失败告警触达运维 2026-09-24 19:34:09 +08:00
Lemon-miaow 22becd6859 fix(reaper): 认领时重置活跃时钟与预警,回收只复用本次认领后的 reaper 归档 2026-09-24 19:31:23 +08:00
Lemon-miaow 47890ca913 feat(offsite): 世界归档与数据库备份加密同步到异地 S3,reaper 确认异地副本后才删除世界 2026-09-24 19:25:16 +08:00
Lemon-miaow d17524cd67 feat(watchdog): 主机侧巡检定时器按异常邮件通知平台所有者,operator 增加 phase 与 build_info 指标、卡死存活探针与告警规则 2026-09-24 18:06:19 +08:00
Lemon-miaow 215bfd78d7 fix(images): 服务器镜像在创建时固定到仓库 digest,更换镜像需确认备份,安装器重建前先固定旧服并推送不可变版本标签 2026-09-24 16:57:38 +08:00
Lemon-miaow 857b6da886 fix(operator): 删除依赖 rcon-cli 的 preStop 死代码,缩容前由 operator 经 RCON 执行 save-all flush 2026-09-24 16:39:22 +08:00
Lemon-miaow 5521e498a9 fix(platform): minecraft 命名空间强制 PodSecurity baseline,reaper 世界根目录改走静态 hostPath PV 2026-09-24 16:28:12 +08:00
Lemon-miaow c1796bea17 feat(servers): 新建服务器默认空闲 10 分钟自动停服,面板可调,converge 回填旧服,系统服不休眠 2026-09-24 16:12:53 +08:00
Lemon-miaow bd909d5858 fix(operator): 读不出玩家数时暂停空闲自动停服,支持 1.12/Bukkit/EssentialsX 与颜色码 2026-09-24 16:06:55 +08:00
Lemon-miaow 857c73a66a fix(audit): 审计按账号 id 归属并记录来源 IP/UA,登录失败与限速入审计和指标,写入失败计数告警 2026-09-24 16:02:44 +08:00
Lemon-miaow c4e4953f3d fix(auth): 公开登录门按来源限速并设全站发信上限,冷却表定期清理 2026-09-24 15:51:42 +08:00
Lemon-miaow 15f729ffea fix(auth): 邮箱验证码按账号累计错误次数封顶并通知 2026-09-24 15:37:01 +08:00
Lemon-miaow c7db7d4126 feat(db): 控制面 PG 定时备份、迁移前快照与原子恢复 2026-09-24 15:19:42 +08:00
Lemon-miaow abfe60d62d fix(api): wake 与回档/备份/改文件按服务器互斥 2026-09-24 14:39:02 +08:00
Lemon-miaow 7819e5de50 feat(netpol): 锁定游戏服出站并为 registry 加入站围栏 2026-09-24 14:25:17 +08:00
Lemon-miaow b9ebc872ad docs(troubleshooting): the [INERT] legend points at a field that no longer exists
The legend said "§12 lists the one field this still applies to", but the last
inert field (spec.storage.retainOnDelete) was removed rather than implemented
(§13), and §12 has said "every field below is read by a controller" since.
Reworded so a reader whose change looks ignored follows the condition question
instead of hunting for a dead field.
2026-09-24 10:35:35 +08:00
Lemon-miaow c57daaf861 feat(cli): add felis converge for fields a newer desired spec never delivered (tracker #1)
Provisioning is create-if-absent, so a field the desired spec gained after an
install (spec.rcon, spec.startup.healthHTTPPort, a derived env key) never
reaches the existing login/lobby CR while every re-run of setup reports
success — the reported 'configuration updates never reach an installed
deployment' symptom. converge is the explicit pass: it fills exactly the
zero-valued whitelist fields and the derived env (including a missing key,
which refreshDerivedEnv deliberately never adds), and never overwrites a
non-zero value. The timing stays with the operator because enabling RCON or
the HTTP readiness gate on a pre-listener image would wedge that server in
Starting until it was marked Failed.

Tests: fills predated fields while operator edits survive / non-zero values
left alone / absent + foreign + unset-image guards. usage table updated so the
router-parity test passes; troubleshooting gains §12b.
2026-09-24 10:35:14 +08:00
Lemon-miaow 854320ac3f feat(submit): give the upload lane a lifecycle — withdraw + admin delete (#76)
Nothing ever removed a submission: users could not retract a pending row, no
route deleted blobs or rows, and the reaper never touches uploads — so every
upload accumulated on the 5 GiB PVC forever and the only cleanup was SQL or
kubectl against the store.

- Blobs.Delete on both transports (local: RemoveAll of the id-namespaced dir,
  id re-validated at the boundary; S3: idempotent object DELETE).
- Store: DeleteSubmission (admin, any status) and DeletePendingSubmission
  (owner+pending CAS — a reviewed row can never be withdrawn out from under
  its build).
- Manager.Delete / Manager.Withdraw delete the ROW first (under the CAS for
  withdraw) and the blob after, so a live row can never point at a reaped
  blob; a cleanup failure names the orphan explicitly instead of failing mute.
- API: DELETE /me/submissions/{id} (withdraw, app tier) and
  DELETE /api/v1/submissions/{id} (admin) both return the row as it was;
  audit events submission.withdraw / submission.delete; openapi documents both
  paths; admin route pinned in the admin-only table.
- Panel: two-step withdraw on a pending row (frees the pending slot and the
  storage budget); two-step delete on every admin row; zh/en copy; wire tests.

Unit: submit (withdraw happy path / wrong owner / reviewed row / no transport /
blob-cleanup failure), local+S3 delete idempotence, api handlers (200/404/409/
503 + route tier); pgint: withdraw CAS + admin delete exactly-once.
go vet/go test/gofmt clean; panel vitest 120 + typecheck green.
2026-09-24 10:24:23 +08:00
Lemon-miaow ad4d256d8f fix(submit): bound the untrusted upload lane — per-user caps + throttles (#75)
A logged-in user could file submissions without bound and stream a 1 GiB
context per submission. The only limits were the single-blob size cap and the
5 GiB uploads PVC (platform/workloads.go); nothing counted a user's rows or
bytes, so one account could fill the volume and every other user's upload
would start failing.

- Create: per-user pending_review cap (default 5) — the review queue cannot
  be parked full of one account's rows. Check-then-insert, documented soft.
- UploadContext: per-user stored-context budget (default 2 GiB) charged
  against the blob store's REAL sizes (new Blobs.Size on local/S3 stores), so
  the sum cannot drift from the volume; the write is capped at the remaining
  budget, so the excess is refused before it is persisted, and a re-upload is
  charged only for its new bytes.
- API: per-user create/upload throttles (30s/15s, cmd/felis-wired) on a
  dedicated cooldown keyspace, reserve→release so a failed attempt never
  burns the window and a burst collapses to one winner; ErrQuotaExceeded →
  403 submission_quota_exceeded (distinct from the 400 an oversize blob
  gets), 429 submission_cooldown for the throttles.
- Panel: zh/en copy for both codes; openapi documents 403/429 on the two
  user routes; pgint covers the pending-queue count.

Unit tests: submit package (cap, budget boundary/exact-fit/replacement,
oversize-vs-quota split) and api handlers (quota 403 both paths, throttle
429 + recovery + failure-release). go vet/go test/gofmt clean; panel
vitest 118 + typecheck green.
2026-09-24 10:14:42 +08:00
Lemon-miaow b14bacbfc6 fix(build): mirror the Trivy Java DB — jar-bearing builds failed closed at the scan gate (#72) 2026-09-24 03:11:01 +08:00
Lemon-miaow 62e8c87d0e docs(diagrams): align §28 claim/link sequences with the audited implementations
The claim-transaction note still described the pre-audit-#4 shape (SELECT
EXISTS plus a quota pre-check outside the transaction); the implemented
ClaimServer serializes on pg_advisory_xact_lock(user_id), takes the row
FOR UPDATE and re-runs the four-dimension gate inside the same
transaction — the diagram's QuotaAvailable step is only a fast path.
The link flow now selects the code FOR UPDATE, treats a same-user
re-verify as idempotent, and lets a live caller take over a retired
(soft-deleted) owner's link — the 409 is only for a different, live user.
Both re-read from internal/api/pgrepo.go and the handler mappings.
2026-09-24 01:26:32 +08:00
Lemon-miaow 397a400d57 fix(update): point the apply guidance at the installer, not felis setup (#54)
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.

Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
2026-09-23 21:01:32 +08:00
Lemon-miaow 72553cb414 docs(troubleshooting): the installer leaves docker stopped — start it before manual pushes
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
2026-09-23 19:35:13 +08:00
Lemon-miaow fa0e8d7d97 docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
2026-09-23 19:03:11 +08:00
Lemon-miaow 168a37542b feat(api,panel): reviewer context download for submissions; dockerfile field documented as audit-only (audit #45) 2026-09-23 16:41:55 +08:00
Lemon-miaow 43df08b52a feat(alerts): ship Felis alert rules with promtool unit tests; document scraping & rules (troubleshooting §14) 2026-09-23 16:17:51 +08:00
Lemon-miaow 94f71eea19 feat(api): serve felis_* metrics on the internal face (build-failure counter's only scrape path) 2026-09-23 16:17:50 +08:00
Lemon-miaow ae6e9256c6 docs(troubleshooting): 8e — apply build-image overrides through the config Secret (restart alone does not) 2026-09-23 15:56:44 +08:00
Lemon-miaow 2f90851c03 fix(panel): mark platform system services read-only in the fleet table
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
2026-09-23 06:51:45 +08:00
Lemon-miaow 72c4aa3895 fix(submissions): surface each linked build's outcome to the submitter
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
2026-09-23 06:31:13 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00
Lemon-miaow ffe5dc14a8 docs(seams): close the deferred entries that are now live-verified 2026-09-23 04:46:02 +08:00
Lemon-miaow 089d4f3a80 docs(audit): idle auto-stop postmortem — ledger #25 and self-checks in §11
Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
2026-09-23 04:14:11 +08:00
Lemon-miaow 8e7c7bbf24 fix(reaper): deliver pre-reap warnings for real — and never fake a delivery
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:

- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
  `warning suppressed — no warner wired` and does NOT stamp; a delivery error
  logs and retries on the next daily run (bounded by the warning window). The
  stamps are no longer burned by notices nobody received.

- A real channel: mail.SendNotice (the second and last message shape the mail
  package sends) plus a mailWarner that resolves the owner's VERIFIED email
  and mails the notice through the configured [smtp] relay. `felis reaper`
  wires it when [smtp] is set (same password_ref convention as felis-api) and
  prints exactly what happens when it is not.

Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.

Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
2026-09-23 03:47:19 +08:00
Lemon-miaow 1d0ec61c9d docs(audit): fifth-batch ledger — quota atomic gate closed (audit #4)
The deferred-seams entry that waited for a real-Postgres harness is struck:
ClaimServer owns the gate now, proven red-then-green by the pgint concurrency
test (two claims, one win, one 403), and the storage-cache zeroing found in the
same pass is recorded with its hermetic test. auditfix19 is live on the VM.
2026-09-23 03:39:25 +08:00
Lemon-miaow 02fd2de502 fix(build): three drill-driven fixes so the lane actually completes on a starter node
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:

- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
  Pod never scheduled on the platform's own starter node: FailedScheduling /
  Insufficient memory, Pending forever. Requests are now a small floor
  (250m / 512Mi, never above a configured cap) while the limits stay the
  safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
  the source owner; a 65532-owned context (the distroless felis image uid)
  fails that under the pod's dropped capabilities ('copying dockerfile:
  chown /kaniko/Dockerfile: operation not permitted'). The fetch container
  now extracts as root — the uid Kaniko already runs as — so the copy
  succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
  step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
  renders --db-repository, and docs/troubleshooting.md §8e now carries the
  verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
  internal registry; --insecure already covers its plain HTTP).

Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
2026-09-22 23:01:25 +08:00
Lemon-miaow f79e5ebb5e feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:

- submit: derived context refs become the internal-face URL
  /api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
  gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
  route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
  image's new fetch-context entrypoint) that streams the blob with the
  namespace-local service-token Secret — never mounted into Kaniko — and
  extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
  reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
  build namespace gets the token Secret through the existing replica mechanism
  (bootstrap.sh + felis setup); the build egress lock opens exactly the control
  namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
  escapes/symlinks/non-gzip).

Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
2026-09-22 22:45:09 +08:00
Lemon-miaow 87a9f4eb25 feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
2026-09-22 22:03:48 +08:00
Lemon-miaow 0a2d654e68 fix(platform): control plane runs system-cluster-critical, so eviction refuses it (#8)
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.

Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
2026-09-22 21:55:54 +08:00
Lemon-miaow fe310743a2 fix(platform): give the control plane a PriorityClass eviction shield (#8)
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
2026-09-22 21:08:56 +08:00
Lemon-miaow 2b87a5a13b fix(install): grant the reaper traverse on the worlds root (#6)
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.

Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
2026-09-22 21:03:17 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
flyemoji a56c326518 chore: stop tracking the docs/changes ledger
docs/changes held 26 per-feature change notes and their index,
written while each feature was built. They were working records, not
documentation: they cite internal milestone numbers and plan steps,
several describe designs that changed before they shipped (the nano
note's config schema and a proxy plugin that was never built), and
nothing in the code, the build or the other docs refers to them. New
notes stopped being added a while ago; the commit messages carry that
record now.

The directory leaves the tree in this commit. Its contents stay
reachable in history, and the files were kept outside the repository
before removal. No code, build or test changes.
2026-09-22 15:02:47 +09:00
flyemoji 9ee8c48fff docs(openapi): list every answer hasjoined gives
The hasJoined contract listed only 200 and 204 and named authlib as
the caller. The handler now answers four more ways, and a proxy
operator reading the contract could not tell a refused login from a
down source.

- 204 also covers a missing or oversized parameter (no source is
  asked), a third-party name that is not a legal Minecraft username,
  and an identity id that does not parse.
- 400 for a request that declares a body. There is no response body,
  and the connection is closed.
- 500 when the bar-list lookup fails, with the usual error body.
- 503 when no source validated and at least one failed, since that
  source's player may be the one logging in.

The three query parameters now carry the 64-byte cap. The profile name
says a third-party player holding a registered Mojang name gets it
back prefixed and cut to 16 characters. The description names
Velocity, drops the "thin login hook" that does not exist, and says
that a non-200, non-204 answer makes Velocity report the auth servers
as down.
2026-09-22 13:43:08 +09:00
flyemoji afdbfac7a8 docs: index the deferred integration seams and correct two stale markers
INTEGRATION-ONLY and KNOWN-LIMITATION are grep-able, but the grep answers the
wrong question. Thirty-four Go sites share the two markers and they carry four
different meanings: "declared, nothing implements it" reads exactly like
"implemented, only its I/O is unreachable from here", and neither reads
differently from a limitation that was accepted on purpose and is not coming
back. docs/deferred-seams.md sorts them, following the bucketed shape
internal/updater/doc.go already uses for its own package rather than starting a
second convention.

Sorting them turned up two markers that had outlived the condition they describe.

config.go called the modpack upload transport a deferred integration after both
backends had shipped -- LocalContextStore and S3ContextStore, selected in
cmd/felis by the shape of user_uploads_context, with the uploads PVC mounted and
the felis-uploads-s3 Secret rendered. What is still deferred is the far end:
Kaniko reading that context from inside the build Pod.

updater/doc.go listed the `felis update` CLI and the off-cluster Velocity jar read
under REMAINING INTEGRATION. Both exist -- cmd/felis/update.go, and
gatherer_host.go, which answers Velocity from the installed jar's manifest and
felis-api from the running binary's build stamp. The two nil seams that bullet
also names are real, but they belong to the in-cluster gatherer only, so the
bullet now says which caller has what and which is still empty.

The index also records the collision that makes a naive grep misleading:
docs/troubleshooting.md uses [INTEGRATION-ONLY] for something else, defined in its
own opening at :19 -- the symptom is produced by the kubelet, kaniko or a live
handshake, so it cannot be reproduced from the repository. Those twelve marks say
where a failure comes from, not that something is unbuilt, and are excluded.

Both code changes are comments. Every file:line the index cites was checked
against the line it points at.
2026-07-28 17:27:01 +09:00
flyemoji 23792d6251 fix(crd): remove spec.storage.retainOnDelete rather than leave it inert
The field validated, shipped in the CRD, and reached no controller. The world
PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate,
StatefulSet deletion does not cascade to template PVCs, and no finalizer exists
anywhere in the operator. So setting it true described what already happened,
and setting it false did nothing at all. False is the worse half: it reads as a
request to delete a world, and was silently ignored.

This departs from spec v4.1 §5, which asks for
"删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half
was ever built. Restoring that line means adding a finalizer whose other listed
duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs,
so the only work it would newly do is delete worlds, on a path that does not
pass the reaper's verified-backup check. The reaper is the one thing in the
system allowed to destroy a world and it earns that by proving a backup first.
A second door without that check is not an improvement.

The spec is a frozen versioned document, so it is left alone and the departure
is recorded in troubleshooting.md §13, beside the behaviour it explains. §12
loses its inert row and its opening sentence, which existed to introduce this
one field: every field in that table is now read by a controller.

Deployed installs need nothing. A CR still carrying retainOnDelete keeps
working, because a v1 CRD prunes unknown keys on the next write and the
behaviour the field claimed to control was never conditional.

go build, go vet and go test ./... pass on Linux with zero failures; the CRD
still parses and storage keeps size and storageClassName.
2026-07-28 17:27:01 +09:00