diff --git a/docs/changes/2026-06-26-foundational-subsystems.md b/docs/changes/2026-06-26-foundational-subsystems.md deleted file mode 100644 index 238483b..0000000 --- a/docs/changes/2026-06-26-foundational-subsystems.md +++ /dev/null @@ -1,41 +0,0 @@ -# Foundational subsystems: the initial Felis import (ledger backfill) - -- **Type:** feature (initial import) — retroactive ledger entry -- **Date:** 2026-06-26 -- **Area:** `apis/`, `internal/` (naming, rcon, store, config, build, backup, operator, - submit, api, platform), `cmd/felis`, `plugins/` -- **Commits:** - - `7fbebfe` feat(apis): MinecraftServer CRD types (v1alpha1) — the lifecycle source of truth (§1) - - `708cdfc` feat(core): naming, RCON, store (Postgres + embedded migrations), config, image-build libraries - - `43ab921` feat(backup): archive-based world backup/restore + the retention/idle reaper - - `78b8cf6` feat(operator): MinecraftServer controller and reconcilers - - `d39605e` feat(submit): user modpack build + admin-approval pipeline (see [modpack-submission-lane](2026-06-26-modpack-submission-lane.md)) - - `b508fcc` feat(api): dual-faced felis-api — permissions/LuckPerms, modpack lane, admin fleet read - - `47fcd90` feat(platform): node orchestration + the `cmd/felis` single-binary entrypoint - - `93f143f` feat(plugins): Velocity proxy + Fabric/Forge/NeoForge/Paper integration mods -- **Tasks:** #23 (permissions), #24 (modpack lane), #25 (fleet read) - -## What it did - -Stood up the whole backend spine in one build-order sweep: the Kubernetes CRD that is -the lifecycle source of truth, the core libraries (deterministic resource naming, the -RCON client, the Postgres store with embedded SQL migrations, config loading, container -image-build helpers), the backup/restore/reaper subsystems, the operator controller -that drives `MinecraftServer` resources, the user-modpack submit+approval pipeline, the -dual-faced (internal/external) felis-api behind a Zero-Trust guard, the platform -orchestrator that wires it all together under `cmd/felis`, and the server-side -integration plugins. - -## Why - -This is the project's first functional import — the substrate every later change edits. -It predates the change-ledger convention (established `fad48ff`, 2026-07-06), so it never -got a contemporaneous detail doc; this entry backfills one. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the -> change-ledger's detail-doc axis (§ Convention). This entry deliberately describes only -> what these eight commits **introduced** on 2026-06-26 — the named subsystems have been -> extended and reworked many times since (auth, passkey, metrics, quotas, updates), and -> that later work lives in its own dated detail docs, not here. Not independently -> re-verified for this doc; each subsystem was verified at its original commit and the -> current tree builds green at `9911b8c` (WSL oracle, go1.26.4). diff --git a/docs/changes/2026-06-26-modpack-submission-lane.md b/docs/changes/2026-06-26-modpack-submission-lane.md deleted file mode 100644 index 56fa07e..0000000 --- a/docs/changes/2026-06-26-modpack-submission-lane.md +++ /dev/null @@ -1,30 +0,0 @@ -# Modpack submission lane: build/approval pipeline + storage backends (ledger backfill) - -- **Type:** feature — retroactive ledger entry -- **Date:** 2026-06-26 – 2026-07-02 -- **Area:** `internal/submit` (build/approval pipeline, storage backends), `internal/api` (submission endpoints) -- **Commits:** - - `d39605e` feat(submit): user modpack build + approval pipeline — an uploaded modpack stays `pending_review` and is never built until an admin approves; approval is a single-winner compare-and-swap handing off to the image-build Job, keeping the mandatory vulnerability scan in front of any push - - `598f3d3` feat(submit): local + S3 backends for modpack upload contexts, installer-selectable -- **Tasks:** #24 (§8 user-submitted modpack approval lane) - -## What it did - -Built the user-directed extension over the image-build subsystem: a player uploads a -modpack context, it sits in `pending_review`, and an admin's approval is the single-winner -gate that hands off to the build Job — with the vulnerability scan always ahead of any -registry push. `598f3d3` makes the upload-context store pluggable (local filesystem or S3), -selectable at install time. - -## Why - -Untrusted user content must never build or push unreviewed, and the compare-and-swap -approval guarantees exactly one build per submission even under a double-click or retry. -The storage-backend choice lets a single-node demo use local disk while a real deployment -uses S3, without a code change. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The approval -> compare-and-swap and endpoints were unit-tested at their commits; the S3 path is -> integration-configurable. The panel-side submission/approval UI is the collaborator's -> frontend work and is tracked only by its INDEX rows. Not independently re-verified for -> this doc; current tree green at `9911b8c`. diff --git a/docs/changes/2026-06-27-cloudflare-tunnel-access-edge.md b/docs/changes/2026-06-27-cloudflare-tunnel-access-edge.md deleted file mode 100644 index 9518ec2..0000000 --- a/docs/changes/2026-06-27-cloudflare-tunnel-access-edge.md +++ /dev/null @@ -1,38 +0,0 @@ -# Cloudflare Tunnel + Access edge (cfsetup) + NodePort fencing (ledger backfill) - -- **Type:** feature + fix — retroactive ledger entry -- **Date:** 2026-06-27 – 2026-07-01 -- **Area:** `internal/cfsetup` (pure core + integration runner), `cmd/felis` (TUI edge flow), edge nftables fence -- **Commits:** - - `53a7664` feat(cfsetup): recommended Cloudflare Tunnel + Access edge (§14) — domain- and IdP-agnostic; the load-bearing `validateFailClosed` allowlist refuses any policy that could be public; fail-shut 404 catch-all; the raw game host is never proxied - - `ba13839` feat(breakglass): optional Tunnel + Access setup in the sudo TUI, an independent peer of Owner provisioning - - `a531f5e` fix(cfsetup): keep the connector install in the host apply layer only (drop the duplicate `StartConnector`) - - `2810fe8` fix(cfsetup): repoint a stale DNS record when routing a tunnel hostname - - `7d3be64` feat(cfsetup): start the tunnel connector as a setup step - - `346ec68` refactor(deploy): rework the cloudflare-edge walkthrough — restructured the edge TUI flow and added a tested `cfsetup` integration-runner path (with TUI height-measure/root tests) - - `e058a64` feat(edge): close the panel NodePort to the public after the tunnel is up — nftables at prerouting `raw` (-300), before kube-proxy's NodePort DNAT, gated on the connector actually serving; loopback accepted first so the connector origin hop is untouched -- **Tasks:** #37 (fence panel NodePort to public after tunnel) - -## What it did - -Stood up the optional one-click Zero-Trust edge: a Cloudflare Tunnel routing only the web -hostnames plus a fail-closed Access application, provisioned from the sudo TUI against the -operator's own Cloudflare account. `e058a64` then closes the Access-bypass hole where a -direct `https://:/` with the right Host header reached the origin -behind Access, by fencing the NodePort at the nftables raw hook so the packet is caught on -its original destination port — but only once the connector is confirmed serving, so -fencing never severs the only web path to a live origin. - -## Why - -Access is only a security boundary if the origin cannot be reached around it. The -fail-closed policy guard (`validateFailClosed`) and the NodePort fence are the two -load-bearing safety properties: a policy that could be public aborts the run with nothing -created, and a routable-but-unfenced NodePort would defeat the whole edge. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The policy guard, -> ingress generation, request bodies, gating, and the nftables ruleset shape / conn-count -> gate are unit-tested; the live cloudflared/Cloudflare-API and `nft` calls are -> INTEGRATION-ONLY (need a real account). KNOWN-LIMITATION: the fence targets nftables; -> firewalld-native coordination is deferred. Not independently re-verified for this doc; -> current tree green at `9911b8c`. diff --git a/docs/changes/2026-06-27-console-auth-passwordless.md b/docs/changes/2026-06-27-console-auth-passwordless.md deleted file mode 100644 index c4fa9ad..0000000 --- a/docs/changes/2026-06-27-console-auth-passwordless.md +++ /dev/null @@ -1,32 +0,0 @@ -# Console auth: local-password login → passwordless migration (ledger backfill) - -- **Type:** feature + refactor — retroactive ledger entry -- **Date:** 2026-06-27 – 2026-07-04 -- **Area:** `internal/api` (auth handlers, sessions), `internal/store` (users schema) -- **Commits:** - - `af14f02` feat(api): local-password authentication backend — login/logout/change-password on `op.console`; HttpOnly+Secure+SameSite=Lax host-only server-side sessions (SHA-256, 12h TTL); anti-enumeration uniform bcrypt; JSON-only credential writes (415 otherwise); fails closed unless `local_auth_enabled` - - `0c1cc59` feat(auth): migrate console login to passwordless - - `3b43f05` refactor(api): drop the dead login concurrency limiter and reconcile passwordless comments - - `c20b12c` refactor(api): drop the dead password-era `ResetMailer`, reconcile passkey-unbind docs -- **Tasks:** #27 (B1 thin thread), #79/#80/#81 (residue sweep + primitive adjudication) - -## What it did - -Shipped the staff local-password door (`af14f02`) as the primary web login when -Zero Trust is not in front of the API, then migrated the console to passwordless -(`0c1cc59`) once email-OTP + passkey were the intended factors. The two refactors -(`3b43f05`, `c20b12c`) then swept the password-era residue — the now-dead login -concurrency limiter and the `ResetMailer` — so no unused password machinery lingered in -the compile path, and reconciled the stale comments that referenced it. - -## Why - -`op.console` needs a real login even in deployments without a Cloudflare-Access edge; the -password backend was that. Once the passwordless factors landed, keeping the old password -scaffolding around was a bug farm — the sweep is the closeout evidence that the migration -was complete, not half-done. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. `af14f02` was -> covered by Go unit tests (content-type guard, anti-enumeration, forced-change lockdown) -> at its commit. Not independently re-verified for this doc; current tree green at -> `9911b8c` (WSL oracle, go1.26.4). diff --git a/docs/changes/2026-06-27-deploy-bootstrap-installer.md b/docs/changes/2026-06-27-deploy-bootstrap-installer.md deleted file mode 100644 index 9e570a8..0000000 --- a/docs/changes/2026-06-27-deploy-bootstrap-installer.md +++ /dev/null @@ -1,38 +0,0 @@ -# Deploy: one-line bootstrap installer + demo bring-up (ledger backfill) - -- **Type:** feature + fix — retroactive ledger entry -- **Date:** 2026-06-27 – 2026-07-03 -- **Area:** `deploy/` (bootstrap.sh, Dockerfiles, demo-up.sh), image build context -- **Commits:** - - `58fa4b0` feat(deploy): one-line bootstrap installer + distroless felis image (auto-detects apt/dnf, installs Docker/k3s/PostgreSQL, opens pg_hba to the pod CIDR, runs migrations, applies the control-plane bundle, leaves Web disabled pending `felis setup`) - - `94a3b7b` fix(deploy): harden bootstrap for RHEL-family Linux - - `deaa2f8` feat(deploy): zypper support (openSUSE/SLES) - - `318a724` feat(deploy): pacman support (Arch) - - `e5f1682` refactor(deploy)!: TUI (breaking walkthrough restructure) - - `28c3eee` refactor(deploy): improved TUI walkthrough - - `c14ed17` fix(docker): keep embedded `panel/` and `deploy/` in the image build context - - `d9e866f` fix(deploy): make the lobby image actually build (re-include `plugins/paper`, build on `gradle:8.14-jdk21`) - - `b84debf` feat(deploy): one-shot `demo-up.sh` — bootstrap → build/import limbo+lobby images → wire `[velocity]` image refs → `felis setup`, ending in the interactive Owner TUI -- **Tasks:** #26 (Phase A bootstrap verified end-to-end on the Demo VM) - -## What it did - -Made a bare Linux box a running Felis with one command. `bootstrap.sh` auto-detects the -host package manager across the four major families (apt/dnf/zypper/pacman), installs -whatever is missing (Docker, k3s, PostgreSQL, cloudflared), builds+imports the distroless -felis image, opens `pg_hba` to the pod CIDR, runs migrations, and applies the rendered -control-plane bundle. `demo-up.sh` wraps that plus the login-limbo/lobby image build and -`felis setup` into a single command, stopping only at the Owner-creation TUI it cannot -automate. - -## Why - -The spec calls for a self-hostable single-node deployment a SysAdmin can stand up without -a Kubernetes background. The package-manager fan-out and the demo wrapper are what make -"one line" true across real distros rather than only on the author's box. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the -> change-ledger's detail-doc axis. `deploy/` is shell + Dockerfiles (not Go-oracle -> verifiable); `d9e866f` records a real build+boot check (limbo `/healthz` 200, lobby -> reaches "Done"). Not independently re-verified for this doc; current tree green at -> `9911b8c`. diff --git a/docs/changes/2026-06-27-felis-cli-break-glass-setup.md b/docs/changes/2026-06-27-felis-cli-break-glass-setup.md deleted file mode 100644 index a3c5a9c..0000000 --- a/docs/changes/2026-06-27-felis-cli-break-glass-setup.md +++ /dev/null @@ -1,35 +0,0 @@ -# felis CLI: break-glass recovery console + first-run setup (ledger backfill) - -- **Type:** feature + fix — retroactive ledger entry -- **Date:** 2026-06-27 – 2026-06-30 -- **Area:** `cmd/felis` (break-glass/setup TUI, apply, migrate), `internal/api` (audit, owner store), `deploy/` -- **Commits:** - - `e108a37` feat(cli): break-glass emergency console TUI — root-only (`euid==0`), provisions/resets the Owner directly against Postgres, enables local login, prints a durable one-time-password summary - - `2d0bbb0` feat(cli): attribute break-glass recovery to the SysAdmin who runs it — bootstrap / recovery (bcrypt) / root-override, each audited with an honest `verified` flag and payload - - `a94b001` feat(deploy): break-glass Operator account provisioning - - `eb5875a` feat(felis): Operator break-glass op behind an operation menu - - `f5d00f3` feat(cli): `felis apply` for direct CRD creation - - `9c46632` feat(cli): `felis setup` first-run console (shared `runConsoleTUI` model, reclaim protection, cfsetup idempotency, `[auth].admin_hostname` respect) - - `7d91373` fix(migrate): honor `-config` placed after the `up` verb (flag.Parse stops at the first non-flag token) -- **Tasks:** #27 (B1 login→change-pw→TUI reset) - -## What it did - -Built the local-root recovery and first-run surface that bypasses web Zero Trust by -design. `felis breakGlass` mints or resets the Owner when the web login is unreachable; -`2d0bbb0` makes it accountable by recording *which* SysAdmin broke the glass across three -audited modes. `felis setup` is the non-emergency first-run twin sharing the same console -model. `felis apply` writes a `MinecraftServer` CRD directly, and `7d91373` fixes the -`migrate` flag parse so a configured DB path after `up` is honored. - -## Why - -An operator with root on the node and a kubeconfig must always be able to recover the -platform — that is break-glass's whole job, so it never refuses. Attribution -(`2d0bbb0`) closes the gap that root is machine authority, not a human identity: the root -gate is necessary but not sufficient for the audit trail. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The core logic was -> covered by Go unit tests over a fake owner store at each commit (auth match/non-match, -> the three audit modes, headless TUI drive). The bubbletea TUI glue is untested by house -> convention. Not independently re-verified for this doc; current tree green at `9911b8c`. diff --git a/docs/changes/2026-06-27-player-onboarding-b2.md b/docs/changes/2026-06-27-player-onboarding-b2.md deleted file mode 100644 index 54abcaf..0000000 --- a/docs/changes/2026-06-27-player-onboarding-b2.md +++ /dev/null @@ -1,36 +0,0 @@ -# Player onboarding data layer §B2: email-OTP, account-link, QR, Bind-Code (ledger backfill) - -- **Type:** feature + fix — retroactive ledger entry -- **Date:** 2026-06-27 – 2026-07-03 -- **Area:** `internal/api` (onboarding/auth-bind handlers), `internal/store` (migrations 0004–0006) -- **Commits:** - - `dbe34a1` feat(api): player email-OTP verification (§B2) — `POST /account/email/{start,verify}`; 6-digit code, SHA-256-at-rest, 10-min TTL, 5-attempt cap enforced in the repo - - `1f8b9bb` feat(api): record account-link auth source (`mojang|thirdparty`) for the dual-Yggdrasil split (§10) - - `116595f` feat(api): QR scan-login completion poll on the internal face (`GET /internal/account/link/status/{mc_uuid}`) — read-only, reuses `UserByMCUUID`, no migration - - `fe2ece0` feat(api): public Bind-Code onboarding (`POST /auth/bind`) — the one pre-account entrypoint of `console.`; refuses a staff-UUID code with 403 without consuming it, so the public door provably never yields an admin principal - - `55592ed` feat(auth): public auth-bind endpoint wiring - - `6c3999a` fix(api): rate-limit email-OTP sends to close the email-bomb vector - - `879b177` fix(api): make OTP-start throttle atomic to close the concurrent-burst bypass -- **Tasks:** #29 (B2 data layer), #32 (OTP rate-limit), #35 (atomic throttle), #39 (console access model) - -## What it did - -Built the Go-verifiable data layer of forced web onboarding: prove control of an email -(OTP), record which Yggdrasil authenticated an in-game UUID, let a phone already signed in -to the panel complete a QR device-code link, and let an account-less player redeem a -one-time Bind Code minted in the Login Lobby to create+link+session in one public step. -The two fixes bound the OTP abuse surface — a per-target send rate limit and an atomic -reserve that closes the check-then-act race on the attempt counter. - -## Why - -The spec forces onboarding through the web so every account is provably email-controlled -and UUID-linked before it can operate anything. The `op.console` redline in `fe2ece0` — a -staff-UUID code is refused without being consumed — is what keeps the public console door -from ever minting an admin principal. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The account/session -> logic, single-use codes, and the op.console redline were covered by handler tests + the -> OpenAPI parity gate at each commit; the identity guarantee behind a Bind Code lives in -> velocity/Java (CODE-ONLY) and is not verifiable from this repo. Not independently -> re-verified for this doc; current tree green at `9911b8c`. diff --git a/docs/changes/2026-06-27-username-reclaim-b3.md b/docs/changes/2026-06-27-username-reclaim-b3.md deleted file mode 100644 index f15ae1b..0000000 --- a/docs/changes/2026-06-27-username-reclaim-b3.md +++ /dev/null @@ -1,32 +0,0 @@ -# §B3 username-collision reclaim + account migration (ledger backfill) - -- **Type:** feature — retroactive ledger entry -- **Date:** 2026-06-27 – 2026-07-05 -- **Area:** `internal/api` (internal-face reclaim/blacklist, account migrate), `internal/store` (migration 0006) -- **Commits:** - - `a29571d` feat(api): reclaim squatted usernames for Mojang-priority players (§B3, 正版优先) — `POST /internal/player/reclaim` bars the squatter UUID + stashes its data (30-day hold) in one transaction, idempotent, returns the *first* reclaim's expiry; `GET /internal/player/blacklist/{mc_uuid}` is the login-gate check - - `fdb6efb` feat(account): migrate a live account's owned servers to a new account (§B3 inherit) -- **Tasks:** #30 (B3 game-login + username-collision reclaim) - -## What it did - -Built the data layer of the Mojang-priority collision flow: when the configured -third-party Yggdrasil and official Mojang issue the same username under different UUIDs, -the non-genuine squatter is displaced in favour of the real Mojang owner. Both tables are -keyed by `mc_uuid`, so the genuine player — identical username, *different* UUID — is -never caught by the bar. `fdb6efb` adds the inherit half: migrating an existing account's -owned servers onto a new account. - -## Why - -Two players cannot hold one username across two Yggdrasils; the spec resolves it in the -genuine Mojang owner's favour with a 30-day data hold for the displaced squatter, told the -truth about how long their data is kept (the first hold's window, never a fresh `now()+30d` -on retry). - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Handlers + the -> in-memory repo contract were unit-tested at each commit; the Postgres SQL path is -> integration-only, and the velocity collision-routing / limbo prompt / authlib -> dual-backend are code-only (Java) and out of this data-layer slice. Not independently -> re-verified for this doc; current tree green at `9911b8c`. Related: the operator-facing -> `/felis migrate` command has its own doc ([felis-migrate-command](2026-07-05-felis-migrate-command.md)). diff --git a/docs/changes/2026-06-30-felis-api-hardening.md b/docs/changes/2026-06-30-felis-api-hardening.md deleted file mode 100644 index a7e8af4..0000000 --- a/docs/changes/2026-06-30-felis-api-hardening.md +++ /dev/null @@ -1,39 +0,0 @@ -# felis-api security + robustness hardening (audit sweep) (ledger backfill) - -- **Type:** fix — retroactive ledger entry -- **Date:** 2026-06-30 – 2026-07-01 -- **Area:** `internal/api` (login, request-id, listeners, SSE relays, quota/claim, MyServers), `internal/operator` -- **Commits:** - - `7a51c1d` fix(api): bound concurrent login bcrypt to shed CPU-pin floods (429 `auth_busy` before the compare; a cap, not a per-account lockout) — *audit #2* - - `164ac44` fix(api): validate inbound `X-Request-Id` before echo + audit persist (≤64 bytes, log-safe charset) — *audit-integrity* - - `c6c0772` fix(api): read/idle timeouts on all three listeners via a `newAPIServer` factory (closes Slowloris via `ReadHeaderTimeout`; `WriteTimeout` left unset so SSE isn't severed) — *audit #3* - - `3c1d647` fix(api): per-principal SSE stream cap (429 `too_many_streams`) — *audit #1, blast-radius bound* - - `d6e3189` fix(api): per-write deadline on SSE relay to sever a stalled reader (the real leak close behind the cap) — *audit #1* - - `8f41a00` fix(api): clear the SSE write deadline on return so it can't leak onto a reused keep-alive connection — *audit #1* - - `6368ab1` fix(api): `COALESCE` the MyServers `owned` flag so an ownerless row doesn't 500 the listing - - `2a4a81b` fix(api): don't burn the wake cooldown when refused at capacity - - `9873904` fix(operator): populate `Status.Players` from an RCON `list` probe (so the panel doesn't report 0/0) -- **Tasks:** #33 (wake cooldown), #34 (Status.Players), #41–#46 (audit #1–#4) - -## What it did - -A hardening sweep across the API's abuse and robustness surface: bound the two unbounded -CPU/goroutine amplifiers (concurrent bcrypt, per-principal SSE streams), close the SSE -relay's real stalled-reader leak with a per-write deadline (and clear it so it can't leak -onto a pooled connection), validate the caller-supplied request id before it reaches the -audit trail, set listener timeouts to close Slowloris, and fix two functional bugs — the -ownerless-row 500 and the wake cooldown burned on a capacity refusal. - -## Why - -Each is a specific, demonstrated failure mode: a login flood pins every core in bcrypt; a -stalled SSE reader leaks a relay goroutine + its upstream kube-apiserver follow *for the -life of the process*; an unvalidated `X-Request-Id` is a CR/LF log-forgery vector. The -`WriteTimeout`-left-unset detail is load-bearing — a blanket write timeout would sever the -healthy long-lived console/build-log streams the platform depends on. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Each fix shipped a -> targeted test at its commit — notably `d6e3189`/`8f41a00` use a deadline-aware -> `ResponseWriter` that fails closed if the guard is removed. The quota-claim TOCTOU -> (audit #4) is a documented KNOWN-LIMITATION (`2c56d17`), closeable only against a real -> Postgres. Not independently re-verified for this doc; current tree green at `9911b8c`. diff --git a/docs/changes/2026-06-30-felis-metrics.md b/docs/changes/2026-06-30-felis-metrics.md deleted file mode 100644 index 00c4e6b..0000000 --- a/docs/changes/2026-06-30-felis-metrics.md +++ /dev/null @@ -1,25 +0,0 @@ -# felis_* Prometheus metrics (§23) (ledger backfill) - -- **Type:** feature — retroactive ledger entry -- **Date:** 2026-06-30 -- **Area:** `internal/metrics` + the emit sites in build, platform/fleet, and the start lifecycle -- **Commits:** - - `75642d9` feat(metrics): named `felis_*` Prometheus collectors - - `2a93a9e` feat(metrics): record `felis_image_build_failures_total` on failed builds - - `79eae7f` feat(metrics): publish `felis_servers_total` from a fleet snapshot - - `8ac5e64` feat(metrics): observe `felis_start_duration_seconds` across the start lifecycle -- **Tasks:** #17 (§23 felis_* metrics decision) - -## What it did - -Added the named `felis_*` collector set and wired the three emit points that make it -non-empty: a counter incremented on image-build failure, a gauge published from a fleet -snapshot, and a histogram observed across the server start lifecycle. - -## Why - -§23 calls for first-class operational metrics under a stable `felis_` namespace rather than -ad-hoc logging, so an operator can alert on build failures, fleet size, and start latency. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Not independently -> re-verified for this doc; current tree green at `9911b8c` (WSL oracle, go1.26.4). diff --git a/docs/changes/2026-07-01-auto-update-subsystem.md b/docs/changes/2026-07-01-auto-update-subsystem.md deleted file mode 100644 index 76fd06d..0000000 --- a/docs/changes/2026-07-01-auto-update-subsystem.md +++ /dev/null @@ -1,37 +0,0 @@ -# Auto-update subsystem: decision core + sources + gatherer + window API (ledger backfill) - -- **Type:** feature + fix — retroactive ledger entry -- **Date:** 2026-07-01 – 2026-07-05 -- **Area:** `internal/updates` (pure decision core), `internal/updater` (release sources, gatherer), `internal/api` (window admin API) -- **Commits:** - - `c01f133` feat(updates): pure I/O-free decision core — each tracked component is Pinned (Minecraft, left alone), Notify, or Scheduled (apply only inside a SysAdmin window); never force-applied, never a downgrade, never an auto-applied prerelease - - `3673af6` feat(api): admin API for the maintenance window (`GET`/`PUT /updates/window`), stored as JSON under `platform_settings` — API + persistence only, nothing consumes it yet - - `7464fa7` fix(updates): tag `Window` JSON so the persisted window round-trips (the obvious decode is correct by construction; a zero window fails closed to notify-only) - - `96b3cc9` feat(updater): wire `updates.Run` to a caller with PaperMC v3 release discovery - - `7d27640` feat(updater): GitHub Releases source, routing felis-api/k3s/cloudflared - - `7db57b9` feat(updater): `VersionGatherer` extraction core + CLI gather seam -- **Tasks:** #38 (auto-update: Felis/k3s/components/Velocity, pin Minecraft) - -## What it did - -Built the auto-update spine as a pure decision core plus the release-discovery sources -(PaperMC, GitHub Releases) and the version gatherer, with a SysAdmin-set maintenance -window read/written through an admin API. Version parsing tolerates the real feeds (leading -`v`, k3s `+k3s1` suffix, calendar versions, prerelease tails) and orders by SemVer -precedence. - -## Why - -The red lines are `不要强制自动更新` (never force auto-update) and `能不动的就别动` -(Minecraft stays pinned). The design encodes them structurally: a component may be applied -*only* inside a window the operator explicitly set, and Minecraft is Pinned so it is never -touched. `7464fa7`'s fail-closed zero-window (decodes to notify-only, never a rogue apply) -is the safety property for the not-yet-built runner. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The load-bearing -> invariants (pinned never changes, no downgrade, no auto-prerelease, apply-only-in-window) -> and the JSON round-trip contract were unit-tested at their commits. This subsystem is -> deliberately **report-only / integration-deferred**: the concrete Notifier/Applier, -> the `felis update` CLI + CronJob, and the current-version producing seams are declared -> but not wired (see `internal/updater/doc.go`, `openapi.yaml`). Not independently -> re-verified for this doc; current tree green at `9911b8c`. diff --git a/docs/changes/2026-07-01-passkey-enrollment.md b/docs/changes/2026-07-01-passkey-enrollment.md deleted file mode 100644 index 772bd38..0000000 --- a/docs/changes/2026-07-01-passkey-enrollment.md +++ /dev/null @@ -1,37 +0,0 @@ -# Passkey (WebAuthn) enrollment subsystem + hardening (ledger backfill) - -- **Type:** feature + fix — retroactive ledger entry -- **Date:** 2026-07-01 – 2026-07-02 -- **Area:** `internal/passkey` (go-webauthn adapter), `internal/api` (enrollment handlers/audit), `internal/store` (migrations 0007–0009) -- **Commits:** - - `f2c916d` feat(api): passkey enrollment persistence layer - - `742f15f` feat(api): passkey enrollment endpoints - - `0261204` feat(passkey): go-webauthn enrollment verifier adapter (Oracle-verified against a virtual authenticator) - - `fce0fce` feat(passkey): wire the enrollment verifier into felis-api - - `7278cd7` feat(passkey): require + record user verification at enrollment (`UserVerification=required`; capture `user_verified`/`backup_eligible`/`backup_state` — migration 0009) — *fix (d)* - - `cdbb5ab` fix(api): record credential id in the passkey-register audit event so bind/unbind are symmetric — *fix (a)* - - `9953275` fix(api): bound `webauthn_challenges` growth by superseding *all* prior rows per (user, purpose) — *fix (b)* - - `20e31fb` fix(store): cascade-delete passkeys + challenges on user removal (recreate both FKs `ON DELETE CASCADE`, scoped to the passkey tables only) — *fix (c)* - - `54bc6ef` fix(api): clear bound passkeys on password change to close a takeover foothold — *fix (e)* -- **Tasks:** #36 (passkey bind with email-OTP fallback), #48–#52 (fixes a–e) - -## What it did - -Built the WebAuthn *enrollment* half — persistence, the go-webauthn crypto adapter, and -the register-begin/finish endpoints — then hardened it through the five-fix batch (a–e): -symmetric audit, a bounded challenge table, cascade cleanup, enforced+recorded user -verification, and unbinding every passkey on a password reset so a passkey planted through -a transiently-hijacked session cannot survive as a standing login foothold. - -## Why - -Passkeys are the phishing-resistant factor with email-OTP as the fallback. The hardening -batch closes the seams that make enrollment safe to *rely on*: without UV enforcement a -passkey proves possession but not user; without the password-reset clear, a planted -passkey outlives the very remediation meant to evict an attacker. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The adapter crypto -> was verified against a virtual authenticator (virtualwebauthn), and each fix shipped -> with a targeted test (UV-negative rejection, challenge-growth bound, cascade, symmetric -> audit) at its commit. Not independently re-verified for this doc; current tree green at -> `9911b8c`. The assertion/login half is a separate doc ([passkey-login](2026-07-01-passkey-login.md)). diff --git a/docs/changes/2026-07-01-passkey-login.md b/docs/changes/2026-07-01-passkey-login.md deleted file mode 100644 index 978d7f7..0000000 --- a/docs/changes/2026-07-01-passkey-login.md +++ /dev/null @@ -1,36 +0,0 @@ -# Passkey (WebAuthn) login: assertion, discoverable, clone-detection (ledger backfill) - -- **Type:** feature — retroactive ledger entry -- **Date:** 2026-07-01 – 2026-07-05 -- **Area:** `internal/passkey` (assertion crypto), `internal/api` (login/assertion, unbind, UA-guard), `internal/store` (migrations 0013/0014) -- **Commits:** - - `e035142` feat(passkey): WebAuthn login/assertion crypto adapter (BeginLogin/FinishLogin over go-webauthn, Oracle-verified against a virtual authenticator; surfaces the signature counter as a ceremony fact) - - `ec468ba` feat(auth): discoverable (usernameless) passkey login — the from-zero door the username-first assertion couldn't key on - - `0dbd557` fix(store): renumber the discoverable-login migration 0013 → 0014 - - `9e1df12` feat(passkey): advance `sign_count`, reject clone-warned assertions - - `4f59d51` feat(auth): owner-tier passkey-unbind remediation endpoint - - `a63f49d` feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey — a backend-only UA interstitial (the SPA is untouched); asset/API/health requests pass through, an `ua_ack` cookie lets a determined user continue -- **Tasks:** #40 (from-zero discoverable login), #67 (WeChat/QQ UA-guard in `internal/panel`) - -## What it did - -Built the assertion (login) half of the ceremony: the crypto adapter, then discoverable -credentials so a user with no typed identifier can still log in (the enrollment -identifier problem the earlier deferral doc named), clone detection via the advancing -signature counter, and the owner-tier unbind remediation. `a63f49d` guards the flow at the -transport edge — WebAuthn is unusable inside the WeChat/QQ WebViews, so those UAs get a -bilingual "open in your system browser" page instead of the passkey SPA. - -## Why - -Enrollment without a login path is half a feature. Discoverable credentials resolve the -blocker recorded in the earlier deferral (`users.email` is nullable/non-unique and a -player's username is their Minecraft UUID, so username-first assertion had nothing to key -on). The UA-guard stops the most common real-world dead end: a passkey prompt that can -never succeed inside an in-app browser. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The assertion crypto -> was verified against a virtual authenticator (enrollment→assertion chain, origin-mismatch -> and unbound-credential rejection); `9e1df12`'s clone policy and the UA-guard pass-through -> were unit-tested at their commits. Not independently re-verified for this doc; current -> tree green at `9911b8c`. diff --git a/docs/changes/2026-07-02-system-servers-login-limbo-lobby.md b/docs/changes/2026-07-02-system-servers-login-limbo-lobby.md deleted file mode 100644 index 7a17dea..0000000 --- a/docs/changes/2026-07-02-system-servers-login-limbo-lobby.md +++ /dev/null @@ -1,37 +0,0 @@ -# System servers: login-limbo + lobby (always-on gate) (ledger backfill) - -- **Type:** feature — retroactive ledger entry -- **Date:** 2026-07-02 -- **Area:** `internal/config`, `internal/naming`, `internal/api` (CRD readiness), `internal/operator`, `internal/platform`, `cmd/felis`, `plugins/limbo`, `deploy/limbo` + `deploy/lobby` -- **Commits:** - - `9bed51b` feat(config): `[velocity] login_image/lobby_image` — setup provisions the always-on system services only when set (empty = fail-loud skip; no official LOOHP/Limbo image exists) - - `9ef817f` feat(naming): reserved system-server names + service-token identifiers (single source of truth for the internal-API credential Secret) - - `159107b` feat(api): HTTP readiness knob on `MinecraftServer` + user-server fallback defaults to the login gate - - `dc23cb5` feat(operator): system-server pod HTTP readiness probe + login-only `FELIS_SERVICE_TOKEN` env (keyed off the reserved name so it can never leak into a user pod; sourced via `secretKeyRef`, never inlined) - - `3fdb3d0` feat(platform): internal-API base-URL helper + single-sourced token Secret - - `f554d52` feat(cli): provision the reaper-exempt login/lobby servers + replicate the service-token Secret into the minecraft namespace - - `241fe21` feat(limbo): felis-limbo in-game login flow (join → blacklist check → mint bind code → open book to `console.` → poll link-status → BungeeCord transfer to lobby; fail-closed) - - `c7315e4` feat(deploy): login-limbo + lobby images with game-port pinning (server-port pinned to GamePort 25565 on every start) -- **Tasks:** #53–#68 (system-server plumbing L1–L4, limbo plugin, operator env injection) - -## What it did - -Stood up the always-on authentication gate: reserved, reaper-exempt login/lobby -`MinecraftServer`s provisioned by setup, an HTTP readiness path for the RCON-less LOOHP/Limbo -loader (which reports "started" only after the first tick), and the felis-limbo plugin that -runs the whole onboarding *inside* Limbo before transferring an admitted player to the -lobby. A fresh connection always lands on the login gate, never a user backend, so -authentication is always in front. - -## Why - -The spec requires that a player authenticate before reaching any real server. That needs a -purpose-built always-on front server (Limbo) that speaks to the internal API — hence the -login-only service-token injection (keyed to the reserved name so it can never reach a user -pod) and the HTTP readiness knob for a loader that has no RCON. - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The Go layer -> (config/naming/readiness/operator env/platform) was unit-tested at each commit; the -> felis-limbo plugin is Java verified against a real Limbo jar via podman (#65), and the -> images carry a real build+boot check (#63, limbo `/healthz` 200 on 25565). Not -> independently re-verified for this doc; current tree green at `9911b8c`. diff --git a/docs/changes/2026-07-05-break-glass-halt.md b/docs/changes/2026-07-05-break-glass-halt.md deleted file mode 100644 index c0fe819..0000000 --- a/docs/changes/2026-07-05-break-glass-halt.md +++ /dev/null @@ -1,83 +0,0 @@ -# Break-glass "halt a running server" op (#31 B4) - -- **Type:** feature (addition) -- **Date:** 2026-07-05 -- **Area:** `cmd/felis` — break-glass recovery console (Go, oracle-verifiable) -- **Commit:** `c2ee21a` — feat(breakglass): add halt-a-server op to the recovery console (§B4) -- **Task:** #31 Phase B4 (felis TUI break-glass ops) - -## What it does - -Adds a **"Halt a running server"** operation to the root-gated break-glass console. -The operator picks a server from the live fleet and the console flips that -`MinecraftServer` CRD's `spec.desiredState` to `Stopped`, letting the operator -reconcile it into a graceful shutdown. It is the emergency "stop this now" lever for -when the panel is unreachable but the box still has `root` + a kubeconfig. - -## Why - -The break-glass console already provisions the Owner and adds Operators, but there -was no local, panel-independent way to **stop** a misbehaving server (runaway, -compromised, resource-pinning). Halting is a reversible state nudge — the safest -possible break-glass power — so it belongs in the same root-gated recovery surface. - -## Design decisions - -- **CRD write, not pod kill.** The console flips `spec.desiredState=Stopped` with a - **spec-only merge patch** (`client.MergeFrom`), never a full-object `Update`. The - operator writes `status` on the same object continuously; a merge patch of - `spec.desiredState` touches a disjoint field and cannot race/clobber the operator's - status writes. A halt is therefore exactly the CRD write the operator already knows - how to honour. -- **Authority = root + kubeconfig.** The accountable actor is the OS user who - escalated to root (`osUser`), recorded for attribution — not proof. The root gate - plus kubeconfig possession *is* the authority, so (unlike the owner/operator paths) - no credential-minting auth sub-flow is needed for a reversible state change. -- **System servers allowed but named.** Halting the `login`/`lobby` system servers - takes the shared front door down (login has no fallback). Break-glass is deliberately - full power, so the console **warns** rather than forbids: a `⚠ system` tag in the - picker and an explicit `WARNING` line in the post-exit summary. -- **Audit is best-effort.** `performHalt` mirrors `performBreakGlass`: the halt - succeeds even if the audit sink is down (break-glass must work with logging broken); - any audit error rides back in the outcome and is surfaced as a summary `WARNING`. -- **Already-stopped is a no-op** reported distinctly ("was already stopped" vs "is now - stopping"), so the console never claims a stop it didn't perform. -- **Namespace from config.** The target namespace is `cfg.K8s.Namespace`, threaded - through the console constructors — never hardcoded. - -## Files - -| File | Change | -|---|---| -| `cmd/felis/halt.go` | **new** — pure core (no bubbletea): `listServersForHalt`, `haltServer` (merge patch), `isSystemServer`, `performHalt`, `auditHalt` | -| `cmd/felis/halt_test.go` | **new** — table tests against a controller-runtime **fake client** (applies patches for real): running→stopped persists, already-stopped no-op, missing→error, system flag, list projection + desired-state fallback, audit success, audit-failure-still-halts | -| `cmd/felis/tui_halt.go` | **new** — bubbletea/huh shell mirroring `ownerModel` (load → pick → work → done), empty-fleet guard, `⚠ system` picker labels, outcome card | -| `cmd/felis/tui_menu.go` | `bgHaltServer` enum + "Halt a running server" menu option | -| `cmd/felis/tui_root.go` | `namespace` field; `bgHaltServer` dispatch to `newHaltModel`; `haltResultMsg` terminal handling | -| `cmd/felis/breakglass.go` | halt fields on `breakGlassResult`; `namespace` threaded through `runBreakGlassTUI`/`runSetupTUI`/`runConsoleTUI`; post-exit halt summary (stopping / already-stopped, system + audit warnings, restart hint) | -| `cmd/felis/setup.go` | pass `cfg.K8s.Namespace` into `runSetupTUI` | -| `cmd/felis/tui_root_test.go` | pass `"minecraft"` namespace into `newRootModel` test call | - -## Verification - -WSL oracle (go1.26.4, FedoraLinux-44), authoritative for Go: - -``` -go build ./... → BUILD_OK -go vet ./cmd/felis/... → VET_OK -go test ./... → all 20 packages ok, ALL_GREEN -``` - -The core (`halt.go`) is fully unit-tested against a real `fake.Client`, which applies -the merge patch, so the test asserts the **persisted** `spec.desiredState`, not merely -that `Patch` was called. `tui_halt.go` is thin bubbletea glue (untested by house -convention, mirrors the existing `tui_owner.go`). - -## Self-review outcome - -- **ponytail (over-engineering):** lean — no one-impl interface, every field consumed, - audit seam justified. Nothing cut. -- **correctness:** caught and fixed a misleading restart hint — the summary originally - pointed at `felis apply`, but that command is **create-only** (errors "already - exists" on an existing server); corrected to "restart from the panel, or set - `spec.desiredState` back to Running." diff --git a/docs/changes/2026-07-05-felis-migrate-command.md b/docs/changes/2026-07-05-felis-migrate-command.md deleted file mode 100644 index 0624904..0000000 --- a/docs/changes/2026-07-05-felis-migrate-command.md +++ /dev/null @@ -1,80 +0,0 @@ -# `/felis migrate` in-game command (§B3 inherit, Velocity side) - -- **Type:** feature (addition) -- **Date:** 2026-07-05 -- **Area:** `plugins/velocity` + `plugins/shared` — Velocity proxy plugin (Java, compile-verified) -- **Commit:** `c1aa38b` — feat(velocity): add /felis migrate to open an account migration (§B3 inherit) -- **Task:** completes the code-only gap named in `internal/api/handlers_account_migrate.go` - -## What it does - -Adds the in-game `/felis migrate` command that a player runs to **open an account -migration** — the first step of handing their owned servers to another account (spec -§B3 "inherit", scenario A). The command posts the player's Mojang-verified UUID to the -backend, which puts that account into migrate mode (`state=initiated`). The player then -finishes the migration on the web console (prove it's them, name the receiving account, -redeem a one-time code). - -The Go backend (`handleMigrateStart` and the web-driven steps 2–4) already existed and -was tested; its header comment explicitly named **"the `/felis migrate` command that -calls handleMigrateStart"** as the code-only gap. This change closes that gap. - -## Why - -Without the in-game command, the migration flow had no entry point — the backend -handler was reachable only in theory. `/felis migrate` is the trustworthy initiator: -Velocity has already established the caller's online-mode UUID, so the sensitive proof -can be deferred to the web step-up while the in-game command just opens the migration. - -## Design decisions - -- **Mirrors the existing command suite verbatim.** `doMigrate` follows `doClaim`; - `migrateError` follows `claimError`; `migrateStart` follows `claim`/`opLoginApprove`. - No new imports, types, or idioms — every construct already appears in the same files. -- **Identity-bound + out-of-limbo, but server-independent.** Like `claim`, it requires - a real player past the login limbo (`requirePlayer` + `ensureOutOfLimbo`). Unlike - `claim`, it acts on the caller's *account*, not the server they stand on, so there is - **no** `registry`/current-server check. -- **Expects HTTP 201.** `migrateStart` posts to - `/api/v1/internal/account/migrate/start` and expects **201 Created** (`handleMigrateStart` - returns `StatusCreated`) — not 200 like the other calls. A 201 that does not affirm - `started:true` is treated as a contract breach, not a refusal. -- **Error mapping matches the handler's refusals:** 404 `not_linked` → "Link your - account on the web console before migrating"; 409 `account_retired` → "This account - can't start a migration (already migrated or retired)"; transport (0) and default → - generic retry text. -- **Points the player to the console on success.** The command only *opens* the - migration, so on success it prints the player web console URL - (`https://console.`, derived from config — never a hardcoded domain) and - a one-line description of the remaining steps. A proxy-side `logger.info` records the - initiating username against the UUID (the backend audit only has the UUID). - -## Files - -| File | Change | -|---|---| -| `plugins/shared/.../link/FelisApiClient.java` | **+`migrateStart(UUID)`** — POST mc_uuid, expect 201, affirm `started:true` | -| `plugins/velocity/.../FelisVelocityPlugin.java` | `migrate` literal in the Brigadier tree; **`doMigrate`** handler; **`migrateError`** mapper; `/felis migrate` help line | - -## Verification - -Java is not oracle-verifiable via the Go suite, but it **is** compile-verifiable via -the podman gradle toolchain established in #63/#65: - -``` -podman run --rm -v plugins:/work -w /work/velocity \ - docker.io/library/gradle:jdk17 gradle --no-daemon compileJava -→ BUILD SUCCESSFUL in 19s (compiled against real velocity-api:3.3.0-SNAPSHOT) -``` - -The change compiles clean against the real Velocity API jar (including the shared -`FelisApiClient` compiled straight into the velocity module). The backend contract it -speaks to (`handleMigrateStart`) is covered by `handlers_account_migrate_test.go` on -the Go side. - -## Self-review outcome - -- **ponytail (over-engineering):** lean — pure mirror of three existing, compiling - methods; no speculative abstraction. Nothing cut. -- **correctness:** the one contract divergence (201 vs 200) was verified against the Go - handler source before writing. diff --git a/docs/changes/2026-07-05-operator-idle-quota-readiness.md b/docs/changes/2026-07-05-operator-idle-quota-readiness.md deleted file mode 100644 index 76ef2bc..0000000 --- a/docs/changes/2026-07-05-operator-idle-quota-readiness.md +++ /dev/null @@ -1,29 +0,0 @@ -# Operator: idle auto-stop, quotas, startup/readiness timeouts, /readyz (ledger backfill) - -- **Type:** feature + fix — retroactive ledger entry -- **Date:** 2026-07-05 -- **Area:** `internal/operator` (idle stop, timeouts), `internal/api` (quotas, /readyz) -- **Commits:** - - `91bfa27` feat(operator): idle auto-stop (§8) - - `e574749` feat(api): enforce CPU/memory/storage quotas (§9.3, §22) - - `7f7e459` fix(operator): enforce startup and readiness timeouts (§5, §8) - - `7becb38` fix(api): implement `/readyz` with real DB + K8s API + CRD checks (§7) -- **Tasks:** §5/§7/§8/§9.3/§22 operator + resource-governance spec items - -## What it did - -Rounded out the operator's lifecycle governance: stop idle servers automatically, enforce -per-resource CPU/memory/storage quotas at claim/create, bound how long a server may sit in -startup/readiness before the operator gives up, and make `/readyz` a real dependency check -(DB, Kubernetes API, and the CRD) rather than a static 200. - -## Why - -An orchestrator that never reclaims idle capacity or bounds startup will accumulate stuck -and wasteful workloads; a `/readyz` that always returns 200 tells the load balancer a -broken control plane is healthy. These are the spec's resource-governance and -readiness-correctness requirements (§5/§7/§8/§9.3/§22). - -> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Covered by Go unit -> tests at each commit. Not independently re-verified for this doc; current tree green at -> `9911b8c` (WSL oracle, go1.26.4). diff --git a/docs/changes/2026-07-07-break-glass-backup-peer.md b/docs/changes/2026-07-07-break-glass-backup-peer.md deleted file mode 100644 index 89e8bc5..0000000 --- a/docs/changes/2026-07-07-break-glass-backup-peer.md +++ /dev/null @@ -1,117 +0,0 @@ -# Break-glass "back up a world now" console peer (§B4 "Sync", phase 2b) - -- **Type:** feature (addition) -- **Date:** 2026-07-07 -- **Area:** `cmd/felis` (Go, oracle-verified); `internal/api` + `docs/openapi.yaml` - (the internal endpoint's `os_user` accountability extension) -- **Commit:** `fc748d3` -- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation. Per the user's - **"两者都要"** decision the feature was built in two halves: the felis-api endpoint - that does the real backup-Job orchestration (phase 1 `7a7c0d5` external face, phase - 2a `f2fc57c` internal face) and a break-glass menu peer that calls it while the API - is alive. **This change is that peer — phase 2b — the last build step of "Sync".** - (S3 remains open in B4; this does not close the phase.) - -## What it does - -Adds a **"Back up a world now (Sync)"** operation to the root-gated break-glass -console. The operator picks a server from the live fleet; the peer resolves the -`felis-api-internal` ClusterIP Service and the service token from the control -namespace, then POSTs the internal backup endpoint to snapshot that server's world -while felis-api is alive. It is the on-node counterpart to the halt op: an emergency -"snapshot this now" lever for when the panel is unreachable but the box still has -`root` + a kubeconfig and the API is running. - -It also extends the internal backup endpoint (phase 2a) to accept an optional -`{"os_user":"..."}` body so the audit row names the operator at the keyboard rather -than the generic `break-glass`. - -## Why - -The console cannot render the backup Job itself — it lacks the deployment coordinates -(`FELIS_IMAGE`, `FELIS_BACKUP_PVC`) that only felis-api holds. So, unlike the halt peer -(which writes the `MinecraftServer` CRD directly), the backup peer must go **through** -the API. The internal face exists precisely so an on-node machine caller with the -service token — no browser session, no Cloudflare-Access Principal — can reach that -orchestration. This change is the client that knocks on that door. - -## Design decisions - -- **Goes through the API, does not orchestrate locally.** Mirrors the phase-2a - rationale: deployment coordinates live only in felis-api. The peer's job is to - resolve the endpoint, authenticate with the token, and translate the HTTP result - into a friendly outcome card — not to build a Job. -- **ClusterIP resolution, not DNS.** `resolveInternalAPI` `Get`s the - `platform.APIInternalServiceName` Service and dials its `Spec.ClusterIP:APIInternalPort` - directly, erroring on an empty or `None` (headless) ClusterIP. The on-node console's - host resolver is not CoreDNS, so the in-cluster Service DNS name would not resolve - from the host; the ClusterIP is routable from the node and is what the `2ba9948` - dedicated ClusterIP Service exists to provide. -- **`os_user` attribution, parity with halt.** The console sends the escalated OS user - in the request body; the endpoint makes it the audit actor. The body is decoded - whenever `ContentLength != 0` — **not** gated on `Content-Type` (`decodeJSON` checks - only for unknown/trailing fields, not the header) — so a console that forgets the - header still records the operator. Absent/blank falls back to `break-glass`. The - console does **not** double-audit: the API audits at the boundary, single-sourced. -- **Stopped-gate stays server-side.** The world PVC is RWO, so the server must be - stopped. The peer does not pre-check this; it lets the API's stopped-gate return - `409 not_stopped` and renders that as a "must be stopped — halt it first" card. The - safety check is single-sourced in the API, never duplicated (and possibly drifting) - in the console. -- **A running pick ends the session with exit 1 — deliberately accepted.** The picker - lists all servers (a running pick is easy to hit), and a `409` surfaces as an error - through `backupResultMsg{err}` → root sets `m.err` → `breakglass.go` prints the - friendly card to stderr and exits non-zero. A dedicated `not_stopped` non-error path - was considered and **rejected**: every break-glass op ends the session anyway (all - `tea.Quit`), so `409`-vs-success differs only in exit code — marginal for an - interactive TUI. No error is swallowed; the friendly card is shown either way. Adding - a soft-landing state machine for one status code is complexity the interactive - surface does not earn. -- **Core/shell split, mirrors halt.** All decision logic (`resolveInternalAPI`, - `requestBackup`, `backupErrorFromResponse`, `performBackupNow`) lives in `backupnow.go` - and is unit-tested against a controller-runtime fake client + `httptest`. - `tui_backupnow.go` is thin bubbletea/huh glue (untested by house convention, mirrors - `tui_halt.go`). The picker **reuses** `listServersForHalt`/`haltableServer` rather - than cloning a second server-listing path. - -## Files - -| File | Change | -|---|---| -| `cmd/felis/backupnow.go` | **new** — pure core: `resolveInternalAPI` (Service ClusterIP + token secret), `requestBackup` (POST + Bearer + `os_user` body, status→outcome), `backupErrorFromResponse` (409/503/404/error-body mapping), `performBackupNow` | -| `cmd/felis/backupnow_test.go` | **new** — table tests against a fake client + `httptest`: happy resolve, headless/empty-token errors; 202 asserts Bearer + `os_user` + path; 409/503/404 + transport-failure mapping | -| `cmd/felis/tui_backupnow.go` | **new** — bubbletea/huh shell mirroring `tui_halt.go`: load → pick → work → outcome card; empty-fleet guard; friendly error card | -| `cmd/felis/tui_menu.go` | `bgSyncBackup` enum + "Back up a world now (Sync)" menu option after the halt option | -| `cmd/felis/tui_root.go` | `bgSyncBackup` dispatch to `newBackupModel`; `backupResultMsg` terminal handling into `breakGlassResult` | -| `cmd/felis/breakglass.go` | `backedUp`/`backupServer`/`backupStatus` result fields; cancel guard; post-exit backup summary | -| `internal/api/handlers_backups.go` | `handleInternalBackup` decodes the optional `os_user` body (gated on `ContentLength`, not `Content-Type`) and passes it as the audit actor; defaults `break-glass` | -| `internal/api/handlers_backup_now_test.go` | **+subtest** in `TestInternalBackup`: `os_user` body attributes the audit to the operator | -| `docs/openapi.yaml` | document the `internalBackupNow` optional `os_user` request body | - -## Verification - -WSL oracle (go1.26.4, FedoraLinux-44, authoritative for Go): - -``` -go build ./... && go vet ./... && go test ./... → ALL_GREEN -``` - -`cmd/felis` and `internal/api` both re-ran (not cached), so the new `backupnow_test.go` -and the added `TestInternalBackup` subtest executed. `TestOpenAPIMatchesServedRoutes` -still passes: the `os_user` body is an addition to an already-documented operation, so -the served⇔documented route match is unchanged. The core is covered against a real -`fake.Client` (resolves the Service/Secret) + `httptest.Server` (asserts the wire -request and maps every status), so the tests exercise persisted/observable behaviour, -not merely that a call was made. `tui_backupnow.go` is thin glue, untested per the -`tui_halt.go` convention. - -## Self-review outcome - -- **ponytail (over-engineering):** lean — no new abstraction beyond the four core - functions two call sites (test + TUI) already justify; the picker reuses halt's - server-list core rather than cloning it; the deliberate rejection of a `not_stopped` - soft-landing path kept the state machine at four steps. Nothing cut. -- **correctness:** the `os_user` decode is gated on `ContentLength`, matching how - `decodeJSON` actually works (no `Content-Type` check), so a header-less console still - attributes correctly; the stopped-gate and audit stay single-sourced server-side, so - the peer cannot drift from the endpoint on the RWO safety check or the audit record. diff --git a/docs/changes/2026-07-07-internal-api-clusterip-service.md b/docs/changes/2026-07-07-internal-api-clusterip-service.md deleted file mode 100644 index c0ceb56..0000000 --- a/docs/changes/2026-07-07-internal-api-clusterip-service.md +++ /dev/null @@ -1,82 +0,0 @@ -# Separate ClusterIP Service for the felis-api internal face (8081) - -- **Type:** bug fix (latent networking gap) + enabling change -- **Date:** 2026-07-07 -- **Area:** `internal/platform` — Go (struct render oracle-verified; packet path **not** - verifiable in this environment — see Verification) -- **Commit:** `2ba9948` -- **Task:** #31 Phase B4 — surfaced while wiring the break-glass backup console peer - (phase 2b): the peer needs a routable path to the internal face, and that path was - broken for the login pod too. - -## What it does - -Renders a **new ClusterIP-only Service `felis-api-internal`** (control namespace) -that fronts the felis-api pod's internal port 8081, and repoints -`InternalAPIBaseURL` (the URL baked into the login pod's `FELIS_API_BASE_URL`) at -that Service name. Adds exported `APIInternalServiceName` / `APIInternalPort` so the -on-node break-glass console can resolve the Service's ClusterIP and dial it. - -## Why (the latent bug) - -The login limbo pod is configured with -`FELIS_API_BASE_URL = http://felis-api..svc.cluster.local:8081` (setup.go) and -dials the internal face with the service token to mint bind codes and poll link -status. But the only Service named `felis-api` is the **external** face: a NodePort -Service that declares **only** port 443. A Service answers only on its declared -ports, so `felis-api:8081` had no backend — **every login-pod call to the internal -API silently failed to connect.** `deploy/limbo/README.md` even documented the -"login-pod → felis-api internal-port (8081) path" as reachable; it was not. - -## Design decisions - -- **A separate Service, not a second port on `felis-api`.** A `Type: NodePort` - Service allocates a node port for **every** declared port, with no per-port - opt-out. Folding 8081 into the NodePort `felis-api` Service would therefore publish - the internal face — which is service-token-only, explicitly **no Zero Trust** — on - every node's external IP. That violates the two-face security posture. A distinct - `ClusterIP` Service exposes 8081 **in-cluster only**: reachable by the login pod via - cross-namespace DNS, and by the on-node console via the ClusterIP (kube-proxy - programs ClusterIPs into the node's routing). -- **Repoint `InternalAPIBaseURL` to the new Service name.** The helper single-sources - the name the login pod is told to call; pointing it at `felis-api-internal` keeps - the login pod and the Service in agreement by construction. -- **Export the name + port for the console.** The break-glass backup peer (phase 2b) - resolves `APIInternalServiceName`'s ClusterIP at runtime and dials - `http://:APIInternalPort` — it cannot use the cluster-DNS form because - the host's resolver is not CoreDNS. - -## Files - -| File | Change | -|---|---| -| `internal/platform/workloads.go` | **+`apiInternalService`** (ClusterIP, 8081→`internal`), wired into `Workloads()`; **+exported `APIInternalServiceName`/`APIInternalPort`**; `InternalAPIBaseURL` repointed at the internal Service, comment corrected | -| `internal/platform/workloads_test.go` | **+`TestAPIInternalService_ClusterIP`** — ClusterIP (never NodePort), 8081→`internal`, no nodePort, selects the api pods, name distinct from `felis-api` | -| `deploy/limbo/README.md` | document the `felis-api-internal` Service; correct the reachability note | -| `docs/troubleshooting.md` | §6 note: internal calls reached via `felis-api-internal`; a *connect* failure (not 401) points at that Service | - -## Verification - -WSL oracle (go1.26.4, authoritative for Go): - -``` -go build ./... && go vet ./... && go test ./... → ALL GREEN -``` - -`TestAPIInternalService_ClusterIP` freezes the Service's shape. **This is a -code-level fix only.** `go build/vet/test` verifies the Service *struct* renders -correctly; it verifies **nothing** about packets flowing — not the login pod's -in-cluster call, not the console's host→ClusterIP dial (which relies on kube-proxy's -OUTPUT-chain DNAT, present on k3s but unverified here), not that 8081 is programmed -on a live cluster. Per the project's "Java/K8s code-only" reality, the runtime path -is **pending real-cluster verification**; the manifest-level defect (a DNS name with -no backing port) is fixed and asserted. - -## Self-review outcome - -- **ponytail (over-engineering):** one Service + two exported identifiers, all - load-bearing (the login pod and the console both need the routable 8081). No new - abstraction; `apiInternalService` mirrors `apiService`/`registryService`. -- **correctness / security:** the ClusterIP-not-NodePort choice is the crux — it keeps - the no-Zero-Trust internal face off every node's external interface, which a second - port on the NodePort Service could not. diff --git a/docs/changes/2026-07-07-internal-backup-endpoint.md b/docs/changes/2026-07-07-internal-backup-endpoint.md deleted file mode 100644 index fda2379..0000000 --- a/docs/changes/2026-07-07-internal-backup-endpoint.md +++ /dev/null @@ -1,83 +0,0 @@ -# Internal-face break-glass world backup endpoint (§B4 "Sync", phase 2a) - -- **Type:** feature (addition) -- **Date:** 2026-07-07 -- **Area:** `internal/api`, `docs/openapi.yaml` — Go, oracle-verified -- **Commit:** `f2fc57c` -- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation. Per the user's - **"两者都要"** decision the feature is built in two halves: the felis-api endpoint - that does the real backup-Job orchestration (phase 1, `7a7c0d5`) and a break-glass - menu peer that calls it while the API is alive (phase 2b, follow-up). **This change - is phase 2a: the second, internal face of that endpoint** — the door the console - peer will knock on. - -## What it does - -Adds `POST /api/v1/internal/servers/{name}/backup`, an **internal-face** twin of the -external `POST /api/v1/servers/{name}/backup`. The on-node break-glass console (root -on the host, holding the service token) POSTs here to snapshot a stopped world while -felis-api is alive. Same 202 `backing_up` / 409 `not_stopped` / 503 -`backup_unavailable` / 404 / 400 `bad_name` surface as the external face. - -## Why - -The console cannot render the backup Job itself: it lacks the deployment coordinates -(`FELIS_IMAGE`, `FELIS_BACKUP_PVC`) that only felis-api holds — the same reason the -endpoint exists at all (phase 1). But the external face requires a Cloudflare-Access -Principal the console does not have. The internal face authenticates with the service -token (a trusted machine caller, no Principal), so the console can reach the same -orchestration without a browser session. - -## Design decisions - -- **No owner gate on the internal face.** The external handler enforces owner-or-admin - from the Principal; the internal handler has none — the service token IS the - authorization (the operator already has root on the node), so a server owned by - someone else still backs up. This mirrors how the other internal-face handlers - (op-login approve, QR poll) trust the token rather than a Principal. -- **Shared `enqueueBackup` tail.** The RWO stopped-gate, the optional-Backuper 503, - the async hand-off, and the audit+202 were refactored out of `handleBackupNow` into - a single `enqueueBackup(w, r, name, rec, actor, source)` that both faces call. The - two faces differ **only** in how the caller is authorized and in the audit - actor/source — the security-critical stopped-gate is single-sourced so the faces - cannot drift apart. -- **Audit attributed to break-glass/internal.** The internal handler audits directly - via `Repo.Audit` with `Actor:"break-glass", Source:"internal"` (the `a.audit` - helper hardcodes `Source:"external"`), so a console-initiated backup is - distinguishable in the audit log from an owner's self-service one. -- **Console does not double-audit.** Unlike the halt peer — which writes the CRD - directly and audits locally — the backup peer goes through the API, and the API - audits at the boundary. Auditing is single-sourced there; the console will not - emit its own row. - -## Files - -| File | Change | -|---|---| -| `internal/api/handlers_backups.go` | **+`handleInternalBackup`**, **+`enqueueBackup`**; `handleBackupNow` tail now calls `enqueueBackup(..., p.Email, "external")` | -| `internal/api/api.go` | register `POST /api/v1/internal/servers/{name}/backup` on the internal-face route table | -| `docs/openapi.yaml` | document the `internalBackupNow` operation (`x-felis-face: [internal]`, `serviceToken` security) | -| `internal/api/handlers_backup_now_test.go` | **+`TestInternalBackup`** — no-owner-gate, break-glass/internal audit, stopped-gate/503/404/400 | - -## Verification - -WSL oracle (go1.26.4, authoritative for Go): - -``` -go build ./... && go vet ./... && go test ./... → ALL GREEN -``` - -`TestOpenAPIMatchesServedRoutes` gates the new route against `docs/openapi.yaml` in -both directions (served⇔documented) and passes. `TestInternalBackup` (5 subtests) and -the existing `TestBackupNow` (10) both pass — the external refactor is -behaviour-preserving (same audit actor `p.Email`/source `external`). - -## Self-review outcome - -- **ponytail (over-engineering):** the internal face is not a copy of the external - handler — the shared tail (`enqueueBackup`) collapses the duplication, and the two - handlers hold only their distinct auth + audit-attribution. No new abstraction - beyond the one shared function two callers already justify. -- **correctness:** the no-owner-gate difference is deliberate and matches the other - service-token handlers; the stopped-gate is unchanged and now single-sourced, so the - external and internal faces cannot diverge on the RWO safety check. diff --git a/docs/changes/2026-07-07-on-demand-world-backup.md b/docs/changes/2026-07-07-on-demand-world-backup.md deleted file mode 100644 index 71f55b1..0000000 --- a/docs/changes/2026-07-07-on-demand-world-backup.md +++ /dev/null @@ -1,115 +0,0 @@ -# On-demand world backup (§B4 break-glass "Sync"; felis-api endpoint + Job executor) - -- **Type:** feature (addition) -- **Date:** 2026-07-07 -- **Area:** `internal/backupjob` (new pkg), `internal/api`, `cmd/felis`, `docs/openapi.yaml` — Go, oracle-verified -- **Commit:** `7a7c0d5` -- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation, resolved with the user as **immediate/on-demand world backup**. Per the user's "两者都要" decision this is built in two halves: **(this change) the felis-api endpoint that does the real backup-Job orchestration**, and (a follow-up) a break-glass menu peer that calls it while the API is alive. - -## What it does - -Adds `POST /api/v1/servers/{name}/backup`: an owner or admin snapshots a **stopped** -server's world into the archive store on demand, recorded as a first-class -`world_backups` row (reason `manual`) — restorable later by the existing restore path -and expired by the reaper's retention pass, so it never leaks as an orphan archive. - -The backup runs asynchronously as a one-shot Kubernetes Job (the new -`internal/backupjob` package), mirroring how restore and image builds hand off to -Jobs. The handler answers **202 `backing_up`**. - -## Why - -felis-api cannot archive a world in-process: the world PVC is **RWO** and owned by the -operator's StatefulSet, so the API has nothing to mount at request time — the same -constraint that already makes `internal/restore` a Job. The break-glass console (which -runs direct-to-Postgres) likewise lacks the deployment coordinates (`FELIS_IMAGE`, -`FELIS_BACKUP_PVC`) needed to render the Job. Both point to the same home: the -orchestration belongs in felis-api, which holds those coordinates; other callers -invoke the endpoint. - -## Design decisions - -- **Backup Job self-records its `world_backups` row.** Unlike the restore Job — which - is deliberately DB-blind because it processes a potentially poisoned archive — the - backup Job **does** mount the felis config Secret and inserts its own backup row, - exactly like the reaper (the only other component holding both a world mount and the - database). This avoids the archive-then-async-record split that would otherwise leak - orphan archives on a crash. The security review for that one departure lives in - `internal/backupjob/jobspec.go` and is frozen by `jobspec_test.go`. Rationale: a - backup only **reads** a world the operator already owns and tars it (bytes, never - executed), so restore's poisoned-input threat does not apply; its blast radius (DB + - two PVCs) is a strict subset of the reaper's, and it never deletes a PVC nor calls - the K8s API (SA token stays un-mounted). -- **World mounted read-only, backup PVC read-write** — the mirror image of restore. -- **Stopped-gate (409 `not_stopped`).** The world PVC is RWO and held by a running - server, so a backup Job cannot double-mount it; the handler refuses unless the server - is fully stopped (`info.Ready || DesiredState != Stopped`). This also guarantees a - quiescent, non-torn archive. Mirrors `handleRestoreBackup`'s gate. -- **Authorization is restore's front half, minus the former-owner match.** Backup is - initiated by the **current** owner and records **their** ownership, so there is no - prior owner's data to leak — the leak guard that restore needs does not apply here. - An admin may back up an unowned (released) world; the recorded former owner is then - empty, exactly as the reaper records for an unowned reap. -- **Unique Job name per request.** Each backup Job is named `backup--`, - not a deterministic `backup-`. A deterministic name would collide with a - just-finished Job still inside its `TTLSecondsAfterFinished` window (10m), and the - `AlreadyExists → 202` path would then silently produce **no** archive — the exact - window a user (or the console "立即备份" button) retries in. Unique names make every - request produce its own archive; `ErrAlreadyExists` remains only as a defensive - no-op on the ~impossible suffix collision. Ceiling (documented in `backup.go`): two - truly simultaneous taps may schedule two backup Pods — both mount the world PVC - read-only, so neither corrupts anything; single-flight-on-running is the upgrade - path if a double-tap storm ever appears. -- **One retention clock.** The entrypoint reuses the reaper's `reaperConfig` derivation - so a manual backup expires on the same schedule as an inactivity backup — one policy, - not two. The `"bk-"+hex` id scheme also matches, so manual and inactivity backups are - indistinguishable downstream. -- **Fail-safe on record failure.** If the row insert fails, the entrypoint deletes the - just-written archive so a failed backup leaves no unrecorded bytes. -- **Optional executor, honest 503.** Wired only when `FELIS_IMAGE` + `FELIS_BACKUP_PVC` - are supplied (same gate as restore); otherwise `API.Backuper` is nil and the endpoint - returns 503 `backup_unavailable`, so the authorization boundary is exercised before - the Job executor is deployable. - -## Files - -| File | Change | -|---|---| -| `internal/backupjob/jobspec.go` | **new** — `BackupJob` renderer + `BackupJobName`; weak SA, token off, hardened container, world RO / backup RW, config-Secret mount | -| `internal/backupjob/backup.go` | **new** — `Backuper` (idempotent enqueue) + `Config`/`withDefaults` | -| `internal/backupjob/k8sjobs.go` | **new** — controller-runtime `CreateBackupJob` (AlreadyExists → idempotent) | -| `internal/backupjob/jobspec_test.go` | **new** — freezes the Job's security shape incl. the deliberate config-Secret mount | -| `internal/backupjob/backup_test.go` | **new** — asserts each `Backup` call mints a unique Job name (repeat-tap must not silently no-op) | -| `cmd/felis/backup.go` | **new** — `felis backup` in-Pod entrypoint: archive + self-record + orphan-cleanup | -| `cmd/felis/run.go` | dispatch `case "backup"` + usage line | -| `internal/api/backuper.go` | **new** — the narrow `Backuper` port | -| `internal/api/handlers_backups.go` | **+`handleBackupNow`** | -| `internal/api/api.go` | `Backuper` field + `POST /servers/{name}/backup` route | -| `internal/api/handlers_backup_now_test.go` | **new** — `fakeBackuper` + handler subtests | -| `internal/api/backuper_wire_test.go` | **new** — compile-time `Backuper = (*backupjob.Backuper)(nil)` | -| `cmd/felis/api.go` | wire `backuper` under the `FELIS_IMAGE`+`FELIS_BACKUP_PVC` gate; `backupConfig` helper | -| `docs/openapi.yaml` | document the `backupNow` operation | - -## Verification - -WSL oracle (go1.26.4, authoritative for Go): - -``` -go build ./... && go vet ./... && go test ./... → ALL GREEN -``` - -The `internal/api` OpenAPI served-route contract test (`TestOpenAPIMatchesServedRoutes`) -initially failed — the new route was served but undocumented — and passes after adding -the `backupNow` operation to `docs/openapi.yaml`. `internal/backupjob` and the new -handler subtests pass. The controller-runtime `K8sJobs` binding is integration-only -(needs a live cluster) and is exercised only by the interface conformance test. - -## Self-review outcome - -- **ponytail (over-engineering):** the backup Job is a near-mirror of the restore Job, - not a shared parameterization — deliberate, because its security shape differs (it - holds DB creds) and must be asserted independently, not hidden behind a shared knob. - No speculative config; `Config.withDefaults` fills only real deployment values. -- **correctness:** the RWO stopped-gate and the self-recording atomicity were traced to - the reaper and restore before writing; the former-owner asymmetry vs restore is - justified above. diff --git a/docs/changes/2026-07-07-test-quality-mutation-audit.md b/docs/changes/2026-07-07-test-quality-mutation-audit.md deleted file mode 100644 index 440f4fd..0000000 --- a/docs/changes/2026-07-07-test-quality-mutation-audit.md +++ /dev/null @@ -1,155 +0,0 @@ -# Test-quality integrity audit — do the verifications verify FUNCTION, or just go green? - -- **Type:** audit / verification evidence (no code changed) -- **Date:** 2026-07-07 -- **Method:** mutation testing on the WSL oracle (go1.26.4) + per-function coverage backbone -- **Scope:** the load-bearing safety invariants the backfilled change ledger *claims* were tested -- **Tree state:** every mutation reverted; authoritative Windows-git working tree clean at `5a7cd5a` -- **Point verified:** each gate is broken at `HEAD` (`5a7cd5a`), not per-commit — this is the - right reading of "does the verification verify the FUNCTION": the current test pins the - current implementation. A per-commit sweep would audit history hygiene, a different question. - -## Why this audit exists - -The ledger backfill asserts, per subsystem, that a set of load-bearing safety -properties are "unit-tested". A passing suite proves the tests are GREEN; it does -not prove they would go RED if the behaviour broke. Those are different claims — -"passing ≠ verifying". This audit closes that gap the only way that earns the word -*verified*: **break the implementation, confirm the specific test turns red.** A -subagent (or a human) *reading* a test and judging it "looks thorough" reproduces -the exact error being audited (looks-right ≠ verifies), so reading was used only to -locate the gate line; the verdict is always the mutation result. - -## Result: 18 / 18 crown-jewel invariants mutation-verified - -Each row is a one-line break of the implementation, run against its own package on -the oracle. **CAUGHT = the suite went red** = the test genuinely pins the behaviour. - -| # | Invariant (claimed tested) | Impl gate mutated | Verdict | -|---|---|---|---| -| 1 | Pinned component is NEVER changed | `plan.go` pin branch → fall through | CAUGHT | -| 2 | A downgrade is NEVER proposed | `plan.go` `lv.After(current)` → `true` | CAUGHT | -| 3 | A prerelease is NEVER auto-applied | `plan.go` `!lv.IsPrerelease()` → `true` | CAUGHT | -| 4 | Apply ONLY inside the SysAdmin window | `plan.go` `Window.Contains(now)` → `true` | CAUGHT | -| 5 | `After` is strict (no equal-version churn) | `version.go` `> 0` → `>= 0` | CAUGHT | -| 6 | Clone-warned assertion refused fail-closed | `handlers_passkey.go` `if va.CloneWarning` → `if false` | CAUGHT | -| 7 | Approval CAS builds exactly once | `submit.go` `if !won` → `if false` | CAUGHT | -| 8 | An `everyone` base is not fail-open | `cfsetup.go` `if !includeHasEveryone` → `if true` | CAUGHT | -| 9 | Scoped-identity recognition actually admits | `cfsetup.go` `scoped = true` → `scoped = false` | CAUGHT | -| 10 | SSE per-principal stream cap holds | `api.go` cap-disable threshold | CAUGHT | -| 11 | OTP atomic reserve → one winner per burst | `api.go` `Sub(last) < window` → `< 0` | CAUGHT | -| 12 | Idle server is auto-stopped | `reconciler.go` `AutoStopEnabled &&` → `false &&` | CAUGHT | -| 13 | Startup/readiness timeout fires | `reconciler.go` `>= timeout` → `>= timeout + 1h` | CAUGHT | -| 14 | `/readyz` 503s when DB/K8s is down | `handlers_internal.go` dep-check `err != nil` → `false` | CAUGHT | -| 15 | NodePort fence only fires once connector serves | `tui_edge_apply.go` `connectorConnCount` parse-fail `return 0` → `1` | CAUGHT | -| 16 | CRITICAL-CVE build is NEVER admitted | `build.go` scan-gate `JobFailed`→`StatusFailed` → `StatusSucceeded` | CAUGHT | -| 17 | A user can NEVER claim a reserved system name | `naming.go` `reserved[name]` → `reserved["__nomatch__"]` | CAUGHT | -| 18 | Service token reaches ONLY the login pod | `builders.go` `Name == SystemLoginServer` → `true` | CAUGHT | - -Rows 15–18 close the gap a review of this audit surfaced: the first pass verified a -*subset* and worded the verdict as the whole set. They are the four remaining -load-bearing safety properties the ledger docs name as "unit-tested" (§ *Documented-tested -claim reconciliation* below). Each mutation produced a real `--- FAIL` on the specifically -named test — e.g. #15 reddened `TestConnectorConnCount/garbage_is_not_a_healthy_tunnel`, -#16 `TestSyncFailedDoesNotAdmitImage`, #18 `TestBuildEnvWithholdsServiceTokenFromUserServers` -— i.e. an assertion failure, not a compile break. - -Not one crown-jewel test was vacuous. The `cfsetup` fail-closed test additionally -feeds five distinct *violating* policies (bare-everyone, everyone-OR-identity, -unrecognized `ip` type, empty rule, wrong decision) and asserts each is rejected — -strong negative-path coverage, confirmed by mutations #8–#9. - -## Coverage backbone — what no oracle test executes (failure-mode B) - -Coverage triages code that no test even runs (so it cannot be verified). It does NOT -itself earn "verified" — high coverage with weak asserts is the same green-number -trap. Per-function scan of the security packages: - -**Integration-only by design (0% on the oracle — honest, NOT a gap).** The real -adapters run only against live infra; unit tests exercise the ports through fakes: -- `pgrepo.go` — all SQL, **including `QuotaAvailable`/`QuotaCheck` (the quota TOCTOU - atomic claim)**. This matches task #45's own "ENV-blocked" note: the atomic claim - is a Postgres `INSERT … WHERE`, verifiable only against a real DB. -- `k8scluster.go`, the K8s console/log-stream adapters — real Kubernetes/RCON I/O. -- `tui_edge_apply.go` `verifyConnectorServing` + the nftables fence apply — shell out to - live `cloudflared`/`nft`. **Correction from the first pass:** the doc splits this from the - *pure* `connectorConnCount` decision gate, which IS unit-tested and is now mutation-proven - (#15). The first pass wrongly folded the whole fence into "integration-only"; only the live - calls are. The gate that decides *whether* to fence is verified. - -**Genuine coverage gap (untested at the HTTP layer — "not verified").** These are -*missing* tests, not fake-passing ones: -- `handlers_users.go` — the P5 SysAdmin account-management suite: `handleCreateUser`, - `handleGetQuotas`, `handleSetQuotas`, `handleListUsers`, `handleGetUser`, - `handlePatchUser`, `handleDeleteUser`, `handleDisableUser`, `handleLinkAccount`, - `handleUnlinkAccount`, `handleListUserSessions`, `handleRevokeUserSessions`, and - `validateUsername`. All 0%; no `handlers_users*_test.go` exists. (The adjacent - `DELETE …/passkeys` remediation handler *is* tested by `TestUnbindUserPasskeys`.) -- `handleReady` — the internal-face "server is up" push (distinct from the tested - `handleReadyz`); 0%. - -These handlers are owner/operator-role-gated, so the blast radius is bounded, but -`validateUsername` is load-bearing input validation and is the highest-value target -for a follow-up test. **Recommendation:** add an `handlers_users_test.go` covering -create/quota/link + `validateUsername` negative paths. Filed as a proposed change, -not made here (this is a read-and-verify audit — no test/impl was modified). - -## Cheap tells (static pre-pass) - -- 3 `t.Skip` sites, all benign: RNG-collision reruns (a 1-in-10^6 OTP code clash), - not coverage-gating skips. -- No test file falls below 2 assertions per test function. - -## Documented-tested claim reconciliation - -To avoid the subset-verified/whole-worded trap a second time, every "unit-tested" -string in the ledger docs was enumerated (`grep -niE "unit-tested" docs/changes/*.md`) -and mapped to a verdict — verified fail-open gates get a mutation; behavioural/contract -claims are scoped, not silently dropped: - -| Doc claim | Verdict | -|---|---| -| modpack: approval CAS builds once | mutation #7 | -| modpack: **scan in front of any push** | mutation #16 | -| cloudflare: `validateFailClosed` refuses public policy | mutations #8–#9 | -| cloudflare: **conn-count fence gate** | mutation #15 | -| system-servers: **naming reservation** | mutation #17 | -| system-servers: **service-token → login pod only** | mutation #18 | -| operator: idle stop / startup+readiness timeout / `/readyz` | mutations #12 / #13 / #14 | -| auto-update: pin / no-downgrade / no-prerelease / window / strict-`After` | mutations #1–#5 | -| passkey: clone-warned assertion refused | mutation #6 | - -**Scoped, NOT individually mutation-proven** (behavioural/contract-level, not fail-open -safety gates — they rest on the green suite + the coverage backbone, and are called out here -rather than folded into the verdict): -- console-auth: content-type guard, anti-enumeration, forced-change lockdown. Anti-enumeration - is the one with security weight; the current public login door is email-OTP/passkey, and its - anti-enumeration behaviour is a candidate for a future mutation pass. -- break-glass setup: owner-auth match/non-match over the fake store. -- username-reclaim + auto-update JSON round-trip: in-memory repo contract / serialization. - -## Toolchain honesty - -Only Go runs on the oracle. The felis-limbo plugin (Java) is podman-verified against -a real Limbo jar (#65); the limbo/lobby images carry a build+boot check (#63). Those -completions were never a green-Go-tests claim and are not audited as if they were. - -## Verdict - -The commit history's verification claims are **accurate**: all 18 load-bearing *fail-open -safety gates* the ledger names as tested — spanning every subsystem, reconciled one-for-one -against the docs' "unit-tested" claims above — are mutation-proven to pin behaviour, not -merely to pass. No crown-jewel test was vacuous. The shortfalls are (a) integration seams -unrunnable on the oracle by design (honestly classified — including the live -`cloudflared`/`nft` fence-apply, whose *decision* gate is nonetheless verified); (b) one -untested cluster of admin user-management handlers — a missing test, not a false green; and -(c) a residue of behavioural/contract-level "unit-tested" claims (console anti-enumeration, -break-glass owner-auth, repo/JSON contracts) that rest on the green suite plus coverage and -are scoped above rather than individually mutation-proven — the honest boundary of this pass. - -**Method note.** The first pass mutation-verified 14 gates but worded its verdict as "every" -invariant; a review caught that 4 documented safety gates (fence, scan, naming, service-token) -were named-as-tested yet unverified, and one (the fence) was mis-classified as integration-only. -Those four are now mutation-proven (#15–#18) and the classification corrected. The lesson is -the audit's own thesis turned on itself: *reading a scope and judging it complete* reproduces -the *looks-right ≠ verifies* error — only the enumerate-and-mutate reconciliation earns the word. diff --git a/docs/changes/2026-07-08-adversarial-input-audit.md b/docs/changes/2026-07-08-adversarial-input-audit.md deleted file mode 100644 index 99080d5..0000000 --- a/docs/changes/2026-07-08-adversarial-input-audit.md +++ /dev/null @@ -1,129 +0,0 @@ -# Adversarial input-validation audit — every dangerous sink fails closed - -- **Type:** negative-path / input-validation audit (no production code change) -- **Date:** 2026-07-08 -- **Area:** `internal/api`, `internal/naming`, `internal/submit`, `internal/rcon` -- **Task:** a different question from the #82 / round-2 mutation audits. Those asked - *do the TESTS catch a gate regression?* This asks *does the CODE reject hostile - INPUT, or does bad data PASS?* — feed the real endpoints malformed, boundary, and - hostile bodies (故意加错误数据) and confirm they fail closed (4xx) rather than letting - the garbage reach a sink. - -## Method — sink-first, not fuzz-everything - -The low-hanging garbage (oversized body, unknown field, wrong content-type) is already -caught by the universal body guards, so a blanket "fuzz all ~80 handlers" would burn -effort where the answer is known. The real "can bad data PASS?" risk lives at the -**sinks** — the few places a request string is concatenated into an RCON command, used -as a K8s object name, joined into a filesystem/archive path, put in a SQL query, or -accepted as an enum/quantity **without a validator in front**. So the audit traces each -dangerous sink class from its handler entry to the sink, and for the crown-jewel class -(text → RCON) **mutation-verifies** the guard is non-vacuously pinned: loosen the guard -in source, run the package tests, confirm the specifically-named negative test reddens -(`--- FAIL: `), then revert. Oracle: WSL Fedora-44, go1.26.4. - -## The universal belt (caught before any sink) - -`decodeJSON` (`internal/api/util.go`) wraps every body in -`http.MaxBytesReader(w, r.Body, 1<<20)` (1 MiB cap), sets `DisallowUnknownFields()`, and -rejects trailing data after the first JSON value — all → `400 bad_request`. -`requireJSONContentType` returns `415 unsupported_media_type` on credential writes -(a CSRF belt). So oversized, unknown-field, multi-document, and wrong-type bodies never -reach a handler body at all. - -## The dangerous sinks — each traced fail-closed - -### 1. Text → RCON (the lead). Two vectors, both fenced. - -**(a) Structured commands** — `handlers_access.go` (LuckPerms permission/group, -whitelist/ban/kick). Every operand is validated against an anchored allow-list charset -*before* it is concatenated: `mcNameRe = ^[A-Za-z0-9_]{1,16}$` (player), -`lpNodeRe = ^[A-Za-z0-9_.*-]{1,64}$` (node), `lpCtxRe = ^[A-Za-z0-9_-]{1,48}$` -(world/group). No space, separator, or control character can appear in a validated -operand, and **there is no free-text field anywhere** — a ban/kick deliberately carries -no reason string (that would be the one splice vector). Go's `$` is `\z` (absolute end, -not `\Z`), so even a single trailing `\n` is rejected. -*Mutation-verified:* loosening `mcNameRe` to admit a space -(`^[A-Za-z0-9_ ]{1,16}$`) reddens -`TestAccessInjectionRejected/{whitelist,ban,kick,permission}_player_space` — the guard -is real, not vacuous. - -**(b) Free-text passthrough** — `handlers_console.go` `handleCommand`, POST -`/servers/{name}/command`. This is the *one deliberate* free-text → RCON vector, and it -is **owner/admin-gated** (403 for a stranger, 404 for an unknown server). Its input -fence: trim + strip a single leading `/`, reject empty, cap at 1024 bytes, and -`strings.IndexFunc(command, func(c rune) bool { return c < 0x20 }) >= 0 → 400` — every -C0 control (incl. `\n`) is rejected so one request cannot splice a second command. -*Mutation-verified:* disabling the scan (`c < 0x20` → `c < 0x00`) reddens -`TestConsoleCommand/control_character_(newline)_->_400,_no_RCON_call`, whose input is -literally `{"command":"say hi\nop attacker"}` and whose assertion is 400 **and** -`console.calls == 0`. (Severity note: because this vector is owner-gated by design, the -scan is an audit-integrity measure — one request = one command — not a privilege -boundary; a splice on your *own* server escalates nothing, since the owner may already -run any RCON command. The fence exists regardless.) - -### 2. Break-glass / internal-face (the newest code — scrutinised specifically) - -This surface runs under a "service-token-authed / local-root, inputs trusted" posture, -the classic place a field-level guard gets skipped. Traced end to end: - -- **Server name** — every internal handler (`handleReady`, `handleJoinEvent`, - `handleInternalWake`, `handleInternalClaim`, `handleInternalMenuStatus`) validates the - path name with `naming.ValidateServerName` before use. -- **`mc_uuid`** (join/wake/claim bodies) — checked non-empty, then flows *only* to - DB-parameterized calls (`RecordJoin`, `UserByMCUUID`, `UUIDInAllowlist`). The - "allowlist" is a **DB table**, not a live RCON `whitelist add` — there is no - `mc_uuid` → RCON path. -- **Break-glass "OP-create"** (`performAddOperator` → `provisionOperator` → - `InsertOperator(ctx, id, username, email)`) is a **parameterized DB INSERT** creating a - *panel staff account*, **not** a Minecraft `op` RCON command. The hypothesised - name → RCON `op` sink was checked and **does not exist** in this shape; the username is - `TrimSpace`d and reaches only `$N`-parameterized SQL, from a local-root caller. -- **`os_user`** (internal backup attribution) — `TrimSpace`d, sets only the audit actor - (a DB row); shown non-vacuous in the round-2 backup/restore audit (`b7b4a3b`). - -### 3. SQL injection — dismissed. - -`pgrepo.go` uses uniform `$1/$2/$3` parameterization throughout -(`QueryRowContext`/`ExecContext(ctx, q, args…)`); no request string is `Sprintf`'d into a -query. - -### 4. Path traversal (submit) — dismissed. - -`internal/submit` validates the submission id (rejects `..`, path separators, uppercase, -space, empty), and the on-disk blob name is a **fixed** constant (`contextBlobName`) — no -attacker-supplied filename is ever joined. The hostile-id matrix -`{"../evil","sub/../../etc","SUB-UPPER","has space","","a/b"}` is test-pinned in both the -local and S3 backends. The one free-form field a submission carries (`DisplayName`) is -charset-constrained by `displayNameRE` and rejects control chars -(`submit_test.go` "control chars" case). - -### 5. K8s object names — validated at every cluster write. - -`ValidateServerName` / `ValidateSystemServerName` (`^[a-z0-9-]{3,32}$`, no -leading/trailing dash, reserved-name set) and `ValidateHostname` -(`dnsLabelRE`, single label under the configured root domain) gate every create/patch. -`cluster.go` documents the invariant: "there is no free-form YAML path — every field is a -typed, validated value," so no raw CRD field can be smuggled through a create/patch body. - -### 6. Numeric / enum — fail-closed. - -`resolveResources` routes **every** quantity (memory, `resources.cpu/memory`, requests) -through `parsePositiveQuantity`, which rejects `q.Sign() <= 0` (negative *and* zero) with -a field-named 400, plus a request>limit guard; `parseStorageSize` carries the same guard. -Enums are closed sets: `parseAutostartPolicy` (ownerOnly/public/allowlist), the -access-action switch, and the image `ImageAdmitted` allow-list (no free image string). - -## Verdict - -**Bad data does not pass.** Every dangerous sink is fail-closed — including the -break-glass / internal-face surface, where the hypothesised `mc_uuid`/OP-create → RCON -paths were traced and found not to exist (parameterized DB, not RCON). Both text → RCON -vectors — structured (`handlers_access`) and free-text (`handleCommand`) — are -mutation-pinned by their named negative tests. No gap was found and no production code -changed; the honest result of "故意加错误数据" is that the input surface rejects it. - -Scope is deliberately bounded to the dangerous sinks and the newest (break-glass) code, -not an exhaustive fuzz of all ~80 handlers — the claim proven is "every place user input -reaches a dangerous sink validates before the sink," by trace plus two mutations, not -"every handler was fuzzed." diff --git a/docs/changes/2026-07-08-b4-close-s3-deferred.md b/docs/changes/2026-07-08-b4-close-s3-deferred.md deleted file mode 100644 index 316ccce..0000000 --- a/docs/changes/2026-07-08-b4-close-s3-deferred.md +++ /dev/null @@ -1,57 +0,0 @@ -# §B4 phase close — S3 archive backend deferred by design (decision record) - -- **Type:** decision / scope record (no code changed) -- **Date:** 2026-07-08 -- **Area:** `internal/config`, `internal/backup` — the archive-store backend selection -- **Task:** #31 Phase B4 break-glass ops — **closes the phase**, superseding the phase-2b - note in [break-glass-backup-peer](2026-07-07-break-glass-backup-peer.md) ("S3 remains - open in B4; this does not close the phase"). - -## Decision - -The three operational B4 break-glass ops are built, wired into the recovery console menu, -and oracle-verified: - -| Op | Menu enum | Commit | -|---|---|---| -| Provision/reset Owner | `bgProvisionOwner` | (Phase B1 lineage) | -| Add Operator ("OP create") | `bgAddOperator` | recovery-console menu | -| Halt a running server | `bgHaltServer` | `c2ee21a` | -| Back up a world now ("Sync") | `bgSyncBackup` | `7a7c0d5` / `f2fc57c` / `fc748d3` | - -The fourth B4 line item — **S3 archive backend (`tarS3`)** — is **deferred by design**, not -left as a silent gap. It is closed as a documented deferral and **#31 is done**. - -## Why deferring is safe (not a loose end) - -- **Fail-closed at config load, frozen by a test.** `config.Load` rejects - `store = "tarS3"` (and `volumeSnapshot`, `longhorn`) with an error that points the - operator at the `tarLocal` remediation. `TestLoadRejectsUnimplementedArchiveStore` - freezes exactly this: a config naming an unimplemented backend fails at load, so - felis-api can never boot green while the reaper CronJob fails every run and restore - silently 503s. tarS3 cannot be selected into a broken state. -- **Peer to two other deferred backends.** `tarS3` sits beside `volumeSnapshot` and - `longhorn` as recognized-but-unimplemented store names. The spec's own phasing is - tarLocal-first ("起步 `tarLocal` … 要异地/跨集群 → `tarS3`"): the baseline single-node - path is `tarLocal` (tar → backup PVC), which is implemented, tested, and the backend - every built backup/restore path uses today. -- **Offsite/cross-cluster DR is the only capability gap**, and it is opt-in future work, - not a correctness hole in the shipped baseline. - -## The build path, when offsite DR is wanted - -`minio-go/v7` is already vendored (the modpack upload lane's `internal/submit/s3store.go`), -so tarS3 adds no dependency. A future build is bounded: - -1. `internal/backup/tars3.go` — a `WorldArchiver` reusing the existing package-level - `writeTarGz`/`readTarGz`/`pruneToManifest`, streaming the tar to an object via - `PutObject` (size −1, multipart) and reading it back via `GetObject`, mirroring - `s3store.go`'s fakeable-client testability. -2. Store-selection factory in the backup/reaper entrypoint (`store = "tarS3"` → construct - the minio-backed archiver) + remove tarS3 from the config fail-closed list (update - `TestLoadRejectsUnimplementedArchiveStore` to keep only volumeSnapshot/longhorn). -3. Inject the S3 Secret into the backup/reaper Job Pods (integration wiring, like the - Kaniko S3-context credential path). - -The live-S3 upload + Secret-into-Pod would be integration-only verified, exactly as the -modpack S3 backend and the pgrepo SQL are. diff --git a/docs/changes/2026-07-08-backup-restore-mutation-audit.md b/docs/changes/2026-07-08-backup-restore-mutation-audit.md deleted file mode 100644 index b518749..0000000 --- a/docs/changes/2026-07-08-backup-restore-mutation-audit.md +++ /dev/null @@ -1,93 +0,0 @@ -# Backup/restore data-safety mutation audit (round 2) — 7 gates pinned, 1 gap closed - -- **Type:** test-quality audit + one test added (code change: `internal/api/handlers_backups_test.go`) -- **Date:** 2026-07-08 -- **Area:** `internal/api` (`handlers_backups.go`), `internal/backupjob` -- **Task:** continues the #82 test-quality integrity audit onto the on-demand - backup/restore surface, which postdates the 18-gate round-1 audit - ([mutation audit `4626ab5`](2026-07-07-test-quality-mutation-audit.md)). The - backup/restore endpoints (`7a7c0d5` / `f2fc57c` / `fc748d3`) were not in that pass. - -## Method - -Same as round 1: apply a one-line mutation to a fail-open gate in the source, run the -package tests, confirm the **specifically-named** test reddens with an *assertion* -failure (`--- FAIL: `), then revert. A build break (`declared and not used`, -`undefined`) is not a valid verdict, so mutations are operator-flips that keep every -operand referenced (`!=`→`==`, drop a `!`, a literal→`true`, or `if false && ` to -disable a gate without orphaning its variables). Airtightness: each mutation is re-run -with `-run` scoped to the intended subtest and `grep -- "--- FAIL: "`, so a -reddening sibling can't be mistaken for the gate under test. Oracle: WSL Fedora-44, -go1.26.4. - -## Fail-open gates mutation-verified (all CAUGHT at the named subtest) - -| # | Gate (file:line) | What it guards | Mutation | Subtest that reddened | -|---|---|---|---|---| -| A | `handlers_backups.go:282` enqueueBackup stopped-gate | RWO double-mount / torn archive while the world is up | `!=`→`==` | `TestBackupNow/starting_server_->_409_not_stopped` | -| B | `:151` restore stopped-gate | restore Job can't mount a live world's RWO PVC | `!=`→`==` | `TestRestoreBackup/starting_server_->_409_not_stopped` | -| C | `:136` restore former-owner match | a fresh claimant resurrecting the previous owner's world | `!=`→`==` | `TestRestoreBackup/current_owner_who_is_not_former_owner_->_403` | -| D | `:116` restore cross-server guard | restoring server A's backup onto server B | `!=`→`==` | `TestRestoreBackup/restore_by_backup_id_cross-server_->_403` | -| E | `:214` backup owner-or-admin authz | a stranger backing up someone else's world | drop `!` | `TestBackupNow/non-owner_->_403,_no_backup` | -| F | `:26` list cross-user scope | a user seeing other tenants' backups | `p.IsAdmin()`→`true` | `TestListBackups/user_sees_only_own_former-owned_present_backups` | - -## The gap this audit found — and closed - -**Restore's owner-or-admin gate (`handlers_backups.go:85`) was not pinned by any test.** -It is the twin of gate E, but the two are *not* symmetric. Disabling it -(`if false && !a.isOwnerOrAdmin(p, rec)`, which keeps `rec` referenced so the package -still builds) reddened **nothing** — `go test ./internal/api/` stayed `ok`. The same -disable applied to backup's L214 (gate E) reddened `non-owner` immediately, proving the -technique valid and the asymmetry real. - -Root cause: the former-owner gate at L136 backstops every non-owner case the suite -exercised (a stranger and a wrong-backup current owner both fail L136 *and* L85, so -L136's 403 masks a broken L85). The one case only L85 catches went untested: a -**superseded former owner** — a user who owned a server, took this backup -(`FormerOwner=them`), then released it to a *new* owner. They still pass L136 (they *are* -the former owner) but must be stopped by L85, or they could roll the new owner's live -server back onto their old world (cross-tenant clobber). The handler comment names this -the "must re-claim first" rule (`handlers_backups.go:77-79`). - -**Fix (code):** added `TestRestoreBackup/former owner after release -> 403, no restore`, -the mirror of the existing L136 test. Verified both directions: green on the clean tree, -and it is the sole subtest that reddens when L85 is disabled — so it now pins the owner -gate specifically, not L136. No production code changed; `handlers_backups.go` is a pure -test addition away from where it was. - -## Enumeration — covered vs. scoped (so "the gates" means all of them) - -- **Fail-open data-safety gates — all pinned:** A–F above, plus L85 (now closed). 7/7. -- **Accountability, not fail-open (verified non-vacuous):** the `os_user` attribution at - `handlers_backups.go:254` — a supplied operator name overrides the default `break-glass` - audit actor. Mutating `u != ""`→`u == ""` reddens - `TestInternalBackup/os_user_body_attributes_the_audit_to_the_operator`, so the - attribution test isn't vacuous. A failure here degrades the audit actor; it grants no - bypass, so it is out of the fail-open bucket. -- **Contract/behavioral (tested, out of mutation scope):** nil `Backuper`/`Restorer` → 503; - a failed backup/restore → 500 **not** audited; `backup_ref` never serialized to the wire; - the internal-face actor defaults to `break-glass`. Each has a direct test; none is a - fail-open safety gate. -- **`internal/backupjob` (glanced, not mutated):** orchestration only — each backup gets a - unique Job name (`BackupJobName` + random suffix) so a repeat "立即备份" tap can't collide - with a just-finished Job still inside its TTL; `ErrAlreadyExists` is a defensive no-op; - `Backup` returns once the Job is created (the async 202 is honest). Unit-tested against a - fake `Jobs`; the controller-runtime `k8sjobs.go` and the `jobspec.go` Pod shape (weak SA - with its token un-mounted, config Secret mounted for the self-recorded row, read-only - world mount) are integration-verified per the package doc — not fail-open API gates. - -## Coverage nuance (documented, not a gate failure) - -The stopped-gate is `if info.Ready || info.DesiredState != DesiredStopped`. The `!=`→`==` -mutation pins the `DesiredState` operand (both A and B reddened), but `info.Ready` is not -*independently* pinned: no test sets `Ready=true` together with `DesiredState=Stopped` — -the stopping-but-still-up race. Low risk because in practice `Ready` drops as -`DesiredState` leaves `Stopped`, but the belt-and-suspenders `Ready` operand rides on -coverage of the operand beside it rather than its own case. - -## Verdict - -The backup/restore data-safety surface is a coherent unit, and this closes it: **7/7 -fail-open gates pinned** (6 pre-existing, 1 added this round), one accountability gate -shown non-vacuous, one coverage edge documented. Not extended to every handler — that -would be an unbounded "continue the audit." diff --git a/docs/changes/2026-07-12-felis-nano-hasjoined-resolver.md b/docs/changes/2026-07-12-felis-nano-hasjoined-resolver.md deleted file mode 100644 index 5c18098..0000000 --- a/docs/changes/2026-07-12-felis-nano-hasjoined-resolver.md +++ /dev/null @@ -1,110 +0,0 @@ -# Felis-nano — the federating `hasJoined` multiplexer (step 1: the Go resolver) - -- **Type:** feature (new endpoint) — the verifiable "brain" of Felis-nano -- **Date:** 2026-07-12 -- **Area:** `internal/api` (`handlers_hasjoined.go` + test), `internal/api/api.go` - (route + `AuthSources` field), `docs/openapi.yaml`, `go.mod` -- **Task:** Felis-nano provides a MultiLogin-like capability — one Velocity proxy that - accepts logins verified by **several** Yggdrasil auth servers at once (Mojang + N - third-party roots), 正版优先 (Mojang-first). This change builds **step 1**: the Go - `hasJoined` multiplexer that does the federating verification. It is the only part of - the plan that produces immediate verifiable hard evidence (a unit-tested HTTP endpoint); - the two delivery shells that point Velocity at it (a JVM `-Dmojang.sessionserver` flag, - and a thin reflection-hook plugin for third-party servers) are later steps. - -## What Velocity asks for, and what this answers - -On a Minecraft login Velocity's authlib computes the `serverId` hash and issues -`GET /session/minecraft/hasJoined?username=&serverId=[&ip=]` against -whatever URL its `mojang.sessionserver` system property names. A 200 with a game profile -means "verified"; a 204 means "not verified" and authlib rejects the login. Vanilla points -this at Mojang alone. Felis-nano points it **here**, and this endpoint fans the same query -out to the configured Yggdrasil roots **in priority order**, returning the first source -that validates. Each upstream Yggdrasil runs its own `serverId`-hash check — the -multiplexer only relays, it computes no hashes. - -## The one non-negotiable transform — per-source UUID namespacing - -A third-party Yggdrasil's UUIDs are **self-asserted**: nothing stops a malicious source -from answering with a *genuine Mojang player's* UUID. If that UUID were emitted as-is, the -third-party could impersonate any Mojang player with full UUID fidelity — and the reclaim/ -blacklist layer could never catch it, because its whole invariant is "the genuine Mojang -player has a **different** UUID from any squatter." That invariant would simply be false. - -So the resolver rewrites every non-identity source's profile into a per-source namespace -**before it leaves the resolver** — the single entry point every login crosses: - -``` -canonical = UUIDv3(felisAuthNS, tag + ":" + nativeID) // third-party -canonical = the source's UUID verbatim // Mojang (Identity: true) -``` - -MD5 (UUIDv3) preimage resistance means no third-party can mint a value inside Mojang's -UUID space; the per-`tag` prefix means two sources can't collide onto one identity. Every -downstream key — `account_links`, `username_blacklist`, owner checks — then sees exactly -one canonical UUID per real identity, so the reclaim invariant is true **by construction**, -not by assumption. - -## Fail-closed details that bite if wrong - -- **Bar gate at the chokepoint.** The canonical UUID is checked against - `Repo.IsUsernameBlacklisted` *before* the profile is returned, so a reclaimed squatter - stays out even on a consumer that has no limbo plugin. Keyed on the **dashed** canonical - (`.String()`) — the exact form `Repo.ReclaimUsername` stores. A DB error there fails - closed (non-200 → authlib rejects), matching the existing `handleCheckBlacklist` pattern. -- **Emit undashed.** authlib's `GameProfile` expects the 32-hex undashed `id` - (`hex.EncodeToString(u[:])`); the DB/reclaim/blacklist keys are dashed. The resolver - **checks** on the dashed string and **emits** the undashed one. Mixing the two forms is a - silent gate miss — pinned by the tests below. -- **`properties` relayed verbatim** (`[]json.RawMessage`) so a source's signed textures - survive the multiplexer untouched. -- **Inert by default.** `AuthSources` is nil until `cmd/felis` wires configured sources, - so the endpoint 204s every login until deliberately configured — it ships off. -- **Public internal-face route.** authlib sends no service token, so the route is mounted - `Public: true` on the internal face (like `/healthz`); no third face is introduced. The - OpenAPI parity test enforces `x-felis-face: [internal]` + `x-felis-tier: public`. - -## Files - -| File | Change | -|---|---| -| `internal/api/handlers_hasjoined.go` | **new** — `handleHasJoined` + `resolveHasJoined` + `AuthSource`/`sessionProfile` types + `felisAuthNS` | -| `internal/api/handlers_hasjoined_test.go` | **new** — `TestHasJoined`, 7 subtests over `httptest` fake Yggdrasil roots | -| `internal/api/api.go` | `AuthSources []AuthSource` field (nil = inert) + `GET /session/minecraft/hasJoined` `Public` internal route | -| `docs/openapi.yaml` | `/session/minecraft/hasJoined` path — `x-felis-face: [internal]`, `x-felis-tier: public`, `security: []` | -| `go.mod` | promote `github.com/google/uuid` indirect→direct (first direct importer) | - -## Verification evidence - -Oracle: WSL Fedora-44, go1.26.4. `go build ./...` → `BUILD-OK`. Full `internal/api` -package green (`ok felis.lolicon.best/internal/api`), `go vet ./internal/api/` clean. The -full package (not a `-run` filter) was run because this change edits two shared surfaces — -the `API` struct and the `internalAPIRoutes()` table — where a route that isn't under -`/api/v1/` is exactly what a route-table-driven invariant test would trip; nothing -reddened. - -`TestHasJoined` — 7 subtests, all PASS: - -1. `mojang identity passthrough` — Mojang UUID unchanged, `properties` relayed. -2. **`thirdparty UUID rewritten, never emitted as-is`** — the security invariant: an evil - source returns real Notch's Mojang UUID; the resolver emits neither that UUID nor any - Mojang-space value, but the deterministic `UUIDv3(felisAuthNS, "littleskin:"+id)`. -3. `mojang priority wins over thirdparty` — Mojang-first ordering. -4. `fallthrough to thirdparty when mojang 204s` — priority scan continues past a 204. -5. `no source validates -> 204`. -6. `barred canonical UUID -> 204` — the reused reclaim bar gate holds at the resolver. -7. `missing username -> 204` — no source touched on a malformed query. - -`TestOpenAPIMatchesServedRoutes` PASS — the new route's served facets match its -`docs/openapi.yaml` entry in both directions. - -## Deferred (not in this change) - -- **Source configuration** (step 2): a `tag`/`type`/`url`/`priority` schema and - `cmd/felis` wiring that populates `AuthSources`. Until then the endpoint is inert. -- **Delivery shells** (step 3): Shell 1 = the `-Dmojang.sessionserver` JVM flag on - Felis-managed proxies; Shell 2 = the thin reflection-hook Velocity plugin for - third-party servers, plus a Velocity verification runbook. The user has a real - server to test Shell 2 against. -- **Name-match / textures-signature enforcement** — deliberately out of scope: identity - is `tag:nativeID`, not the name, and textures are the cosmetic bucket relayed verbatim. diff --git a/docs/changes/INDEX.md b/docs/changes/INDEX.md deleted file mode 100644 index f84f5c1..0000000 --- a/docs/changes/INDEX.md +++ /dev/null @@ -1,255 +0,0 @@ -# Felis change ledger - -The index of every functional change to Felis — what it did and which commit records -it. This is the durable, in-repo map that `git log` alone doesn't give: it links -substantial changes to their detail docs and flags work that is built and verified but -not yet committed. - -## Convention - -- **Every functional change** (a feature addition, a behaviour change, a bug fix) gets: - 1. a dated detail doc in this directory — `docs/changes/YYYY-MM-DD-.md`, covering - _what it did, why, the files touched, and the verification evidence_; and - 2. a row in the ledger below, carrying its **commit record** (the short SHA). -- A change that is **built and verified but not yet committed** (e.g. while PGP signing - is locked) sits in **Pending** with `commit: pending`, and moves into the ledger with - its real SHA once committed. -- Pure-cosmetic or non-functional commits (docs, style) still appear in the ledger table - for completeness, but do not require a dedicated detail doc. -- The ledger table is generated losslessly from git history and can be regenerated: - ``` - git log --reverse --pretty=format:'| %h | %ad | %s |' --date=short - ``` - -## Pending (built + verified, not yet committed) - -_None._ - -## Detail docs - -Depth docs for substantial changes, keyed to the commit(s) they cover. The committed -ledger table below stays a lossless mirror of `git log` (so it can be regenerated); this -section is where a row's detail doc, when it has one, is found. Most rows — panel/UI, -docs, chore, style — have no detail doc by convention and are recorded by their table row -alone. Entries marked *(backfill)* were reconstructed retroactively on 2026-07-07 from git -history to close the ledger's detail-doc axis for the pre-convention functional commits; -each carries a backfill note stating it was not independently re-verified. Frontend/`panel` -commits are the collaborator's UI work and are not given detail docs here. - -| Detail doc | Commit(s) | Scope | -|---|---|---| -| [foundational-subsystems](2026-06-26-foundational-subsystems.md) *(backfill)* | `7fbebfe` `708cdfc` `43ab921` `78b8cf6` `d39605e` `b508fcc` `47fcd90` `93f143f` | initial import: CRD, core libs, backup, operator, submit, api, platform, plugins | -| [modpack-submission-lane](2026-06-26-modpack-submission-lane.md) *(backfill)* | `d39605e` `598f3d3` | §8 modpack build/approval pipeline + local/S3 backends | -| [deploy-bootstrap-installer](2026-06-27-deploy-bootstrap-installer.md) *(backfill)* | `58fa4b0` `94a3b7b` `deaa2f8` `318a724` `e5f1682` `28c3eee` `c14ed17` `d9e866f` `b84debf` | one-line bootstrap installer + demo bring-up | -| [console-auth-passwordless](2026-06-27-console-auth-passwordless.md) *(backfill)* | `af14f02` `0c1cc59` `3b43f05` `c20b12c` | local-password login → passwordless migration + residue sweep | -| [felis-cli-break-glass-setup](2026-06-27-felis-cli-break-glass-setup.md) *(backfill)* | `e108a37` `2d0bbb0` `a94b001` `eb5875a` `f5d00f3` `9c46632` `7d91373` | break-glass recovery console + first-run setup + apply/migrate | -| [cloudflare-tunnel-access-edge](2026-06-27-cloudflare-tunnel-access-edge.md) *(backfill)* | `53a7664` `ba13839` `a531f5e` `2810fe8` `7d3be64` `346ec68` `e058a64` | §14 Tunnel + fail-closed Access edge + NodePort fence | -| [player-onboarding-b2](2026-06-27-player-onboarding-b2.md) *(backfill)* | `dbe34a1` `1f8b9bb` `116595f` `fe2ece0` `55592ed` `6c3999a` `879b177` | §B2 email-OTP, account-link, QR, Bind-Code + OTP throttle | -| [username-reclaim-b3](2026-06-27-username-reclaim-b3.md) *(backfill)* | `a29571d` `fdb6efb` | §B3 Mojang-priority reclaim + account migration | -| [felis-metrics](2026-06-30-felis-metrics.md) *(backfill)* | `75642d9` `2a93a9e` `79eae7f` `8ac5e64` | §23 felis_* Prometheus collectors | -| [felis-api-hardening](2026-06-30-felis-api-hardening.md) *(backfill)* | `7a51c1d` `164ac44` `c6c0772` `3c1d647` `d6e3189` `8f41a00` `6368ab1` `2a4a81b` `9873904` | audit #1–#3 + robustness fixes | -| [passkey-enrollment](2026-07-01-passkey-enrollment.md) *(backfill)* | `f2c916d` `742f15f` `0261204` `fce0fce` `7278cd7` `cdbb5ab` `9953275` `20e31fb` `54bc6ef` | WebAuthn enrollment + hardening a–e | -| [passkey-login](2026-07-01-passkey-login.md) *(backfill)* | `e035142` `ec468ba` `0dbd557` `9e1df12` `4f59d51` `a63f49d` | WebAuthn assertion/discoverable login + UA-guard | -| [auto-update-subsystem](2026-07-01-auto-update-subsystem.md) *(backfill)* | `c01f133` `3673af6` `7464fa7` `96b3cc9` `7d27640` `7db57b9` | update decision core + sources + gatherer + window API (report-only) | -| [system-servers-login-limbo-lobby](2026-07-02-system-servers-login-limbo-lobby.md) *(backfill)* | `9bed51b` `9ef817f` `159107b` `dc23cb5` `3fdb3d0` `f554d52` `241fe21` `c7315e4` | always-on login-limbo + lobby auth gate | -| [operator-idle-quota-readiness](2026-07-05-operator-idle-quota-readiness.md) *(backfill)* | `91bfa27` `e574749` `7f7e459` `7becb38` | idle auto-stop, quotas, timeouts, /readyz | -| [break-glass-halt](2026-07-05-break-glass-halt.md) | `c2ee21a` | §B4 break-glass halt-a-server op | -| [felis-migrate-command](2026-07-05-felis-migrate-command.md) | `c1aa38b` | §B3 `/felis migrate` account migration | -| [on-demand-world-backup](2026-07-07-on-demand-world-backup.md) | `7a7c0d5` | §B4 Sync phase 1 — external backup endpoint + Job | -| [internal-backup-endpoint](2026-07-07-internal-backup-endpoint.md) | `f2fc57c` | §B4 Sync phase 2a — internal-face backup endpoint | -| [internal-api-clusterip-service](2026-07-07-internal-api-clusterip-service.md) | `2ba9948` | felis-api internal-face ClusterIP Service | -| [break-glass-backup-peer](2026-07-07-break-glass-backup-peer.md) | `fc748d3` | §B4 Sync phase 2b — console backup peer | -| [adversarial-input-audit](2026-07-08-adversarial-input-audit.md) | `667c6d3` | adversarial input-validation audit — sink-first negative-path, two text→RCON guards mutation-pinned, no code change | -| [felis-nano-hasjoined-resolver](2026-07-12-felis-nano-hasjoined-resolver.md) | `ff550c4` | Felis-nano §B3 — federating hasJoined multiplexer + per-source UUID namespacing | - -## Committed change ledger - -Oldest first (project build order). Commit = short SHA on `main`. Frontend/`panel` -commits are the collaborator's UI work; backend (Go/Java/K8s) is tracked here as the -primary record. - -| Commit | Date | Change | -|---|---|---| -| 5b7b38d | 2026-06-26 | chore: add Go module manifest and ignore rules | -| 5a30aa5 | 2026-06-26 | docs: add OpenAPI 3.1 served-route contract | -| 7fbebfe | 2026-06-26 | feat(apis): add MinecraftServer CRD types (v1alpha1) | -| 708cdfc | 2026-06-26 | feat(core): add naming, RCON, store, config, and image-build libraries | -| 43ab921 | 2026-06-26 | feat(backup): add backup, restore, and reaper subsystems | -| 78b8cf6 | 2026-06-26 | feat(operator): add MinecraftServer controller and reconcilers | -| d39605e | 2026-06-26 | feat(submit): add user modpack build and approval pipeline | -| b508fcc | 2026-06-26 | feat(api): add felis-api service with permissions, modpack lane, and fleet read | -| 47fcd90 | 2026-06-26 | feat(platform): add node orchestration and the felis entrypoint | -| eee00c2 | 2026-06-26 | feat(panel): add three-sided web console (User, Admin, SysAdmin) | -| 93f143f | 2026-06-26 | feat(plugins): add Velocity proxy and Fabric/Forge/NeoForge/Paper integration mods | -| ce0ba76 | 2026-06-27 | chore(api): add kubebuilder object-generation markers to v1alpha1 | -| 7d91373 | 2026-06-27 | fix(migrate): honor -config flag placed after the up verb | -| 99de43f | 2026-06-27 | chore: ignore plugin build artifacts and editor config | -| 58fa4b0 | 2026-06-27 | feat(deploy): add one-line bootstrap installer and container image | -| af14f02 | 2026-06-27 | feat(api): local-password authentication backend | -| e108a37 | 2026-06-27 | feat(cli): break-glass emergency console TUI | -| 885c4a9 | 2026-06-27 | feat(panel): local-password login and forced password change | -| 2d0bbb0 | 2026-06-27 | feat(cli): attribute break-glass recovery to the SysAdmin who runs it | -| dbe34a1 | 2026-06-27 | feat(api): add player email OTP verification (spec §B2 onboarding) | -| 1f8b9bb | 2026-06-27 | feat(api): record account-link auth source (mojang/thirdparty) | -| a29571d | 2026-06-27 | feat(api): reclaim squatted usernames for Mojang-priority players (spec §B3) | -| 53a7664 | 2026-06-27 | feat(cfsetup): recommended Cloudflare Tunnel + Access edge setup | -| ba13839 | 2026-06-27 | feat(breakglass): optional Cloudflare Tunnel + Access setup in the TUI | -| a5a6482 | 2026-06-27 | feat: dev mock | -| dd2fc6f | 2026-06-27 | feat(panel): i18n | -| 7b458da | 2026-06-27 | feat(panel): light/dark theme | -| 51c9eab | 2026-06-27 | docs: CONTRIBUTOR.md | -| 74e7e6e | 2026-06-27 | docs: CONTRIBUTING.md | -| b81b334 | 2026-06-27 | Merge branch 'main' of https://github.com/MliroLirrorsIngenuity/Felis | -| 9d13787 | 2026-06-27 | refactor(panel): dashboard | -| f3521e3 | 2026-06-27 | refactor(panel): uniform margins | -| 18b4be0 | 2026-06-27 | refactor(panel): uniform title icon styles | -| 665841b | 2026-06-27 | fix(panel): remove internal spec references from user-facing text | -| cf88bcc | 2026-06-27 | fix(panel): extract hardcoded security note into i18n keys | -| 061482d | 2026-06-27 | fix(panel): extract hardcoded Chinese text to i18n keys | -| cfe126f | 2026-06-27 | feat(panel): add RCON command input to server console | -| ae91133 | 2026-06-27 | style(panel): refine button styles with shadow, active scale, toned-down colors | -| e189265 | 2026-06-27 | refactor(panel): compact server card layout, denser grid | -| f4df3e2 | 2026-06-27 | feat(panel): pagination for server lists | -| b512a18 | 2026-06-27 | refactor(panel): adjust margins | -| a49443c | 2026-06-27 | refactor(panel): simplify sidebar | -| 86f2ae4 | 2026-06-28 | feat(panel): sidebar foot shows current account + sign-out; reorder Account cards | -| f5d00f3 | 2026-06-28 | feat(cli): implement felis apply command for direct CRD creation | -| 832b200 | 2026-06-28 | fix(panel): reactive system theme detection | -| c9cd9dc | 2026-06-28 | style(panel): unify dialog animation to fade and scale from center | -| 94a3b7b | 2026-06-28 | fix(deploy): harden bootstrap for RHEL-family Linux | -| 9c46632 | 2026-06-28 | feat(cli): add felis setup first-run console with reclaim protection and cfsetup idempotency | -| 5450c26 | 2026-06-28 | chore: normalize line endings and apply formatting | -| deaa2f8 | 2026-06-28 | feat(deploy): add zypper support for openSUSE/SLES | -| 318a724 | 2026-06-28 | feat(deploy): add pacman support for Arch Linux | -| e5f1682 | 2026-06-29 | refactor(deploy)!: TUI | -| 28c3eee | 2026-06-30 | refactor(deploy): improved TUI walkthrough | -| 116595f | 2026-06-30 | feat(api): add QR scan-login completion poll on the internal face | -| a94b001 | 2026-06-30 | feat(deploy): add break-glass Operator account provisioning | -| 346ec68 | 2026-06-30 | refactor(deploy): improved cloudflare walkthrough | -| 563041a | 2026-06-30 | feat(panel): add fail-closed role-switcher view-mode logic | -| 50b8487 | 2026-06-30 | feat(panel): wire role-switcher into the app shell | -| 75642d9 | 2026-06-30 | feat(metrics): add named felis_* Prometheus collectors | -| 2a93a9e | 2026-06-30 | feat(metrics): record felis_image_build_failures_total on failed builds | -| 79eae7f | 2026-06-30 | feat(metrics): publish felis_servers_total from a fleet snapshot | -| 8ac5e64 | 2026-06-30 | feat(metrics): observe felis_start_duration_seconds across the start lifecycle | -| 676407d | 2026-06-30 | docs(diagrams): align §28 sequence diagrams with implemented routes | -| ac02c69 | 2026-06-30 | docs(troubleshooting): add operator failure-mode checklist | -| eb5875a | 2026-06-30 | feat(felis): add Operator break-glass op behind an operation menu | -| c14ed17 | 2026-06-30 | fix(docker): keep embedded panel/ and deploy/ in the image build context | -| 2a4a81b | 2026-06-30 | fix(api): don't burn wake cooldown when refused at capacity | -| 6c3999a | 2026-06-30 | fix(api): rate-limit email-OTP sends to close the email-bomb vector | -| 9873904 | 2026-06-30 | fix(operator): populate Status.Players from an RCON list probe | -| 879b177 | 2026-06-30 | fix(api): make OTP-start throttle atomic to close concurrent-burst bypass | -| 29f5341 | 2026-07-01 | docs(api): correct cooldownLimiter doc for its OTP reuse | -| 7507cfa | 2026-07-01 | Revert "feat(panel): wire role-switcher into the app shell" | -| f2c916d | 2026-07-01 | feat(api): add passkey enrollment persistence layer | -| 742f15f | 2026-07-01 | feat(api): add passkey enrollment endpoints | -| d2de11a | 2026-07-01 | feat(panel): fleet | -| 0261204 | 2026-07-01 | feat(passkey): add go-webauthn enrollment verifier adapter | -| fce0fce | 2026-07-01 | feat(passkey): wire enrollment verifier into felis-api | -| 2810fe8 | 2026-07-01 | fix(cfsetup): repoint stale DNS record when routing a tunnel hostname | -| 7d3be64 | 2026-07-01 | feat(cfsetup): start the tunnel connector as a setup step | -| a531f5e | 2026-07-01 | fix(cfsetup): keep connector install in the host apply layer only | -| e058a64 | 2026-07-01 | feat(edge): close the panel NodePort to the public after the tunnel is up | -| c01f133 | 2026-07-01 | feat(updates): add pure decision core for component self-update | -| fe2ece0 | 2026-07-01 | feat(api): add public Bind-Code onboarding for the player console | -| 3673af6 | 2026-07-01 | feat(api): add admin API for the SysAdmin-set auto-update maintenance window | -| 7464fa7 | 2026-07-01 | fix(updates): tag Window JSON so the persisted maintenance window round-trips | -| e035142 | 2026-07-01 | feat(passkey): add WebAuthn login/assertion crypto adapter | -| f34711c | 2026-07-01 | docs(api): record passkey login-handler deferral rationale | -| 7a51c1d | 2026-07-01 | fix(api): bound concurrent login bcrypt to shed CPU-pin floods | -| 164ac44 | 2026-07-01 | fix(api): validate inbound X-Request-Id before echo and audit persist | -| c6c0772 | 2026-07-01 | fix(api): set read/idle timeouts on the felis-api listeners | -| 3c1d647 | 2026-07-01 | fix(api): cap concurrent SSE streams per principal | -| d6e3189 | 2026-07-01 | fix(api): bound SSE relay writes with a deadline to sever stalled readers | -| 2c56d17 | 2026-07-01 | docs(api): record the quota-claim TOCTOU as a KNOWN-LIMITATION (audit #4) | -| 8f41a00 | 2026-07-01 | fix(api): clear the SSE write deadline on return so it can't leak to a reused connection | -| 15c58d9 | 2026-07-01 | feat(panel): player management | -| 8ae65ae | 2026-07-02 | feat(panel): backup management | -| a15ff55 | 2026-07-02 | refactor(panel): optimize player list layout and horizontal operations | -| 149f01a | 2026-07-02 | fix(panel): change console button to outline variant on my servers page | -| 4ecaf3c | 2026-07-02 | refactor(panel): set defaultOpen parameter of whitelist card to false | -| a9dbc8b | 2026-07-02 | feat(panel): add search and status filtering to my servers page | -| fa7bab5 | 2026-07-02 | style(panel): refine search and filter layout to align with header | -| 70d17a0 | 2026-07-02 | feat(panel): align my servers page search layout with fleet table | -| 5a8eff1 | 2026-07-02 | feat(panel): remove developer comment footer cards from my servers and server admin pages | -| 92770ea | 2026-07-02 | style(panel): adjust pagination padding to pt-3 for balanced spacing | -| 6e43a46 | 2026-07-02 | fix(panel): pin sidebar navigation and enable independent content scroll | -| 3b4298d | 2026-07-02 | refactor(panel): unify servers cockpit layout, resolve duplicate pages and adjust spacing | -| c0d333b | 2026-07-02 | feat(panel): support full server config edit dialog with status prefilling | -| 6368ab1 | 2026-07-02 | fix(api): coalesce MyServers owned flag so ownerless rows do not 500 | -| cdbb5ab | 2026-07-02 | fix(api): record credential id in passkey-register audit event | -| 9953275 | 2026-07-02 | fix(api): bound webauthn_challenges growth by superseding all prior rows | -| 20e31fb | 2026-07-02 | fix(store): cascade-delete passkeys and challenges on user removal | -| 7278cd7 | 2026-07-02 | feat(passkey): require and record user verification at enrollment | -| 54bc6ef | 2026-07-02 | fix(api): clear bound passkeys on password change to close a takeover foothold | -| 19f500b | 2026-07-02 | style(panel): update destructive red color and rename wake to start | -| e0bc288 | 2026-07-02 | feat(panel): implement image build pipeline and admin whitelist with mock dev api | -| 8594622 | 2026-07-02 | feat(panel): implement email OTP verification and passkey registration management | -| 9bed51b | 2026-07-02 | feat(config): add [velocity] login_image/lobby_image for system servers | -| 9ef817f | 2026-07-02 | feat(naming): system-server names, validation, and service-token identifiers | -| 159107b | 2026-07-02 | feat(api): HTTP readiness knob on MinecraftServer and login-gate fallback default | -| dc23cb5 | 2026-07-02 | feat(operator): system-server pod readiness probe and login service-token env | -| 3fdb3d0 | 2026-07-02 | feat(platform): internal API base-URL helper and single-sourced token secret | -| f554d52 | 2026-07-02 | feat(cli): provision login/lobby system servers with login env and token replica | -| a63f49d | 2026-07-02 | feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey | -| 241fe21 | 2026-07-02 | feat(limbo): felis-limbo in-game login flow over the shared account-link client | -| c7315e4 | 2026-07-02 | feat(deploy): login-limbo and lobby images with game-port pinning | -| 191640c | 2026-07-02 | feat(panel): implement admin submission approval and reject queue | -| 598f3d3 | 2026-07-02 | feat(submit): local + S3 backends for modpack upload contexts, installer-selectable | -| adf0d99 | 2026-07-02 | feat(panel): implement user-side modpack submissions with drag & drop context upload | -| d9e866f | 2026-07-03 | fix(deploy): make the lobby image actually build | -| b84debf | 2026-07-03 | feat(deploy): one-shot demo bring-up wrapper | -| 73d6ec1 | 2026-07-03 | feat(mock): add mock submissions for owner account | -| 5427bc7 | 2026-07-03 | feat(servers): support claiming servers directly from ServersPage list | -| 55592ed | 2026-07-03 | feat(auth): support public auth bind endpoint | -| 804459c | 2026-07-03 | feat(panel): implement admin maintenance window settings page | -| dd7dff6 | 2026-07-03 | feat(panel): support importing parameters from submission with owner-restricted unapproved entries | -| a8701c1 | 2026-07-03 | fix(panel): prevent automatic wake during server claim in mock api | -| f749c2c | 2026-07-03 | style(panel): resolve double borders and uneven padding in server console | -| 5f402b1 | 2026-07-03 | fix(panel): force dark mode and pure black bg on server console card | -| b7d8000 | 2026-07-03 | feat(panel): implement dedicated LuckPerms permissions and groups management sub-page | -| e60784b | 2026-07-04 | fix(panel): eliminate page collapse and scroll shifts during LuckPerms query reload | -| 439f19e | 2026-07-04 | fix(panel): prevent page collapse and scroll shifts in players and bans management sections during reload | -| 83e57b4 | 2026-07-04 | fix(panel): implement two-step confirmation for claiming a server to prevent accidental operations | -| 3347cc0 | 2026-07-04 | feat(panel): implement user management administration panel with sessions and minecraft link support | -| 67b4e19 | 2026-07-04 | fix(panel): override generic already_exists error message during user creation and profile editing | -| 2c95da8 | 2026-07-04 | fix(panel/i18n): add missing users_col_user key to translation files | -| 627883e | 2026-07-04 | fix(panel): refine reset password messages and fix empty email placeholder in mock api response | -| 0c1cc59 | 2026-07-04 | feat(auth): migrate console login to passwordless | -| 3b43f05 | 2026-07-04 | refactor(api): drop dead login concurrency limiter and reconcile passwordless comments | -| 4f59d51 | 2026-07-04 | feat(auth): add owner-tier passkey-unbind remediation endpoint | -| c20b12c | 2026-07-04 | refactor(api): drop dead password-era ResetMailer, reconcile passkey-unbind docs | -| 0a2accd | 2026-07-04 | chore: stop tracking Autohand-generated AGENTS.md | -| 96b3cc9 | 2026-07-04 | feat(updater): wire updates.Run to a caller with PaperMC v3 release discovery | -| 9896fe1 | 2026-07-05 | docs(updater): correct PaperMC UA/fixture overclaims, re-tier the boundary | -| 7d27640 | 2026-07-05 | feat(updater): add GitHub Releases source and route felis-api/k3s/cloudflared | -| bd49313 | 2026-07-05 | refactor(panel): 抽取 10 个公共组件,消除 ~150 处重复代码 | -| 91bfa27 | 2026-07-05 | feat(operator): implement idle auto-stop (spec §8) | -| e574749 | 2026-07-05 | feat(api): enforce CPU/memory/storage quotas (spec §9.3, §22) | -| 7db57b9 | 2026-07-05 | feat(updater): add VersionGatherer extraction core and CLI gather seam | -| ec468ba | 2026-07-05 | feat(auth): add discoverable (usernameless) passkey login | -| 154002e | 2026-07-05 | docs(auth): cite MultiLogin reference for UUID-keyed reclaim split | -| 0dbd557 | 2026-07-05 | fix(store): renumber discoverable-login migration 0013 -> 0014 | -| 9e1df12 | 2026-07-05 | feat(passkey): advance sign_count, reject clone-warned assertions | -| 7f7e459 | 2026-07-05 | fix(operator): enforce startup and readiness timeouts (§5, §8) | -| 7becb38 | 2026-07-05 | fix(api): implement /readyz with real DB + K8s API + CRD checks (§7) | -| 9079a2c | 2026-07-05 | feat(panel): implement email otp and passkey login interface | -| cfe68ae | 2026-07-05 | fix(panel): align status distribution order to put Stopped at the end | -| bbcfaeb | 2026-07-05 | refactor(panel): remove redundant voxel network topology description subtitle | -| fdb6efb | 2026-07-05 | feat(account): migrate a live account's owned servers to a new account (§B3 inherit) | -| abad137 | 2026-07-06 | style(panel): unify vertical spacing below PageHeader across pages | -| c2ee21a | 2026-07-06 | feat(breakglass): add halt-a-server op to the recovery console (§B4) | -| c1aa38b | 2026-07-06 | feat(velocity): add /felis migrate to open an account migration (§B3 inherit) | -| 7a7c0d5 | 2026-07-07 | feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync) | -| f2fc57c | 2026-07-07 | feat(api): add internal-face break-glass world backup endpoint (§B4 Sync) | -| 2ba9948 | 2026-07-07 | fix(platform): front the felis-api internal face on its own ClusterIP Service | -| fc748d3 | 2026-07-07 | feat(breakglass): add "back up a world now" console peer (§B4 Sync) | -| 9911b8c | 2026-07-07 | docs(changes): record the break-glass backup console peer (§B4 Sync phase 2b) | -| 096d597 | 2026-07-07 | docs(changes): backfill detail docs for pre-ledger functional commits | -| 5a7cd5a | 2026-07-07 | docs(changes): fold 346ec68 cloudflare-edge walkthrough into its detail doc | -| 4626ab5 | 2026-07-07 | docs(changes): mutation-audit the ledger's "unit-tested" safety claims | -| 729bd7b | 2026-07-07 | docs(changes): index the mutation audit and two lagging ledger rows | -| c67a4d3 | 2026-07-08 | docs(changes): close §B4 with the S3 archive backend deferred by design | -| 85b8a92 | 2026-07-08 | test(api): pin restore's owner gate against a superseded former owner | -| b7b4a3b | 2026-07-08 | docs(changes): record the round-2 backup/restore mutation audit and index the owner-gate test |