195 Commits
Author SHA1 Message Date
Lemon-miaow b9c97ffac9 docs(audit): batch-46 — upload-lane quotas/lifecycle, tracker #8/#1, release-pipeline smoke
Real-machine closure for everything batched since batch 45: #75/#76 verified
end to end on auditfix77 (cooldowns 429, pending cap 403, 2GiB budget 403 via
the real accounting path, failure-does-not-burn-window, withdraw/admin-delete
row+blob double clean, panel two-step confirms in CDP), tracker #8 and #1
landed and verified (#77 panel redirect, #78 felis converge), plus #79 docs.
Also: pgint's new assertions first run on real PG (17/17) and the release
pipeline's first-ever run — tag v0.1.0-rc1, multiarch assets published and
executed on the target arch, /releases/latest deliberately untouched.
2026-09-24 11:11:07 +08:00
Lemon-miaow b9ebc872ad docs(troubleshooting): the [INERT] legend points at a field that no longer exists
The legend said "§12 lists the one field this still applies to", but the last
inert field (spec.storage.retainOnDelete) was removed rather than implemented
(§13), and §12 has said "every field below is read by a controller" since.
Reworded so a reader whose change looks ignored follows the condition question
instead of hunting for a dead field.
2026-09-24 10:35:35 +08:00
Lemon-miaow c57daaf861 feat(cli): add felis converge for fields a newer desired spec never delivered (tracker #1)
Provisioning is create-if-absent, so a field the desired spec gained after an
install (spec.rcon, spec.startup.healthHTTPPort, a derived env key) never
reaches the existing login/lobby CR while every re-run of setup reports
success — the reported 'configuration updates never reach an installed
deployment' symptom. converge is the explicit pass: it fills exactly the
zero-valued whitelist fields and the derived env (including a missing key,
which refreshDerivedEnv deliberately never adds), and never overwrites a
non-zero value. The timing stays with the operator because enabling RCON or
the HTTP readiness gate on a pre-listener image would wedge that server in
Starting until it was marked Failed.

Tests: fills predated fields while operator edits survive / non-zero values
left alone / absent + foreign + unset-image guards. usage table updated so the
router-parity test passes; troubleshooting gains §12b.
2026-09-24 10:35:14 +08:00
Lemon-miaow 43699b46db fix(panel): route a setup-locked session to the wizard, not a permission error (tracker #8)
A session that still owes forced onboarding gets 403 setup_required from every
protected route, but the panel rendered it as the generic forbidden line — the
one step that unlocks the app read as missing authorization. api.ts now
announces the code on a window event (client module has no router) and the App
shell, inside the Router, navigates to /setup; the wizard resumes from the
surviving session with or without a token. Other 403s are untouched.

Panel tests: +2 (fires on setup_required, silent on any other 403).
2026-09-24 10:32:38 +08:00
Lemon-miaow 854320ac3f feat(submit): give the upload lane a lifecycle — withdraw + admin delete (#76)
Nothing ever removed a submission: users could not retract a pending row, no
route deleted blobs or rows, and the reaper never touches uploads — so every
upload accumulated on the 5 GiB PVC forever and the only cleanup was SQL or
kubectl against the store.

- Blobs.Delete on both transports (local: RemoveAll of the id-namespaced dir,
  id re-validated at the boundary; S3: idempotent object DELETE).
- Store: DeleteSubmission (admin, any status) and DeletePendingSubmission
  (owner+pending CAS — a reviewed row can never be withdrawn out from under
  its build).
- Manager.Delete / Manager.Withdraw delete the ROW first (under the CAS for
  withdraw) and the blob after, so a live row can never point at a reaped
  blob; a cleanup failure names the orphan explicitly instead of failing mute.
- API: DELETE /me/submissions/{id} (withdraw, app tier) and
  DELETE /api/v1/submissions/{id} (admin) both return the row as it was;
  audit events submission.withdraw / submission.delete; openapi documents both
  paths; admin route pinned in the admin-only table.
- Panel: two-step withdraw on a pending row (frees the pending slot and the
  storage budget); two-step delete on every admin row; zh/en copy; wire tests.

Unit: submit (withdraw happy path / wrong owner / reviewed row / no transport /
blob-cleanup failure), local+S3 delete idempotence, api handlers (200/404/409/
503 + route tier); pgint: withdraw CAS + admin delete exactly-once.
go vet/go test/gofmt clean; panel vitest 120 + typecheck green.
2026-09-24 10:24:23 +08:00
Lemon-miaow ad4d256d8f fix(submit): bound the untrusted upload lane — per-user caps + throttles (#75)
A logged-in user could file submissions without bound and stream a 1 GiB
context per submission. The only limits were the single-blob size cap and the
5 GiB uploads PVC (platform/workloads.go); nothing counted a user's rows or
bytes, so one account could fill the volume and every other user's upload
would start failing.

- Create: per-user pending_review cap (default 5) — the review queue cannot
  be parked full of one account's rows. Check-then-insert, documented soft.
- UploadContext: per-user stored-context budget (default 2 GiB) charged
  against the blob store's REAL sizes (new Blobs.Size on local/S3 stores), so
  the sum cannot drift from the volume; the write is capped at the remaining
  budget, so the excess is refused before it is persisted, and a re-upload is
  charged only for its new bytes.
- API: per-user create/upload throttles (30s/15s, cmd/felis-wired) on a
  dedicated cooldown keyspace, reserve→release so a failed attempt never
  burns the window and a burst collapses to one winner; ErrQuotaExceeded →
  403 submission_quota_exceeded (distinct from the 400 an oversize blob
  gets), 429 submission_cooldown for the throttles.
- Panel: zh/en copy for both codes; openapi documents 403/429 on the two
  user routes; pgint covers the pending-queue count.

Unit tests: submit package (cap, budget boundary/exact-fit/replacement,
oversize-vs-quota split) and api handlers (quota 403 both paths, throttle
429 + recovery + failure-release). go vet/go test/gofmt clean; panel
vitest 118 + typecheck green.
2026-09-24 10:14:42 +08:00
Lemon-miaow 3ed8bd7be9 docs(audit): batch-45 — modpack base-image chain (#70/#71/#72), submission-to-server capstone, tooling fixes (#69/#73/#74) 2026-09-24 03:19:21 +08:00
Lemon-miaow 96aa8176cd test(api): make the console-disconnect teardown case wait for the write, not just the read (#74) 2026-09-24 03:16:40 +08:00
Lemon-miaow 5cbfa893a5 docs: correct the modpack feature claim — approval auto-builds; deployment is an explicit server-image choice (#69) 2026-09-24 03:11:33 +08:00
Lemon-miaow 2961beb8fe test(bootstrap): keep the worlds-root warning case hermetic on hosts that already run k3s (#73) 2026-09-24 03:11:10 +08:00
Lemon-miaow b14bacbfc6 fix(build): mirror the Trivy Java DB — jar-bearing builds failed closed at the scan gate (#72) 2026-09-24 03:11:01 +08:00
Lemon-miaow 6e47730501 fix(build): allow kaniko to unpack base-image layers — drop-ALL removed the caps the tar apply needs (#71) 2026-09-24 03:05:07 +08:00
Lemon-miaow ac403b9cd3 fix(build): let kaniko pull base images from the plain-HTTP registry — push-only insecure flags broke every FROM registry.felis.svc build (#70) 2026-09-24 02:56:56 +08:00
Lemon-miaow 7791ed74f2 docs(audit): batch-44 — modpack-scale (200 MiB) context drill + concurrent builds, zero defects 2026-09-24 02:50:31 +08:00
Lemon-miaow 6abce99b49 docs(audit): batch-43 addendum — #68 context-fetch retry, live-verified; long-poll closure 2026-09-24 02:44:46 +08:00
Lemon-miaow b76d0acec7 fix(build): retry the context fetch through a control-plane restart (#68) 2026-09-24 02:27:45 +08:00
Lemon-miaow 580032056e docs(audit): forty-third batch — build-job reaping (#67) + lifecycle/fault-injection soak
Six CR-level run/stop cycles with three chaos injections (operator pod
kill, postgres restart, api pod kill) all converged (Running 23-29s,
Stopped 3s, pods gone per cycle); PG outage keeps the 503-not-401
session semantics; 30-request burst all 200; ownerless wake correctly
403s.  The build lane's missing Job TTL (#67, fixed in 2755e41) was
verified live on auditfix62: a real build Job carries ttl=604800 and an
old Job patched to ttl=30s was reaped, pod and all, within 40s.
2026-09-24 02:23:18 +08:00
Lemon-miaow 2755e41ff3 fix(build): reap finished build Jobs — they accumulated forever
Every other Job family carries a TTLSecondsAfterFinished (fileedit 2m,
backup/restore 10m) but the build lane never set one: each build left a
completed Job + Pod in felis-build indefinitely (5 already on the drill
cluster, oldest 26h), growing etcd and — since completed pods count
against the node's pod budget (110 on stock k3s) — eventually blocking
new builds.  The code even anticipated GC it never got ("JobUnknown
means the Job was not found (e.g. GC'd)").

Set a deliberately long TTL (7 days): the kaniko log is the admin
failure-triage surface (GET /images/build/{id}/logs), so the window
keeps a week of logs while bounding steady-state pods; terminal builds
are idempotent under Sync, so a late log-404 is the only cost.
2026-09-24 02:03:20 +08:00
Lemon-miaow d5e623c7ba docs(audit): batch-42 addendum — static review of the never-run release pipeline
The release workflow has never executed (no tags/releases exist yet).
Its three load-bearing contracts were verified against their other
halves: the asset names vs bootstrap's download_release_binary (and its
'felis ${REF}' convergence check), the local-exporter path vs the
Dockerfile's final COPY, and the version assertion vs cmdVersion's first
line.  The one un-replicable check (file(1) machine type) was probed
with a cross-compiled linux/arm64 binary: 'ARM aarch64' matches.  No
pre-tag fixes needed.
2026-09-24 01:33:56 +08:00
Lemon-miaow aa82321907 docs(audit): forty-second batch — coverage sweep, README_EN (#66), §28 diagram alignment
Module-by-module coverage matrix against the ledger: repo hygiene clean,
no TODO/FIXME debt, CLI dispatch test-guarded, CI syntax-checks every
tracked shell script, all 22 internal packages accounted for (CRD types
and game-image dirs covered indirectly), panel's 17 page routes drilled
in earlier batches.  Fixes this batch: #66 README_EN was missing the
private-repo install workaround and the rerun/upgrade notes (an English
reader 404s on step one); the §28 claim/link sequence diagrams were
realigned with the audited implementations.  A local-only scripts/
sync.sh conflict (plugins/ exclusion vs go:embed) was reproduced and
fixed on the workstation — private per .git/info/exclude, not committed.
2026-09-24 01:27:42 +08:00
Lemon-miaow 62e8c87d0e docs(diagrams): align §28 claim/link sequences with the audited implementations
The claim-transaction note still described the pre-audit-#4 shape (SELECT
EXISTS plus a quota pre-check outside the transaction); the implemented
ClaimServer serializes on pg_advisory_xact_lock(user_id), takes the row
FOR UPDATE and re-runs the four-dimension gate inside the same
transaction — the diagram's QuotaAvailable step is only a fast path.
The link flow now selects the code FOR UPDATE, treats a same-user
re-verify as idempotent, and lets a live caller take over a retired
(soft-deleted) owner's link — the 409 is only for a different, live user.
Both re-read from internal/api/pgrepo.go and the handler mappings.
2026-09-24 01:26:32 +08:00
Lemon-miaow 62c5a8a408 docs(readme): port the private-repo install path and upgrade notes to the English README (#66)
README.md grew a required workaround — the repository is private, so the
plain raw.githubusercontent one-liner returns 404.  The credentialed form
(token handed to curl via --config - so it never touches argv, sudo -E so
the installer inherits it) and the rerun/upgrade notes (the rerun upgrades
felis-api; the channel is not inherited, so main-followers need
FELIS_VERSION_BOOTSTRAP=dev) never made it into README_EN.md, which is
linked as the English entry point.  An English reader following that page
could not install at all.  Bring it back in sync with the Chinese one.
2026-09-24 01:24:27 +08:00
Lemon-miaow 5f9380dfca docs(audit): forty-first batch — loader mod layer closed
The last module family nobody ever built (#64 gradlew exec bit, #65
license metadata), closed with three real dedicated-server boots
(fabric/forge/neoforge: mod loads, /link registers, console refused)
and a permanent JDK-17 gate in CI + release.  Reachability table now
#1–#65; conclusions item 20; remaining-queue note updated.  CI run
35892544563 = 5/5 with the new mods job green on its first run.
2026-09-24 01:14:49 +08:00
Lemon-miaow aa5abfa911 ci(plugins): gate the three loader mods (JDK 17 wrapper builds) in CI and release
Nothing ever compiled these modules — no CI job, no install path — which is
why the gradlew exec-bit bug (previous commit) shipped unnoticed.  Add
plugins/test-mods.sh: runs each module's vendored wrapper under JDK 17
(they target the Java-17 Minecraft lines; paper/limbo stay on the JDK 21
gate).  ci.yml gains a `mods` job (temurin 17 + setup-gradle, wrapper
pinned per module), release.yml gates both plugin gates before shipping.
README: status now records the server-boot verification, the Building
section pins limbo's `-PlimboVersion=<release>` (the `+` default is
unresolvable from the LOOHP repo) and documents both gates.
2026-09-24 00:59:24 +08:00
Lemon-miaow c2fe6a6bf4 fix(plugins): loader mod metadata — AGPL-3.0-only license, real issue tracker
The three loader mods declared `license = "MIT"` (and Forge/NeoForge a
placeholder `issueTrackerURL = https://example.invalid/felis`) while the
repository is AGPL-3.0-only (README.md:73, LICENSE).  The mods were added
2026-06-26, the LICENSE landed 2026-07-12 — stale leftovers that nothing
ever read back: no CI job, no install path.  Fabric loader prints the
license from fabric.mod.json at boot and both mods.toml files are parsed
by their loaders, so the wrong claim is user-visible.  Align both with
reality — boot-verified on real fabric/forge/neoforge dedicated servers
(AUDIT-2026-09-22.md, batch 41).
2026-09-24 00:59:19 +08:00
Lemon-miaow f6048f268b fix(plugins): commit the gradlew exec bit for the loader mods
All three vendored wrappers were tracked as 100644, so the README's documented
'plugins/<loader>/gradlew -p plugins/<loader> build' failed on every fresh
clone with 'Permission denied' (exit 126) — and since no install path or CI job
ever ran them, nothing caught it.

git update-index --chmod=+x for the three files; the VM then built all three
modules for the first time (results in the gate commit and the audit).
2026-09-24 00:47:32 +08:00
Lemon-miaow 2d04c448c2 docs(audit): fortieth batch — java plugin layer into CI, demo-up single origin, #62 screened out
#63 fixed in f5a76cf (demo-up delegates to bootstrap; T1/T2/T3 real-machine);
plugins/test.sh + the new CI plugins job verified on the VM and in CI
(c59b387, run 35888009965, 4/4 jobs); #62 re-checked against all three
install paths and screened out as unreachable/nonexistent; plugin-layer
real-machine E2E via the MC status/login probes recorded.
2026-09-24 00:25:43 +08:00
Lemon-miaow c59b38775b ci(plugins): run the plugin self-tests and build the shipped plugin jars
The three framework-free test mains under plugins/*/test were never run by
anything — not CI, not the plugin builds — and the velocity/paper/limbo jars
were only ever compiled by deploy/bootstrap.sh on a live host. CI gains a
'plugins' job (JDK 21 plus the Gradle 8.14 the plugin Dockerfiles pin) running
plugins/test.sh: the three mains (InviteCardTest's jars fetched from Maven
Central, pinned and digest-checked) and the three production builds.

The first real run surfaced and fixed two untested assumptions: InviteCardTest's
documented javac line omitted examination-api (adventure-api's Component
signatures reference Examinable, so javac needs it too), and limbo's '+' version
default cannot resolve — LOOHP's repository serves no maven-metadata — so the
script resolves the current release off the Limbo CI artifact name (the same
source bootstrap reads) and plugins/README.md stops advertising a bare
'gradle -p plugins/limbo build' that can never work.

Verified in gradle:8.14-jdk21 on the VM: mains OK (32/36/48 checks);
velocity/paper/limbo BUILD SUCCESSFUL.
2026-09-24 00:19:56 +08:00
Lemon-miaow f5a76cf88a fix(deploy): make demo-up a thin wrapper — bootstrap is the one origin for the game stack (#63)
demo-up.sh grew its own image lane in July, before bootstrap could build the
game stack. The copy has since drifted from the installer it duplicates: it
pinned Paper 1.21.8 while bootstrap derives one version from the Limbo login
gate (both hops of a login must speak one protocol), it never built
felis-velocity.jar (a proxy without it silently routes nothing), it imported
images under local tags the [velocity] wiring no longer points at, and it left
the docker daemon running.

It now runs deploy/bootstrap.sh (unless SKIP_BOOTSTRAP=1), checks that the
earlier run really left the full stack behind, and hands over to 'felis setup'
— so a half-installed base is a clear error instead of a proxy that accepts
logins and routes nowhere.

Verified on the VM: healthy base passes to the handoff (exit 0); a hidden
felis-velocity.jar dies naming the jar and the remedy (exit 1, jar restored).
2026-09-24 00:19:49 +08:00
Lemon-miaow 1184d6b787 docs(audit): thirty-ninth batch — reaper multi-node drill on a cloned second node 2026-09-23 23:58:58 +08:00
Lemon-miaow 811f2b8898 docs(audit): thirty-eighth batch — #56–#59 evidence archived, server detail and ops sweep clean 2026-09-23 23:50:47 +08:00
Lemon-miaow 4c842f5d70 docs(audit): thirty-seventh batch — admin interaction sweep clean, #61 LuckPerms write promise scoped 2026-09-23 23:29:13 +08:00
Lemon-miaow b4ef42d947 fix(panel): scope the LuckPerms write promise to names the server can resolve (#61) 2026-09-23 23:28:28 +08:00
Lemon-miaow ecd42ed484 docs(audit): thirty-sixth batch — panel error localization (#60), Run4c admin write-path review clean 2026-09-23 23:12:14 +08:00
Lemon-miaow c9af5481e7 fix(panel): localize every user-reachable API error code (#60) 2026-09-23 23:09:28 +08:00
Lemon-miaow a2ff2a1102 fix(panel): make the LuckPerms page honest about what the server said (#59)
Live with real LuckPerms 5.5.85: every lp command returns an empty RCON body
(list/plugins answer normally; a standalone RCON client sees the same, and
creategroup/permission-set still persist), so the read projection can never
populate and the page asserted "no parent groups / no explicit nodes" for a
state it could not actually read. The raw reply now rides the same disclosure
the rosters carry, a silent entry-less reply shows an explicit notice instead
of the false empty claims, and the write history's placeholder no longer
dresses up a fabricated "[RCON] ..." line as output.
2026-09-23 22:43:29 +08:00
Lemon-miaow 4d4cdd6ea7 fix(api,panel): refuse file operations on a server without a world volume (#58)
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
2026-09-23 22:17:08 +08:00
Lemon-miaow 70c988e702 fix(panel): render the console command's reply (#57)
sendCommand returns the RCON reply and the component's own contract says it
"displays its plain-text reply", but the reply was discarded and the pod log
does not echo command output, so a sent command produced no visible result at
all (verified live with 'list'). The last command echo + reply now render
above the prompt, terminal-style.
2026-09-23 22:08:19 +08:00
Lemon-miaow 0790f8dfd3 fix(panel): surface the RCON reply for player-access mutations (#56)
The whitelist/ban/kick mutations reported a canned success message and threw
away the server's reply, so a refused command still read as done: live, the
vanilla server answers "That player does not exist" for a name it has never
seen (any player who has not joined yet), while the panel said the player had
been whitelisted/banned. api.ts documents these replies as "surfaced verbatim
as confirmation"; now they are. The localized string stays as the fallback
for a silent server.
2026-09-23 22:08:19 +08:00
Lemon-miaow 9c3a1c5be1 docs(audit): thirty-fifth batch — NetworkPolicy live enforcement matrix green (whitelist closed loop), velocity refresh loop verified 2026-09-23 21:43:23 +08:00
Lemon-miaow cbb0f11288 docs(audit): thirty-fourth batch — felis nano sweep (red 13 -> green 60/60), #55 bad-source ladder stall 2026-09-23 21:30:53 +08:00
Lemon-miaow 9dad61f508 fix(nano): screen unusable profiles per source instead of stopping the ladder (#55)
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.

Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.

Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
2026-09-23 21:21:21 +08:00
Lemon-miaow 971ae01caf docs(audit): thirty-third batch — /updates window API+UI sweep green, #53 repo move, #54 update guidance red/green
Records: (1) the /updates maintenance-window API sweep (unset null, 400s/415 for the invalid set, write->read-back->survives API pod restart, 401/403 auth) and the CDP panel sweep (status transitions, validation copy, clear, audit x3, zero console errors); (2) #53 red/green with the live fix53 binary ('MliroLirrorsIngenuity/Felis' -> 'FelisMC/Felis' in the 404 line) plus the token'd check (old path 301 / new path 200) and the fact the new home has no stable release yet; (3) #54 red/green with the live fix54 binary (installer one-liner + single trailer replace 'run: sudo felis setup'), the setup-vs-installer evidence, and the doc/test synchronization. Reachability table gains #53 (2) and #54 (2, docs); stats 54 total; repro-entry notes updated to the new remote, host binary v0.0.0+fix54 and the batch's artifacts.
2026-09-23 21:04:02 +08:00
Lemon-miaow 397a400d57 fix(update): point the apply guidance at the installer, not felis setup (#54)
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.

Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
2026-09-23 21:01:32 +08:00
Lemon-miaow 2c6739ad76 fix(updater,install,docs): follow the move to FelisMC/Felis (#53)
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.

Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
2026-09-23 20:58:34 +08:00
Lemon-miaow cec9a98305 docs(audit): thirty-second batch — S3 storage wizard sweep, #52 mirror lag red/green, cleanup 2026-09-23 20:39:09 +08:00
Lemon-miaow de7fb2c936 fix(setup): converge the workload felis-config mirror on every apply path (#52)
felis setup's in-TUI applies (storage / connection / edge) refreshed only the
control-namespace felis-config Secret; the workload-namespace mirror kept the
render from the previous run's startup pass until the next setup or installer
run. Found live: after 's -> Local' the minecraft copy still carried
user_uploads_context = s3://felis-wizard-uploads while the control copy and
both tomls were local. The 'configure email' path already overwrote both
mirrors, so storage/connection were the odd ones out.

Move the mirror refresh into applyFelisConfigSecret — the single choke point
every apply path calls — best-effort with a warning, since a control-plane
default install may not have the workload namespace at all. The smtp helper
drops its now-duplicate felis-config block.
2026-09-23 20:31:48 +08:00
Lemon-miaow 01988305a8 docs(audit): thirty-first batch — first-install walkthrough, #51 replica refresh red/green 2026-09-23 20:15:46 +08:00
Lemon-miaow 328e570309 fix(setup,install): refresh the workload namespace's felis-config mirror (#51) 2026-09-23 20:07:20 +08:00
Lemon-miaow cf5a790ea8 docs(audit): thirtieth batch — #50 smtp carry hoard, healed by two live re-runs 2026-09-23 19:57:00 +08:00
Lemon-miaow 4d3c85fd06 fix(bootstrap): [smtp] carry stops hoarding the auth_source comment block (#50) 2026-09-23 19:45:26 +08:00
Lemon-miaow 28fe7c43a9 docs(audit): twenty-ninth batch — image durability live drills; #46–#49
- registry hosting + loopback pull path landed (a9b275a/13d64e0/fa0e8d7) and
  drilled live: three installer re-runs, then GC simulations on the control
  plane (rolled felis-api pulled back in 25ms) and a game image (lobby-0,
  182MB in 10ms).
- #46 registry OOM (475MB-layer push killed the 256Mi template; dmesg evidence)
  fixed and re-verified: oom-kill count unchanged across a full rebuild+push.
- #47 AppleDouble ._*.sql embedding broke felis migrate on a Mac-staged tree;
  .dockerignore fix probed live with a planted junk file.
- #48 per-image docker start/stop tripped systemd start-limit-hit mid-batch;
  one wrap per batch, re-run mirrors all four.
- #49 installer re-runs silently reverted operator [registry]/[archive] config;
  carry-forward landed + live-verified into host toml, pod toml and the Secret,
  and the carried pins drove a successful POST /images/build.
- reachability table extended to #49 (① 20 | ② 19 | ③ 3+ | ④ 4 | 决策 3).
2026-09-23 19:35:13 +08:00
Lemon-miaow 72553cb414 docs(troubleshooting): the installer leaves docker stopped — start it before manual pushes
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
2026-09-23 19:35:13 +08:00
Lemon-miaow 20a95da487 docs(update): the plugins note is a rebuild + registry re-mirror now, not a node re-import
The felis-paper/felis-limbo jars are baked into the lobby/limbo images; with
the images hosted in the in-cluster registry, the extra step is pushing the
rebuilt image there (which is also what survives an image GC), not a bare
containerd import. The installer re-run does both.
2026-09-23 19:33:04 +08:00
Lemon-miaow c7e585e21d fix(bootstrap): mirror the image batch under ONE docker start/stop
Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.

push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
2026-09-23 19:23:44 +08:00
Lemon-miaow 5fa8b7412e fix(build): keep macOS ._*/.DS_Store junk out of the image
A Mac-staged tree (BSD tar materializes extended attributes as ._<name>
sidecars) went through the docker build and one landed in
internal/store/migrations/ — //go:embed-ed into the binary, where every
`felis migrate` then died with 'migration "._0004..." has a non-numeric
version'. Observed live wiring up the auditfix42 image: the installer's own
run_migrations failed on it. Exclude the sidecars and .DS_Store from the
build context; deploy/*.yaml and plugins/ have the same exposure.
2026-09-23 19:19:04 +08:00
Lemon-miaow b8e554dac7 fix(bootstrap): carry the operator's [archive] keys across re-runs too
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.

persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
2026-09-23 19:09:57 +08:00
Lemon-miaow 765a8923a4 fix(bootstrap): re-runs keep the operator's [registry] overrides
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.

Found while landing the registry-hosting work, which depends on those same
keys surviving.

- persisted_registry_block carries the operator-owned [registry] keys and the
  [registry.s3] subtable forward, same first-readable-file rule as
  persisted_smtp_block; url/build_namespace stay installer-owned (they must
  match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
  as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
  idempotence (a second re-run writes a byte-identical file).
2026-09-23 19:05:54 +08:00
Lemon-miaow fa0e8d7d97 docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
2026-09-23 19:03:11 +08:00
Lemon-miaow 13d64e0000 feat(bootstrap): host every built image in the internal registry — GC-durable pulls
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:

- Every image the installer builds is now a registry ref
  (registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
  containerd under that exact name (first boot needs no registry round-trip)
  and mirrored into the registry after deploy_bundle (push_image_to_registry:
  push endpoint 127.0.0.1:5000, and only the path after the host matters to the
  registry — a push there lands where kubelet's mirrored pull looks). A ref
  outside the registry is warned about, not silently unmirrored.

- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
  registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
  registry Deployment binds (node containerd cannot dial the Service VIP — live
  drill: "Empty reply"). k3s regenerates containerd config only at agent start,
  so a CONTENT change restarts k3s and an identical file (every re-run)
  restarts nothing.

- import_registry_image caches registry:2 into containerd so the registry
  Deployment can start on a box that cannot reach Docker Hub.

- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
  'felis-paper:demo') at the registry refs — a user server created from those
  rows must not strand when GC collects the bare tag. Only recommended rows
  still holding the old seed are touched; enabled is preserved; a pre-existing
  target row wins over a duplicate.

bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
2026-09-23 19:02:58 +08:00
Lemon-miaow a9b275abbb fix(platform): registry OOM (audit #46) + loopback hostPort — the node-side pull path
Two changes to the registry Deployment, both prerequisite to GC-durable images:

- Dedicated resource template: the control plane's 256Mi memory limit was a
  live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
  mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
  same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.

- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
  reach the Service VIP (live stack: "Empty reply"), so the node-side pull
  path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
  http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
  the plain-HTTP registry off every other interface.

Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
2026-09-23 18:55:23 +08:00
Lemon-miaow 92c06ac8dc docs(audit): twenty-eighth batch — #45 blind-review fix live-verified (context download, byte-exact); reachability 1-45 2026-09-23 17:01:19 +08:00
Lemon-miaow 168a37542b feat(api,panel): reviewer context download for submissions; dockerfile field documented as audit-only (audit #45) 2026-09-23 16:41:55 +08:00
Lemon-miaow edd9d63f5e docs(audit): twenty-seventh batch — alert module live drill (real build failure -> pending -> firing), cleanup, #44 reachability 2026-09-23 16:35:23 +08:00
Lemon-miaow 43df08b52a feat(alerts): ship Felis alert rules with promtool unit tests; document scraping & rules (troubleshooting §14) 2026-09-23 16:17:51 +08:00
Lemon-miaow 94f71eea19 feat(api): serve felis_* metrics on the internal face (build-failure counter's only scrape path) 2026-09-23 16:17:50 +08:00
Lemon-miaow 17ede3c6aa docs(audit): twenty-sixth batch ledger — setup wizard re-run screens, #44, build-pin drift incident & re-verify 2026-09-23 16:04:20 +08:00
Lemon-miaow ae6e9256c6 docs(troubleshooting): 8e — apply build-image overrides through the config Secret (restart alone does not) 2026-09-23 15:56:44 +08:00
Lemon-miaow abb5910d2f fix(cli): setup re-run keeps its already-set-up framing after connect/storage reconfigure 2026-09-23 15:56:44 +08:00
Lemon-miaow 5450ec786f docs(audit): reachability grading for findings #1-#43 (who actually hits each one) 2026-09-23 15:45:19 +08:00
Lemon-miaow 24a6ab3d1e docs(audit): twenty-fifth batch ledger — breakGlass console screens & backup/restore gates (#39–#43, live-verified) 2026-09-23 07:50:27 +08:00
Lemon-miaow ac3a557566 fix(cli): Sync picker hides system servers; keep the two 409 refusals apart
Two defects from the live Sync drill:

- The picker listed the system servers (login/lobby), which the backup API can
  never accept (reserved names, no servers row): the pick died in name
  validation with a raw "server name is reserved" error. backupPickable now
  filters them out; the halt picker keeps them on purpose (break-glass retains
  full power over system servers).
- backupErrorFromResponse mapped every 409 to the stopped gate, so the new
  world-volume refusal would have displayed the wrong reason. The 409 arm now
  keys on the body's error code; a code-less body still reads as the stopped
  gate.

Live (auditfix38): the picker shows only user servers; a world-less pick shows
the API's own "no world volume yet — start it once" text; the not_stopped text
is unchanged.
2026-09-23 07:49:07 +08:00
Lemon-miaow 508a1c02da fix(api): refuse backup/restore before a missing world volume
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
2026-09-23 07:49:00 +08:00
Lemon-miaow 55d515d41f fix(provisioning): keep the Owner seat single; clash on the operator name stays retryable
Two defects live-drilled in the break-glass staff provisioning:

- An Owner reset that typed any username other than the occupied seat took
  UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
  existing seat — possibly the compromised account the reset was meant to
  replace — live; every owner row is undeletable through the panel, so the tier
  could never converge back to one. provisionOwner now refuses with
  ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
  form); bootstrap still mints, and the seat's own username still resets in
  place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
  console keys its rename prompt off api.ErrConflict — the "choose another
  name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
  the contract; PGRepo had drifted). Map the unique violation to ErrConflict
  and pin it in pgint.

Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
2026-09-23 07:48:54 +08:00
Lemon-miaow f6dbfd3625 docs(audit): twenty-fourth batch ledger — live S3 upload-channel drill (0 defects, reverted clean) 2026-09-23 07:05:42 +08:00
Lemon-miaow f378953982 docs(audit): twenty-third batch ledger — reaper node pin (#38 + multi-node gap) 2026-09-23 07:00:33 +08:00
Lemon-miaow daf760220b fix(cli): pin the reaper to its storage node; drop the stale uid-1000 note
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
2026-09-23 06:58:29 +08:00
Lemon-miaow a31eca65c3 docs(audit): twentieth–twenty-second batch ledger — build outcome visibility, files page, fleet system services 2026-09-23 06:53:40 +08:00
Lemon-miaow 2f90851c03 fix(panel): mark platform system services read-only in the fleet table
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
2026-09-23 06:51:45 +08:00
Lemon-miaow 0a36b3fda9 feat(panel): add the server files page for the world-volume repair lever
The backend could list/read/write a stopped server's world volume since the
file-editor slice, but the panel had no entry, so the one repair path for a
server that will not boot (a wrong line in server.properties) was API-only.
New /servers/:name/files page: breadcrumb browser, editor dialog with the
base64 []byte codec, binary files open read-only, the stopped gate is owned
up front (with a stop action) instead of letting every call 409, and a
doorway card on the console. i18n files namespace + wire-shape tests.
2026-09-23 06:42:45 +08:00
Lemon-miaow 72c4aa3895 fix(submissions): surface each linked build's outcome to the submitter
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
2026-09-23 06:31:13 +08:00
Lemon-miaow 4933c075b0 docs(audit): nineteenth-batch ledger — passkey unbind panel entry 2026-09-23 06:25:01 +08:00
Lemon-miaow 11ac4f50e6 feat(panel): expose owner passkey unbind in the user danger zone
DELETE /users/{id}/passkeys shipped as the owner-tier remediation for a
lost or compromised authenticator, but nothing in the panel reached it.
Add the danger-zone action with a confirm dialog; the account keeps its
other doors (email OTP, in-game op-login re-enrollment), so this severs
a credential without locking anyone out. Wire-shape test pins the call.
2026-09-23 06:24:45 +08:00
Lemon-miaow 35d93d7612 docs(audit): eighteenth-batch ledger — #35 world-executor identity defect and the backups-page completion 2026-09-23 06:20:44 +08:00
Lemon-miaow 97a64c8a33 feat(panel): add back up now and recent operations to the backups page
The backups page could list and restore archives but not create one,
and nothing surfaced backup/restore Job outcomes — a failed 202 was
visible only through kubectl. Add a Back up now action (enabled only on
a stopped server, the backend's own gate; a raced 409 is surfaced in
its words) and a Recent operations card fed by GET /servers/{name}/jobs
that shows running/succeeded/failed with the Job's failure message,
re-reads on an interval while a Job is running, and persists across
reloads. Wire-shape tests pin both endpoints.
2026-09-23 06:20:16 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00
Lemon-miaow f21aef3cfa docs(audit): seventeenth-batch ledger — hasJoined multiplexer drill (fake Yggdrasil) 2026-09-23 05:58:26 +08:00
Lemon-miaow 4298cd5de1 docs(audit): sixteenth-batch ledger — internal-face residual endpoints swept clean 2026-09-23 05:54:32 +08:00
Lemon-miaow 2bd25be712 docs(audit): fifteenth-batch ledger — #34 live closure and executor image refresh 2026-09-23 05:52:12 +08:00
Lemon-miaow bb9798e32c fix(api): in-game identity resolution and link takeover ignore dead accounts
UserByMCUUID now resolves only live accounts: claim, menu, wake
authorization, op-login vouch and the QR link-status poll treat a
disabled or soft-deleted link holder exactly like an unlinked UUID
instead of a retired identity. VerifyLinkCode lets a soft-deleted
link be taken over by a fresh in-game code (the deleted account is
gone, e.g. a migrated source), while a disabled holder still 409s so
the lockout is not bypassable; failed attempts still do not consume
the code. Fake repo and pgint coverage pin both branches.
2026-09-23 05:40:29 +08:00
Lemon-miaow 3ffa3f5318 docs(audit): fourteenth-batch ledger — dead-account resurrection (#33) and its live closure 2026-09-23 05:33:54 +08:00
Lemon-miaow 58535890c4 fix(api): dead accounts cannot log in, hold sessions, or keep identity assets 2026-09-23 05:28:32 +08:00
Lemon-miaow ba9d98f7ce docs(audit): thirteenth-batch ledger — CLI, direct build, panel CDP sweep, #31/#32 2026-09-23 05:15:51 +08:00
Lemon-miaow 6907961ce0 fix(panel): the build page trusts the server-side owner tier and drops its mock build seeds 2026-09-23 05:11:49 +08:00
Lemon-miaow d0b1f9694e docs(audit): twelfth-batch ledger — users admin matrix, defect #30 fix on live 2026-09-23 04:57:36 +08:00
Lemon-miaow 1918da29be fix(api): the quota/link admin sub-resources require a live user (404, not FK 500) 2026-09-23 04:56:06 +08:00
Lemon-miaow fbb6b0c180 docs(audit): eleventh-batch ledger — submission negative matrix, internal context fetch, op-login remint 2026-09-23 04:50:11 +08:00
Lemon-miaow ffe5dc14a8 docs(seams): close the deferred entries that are now live-verified 2026-09-23 04:46:02 +08:00
Lemon-miaow d4bb8d344b docs(audit): tenth-batch ledger — configure-email mirror fix (#29) and auditfix25 deployment 2026-09-23 04:44:12 +08:00
Lemon-miaow ed722d55f8 fix(setup): the workload-ns SMTP mirror must carry the target namespace
'smtpSecretManifest' hardcoded namespace=felis, so the 'configure email'
refresh of the minecraft-namespace copies failed before it began: kubectl
refuses a manifest whose namespace conflicts with -n (found live: 'the
namespace from the provided object "felis" does not match the namespace
"minecraft"'), and the felis-config mirror never ran at all because the
smtp apply returned early. A later SMTP change could therefore never reach
the reaper's pre-reap warnings — the exact failure the refresh was added to
close.

Render the Secret with the caller's namespace (felis for the control-plane
apply, the workload namespace for the mirror). Regression test pins both.
2026-09-23 04:42:59 +08:00
Lemon-miaow 311b1a7ec4 docs(audit): ninth-batch ledger — cfsetup #28 and the unused-endpoint sweep
Records the Access-policy upsert fix and the first end-to-end runs of
account/migrate (all four steps + negative matrix + retire assertions),
passkey credential management, the access player-management group, the
updates window, and fleet — plus the environment restore notes.
2026-09-23 04:41:01 +08:00
Lemon-miaow 36b954d347 docs(backupjob): the backup Job name is not deterministic anymore
The comment described a deterministic-name collision that the unique random
suffix made near-impossible; align it with Backuper.Backup and jobspec's
contract (ErrAlreadyExists survives only as the defensive no-op).
2026-09-23 04:41:01 +08:00
Lemon-miaow 30857df5b2 fix(cfsetup): upsert the Access policy — never swallow already-exists over a broader rule set
CreateAccessPolicy treated a Cloudflare "policy_already_exists" as idempotent
success and kept whatever policy was there. On a re-run with a changed
identity — or against a hand-made broader policy — op.console would stay
guarded by something weaker than the fail-closed body this package builds and
guards, while Setup reported success. The fail-closed validation only ever ran
on the policy we built, never on the one that stayed live.

Now it upserts by name: lookup, PUT the guarded body over the existing policy,
POST only when absent (a racing POST re-looks up and PUTs). apiPost/apiPut
share one apiWrite; three httptest cases pin update-over-existing, create-when-
absent, and the race fallback.
2026-09-23 04:30:53 +08:00
Lemon-miaow faa508e87a docs(audit): operator self-healing postmortem — ledger #26/#27 with live drills
RCON-secret deletion lockup and the stale start anchor (with its permanent
Provisioned=False) get their full live evidence trail, plus the sts
accidental-deletion drill. Deployed image note bumped to auditfix24.
2026-09-23 04:29:02 +08:00
Lemon-miaow 82b5a606f7 fix(operator): end the start attempt on success — stale anchor caused false StartupTimeout
Found live while validating the RCON-secret heal: a server that had already
recovered to Ready was marked Failed(StartupTimeout) minutes later, the moment
an unrelated pod rollout briefly dropped readyReplicas. The anchor
(status.startRequestedAt) was never cleared on success, so its 300s budget
kept ticking under a healthy server and any later blip spent it.

markRunningReady now clears the anchor: every start-or-recovery attempt gets
its own budget. It also flips ConditionProvisioned back to True — markFailed
sets it False and nothing ever reset it, leaving a permanent failure flag on
recovered servers that every conditions consumer would read.

Unit tests pin both: anchor cleared on Ready, Provisioned recovers from
Failed to Running.
2026-09-23 04:23:59 +08:00
Lemon-miaow 56f3abdb36 fix(operator): heal a deleted RCON Secret instead of locking the server out
Deleting the per-server RCON Secret used to leave a running pod authenticating
with the lost password while the operator re-minted a fresh one and probed
with it: the RCON gate failed forever (live: 96s+ of RconNotReachable, headed
for ReadinessTimeout) and nothing re-triggered a pod restart — the server only
came back when the pod was deleted by hand.

Two changes pair up:
- Owns(&corev1.Secret{}) so the deletion is noticed at all (a quiet Running
  server emits no other events; the Secret is controller-owned, so the watch
  maps it back to the CR).
- The pod template now carries a fingerprint of the current password
  (RconSecretAnnotation). Re-creation changes the fingerprint, the
  StatefulSet rolls, and the new pod picks the new password up; while the
  Secret is untouched the value is stable so no spurious rolls.

Unit tests pin stability across reconciles and the change-on-recreation roll.
2026-09-23 04:20:49 +08:00
Lemon-miaow 089d4f3a80 docs(audit): idle auto-stop postmortem — ledger #25 and self-checks in §11
Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
2026-09-23 04:14:11 +08:00
Lemon-miaow f650bf892a fix(operator): idle auto-stop couldn't write — patch the spec, and grant the patch
Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:

- The stop used a whole-object Update while the same reconcile loop writes
  status; that risks clobbering a concurrent status write. Switch to the
  reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
  markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
  closed with 403 (visible in the operator log as 'cannot update resource
  "minecraftservers"'). Grant patch and pin it in the RBAC scope test.

With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
2026-09-23 04:08:58 +08:00
Lemon-miaow 1c89a5eeeb fix(operator): wake up for idle auto-stop — the timer had no driver
EmptySince was stamped and then never revisited: player joins/leaves do not
touch the CRD, RCON is only probed inside Reconcile, and a steady Running
server produces no watch events (its status update goes out unchanged and is a
no-op). Live, an empty server with a 30s grace sat Running for minutes with
zero reconciles in the log — the auto-stop existed only on paper.

reconcileRunning now returns a RequeueAfter for idle-enabled servers: exactly
at the deadline while empty, or a 30s probe cadence while occupied so the
moment the last player leaves is noticed. New unit tests pin all three:
deadline requeue, occupied cadence, and no requeue when disabled.
2026-09-23 04:05:31 +08:00
Lemon-miaow c04a3f083e fix(crd): persist status.emptySince — the field was pruned away by the schema
The operator stamps EmptySince to time the idle auto-stop window, but the CRD's
status schema never declared it. Kubernetes pruned the field on every write
(apiserver warning: unknown field "status.emptySince"), so the timer reset to
nil on every read and idle auto-stop could never fire — a defect invisible to
the fake-client unit tests, which do not enforce the CRD schema. Found live:
the stamp was silently dropped the moment it was set.

Schema now declares emptySince (date-time) like its sibling timestamps.
2026-09-23 04:05:22 +08:00
Lemon-miaow 72f0b4a258 docs(audit): sixth-batch ledger — reaper warning path drilled end-to-end on the live cluster
The auditfix20 image (8e7c7bb) was exercised against a real SMTP sink with a
dedicated warntest server: delivery content, tier precedence, dedupe, retry
semantics (bad relay / unverified email / unowned), threshold boundaries, and
zero backup side effects. Ledger #24 records the defect and the evidence; the
deployed image note is bumped to auditfix20.
2026-09-23 03:56:07 +08:00
Lemon-miaow 8e7c7bbf24 fix(reaper): deliver pre-reap warnings for real — and never fake a delivery
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:

- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
  `warning suppressed — no warner wired` and does NOT stamp; a delivery error
  logs and retries on the next daily run (bounded by the warning window). The
  stamps are no longer burned by notices nobody received.

- A real channel: mail.SendNotice (the second and last message shape the mail
  package sends) plus a mailWarner that resolves the owner's VERIFIED email
  and mails the notice through the configured [smtp] relay. `felis reaper`
  wires it when [smtp] is set (same password_ref convention as felis-api) and
  prints exactly what happens when it is not.

Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.

Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
2026-09-23 03:47:19 +08:00
Lemon-miaow 1d0ec61c9d docs(audit): fifth-batch ledger — quota atomic gate closed (audit #4)
The deferred-seams entry that waited for a real-Postgres harness is struck:
ClaimServer owns the gate now, proven red-then-green by the pgint concurrency
test (two claims, one win, one 403), and the storage-cache zeroing found in the
same pass is recorded with its hermetic test. auditfix19 is live on the VM.
2026-09-23 03:39:25 +08:00
Lemon-miaow bb68fefe04 fix(quota): make the claim gate atomic, and stop zeroing the storage cache
Two defects in the §9.3 quota path, both invisible to the hermetic suite:

- Audit #4's TOCTOU was real and documented: QuotaCheck and ClaimServer were
  separate statements, so two concurrent claims by one user for two different
  ownerless servers both read count < max_servers and both won. The gate now
  lives inside ClaimServer, in the SAME transaction as the ownership write,
  under pg_advisory_xact_lock(hashtext(user_id)) — the aggregate read, the
  four-dimension re-check (shared with QuotaCheck via one helper so the two
  cannot drift), and the UPDATE are one serialized decision. The loser gets
  ErrQuotaExceeded, which both claim handlers map to the same 403 the
  sequential path gives; the server row is additionally taken FOR UPDATE so
  same-server races still resolve to exactly one winner.

- The server PATCH path called UpdateServerResources(..., 0) for storage even
  though a resources patch cannot change storage. The cached columns are the
  ONLY input to the quota aggregate, so every resource patch silently dropped
  that server's storage contribution from its owner's cap. The handler now
  reads the current spec and passes storage through.

Red-then-green: the new pgint test drives two real concurrent claims against
max_servers=1 (before: both win; now: exactly one win + one gated 403, and the
DB shows one owned row); the hermetic suite pins the 403 mapping and the
storage-preserving cache write.
2026-09-23 03:37:59 +08:00
Lemon-miaow d829267f1c docs(audit): fourth-batch ledger — PG contract tests land, verified-email and owner-role defects reconciled
#20–#22 recorded with live evidence: the pgint harness caught the
never-shipped verified-email index on its first run; migration 0020 applied
and the 409 drill replayed live; the admin email edit now drops the stale
proof; the owner role is written by both provisioning paths and its staff
doors, guards, and reclaim protections were drilled end to end on auditfix18.
The remaining-work item "PG-level contract tests" is checked off.
2026-09-23 03:32:48 +08:00
Lemon-miaow e0d23780d8 fix(auth): make the owner role real — provisioning, staff doors, panel guards
Found live while verifying the admin email-edit fix: the Owner account could
not load /api/v1/users at all. Root cause: migration 0011 adds the 'owner'
role and gates every user-administration route on it, but NOTHING ever wrote
it. break-glass (UpsertOwner), the setup MC-bind (CompleteOwnerSetup), and the
re-provision path all forced 'admin', so in a fresh install the entire
owner tier — list/create/edit/disable/delete users, quotas, sessions — was
unreachable. The role was a dead letter in the other direction too: staff
predicates that predate the role did not know it.

- UpsertOwner and CompleteOwnerSetup now write role='owner'; the username-
  conflict arm re-asserts it, which is also the documented pre-0011 promotion
  path ("re-provision via break-glass"). InsertOperator stays plain 'admin'.
- Staff doors learn the role: op-login start/finish admit the Owner; the
  player email door refuses it like any staff account; the in-game approver
  check already used staffRole.
- Reclaim protection: IsProtectedAdminLink (and the break-glass bootstrap
  switch AdminExists) count admin OR owner — the Owner must never be displaced
  by a Mojang-priority reclaim.
- Panel guards make migration 0011's claim true now that owner rows exist: an
  owner can never be demoted, deleted, or disabled through the API (only the
  local break-glass console resets the identity); username/email edits still
  work.

Tests: pgint pins both provisioning paths, the protected-link predicate and
the reset/promote semantics; hermetic suites cover the owner-admitting staff
door, the owner-refusing player door, the three panel guards, and break-glass
attribution.
2026-09-23 03:30:29 +08:00
Lemon-miaow d1ec40f738 fix(users): an admin email edit must clear the stale verification
UpdateUser wrote a new address but kept email_verified, so patching a verified
account asserted a proof of an address nobody had proven — and the
pre-session login mails and resolves on exactly that flag, so a typo'd edit
could hand the account's sign-in codes to the wrong mailbox.

Changing the address now clears the flag in the same write; a no-op patch that
passes the same value keeps it. The fake mirrors the semantics, and the pgint
suite pins both halves (same value keeps proof, new value drops it).
2026-09-23 03:19:55 +08:00
Lemon-miaow b6ef27cd2d fix(auth): enforce the verified-email uniqueness that email login assumes
The design has claimed since migration 0010 that at most one account can hold
a PROVEN email address, with ErrEmailTaken as the 409 a second verifier sees.
Neither half ever shipped: no migration created users_verified_email_unique,
and VerifyEmailOTP had no guard at all — the sentinel was defined but never
returned, so two accounts could both verify one address. The damage is not
cosmetic: the pre-session login resolves accounts BY verified email, so the
duplicate decided which identity a mailed sign-in code belonged to.

- Migration 0020 creates the partial unique index (lower(email) WHERE
  email_verified) the comments have been citing — the database-level backstop.
- VerifyEmailOTP now refuses the take-over with ErrEmailTaken BEFORE consuming
  the code (the address, not the code, is the problem), charges no attempt,
  and maps a lost cross-user race (unique violation) to the same answer.
- The verify handler answers 409 email_taken instead of a generic 500.

Covered by the pgint suite (sequential double-verify refused with the code
still live, a direct duplicate write still loses to the index, the refused
account can still prove its own address) and a hermetic 409 case.
2026-09-23 03:19:35 +08:00
Lemon-miaow 2a55a0d265 test(pgint): verify the business stores against a real Postgres
The hermetic suites encode the store contracts against fakes; PGRepo drifted
behind them three times (attempt accounting, a missing JOIN, a missing FOR
UPDATE) while every unit test stayed green. This harness replays the real
embedded migrations onto a throwaway database — its name must contain "pgint"
or the harness refuses to run — and exercises the SQL directly: sessions, the
onboarding email-OTP lifecycle (supersede/expiry/lockout), the pre-session
login consume, the op-login state machine, link and bind-code redemption,
submissions, and builds with the image admission round trip.

Run it after touching SQL under internal/api/pgrepo.go, internal/submit, or
internal/build; CONTRIBUTING.md carries the one-liner.
2026-09-23 03:18:56 +08:00
Lemon-miaow d2c656533e docs(audit): third-batch ledger — probes verified, build lane proven end to end, #7/#11-#15 reconciled
- The control-plane probes (0c8e29b) and the user-modpack build lane
  (f79e5eb + 02fd2de) get their live evidence recorded, including the three
  drill-only defects the lane fixed (Job scheduling, Kaniko Dockerfile
  ownership, Trivy DB egress).
- The stale 'unfixed' rows #7/#11-#15 are reconciled with their commits.
- Remaining-work list re-stated: image durability, PG contract tests, the
  panel's /jobs block, multi-node reaper placement, alerting, and the newly
  found player-visible build status gap.
2026-09-22 23:05:13 +08:00
Lemon-miaow 02fd2de502 fix(build): three drill-driven fixes so the lane actually completes on a starter node
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:

- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
  Pod never scheduled on the platform's own starter node: FailedScheduling /
  Insufficient memory, Pending forever. Requests are now a small floor
  (250m / 512Mi, never above a configured cap) while the limits stay the
  safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
  the source owner; a 65532-owned context (the distroless felis image uid)
  fails that under the pod's dropped capabilities ('copying dockerfile:
  chown /kaniko/Dockerfile: operation not permitted'). The fetch container
  now extracts as root — the uid Kaniko already runs as — so the copy
  succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
  step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
  renders --db-repository, and docs/troubleshooting.md §8e now carries the
  verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
  internal registry; --insecure already covers its plain HTTP).

Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
2026-09-22 23:01:25 +08:00
Lemon-miaow f79e5ebb5e feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:

- submit: derived context refs become the internal-face URL
  /api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
  gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
  route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
  image's new fetch-context entrypoint) that streams the blob with the
  namespace-local service-token Secret — never mounted into Kaniko — and
  extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
  reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
  build namespace gets the token Secret through the existing replica mechanism
  (bootstrap.sh + felis setup); the build egress lock opens exactly the control
  namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
  escapes/symlinks/non-gzip).

Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
2026-09-22 22:45:09 +08:00
Lemon-miaow 0c8e29b05a fix(platform): give every control-plane Deployment real probes (#8 follow-up)
The api, operator and registry Deployments shipped with no liveness/readiness
probes at all: a wedged process stayed 'Running' forever, and the operator had
no health listener to probe in the first place. Kaniko build evidence on a
fresh install showed the only cluster-wide red after a disk-pressure pass was
Deployment status that never reflected health.

- felis-api: readiness /readyz (DB + K8s API round-trip) and liveness /healthz
  on the internal face (:8081), the only listener that serves both endpoints;
  liveness deliberately avoids /readyz so a DB blip cannot restart the api.
- felis-operator: new --health-probe-bind-address (:8081) with controller-
  runtime's /healthz + /readyz (registered ping checks; an unregistered handler
  map would 404), plus the matching container port and probes.
- registry: /v2/ probes on the pinned port, so a broken storage backend stops
  reading as 'Running'.

Tests pin paths, ports, and that each probe targets a declared container port.
2026-09-22 22:22:37 +08:00
Lemon-miaow edefc34a5b docs(audit): night-2 ledger — default-install backup/reaper verified live, eviction shield re-drilled, OOM + k3s SIGKILL chaos, and the prioritized remaining-work list 2026-09-22 22:15:38 +08:00
Lemon-miaow 87a9f4eb25 feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
2026-09-22 22:03:48 +08:00
Lemon-miaow 0a2d654e68 fix(platform): control plane runs system-cluster-critical, so eviction refuses it (#8)
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.

Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
2026-09-22 21:55:54 +08:00
Lemon-miaow fe310743a2 fix(platform): give the control plane a PriorityClass eviction shield (#8)
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
2026-09-22 21:08:56 +08:00
Lemon-miaow 2b87a5a13b fix(install): grant the reaper traverse on the worlds root (#6)
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.

Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
2026-09-22 21:03:17 +08:00
Lemon-miaow fd33fd05e1 fix(install): backups exist on a default install; retention resolves real world dirs (#6)
Three faces of one gap, all on the supported install path:

- Backup/restore answered 503 out of the box: nothing ever rendered the
  archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
  PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
  'felis manifests' names it by default (--backup-pvc= is the explicit
  no-store shape); bootstrap passes it through so the generated felis.toml
  [archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
  now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
  env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
  nothing deleted).
- Even when enabled it could not find a world on a stock install:
  resolveWorldDir now also resolves the exact local-path directory
  <pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
  glob, so a stale deleted PV's bytes can't be archived in place of the
  current world). Reaper Role gains persistentvolumeclaims:get (weaker
  than the delete it already held).

README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
2026-09-22 20:48:07 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow a2df2f242b fix(operator): re-probe RCON every 2s while Starting
The readiness gate is status-driven; at a 5s re-probe cadence the observed
'container Ready but API still 409 not_running' window was 6~10s. Halving the
cadence halves the worst case; probes still only run while unreachable.
2026-09-22 20:31:53 +08:00
Lemon-miaow 2a8f897e61 fix(api): a session-store outage answers 503, not 401
Resolving a session cookie failed identically whether the credential was
missing or Postgres was unreachable: local_auth_enabled read errors fell into
the fail-closed 'disabled' branch and SessionUser errors into 'invalid
session', both surfacing as 401 'authentication required' — a lie that reads
as 'log in again' during an outage. Split the enabled-read into
(enabled, error), tag non-ErrNotFound store failures with errAuthBackend, and
map that to a new 503 auth_unavailable in requireExternal. Fail-closed is
unchanged: missing setting / bad value / missing session stay 401.
2026-09-22 20:30:43 +08:00
Lemon-miaow abce381faa fix(tui): wrap the one-time setup URL so narrow terminals can't truncate it
The setup URL carries a 43-char token and overruns 80 columns; the TUI
renderer clipped it. Break it at the query '=' boundary (token on its own
line) with a shared wrapDisplayURL helper used by both the Owner wizard and
the mc-bind wizard; unit test pins the no-loss concatenation.
2026-09-22 20:28:01 +08:00
Lemon-miaow a415246adc fix(operator): give controller-runtime a logger instead of a goroutine stack
Without SetLogger, the first reconcile prints
'[controller-runtime] log.SetLogger(...) was never called; logs will not be
displayed' followed by a full stack trace (live-observed in felis-operator).
Route it through logr.FromSlogHandler(slog.Default()) so its messages are
ordinary stderr lines; go-logr/logr promoted to a direct dependency.
2026-09-22 20:25:39 +08:00
Lemon-miaow 1efa8a4b08 docs(audit): batch 2026-09-22 evening — auth E2E, concurrency, reaper drill, fixes #16-#19 2026-09-22 20:19:35 +08:00
Lemon-miaow c839454a1f fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
another namespace either (live drill: 'error looking up service account
minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
RoleBinding subject to the Minecraft namespace alongside the CronJob.
2026-09-22 20:18:36 +08:00
Lemon-miaow e4f2cff532 fix(manifests): render the retention reaper CronJob into the Minecraft namespace
A Pod can only mount PVCs from its own namespace; the CronJob referenced the
minecraft-namespace backup PVC while being rendered under ControlNamespace, so
it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim
felis-backups not found'. The reaper Role/RoleBinding were already
minecraft-scoped (the objects it touches live there), so the CronJob was the
odd one out. The minecraft felis-config replica (felis setup, backup Job fix)
supplies its config mount.
2026-09-22 20:14:19 +08:00
Lemon-miaow 0414913bc7 fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
6-way concurrent redeem of one code 500'd on users_username_key (each request
generated a fresh user id but the same uuid-derived username), plus the rarer
two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
locks its code row and handles the conflict.

- SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
  as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
- INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
  cross-code racers converge on the winner's row (role checked, staff still
  refused) instead of a unique-violation 500.
- account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.

Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
db clean; zero unmapped errors.
2026-09-22 20:08:38 +08:00
Lemon-miaow dcc3b7403e fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake)
The interface doc promised 'each joined to its staff username', the fake and
the pending handler both project username and created_at, but the PG query
selected neither — live internal /op-login/pending returned username:"" and
created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
tests can't see PG-only regressions.
2026-09-22 20:04:32 +08:00
Lemon-miaow 52549f7b3a fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
The PG implementation was a single UPDATE ... WHERE code_hash that returned
ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
the pre-session email-login door (and the op-login finish / migration confirm
doors) fell through to writeError's unmapped-error 500, and attempts were never
charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
intended contract; only the PG side had drifted.

Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
UPDATE the newest live row, expiry + attempt cap before the hash compare,
mismatch charges one attempt and returns ErrOTPInvalid without consuming,
match consumes and commits. Verified live on the VM: 5 wrong guesses return
400 and stop at attempts=5 (correct code then also refused, unconsumed);
fresh code redeems; replay returns 400.
2026-09-22 19:51:04 +08:00
Lemon-miaow 9309ff5a7f chore: apply the missed S1016 conversions in handlers_users
The gofmt/staticcheck commit staged handlers_user.go (singular) for the
formatting fix but missed this sibling for its two struct-literal-to-
conversion cleanups.
2026-09-22 17:57:03 +08:00
Lemon-miaow e690b058db fix(restore): wait for the tracking finalizer before recreating
Live verification of the previous commit showed the immediate retry STILL
stranded: deleting a finished Job leaves it terminating (job-tracking
finalizer), so the re-Create collided with the dying object and was
mapped to ErrAlreadyExists a second time. Poll until the name actually
frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck
finalizer ever outlives the budget. Fake-client tests pin both the
replace-finished and coalesce-in-flight branches.
2026-09-22 17:54:01 +08:00
Lemon-miaow 90ccbfede4 fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
2026-09-22 17:50:04 +08:00
Lemon-miaow fd0794d04d chore(deps): pgx v5.9.2, x/net v0.55.0, x/text v0.39.0
govulncheck flagged pgx v5.7.1 (GO-2026-5004, SQL-injection class) as
reachable from pgrepo.go, plus the old x/net and x/text. Bump all three
to the fixed versions; go vet/test stay green.
2026-09-22 17:50:04 +08:00
Lemon-miaow a05edc934c chore: gofmt the tree, clear staticcheck, add a CI gofmt gate
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
2026-09-22 17:49:54 +08:00
flyemoji a56c326518 chore: stop tracking the docs/changes ledger
docs/changes held 26 per-feature change notes and their index,
written while each feature was built. They were working records, not
documentation: they cite internal milestone numbers and plan steps,
several describe designs that changed before they shipped (the nano
note's config schema and a proxy plugin that was never built), and
nothing in the code, the build or the other docs refers to them. New
notes stopped being added a while ago; the commit messages carry that
record now.

The directory leaves the tree in this commit. Its contents stay
reachable in history, and the files were kept outside the repository
before removal. No code, build or test changes.
2026-09-22 15:02:47 +09:00
flyemoji 7b5b28c587 fix(bootstrap): keep the nano build toolchain under /opt/felis
The source build of the nano binary installed Go at /usr/local/go and
replaced whatever version was already there. On a host that also
builds other things, the operator's own toolchain was removed and
swapped for Felis's pinned version without a word.

GOROOT_DIR is now /opt/felis/go, next to the source, the Velocity
install and the JRE Felis already keeps under /opt/felis, and
install_go_toolchain creates the parent before unpacking. A host where
an earlier run put Go at /usr/local/go downloads it once more on the
next re-run and keeps the old tree untouched; removing it is the
operator's call. The harness now requires the toolchain directory to
be under /opt/felis.
2026-09-22 14:59:39 +09:00
flyemoji 0faec2b02a fix(bootstrap): open the nano port to the proxy alone
For a non-loopback bind, configure_nano_firewall opened the nano port
in firewalld to every source, while the summary told the operator to
restrict it to the proxy. hasJoined takes no token, so on a public
host that port is an auth relay anyone can point a proxy at, spending
this host's Mojang egress until Mojang rate-limits it and the
operator's own players stop getting in.

A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets
one rich rule that admits the port from that source only, ipv4 or
ipv6 by the address given. Without it, no port is opened and the
summary prints the rule to add. A re-run closes the port an earlier
installer opened to every source. A rule for a previous
FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand.
Hosts without firewalld are handled as before.

The value goes into the rule text, so it is checked up front for an
address with one prefix length and nothing else. firewalld's own
parser accepts both rule forms and refuses an ipv6 address under the
ipv4 family. The harness covers the rule for each family, the
closed-by-default case, the re-run cleanup, the loopback case and the
CIDR check.
2026-09-22 14:51:49 +09:00
flyemoji 99c31c1d4e fix(config): refuse plaintext auth-source urls to public hosts
An auth_source url could be http:// to any host. Anyone on the path
to a public root, or anyone who can spoof its DNS name, can then
answer hasJoined with a 200 and log in as any player of that source,
including a third-party account linked to staff. The player's IP also
travels in cleartext. Mojang logins are unaffected, since that source
is built in over https.

Config load now refuses http:// unless the host is localhost or a
loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12,
192.168/16, fc00::/7), so a root on the same host or the LAN still
works without TLS. The decision is made on the literal host because
nothing is resolved at load time, so a LAN root named by hostname
needs its IP address or https. The error says what to change.

The new test covers public names and addresses, link-local, 0.0.0.0
and the first address past 172.16/12 (all refused over http, all
accepted over https), and the loopback and private forms that stay
allowed. It fails on the old check.
2026-09-22 14:24:41 +09:00
flyemoji e9f74f3f0f fix(nano): stop trusting an expired free name while mojang is failing
When the premium-name lookup failed, isPremiumName fell back to any
cached answer, however old. An expired "free" is exactly the answer
that may have stopped being true: someone can buy the name after it
was last seen free. For as long as api.mojang.com kept failing (429,
5xx, a timeout), a third-party player holding that name kept it on
every reconnect, and the Velocity registry, keyed on the name, turned
its new owner away as already connected. A hostile source could drive
the host into Mojang's rate limit on purpose to hold names that way.

A failed lookup now always counts as taken, so the player is renamed
with the source's prefix. An expired "taken" already gave that answer,
so only the stale "free" case changes. The cost is cosmetic: during an
outage an ordinary third-party player may get a prefix they do not
need, and their data follows the UUID, not the name.

A new test gives the cache a free entry past its TTL and has Mojang
answer 429. It fails on the old fallback. The two comments that
described the fallback now describe the fail-closed rule.
2026-09-22 14:17:53 +09:00
flyemoji fa7b54f5ab fix(bootstrap): refuse an unbracketed ipv6 nano listen address
validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081
passed. Go refuses that form ("too many colons in address") and needs
[::1]:8081, so the unit crash-looped on every start. A host part that
contains a colon must now be in brackets.

With that, the bare ::1 pattern in nano_listen_is_loopback can no
longer match an address that gets this far, so it goes. [::1] stays.
The harness adds ::1:8081 to the refused addresses, and [::]:8081 and
:8081, both of which Go binds, to the accepted ones.
2026-09-22 14:08:11 +09:00
flyemoji 928a1fdfff docs(bootstrap): credit velocity, not authlib, with the hasjoined call
Two installer comments still said authlib makes the hasJoined request
and sends no token. Velocity reads -Dmojang.sessionserver and sends
the request itself. Comment text only.
2026-09-22 13:55:12 +09:00
flyemoji b58c20311c test(bootstrap): pin the nano listen default to loopback
nano_listen_is_loopback decides whether configure_nano_firewall opens
the port, and hasJoined takes no token. A default that does not
classify as loopback would make every fresh nano host a public auth
relay.

The harness now runs the classifier on four loopback binds and three
routable ones, and feeds it the default resolve_nano_listen applies on
a first install, with no operator value and no existing unit. Setting
that default to 0.0.0.0:8081 or :8081, or counting 0.0.0.0 as
loopback, now fails the harness. Test only.
2026-09-22 13:54:31 +09:00
flyemoji cf65ffdae5 fix(bootstrap): open up a nano-only config dir an older run left 0750
write_nano_config creates a missing /etc/felis as 0755, but it left
an existing one alone. On a nano-only host an older installer made
that directory with a bare mkdir -p, so under a root umask of 027 it
is 0750. The DynamicUser unit cannot search it, so felis-nano cannot
read its config, and a re-run stops at the service check instead of
repairing the directory.

An existing directory is now set to 0755 unless it holds the full
install's secrets.env or bootstrap.done. The full install locks the
directory to 0700 and writes secrets.env right after, so its directory
keeps that mode, and install_nano_service still reports the lockout
rather than this widening it. The mode cases run only where chmod
works; on a filesystem that ignores it the harness skips them.
2026-09-22 13:54:18 +09:00
flyemoji 2458ee1722 docs(bootstrap): say auth_source tags are permanent and order is trust
Both config templates the installer writes, the nano felis.toml and
the comment above [[auth_source]] in the generated felis tomls, now
state two things an operator editing the list needs to know.

A tag is hashed verbatim into every player UUID of its source, with no
case folding, so renaming it gives all of those players new UUIDs and
orphans their data, links and bans. The list is scanned in order and
the first source that validates wins, so order is trust, and a
compromised root has to be removed, not moved down. Comment text only.
2026-09-22 13:53:15 +09:00
flyemoji 6794e66c4d fix(bootstrap): fail a tokenless private clone instead of prompting
A source build against a private repository with no FELIS_GITHUB_TOKEN,
or a wrong one, made git ask for a username on /dev/tty, and a piped
install sat there waiting.

git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git
fails at once with "terminal prompts disabled". Both fetch_source
failures name FELIS_GITHUB_TOKEN in their message: the fresh clone,
and the fetch into an existing checkout, which had no message of its
own before.
2026-09-22 13:53:03 +09:00
flyemoji 34f73ba19f fix(bootstrap): detect a missing terminal by opening /dev/tty
prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which
never fires on Linux: /dev/tty is mode 0666 whether or not the process
has a controlling terminal, and only opening it fails. Without a
terminal the menu was printed, the read failed with "No such device or
address", and the default was taken by accident rather than by the
documented path.

The guard now opens /dev/tty in a subshell and takes the "no terminal
for a prompt" path when that fails.
2026-09-22 13:52:51 +09:00
flyemoji 0758b9c5d7 docs(bootstrap): pass tunables on the sudo line, not by export
The header said to export tunables before running, but its own
`curl ... | sudo bash` entrypoint resets the environment, so an
exported FELIS_INSTALL_MODE or FELIS_NANO_LISTEN never reached the
installer. The header now shows the two forms that do arrive: the
variable named on the sudo line, or export followed by sudo -E.

The nano summary's hint for a proxy on another machine now prints a
sudo line that can be pasted as is, instead of "re-run with
FELIS_NANO_LISTEN=...". Comment and log text only.
2026-09-22 13:52:40 +09:00
flyemoji 02c079c893 fix(bootstrap): verify the go toolchain tarball against a pinned digest
install_go_toolchain downloaded the tarball to a fixed /tmp name and
unpacked it into /usr/local as root, with no digest check. Another
local user could plant that file first, and nothing would notice a
tampered download.

The tarball is now staged in a mktemp -d directory that the exit
cleanup removes, and its sha256 must match before the old toolchain is
touched, so a refusal leaves the host as it was. The default 1.26.4
carries pinned amd64 and arm64 digests next to its version; they are
the ones https://go.dev/dl/?mode=json&include=all publishes. Any other
FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in
the header, because no pin can cover a version chosen at run time.
Where and which version gets installed is unchanged.
2026-09-22 13:52:29 +09:00
flyemoji 3b0fc7a3e0 fix(bootstrap): print the address nano binds in the install summary
summary_nano printed the node's primary IP for every non-loopback bind
and 127.0.0.1 for every loopback one. A bind to a second private
address, or to [::1], handed the operator a hasJoined URL that nothing
listens on.

The host is now the part of FELIS_NANO_LISTEN before the last ':'. The
node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]),
which names no address a proxy could dial. The loopback and
public-bind notes are unchanged.
2026-09-22 13:52:18 +09:00
flyemoji 404d1172a6 fix(bootstrap): refuse a nano listen address without a usable port
FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in
the firewall while nano bound nothing, a bare 127.0.0.1 printed
http://127.0.0.1:127.0.0.1/... in the summary, and the unit
crash-looped either way.

validate_settings now requires a ':' and a decimal port of 1-65535
after the last one. It runs after resolve_nano_listen, so an address
read back from an existing unit is checked too, and the default always
passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
2026-09-22 13:52:08 +09:00
flyemoji 3918a4b11a fix(bootstrap): install only the full control plane under felis setup
prompt_install_mode also runs inside felis setup. Setup then goes on
to the Owner and edge setup, which need the control plane, so choosing
nano there always ended in a setup error.

Under felis setup the mode is now full before any prompt or default is
considered, and an explicit FELIS_INSTALL_MODE=nano stops with a
message pointing at deploy/bootstrap.sh. That leaves the felis setup
branch of acquire_nano_binary unreachable, so it goes.
install_embedded_binary stays, since the full install still uses it.
2026-09-22 13:52:00 +09:00
flyemoji 515c4a6496 fix(bootstrap): keep a nano host's listen address and mode on re-run
Re-running the installer is how a nano host updates. That re-run reset
FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost
its endpoint and every login through it failed. It also offered the
full control plane as the default, which on a nano host means k3s and
Postgres nobody asked for.

The listen address is now settled by resolve_nano_listen, the first
step of main, so the later checks see the result. The operator's value
wins, then the -listen argument of the installed felis-nano unit, then
loopback. The install mode defaults to nano, at the prompt and without
a terminal, when the felis-nano unit exists and the full install's
bootstrap.done marker does not. Only the full install writes that
marker.

The harness reads back the unit it wrote earlier, and checks the mode
default on a nano-only host, a host with the full install, and a fresh
host.
2026-09-22 13:51:52 +09:00
flyemoji 17b4396460 fix(bootstrap): carry auth_source tables with spaced or quoted headers
A re-run copies the operator's [[auth_source]] tables from the existing
felis toml into the new one. The awk program that finds them matched
only the literal header [[auth_source]], so a table written as
[[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid
TOML, was taken for some other section and dropped from the config.

Each section header now decides afresh whether it opens an auth_source
table, through one regex that allows inner whitespace and a single- or
double-quoted key. The single quote is spelled \047, which gawk and
mawk both honour inside a bracket expression. The harness carries each
spelling and checks that the table still stops at the next section.
2026-09-22 13:51:45 +09:00
flyemoji c2a5645c55 fix: keep internal section numbers out of runtime messages
Four messages that reach an operator or an API client cited sections
of a specification nobody outside the project can read:

- the unimplemented archive store error from config load
- the running-server cap refusal, from both the user wake and the
  internal wake
- the missing memory ceiling guard, in the API and in felis apply

The references are gone and the wording is otherwise unchanged. Each
message still says what went wrong and, where there is one, what to
do about it. The test for the archive store message checks for the
tarLocal remediation, which is still there.
2026-09-22 13:44:38 +09:00
flyemoji 9ee8c48fff docs(openapi): list every answer hasjoined gives
The hasJoined contract listed only 200 and 204 and named authlib as
the caller. The handler now answers four more ways, and a proxy
operator reading the contract could not tell a refused login from a
down source.

- 204 also covers a missing or oversized parameter (no source is
  asked), a third-party name that is not a legal Minecraft username,
  and an identity id that does not parse.
- 400 for a request that declares a body. There is no response body,
  and the connection is closed.
- 500 when the bar-list lookup fails, with the usual error body.
- 503 when no source validated and at least one failed, since that
  source's player may be the one logging in.

The three query parameters now carry the 64-byte cap. The profile name
says a third-party player holding a registered Mojang name gets it
back prefixed and cut to 16 characters. The description names
Velocity, drops the "thin login hook" that does not exist, and says
that a non-200, non-204 answer makes Velocity report the auth servers
as down.
2026-09-22 13:43:08 +09:00
flyemoji 1ebd73a309 docs(nano): describe the hasjoined path as it works
The comments around hasJoined still described an authlib client that
is not in the path. Velocity reads -Dmojang.sessionserver and sends
the request itself, and it turns a 204 into its online-mode-only kick,
not authlib's "failed to verify username". The route comment in api.go
also offered "a thin login hook" as an alternative that does not
exist.

Other comments had drifted from the code:

- The [[auth_source]] doc said an empty list ships the multiplexer
  off. Mojang is always prepended, so an empty list means Mojang is
  the only source.
- The premium-name cache said Mojang does not recycle names. A name
  frees up when its owner renames away. The day-long "taken" TTL still
  holds, because a stale "taken" costs a third-party player only a
  prefix.
- The cache bound claimed entries come only from players who
  authenticated somewhere. Any third-party source that validates a
  login adds one, so a hostile source can force the map to clear. That
  costs repeat lookups, or a fail-closed prefix while Mojang is
  unreachable, never an identity.

The rewrite rationale now states what it costs a backend operator. A
chat-session key that a third-party source signed over its native UUID
cannot verify against the canonical UUID, so chat from those players
can only be accepted unsigned.

In the tests, comments that repeated their subtest names are gone.
2026-09-22 13:42:03 +09:00
flyemoji 59ec23d4a8 test(nano): cover the nano delivery path and its loopback default
felis nano serves the same hasJoined handler as felis api, but behind
nanoStubRepo, which implements only the bar-list lookup and embeds a nil
Repo for everything else. Only the full-api path was tested, against a
complete fake store, so a second store call added to handleHasJoined
would pass CI and panic on every nano login. The loopback default of
-listen, the one thing keeping nano from being an open auth relay, was
not pinned either.

The default moves into a nanoDefaultListen constant, and two tests
cover the path. One serves a login through api.HasJoinedHandler with
nanoStubRepo and a fake identity source and expects the profile back.
The other requires the default to parse as a loopback IP. Taking the
bar-list method off the stub makes the first panic on the nil Repo;
defaulting to 0.0.0.0:8081 or :8081 fails the second.
2026-09-22 13:37:51 +09:00
flyemoji e0ad78af98 test(config): make the identity-key test fail when the key is accepted
TestLoadRejectsAuthSourceIdentityKey is the guard against a config line
identity = true making a third-party source's UUIDs trusted as-is. Its
fixture had no prefix, so Load failed on the prefix rule and the test
passed on that error. With the unknown-key check in decodeConfig
disabled, the test still passed.

The fixture now carries a valid prefix, the error must mention unknown
keys and identity, and LoadNano is checked alongside Load. With the
unknown-key check disabled, both loaders now fail the test; the old
version of the test passes against the same change.
2026-09-22 13:36:29 +09:00
flyemoji 30b4e1dfb2 test(nano): cover the bar-list error, bad identity id and ip relay
Three paths in handleHasJoined had no test that fails when they break:

- A bar-list lookup error answers 500. Logging it and carrying on would
  admit a reclaimed squatter during a database outage.
- An identity (Mojang) id that does not parse answers 204. Ignoring the
  parse error would emit the nil UUID for every such login, so they all
  share one player's data.
- The ip parameter is relayed to each source. Dropping it turns off the
  sources' check that the session is used from the player's own address.

One subtest each. Mutants that ignore the bar-list error, ignore the id
parse error, or stop appending ip each fail their subtest.
2026-09-22 13:35:50 +09:00
flyemoji 942e9a5ff8 test(nano): cover the premium-name cache rules
isPremiumName decides on every third-party login whether the player
keeps their name, and none of its rules had a test that fails when the
rule breaks: treating a 429 or 5xx from api.mojang.com as "free",
swapping the free and taken TTLs, flipping the freshness comparison,
answering "free" from an expired taken entry during an outage, or
dropping the clear-at-4096 bound. Each of those leaves a squatter
holding a name its owner has bought, or grows the cache without limit,
with CI green.

TestPremiumNameCache drives isPremiumName against a stub that answers
with a fixed status and counts lookups, and seeds cache entries at chosen
ages. Five mutants of handlers_hasjoined.go, one per rule above, each
fail at least one subtest. It does not test an expired "free" entry
during an outage; what that case should return is still open.
2026-09-22 13:34:43 +09:00
flyemoji 3f7274d29f test(nano): pin the auth namespace and one rewritten uuid as literals
The rewrite test computed its expected UUID from felisAuthNS itself, so
a change to the namespace seed moved both sides together and still
passed. Such a change gives every third-party player a new UUID on next
login, orphaning their playerdata and account links and letting any
squatter barred by the old UUID back in.

The test now also compares felisAuthNS and the rewrite of
littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e-
9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were
computed independently with Python's uuid5/uuid3, not read back from
the code. Prefixing the seed with https:// fails the test.
2026-09-22 13:32:45 +09:00
flyemoji a7fe525bfc test(api): keep the package's tests off the live mojang profile api
mojangProfileAPI defaults to https://api.mojang.com, and only the tests
that call stubMojangNames or setProfileAPI swap it out. A new test that
reaches a third-party login without doing so would query the real
service: its result then depends on network access and on whether
someone owns the name that day, and the shared premium cache can carry
that answer into later tests.

A TestMain now points the lookup at an address nothing listens on
before any test runs, so a forgotten stub always takes the same
fail-closed path. Tests that stub it restore this address, not the live
one, when they finish.
2026-09-22 13:31:53 +09:00
flyemoji 8e9c8ca4e6 fix(nano): say that [server] listen is ignored instead of defaulting it
LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset,
and a test pinned that value, but felis nano never reads it: it binds
the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An
operator moving nano off loopback by writing [server] listen in its
config got connection refused from the proxy and no hint that the key
did nothing.

LoadNano no longer sets the default, and nano prints a line naming the
ignored value and the address it actually binds whenever the key is
set. It is a warning rather than a load error so a full felis.toml
copied onto a nano host keeps starting. The assertion that pinned the
unused default is removed along with it.

The new test runs cmdNano against a config that sets [server] listen
and one that does not, with an unbindable -listen so it returns after
loading. The first must warn and the second must not; with the old
default restored, the second prints a warning about 0.0.0.0:8080.
2026-09-22 13:30:23 +09:00
flyemoji 1d6c73007e fix(nano): drain in-flight logins on shutdown
The installer and the config template tell the operator to run
systemctl restart felis-nano after editing the source list. nano had no
signal handling, so SIGTERM killed it mid-request: a login waiting on an
upstream had its connection reset, and Velocity disconnected that
player with "authentication servers are down". felis api already drains
on shutdown; nano did not.

nano now listens itself, serves until SIGINT or SIGTERM, then shuts the
server down gracefully with a 30-second limit. That outlasts the source
scan of any realistic list, at five seconds per source, and stays well
inside systemd's default 90-second stop timeout.

The new test holds a request inside the handler, cancels the serve
context, and checks that serveNano is still running 200 ms later, that
the held request then gets its answer, and that serveNano returns 0.
Replacing the graceful shutdown with Close fails it.
2026-09-22 13:28:59 +09:00
flyemoji fa3eda5228 fix(nano): quote and cap the request log line
felis nano logged every request with the raw RequestURI and %s. That
text is the caller's: a right-to-left override reordered the line as
displayed, an invalid UTF-8 byte made journald store the entry as a
binary blob that journalctl -f shows as "[N blob data]", and a query
near net/http's one-megabyte limit became a one-megabyte log line.

The URI is now capped at 256 bytes, several times a real hasJoined
query, and printed with %q, so control, bidi and invalid bytes appear
escaped. The handler assembly moved into nanoHandler so the logged
handler can be tested on its own; cmdNano serves it unchanged.

The new test sends a query carrying U+202E, a 0x9b byte and 4 KiB of
padding, and expects a valid UTF-8 line with the override escaped and
no more than twice the cap. Restoring the old unquoted line fails it.
2026-09-22 13:27:28 +09:00
flyemoji 1905cac950 fix(config): refuse auth-source tags padded with whitespace
A third-party player's UUID is hashed from the source tag byte for byte,
so the tag is a permanent namespace: change it and every player of that
source comes back as someone new, with their playerdata, permissions,
account links and reclaim bans left behind. Nothing said so, and a tag
with a stray leading or trailing space, which nobody can see in the
file, loaded as a brand new namespace.

Such a tag is now rejected at load, and the AuthSourceConfig doc states
that the tag is permanent, case included. The charset stays otherwise
open: tightening it would force existing installs to rename, which is
the very thing that rekeys their players.

The new test loads a tag with a trailing space, a leading space and a
trailing tab through LoadNano; all three loaded before this change.
2026-09-22 13:24:23 +09:00
flyemoji 72a2750461 fix(config): refuse mojang as an auth-source tag
Mojang is prepended in code as the first, identity source, and the
config templates say not to list it. Nothing enforced that. A listed
tag = "mojang" loaded, and nano's startup list printed it as if Mojang
had been pointed at that url, while the real Mojang was still asked
first. The listed entry was a separate third-party source: asked again
on every login that got past Mojang, adding up to five seconds when its
url was Mojang's own and it answered 204 each time.

Any case of "mojang" is now rejected at load with a message saying
Mojang is built in and must not be listed. The duplicate-tag check could
not catch this because the built-in source never passes through it.

The new test loads "mojang" and "Mojang" through LoadNano; both loaded
before this change.
2026-09-22 13:23:37 +09:00
flyemoji 2c74080b78 fix(config): reject auth-source urls the resolver cannot query
The url check only looked for an http:// or https:// prefix. Several
shapes passed it and then left the source dead at login time: no host
("https://"), a bad port, surrounding whitespace (sent as %20 and
answered 404), and any query or fragment. The resolver appends
"?username=…&serverId=…" to the url as a string, so an existing query
swallows those parameters and a fragment hides them from the request
entirely. Each loaded green, and every login from that source failed.

The url is now parsed and must be http or https with a host, no query,
no fragment and no surrounding whitespace. Load and LoadNano share the
check. The shipped LittleSkin default and plain http:// endpoints, such
as a same-host root on loopback, still load.

The new test feeds each rejected shape to LoadNano. Against the previous
prefix check, six of the seven load; only ftp:// was refused.
2026-09-22 13:23:02 +09:00
flyemoji 3338d6f0fe fix(nano): drop oversized hasJoined parameters before asking sources
username, serverId and ip were forwarded to every configured source at
whatever length the caller sent, up to the megabyte net/http allows in a
request line. Velocity never sends more than a 16-character name, a
41-character signed SHA-1 serverId and a textual IP address, so only a
direct caller reaches those sizes, and each such request cost one
oversized upstream call per source.

Any of the three over 64 bytes is now answered 204 before a source is
asked, the same as a missing username or serverId. 64 bytes still
leaves room for a 16-character name in multi-byte UTF-8.

The subtest behind this points a source that validates anything at the
handler and sends missing and oversized fields, expecting 204 and zero
upstream requests, then a well-formed login that gets 200. It replaces
the old missing-username case, whose only source was unreachable, so
the test passed even with the guard removed. Dropping the length check
now fails it on the long username; dropping the whole guard fails it on
the first missing field.
2026-09-22 13:21:43 +09:00
flyemoji a4779186a4 fix(nano): refuse hasJoined requests that declare a body
A GET to hasJoined with a Content-Length and no body held its
connection indefinitely. The handler returned, but net/http tries to
drain an unread body before it writes the answer, and nothing bounds
that wait: ReadHeaderTimeout ends with the headers. One such request
per socket pins a goroutine and a descriptor on nano or on felis-api's
internal face.

Velocity never sends a body, so any request that declares one, including
a chunked one, now gets a 400 with Connection: close, which skips the
drain and releases the connection once the answer is written.

The new subtest writes that request over a raw socket and waits three
seconds for an answer. Before the change it times out with no response
at all; now it reads a 400 marked close.
2026-09-22 13:19:55 +09:00
flyemoji ff81295aa9 fix(nano): report failing sources instead of treating them as a no
A source that timed out, answered 5xx or 429, redirected, or sent a 200
without a usable profile was skipped exactly like one that answered 204.
With nobody else validating, the login got a 204 and Velocity told the
player their account is offline-mode. Nothing was logged, so a dead or
mistyped source URL, or an http:// root that now redirects to https since
redirects stopped being followed, failed every one of its players with
no trace.

Each such failure now logs the source tag and the cause; for a 3xx it
names the Location to configure instead. When no source validates and at
least one failed, the answer is 503, which Velocity reports as the auth
servers being down and logs with the status. A source answering 204 is
still a plain no, and a validating source still wins regardless of
failures before it.

The new subtest puts a 503 source, a redirecting source and an
unreachable one each behind a Mojang that answers 204, and expects 503.
Against the previous handler every case returns 204.
2026-09-22 13:18:33 +09:00
flyemoji a0f54df2a6 fix(bootstrap): fail the nano install when the unit does not stay up
install_nano_service printed "enabled and started" straight after
systemctl restart, which returns as soon as the process is forked. An
upgrade that keeps an old felis.toml the new binary rejects (an
[[auth_source]] without a prefix, say) left the unit crash-looping in
auto-restart while the installer reported success, and every login
through the proxy failed.

The install now waits two seconds and asks systemctl is-active. A unit
that exited is in "activating (auto-restart)", which is-active does not
count as active; on real systemd a unit whose process exits 1 under
Restart=on-failure reads activating/auto-restart and is-active returns
non-zero, while a running one reads active/running and returns 0. On
failure the install prints the unit's last 20 journal lines and stops.
This also surfaces a nano unit locked out of an existing 0700 /etc/felis.

The harness runs the extracted function with systemctl stubbed both
ways. Without the check, the dead-unit cases fail.
2026-09-22 13:05:42 +09:00
flyemoji 26f685be0e fix(bootstrap): create the nano config dir world-searchable
felis-nano runs as a systemd DynamicUser, so it can read
/etc/felis/felis.toml only if others may search /etc/felis.
write_nano_config made the directory with a bare mkdir -p, which takes
its mode from root's umask. On a host hardened to umask 027 that is
0750: nano exits on "permission denied", the unit restarts every five
seconds, and no login gets through.

A missing directory is now created 0755 explicitly. An existing one
keeps its mode, because the full install sets it to 0700 to protect its
secrets and widening that from the nano path would expose them. A nano
unit locked out that way is left for the install to report.

The harness runs the extracted function under umask 027 and checks both
cases. Reverting to the bare mkdir fails the first; an unconditional
chmod 0755 fails the second. The mode checks skip on filesystems that
ignore chmod, such as Git Bash on NTFS.
2026-09-22 13:04:34 +09:00
flyemoji 1dd62a9bdc fix(nano): cap upstream response headers at 16 KiB
The hasJoined and name-lookup clients limited the body to 64 KiB but
left headers at the transport default of 1 MiB. A configured root could
answer with a megabyte of headers and stall the body, holding a few MiB
of heap per in-flight login for the full five seconds; enough parallel
logins take down the host, and every source's logins with it.

Both clients now share a transport with MaxResponseHeaderBytes set to
16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends
338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over
the cap fails the request and the resolver moves on to the next one.

The new subtest puts a source with 64 KiB of headers and a valid profile
ahead of an honest one and expects the honest player. Without the cap
the padded source wins.
2026-09-22 13:02:48 +09:00
flyemoji 2180e77cf5 chore: drop tool-name markers from source comments
Seventeen comments opened with a tag naming the tool that wrote them.
The tag goes and each comment keeps its reasoning, now starting as a
plain sentence. None of the reasoning changes.

The AGENTS.md note in .gitignore drops the story of how the file got
into the tree and keeps the one fact a reader needs: its advice to run
go fmt is destructive on this CRLF working tree.

Comments only; no code, build or test changes.
2026-09-22 12:57:19 +09:00
flyemoji 4e98ae6e56 fix(config): reject auth-source tags that contain a colon
A third-party player's canonical UUID is UUIDv3 over tag+":"+nativeID,
and the native id is whatever the source answers. Tags were only
checked for being non-empty and unique, so both "guild" and "guild:eu"
could be configured. The "guild" root could then answer hasJoined with
id "eu:X" and receive exactly the UUID of "guild:eu"'s player X, along
with their playerdata, permissions and account links. Real native ids
are 32 hex digits, so only the shorter tag's source can do this, and
only when the operator has configured such a pair; when they have, it
is a full impersonation.

Reject a ':' in a tag at load. With colon-free tags the join is
unambiguous: two different (tag, id) pairs can no longer produce the
same input, since equal inputs force equal tags and duplicate tags are
already refused. The tag is deliberately not narrowed any further.
It is a permanent UUID namespace, and forcing an operator to rename a
tag that has no colon would move every one of its players to a new
UUID. The hash input and the native id are left exactly as they were,
so no existing player's UUID changes.

Load and LoadNano share validateAuthSources; the new test runs both
against the guild / guild:eu pair and fails on the previous config.go.
2026-09-22 12:53:35 +09:00
flyemoji 1976fca809 fix(nano): always relay properties as an array
sessionProfile tagged properties with omitempty, so an upstream answer
of "properties": [] (or null, or no key at all) reached Velocity with
no properties key. A Yggdrasil root may legitimately answer that way for
a player without a skin. Velocity 3.5.1's GameProfile deserializer
passes the missing key on as null and ImmutableList.copyOf throws, so
that player hangs at login with nothing logged, even though the same
answer sent straight to Velocity is accepted. Mojang always sends
textures, which is why the premium path and the hardware runs never hit
it.

Drop omitempty and replace a nil slice with an empty one before the
response is written. Removing omitempty alone is not enough: a nil
slice marshals as null, which Velocity rejects the same way.

The new subtest feeds the relay [], null and a missing key and expects
"properties":[] every time. The previous handler fails all three.
2026-09-22 12:52:04 +09:00
flyemoji 28d3638952 fix(nano): stop following redirects from upstream Yggdrasil roots
authHTTPClient kept net/http's default redirect policy, so a configured
third-party root that answered hasJoined with a 3xx made this host fetch
whatever URL it named, up to ten hops. That is a blind SSRF into
anything the host can reach, and it includes the multiplexer's own
listener: a root that redirects back to /session/minecraft/hasJoined
re-enters the handler, which queries Mojang and every source again and
gets redirected again, until the outer 5s client timeout fires. With a
50ms Mojang stub, one login produced 86 nested handler calls and 86
Mojang requests from this host's egress IP. The loopback default does
not help, because the redirect target is resolved from this host.

Return the 3xx as the response instead. resolveHasJoined already skips
any non-200 answer and closes its body, so a redirecting source is now
treated like one that is down, and the next source gets its turn. The
same probe now makes one handler call and one Mojang request.
Neither Mojang's nor LittleSkin's hasJoined redirects.

The new subtest puts a redirecting root ahead of an honest one and
checks that the redirect target is never contacted and the honest
source's player is returned. The pre-fix handler fails it.
2026-09-22 12:50:56 +09:00
flyemoji 07bafebf0d fix(bootstrap): keep the operator's auth sources across re-runs
write_felis_toml regenerates felis.host.toml and felis.pod.toml with a
wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin
block in that heredoc. Re-running the installer, which is also what
`felis setup` does, threw away any edit to the list: a root the operator
added stopped admitting logins, and a root they removed came back. The
generated comment invited exactly that edit.

Carry the tables forward the way [smtp] already is: read every
[[auth_source]] table from the existing felis.host.toml (falling back to
felis.pod.toml) and emit the LittleSkin default only when there is no
earlier file at all. An earlier file with no tables stays empty, because
that is a Mojang-only server rather than a missing value; felis-api now
treats an empty list that way.

The file header and the comment above the list now say what survives a
re-run, and point at felis.host.toml, which is what the next run reads.

bootstrap_test.sh extracts the new function from bootstrap.sh and checks
the fresh-install default, an operator's own table carried without the
default or the following section, an empty list staying empty, and the
indented form the setup TUI writes. It passes under dash with gawk and
with mawk; forcing the function to always return the default fails five
of the new cases.
2026-09-22 12:49:25 +09:00
flyemoji 8fe255e38f fix(api): relay Mojang logins when no auth source is configured
felis api wired the hasJoined multiplexer only when felis.toml had at
least one [[auth_source]]. With none, the source list stayed nil and
every hasJoined answer was a 204. That was harmless while nothing
pointed at the route, but the installer now starts Velocity with
-Dmojang.sessionserver aimed at felis-api unconditionally, and the
generated felis.toml tells the operator to delete the LittleSkin block
for a Mojang-only server. Doing exactly that turned every login away,
premium accounts included, and felis-api logged nothing about it.

Always build the list through authSourcesFromConfig, which prepends
Mojang in code, so an empty config is a Mojang-only relay. felis nano
already behaves this way with the same file.

The new test pins authSourcesFromConfig itself: Mojang first, the only
Identity source, and still present when nothing is configured. Marking
a configured source Identity makes it fail. The call site in cmdAPI is
now a single unconditional assignment and has no unit test of its own.
2026-09-22 12:46:00 +09:00
flyemoji 800a9042a1 test: use placeholder domains in setup and system-server tests
Three tests carried the maintainer's production root domain, a personal
mailbox and the public IP of a live demo host as fixture values. None of
them needs the value to be real: the re-domain test only needs two
different roots, and the setup flow only needs a well-formed address.

Swap them for the placeholders the rest of the suite already uses
(mc.example.net, [email protected]), and move the "before" root in the
re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737
documentation range, so the stale install the test models still has an
IP-derived hostname, which is the case the refresh exists for.
2026-09-22 12:43:54 +09:00
flyemoji d9246ddae6 feat(bootstrap): verify Paper and Velocity jars against Fill's digest 2026-08-04 17:36:47 +09:00
flyemoji 3f2b28d0ec fix(bootstrap): ship deploy/paper in the embedded game-stack tar 2026-08-04 16:34:09 +09:00
flyemoji 587f183191 Merge pull request #19 from MliroLirrorsIngenuity/chore/issue-sweep
chore: work the tracker items that need no cluster
2026-07-29 01:52:15 +09:00
232 changed files with 15450 additions and 3175 deletions

No files matched your search

+11
View File
@@ -31,3 +31,14 @@ Dockerfile
*.key *.key
felis felis
felis.exe felis.exe
# macOS materializes extended attributes as ._<name> sidecars (BSD tar uploads,
# Finder copies, network volumes) and leaves .DS_Store behind. Neither is
# source, and one is actively harmful: a ._*.sql beside the migrations is
# //go:embed-ed into the binary and makes every `felis migrate` fail
# ("non-numeric version") — observed live on a Mac-staged tree. Same exposure
# for any tree the other //go:embed patterns walk (deploy/, plugins/).
._*
**/._*
.DS_Store
**/.DS_Store
+45
View File
@@ -36,6 +36,12 @@ jobs:
with: with:
go-version-file: go.mod go-version-file: go.mod
- name: gofmt
run: |
unformatted=$(gofmt -l .)
if [ -n "$unformatted" ]; then
echo "gofmt needed on:"; echo "$unformatted"; exit 1
fi
- run: go vet ./... - run: go vet ./...
- run: go test ./... - run: go test ./...
@@ -92,3 +98,42 @@ jobs:
- run: npm run typecheck - run: npm run typecheck
working-directory: panel working-directory: panel
plugins:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
# The other jobs never touch the Java layer: the plugin jars were only ever
# compiled by bootstrap on a live host, and the three test mains under
# plugins/*/test were run by hand. JDK 21 plus the Gradle major the plugin
# Dockerfiles pin (8.14) is that same toolchain, in CI.
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21'
- uses: gradle/actions/setup-gradle@v4
with:
gradle-version: '8.14'
- run: bash plugins/test.sh
mods:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
# The three loader mods (Minecraft 1.20.1 / 1.20.4, Java-17 lines) compile
# through their vendored Gradle wrappers, which fetch their own Gradle. Until
# this job nothing ever built them: no install path touches them, and their
# gradlew scripts were committed without the exec bit, so the README's
# one-liners failed on a fresh clone.
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '17'
- uses: gradle/actions/setup-gradle@v4
- run: bash plugins/test-mods.sh
+25
View File
@@ -40,6 +40,31 @@ jobs:
- run: go vet ./... - run: go vet ./...
- run: go test ./... - run: go test ./...
# The same reason, for the Java layer the binary EMBEDS: the release asset is
# the tree's plugin sources (bootstrap_asset.go), and a tag whose plugins don't
# compile turns every install of that release into a failed bootstrap. JDK 21
# gates the install-time plugins + codec/invite tests; JDK 17 gates the loader
# mods (their vendored wrappers fetch their own Gradle).
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21'
- uses: gradle/actions/setup-gradle@v4
with:
gradle-version: '8.14'
- run: bash plugins/test.sh
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '17'
- uses: gradle/actions/setup-gradle@v4
- run: bash plugins/test-mods.sh
# Both architectures, because bootstrap's default release channel DOWNLOADS these # Both architectures, because bootstrap's default release channel DOWNLOADS these
# rather than compiling on the target host — an arm64 host with no asset silently # rather than compiling on the target host — an arm64 host with no asset silently
# falls back to a slow source build. Neither stage is emulated: the Dockerfile pins # falls back to a slow source build. Neither stage is emulated: the Dockerfile pins
+2 -3
View File
@@ -41,9 +41,8 @@ plugins/*/bin/
# ---- Local agent / loop state ---- # ---- Local agent / loop state ----
.claude/ .claude/
# Autohand-generated agent guide — kept on disk for local tooling, never tracked. # Generated tooling guide, kept on disk for local use and never tracked. Its advice
# It rode in via 0c1cc59, claims precedence over CLAUDE.md, and tells agents to # to run `go fmt` is destructive on this CRLF working tree.
# run `go fmt` (destructive on this CRLF working tree).
AGENTS.md AGENTS.md
# ---- Internal planning & design docs (excluded from the public remote per # ---- Internal planning & design docs (excluded from the public remote per
+980
View File
@@ -0,0 +1,980 @@
# Felis 生产就绪审计 — 2026-09-22(真机 E2E + 混沌)
分支:全部修复**逐 commit 直接推 `main`**(不再用功能分支)。remote = `[email protected]:FelisMC/Felis.git`(2026-09-23 起;旧 `MliroLirrorsIngenuity/Felis` 仅靠 GitHub 301 兜住)。环境:CentOS Stream 9 / aarch64 / k3s v1.36.4,
IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:fede:69ec`),
面板经 `socat TCP6:443 → 127.0.0.1:30443` 中继(手动启动,重启 VM 后需重开)。
## 已修复并验证(分支内)
| # | 缺陷 | 证据 | 修复 |
|---|------|------|------|
| 1 | **失败恢复后重试被静默吞掉**:restore Job 固定名 + `ErrAlreadyExists` 一律当"幂等成功";失败 Job 占名 10 分钟(TTL),期间重试返回 202 `restoring` 但什么都不跑 | 真机:坏 ref 制造失败 → 立刻合法重试 → Job 原地不动、无新 Pod | `k8sjobs.go`:撞名时检查已完成(成功/失败)→ 删除+**等 finalizer 释放**(有界 10s)+ 重建;进行中仍幂等吸收。单测 3 个。**真机复验:重试 4s 完成恢复** |
| 2 | **备份 Job 全部 FailedMount**:Job 在 minecraft 命名空间挂 `felis-config`,而 bootstrap 只在 felis 命名空间创建该 Secret | 事件:`MountVolume.SetUp failed: secret "felis-config" not found` | `felis setup` 用既有 `ensureSecretReplica` 把 felis-config(`felis.toml`) 复制到 minecraft;VM 上手工复制后备份成功(167MB 归档) |
| 3 | RBAC 缺 `jobs:get/delete`(修复 #1 需要) | Role 检查 | `APIMinecraftRole` jobs → create/get/delete,测试同步更新 |
| 4 | gofmt 9 文件漂移 + CI 无 gofmt 门禁;staticcheck 9 处 | 基线扫描 | 全部修复;CI 加 gofmt job |
| 5 | 依赖漏洞:pgx v5.7.1(GO-2026-5004,可达 pgrepo.go)、x/net、x/text | govulncheck | 升 pgx v5.9.2 / x/net v0.55.0 / x/text v0.39.0 |
| 16 | **`ConsumeLoginEmailOTP` PG 实现与接口契约漂移**:契约/`fakeRepo` 说「同 VerifyEmailOTP 生命周期(扣尝试/锁定/ErrOTPInvalid/ErrOTPLocked)」,PG 却是单条 UPDATE+`ErrNotFound` → 错码/重放/过期在 login 门、op-login finish、migration confirm 三个入口全部 500;且尝试次数永不累计、`otpMaxAttempts` 锁定失效 | 真机:login 门重放正确码 → **HTTP 500**;修复前错码不扣次。修复后复测:5 次错码 400 且 attempts=5(正确码因锁定也 400、码未消费)、新码可用、重放 400 | `pgrepo.go` 改置为 VerifyEmailOTP 同构事务(FOR UPDATE、先锁后比、mismatch 扣次、match 消费),无 users 写副作用。提交 `52549f7` |
| 17 | **`ListPendingOpLogins` PG 少列**:接口注释承诺「joined to its staff username」,handler 输出 `username`/`created_at`,fake 正确填充;PG SQL 未 JOIN 也未取 `created_at` → 真机 pending 列表 username 为空、created_at 为 `0001-01-01` | 真机 internal `/op-login/pending` 响应 | `pgrepo.go` SQL 改为 JOIN users + 取 created_at。提交见分支 |
| 18 | **绑定码并发兑换 500**:`RedeemPlayerBindCode` 无 `FOR UPDATE`(同文件 `VerifyLinkCode` 有),且裸 INSERT。6 路并发同码兑换 → 3×HTTP 500(`users_username_key` 唯一冲突)+ 1×400 + 2×200;跨码并发同 UUID 同样会撞 | 真机四组并发测试 + API 日志 `unmapped error ... duplicate key` | 同码:码行 `FOR UPDATE`(输家干净地 400 invalid_code);跨码:`INSERT users ... ON CONFLICT (username) DO NOTHING`+重读、`account_links ON CONFLICT (mc_uuid) DO NOTHING`(两路汇聚同一 user)。复测:同码×6=1×200+5×400、双码×2=2×200 同 ID、DB 干净、日志零 unmapped |
| 19 | **reaper CronJob 渲染位置错误,永远无法调度**:`felis manifests` 把 CronJob 渲染到控制 ns,却引用 minecraft ns 的备份 PVC(Pod 不能跨 ns 挂 PVC:真机 `FailedScheduling: persistentvolumeclaim "felis-backups" not found`);修正 ns 后又发现 ServerAccount 也不能跨 ns 使用(`serviceaccount "felis-reaper" not found`)。而 reaper 的 Role/RoleBinding 本就在 minecraft ns | 真机三层取证(PVC/SA/调度)| CronJob 与 SA、RoleBinding subject 全部移到 `MinecraftNamespace`(提交 `e4f2cff`+`c839454`)。**修复后完整演练通过**(见下) |
| 6 | **默认安装无备份能力 + 回收链路不可达/不可用**:(a) 没有任何环节渲染归档 PVC,`FELIS_BACKUP_PVC` 永远为空 → backup/restore 恒 503;(b) bootstrap 从不下传 reaper 三旗标 → 官方安装路径根本无法启用回收;(c) 即便启用,stock k3s 的 `<root>/<pvc>` 布局不存在,且 k3s storage root 是 `0700 root:root`、reaper pod 以 uid 1000 运行 → 真机 `lstat /worlds/…: permission denied`(fail-closed 跳过,但纯空转);(d) README 口径(“定时备份”)不实 | 真机全链路:默认渲染 PVC+env → 建服→marker→备份 202→jobs 端点 running→succeeded→改 marker→restore→读回原值 ✅;再建 resolvecheck 世界(20d idle,marker)→ CronJob 手动 Job:`evaluated=2 reaped=1`,归档含 marker、PVC+宿主目录回收、`world_backups` 得 `inactive_15d` 行、servers 行/CR 保留 ✅ | `platform/workloads.go` 渲染归档 PVC(minecraft ns、RWO 10Gi、默认 SC),`felis manifests --backup-pvc` 默认 `felis-backups`(`=` 空为显式关闭),bootstrap 统一下传 env/旗标并写 [archive] local_path;`resolveWorldDir` 新增精确 local-path 目录解析(读 PVC `spec.volumeName`,非 glob,杜绝陈旧 PV 目录顶替);ReaperRole 增 `pvc:get`;bootstrap 设 `FELIS_WORLDS_HOST_PATH` 时给 uid 1000 授 traverse(setfacl/o+x);README/故障手册改写。提交 `fd33fd0`+`2b87a5a` |
| 8 | **磁盘打满灾难链**:DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后数分钟才恢复调度。恢复依赖人工 `docker save | k3s ctr images import -` | 三轮真机 drill(原缺陷复现 + 两轮修复验证):①自定义 1e6 类:游戏 pod 先走、控制面随后仍被驱逐(kubelet 日志逐条列出 ranked/evicted),随后镜像被 GC → ErrImagePull;②内建 `system-cluster-critical`(2e9):填盘至 1.7G free,kubelet 对 felis-api/operator/registry 全部报 **“cannot evict a critical pod”**,三者在整个 DiskPressure 期间保持 Running;游戏 pod(0)被驱逐;③释放磁盘后:`DiskPressure` 经 ~5 分钟(`eviction-pressure-transition-period`)转 False——即“恢复慢”的主因;被 GC 的游戏镜像按 runbook 重导入后 25s 恢复 ✅ | 控制面四类 pod 挂内建 `system-cluster-critical`(用户自定义类值上限 1e9,达不到 2e9 临界阈值;preemption 保留=管理面可调度,已记录权衡);troubleshooting §13b 固化事件链、5 分钟条件过渡与镜像恢复路径。提交见下 |
| 9 | **升级策略 Recreate**:单副本 + Recreate:任何控制面升级=停机;坏升级(错 tag)中断约 95s 且需人工 `rollout undo`(无自动回滚) | 复核部署模板(Recreate + 单副本 + 无 leader election 的注释理由成立) | 不改策略(Recreate 是对无 leader election 的正确取舍),改为固化 runbook:troubleshooting §15 = 升级即重跑安装器;停机窗口=rollout 时长;坏镜像在安装器 180s 等待内以 `kubectl describe` 诊断呈现;回滚 `kubectl rollout undo`(镜像被 GC 时先按 §13b 重导入)。提交见下 |
| 10 | **备份语义**:归档包含整个 /data(jar、libraries、cache),167MB;是否符合 world backup 定位待评估 | 真机 tar 清单(cache/、config/、eula.txt、server.properties、jar)+ 恢复语义(overlay+prune) | 评估结论:**整卷备份是正确语义**(回滚=整服状态回滚,含配置/插件),保留行为;文档化:troubleshooting §10「backup contains」+ OpenAPI/README 口径改为“整服数据卷”。提交见下 |
| 20 | **验证邮箱唯一性只存在于注释里**:errors.go/repo.go 都声称迁移 0010 建了 `users_verified_email_unique` 部分唯一索引、`VerifyEmailOTP` 会返回 `ErrEmailTaken`;实际上**索引从未创建**、`ErrEmailTaken` 全仓库从未被返回 → 两个账号可同时验证一个邮箱,而登录门正是按 verified email 解析账号 → 同一邮箱的登录码归谁由数据库任意决定 | pgint 首跑即红:第二账号验证同邮箱返回 `nil`;`grep users_verified_email_unique internal/store/migrations/*.sql` 零命中。真机复验(auditfix17 邮箱 409 演练):同码二次 verify 仍 409(码未被消费) | 新增迁移 `0020_verified_email_unique.sql`(`lower(email) WHERE email_verified`);`VerifyEmailOTP` 在消费码**前**查「他人已验证」→ `ErrEmailTaken`(不消费码、不扣次数),并把并发唯一冲突映射为同一答案;handler 新增 409 `email_taken`。提交 `b6ef27c`;PG 契约测试基建 `2a55a0d`(`-tags pgint`,见 CONTRIBUTING) |
| 21 | **admin 改邮箱不清 verified**:`UpdateUser` 写入新地址但保留 `email_verified=true` → 面板改错一个字符就能让登录码寄到别人邮箱(该 flag 正是邮箱登录解析/寄码的依据);fake 同错 | pgint 契约测试 | 改地址时同写 `email_verified = email_verified AND email IS NOT DISTINCT FROM 新值`(同值 no-op 保留证明;换值即清除);fake 同步。提交 `d1ec40f`。真机验证:PATCH→`f`、PATCH 回原值仍 `f`、玩家重验证→`t` |
| 22 | **`role='owner'` 是死信**:迁移 0011 加了 owner 角色并把全部用户管理路由压在 `IsOwner()` 上,但**没有任何代码写过 `owner`**——breakGlass(`UpsertOwner`)、setup MC 绑定(`CompleteOwnerSetup`)、重置路径一律写 `admin` → 全新安装的整个 owner 层(用户列表/创建/编辑/禁用/删除/配额/会话)不可达;且新角色没进各处 staff 判定(op-login 只认 admin、玩家邮箱门只拒 admin、reclaim 保护与 AdminExists 只认 admin);面板一旦有 owner 行还可被降级/删除/禁用 | 真机:owner 会话加载 `/api/v1/users` 403;promote 后 200。op-login:修复前 owner start 中立不发码,修复后铸请求+finish 得 `role=owner` 会话 | `UpsertOwner`/`CompleteOwnerSetup` 改写 `role='owner'`(冲突臂重断言,即 0011 文档的 promote 路径);op-login 双端改 `staffRole`;玩家邮箱门改 `role != 'user'`;`IsProtectedAdminLink`/`AdminExists` 计入 owner;面板新增守卫:owner 行不可降级/删除/禁用(用户名/邮箱编辑仍可)。提交 `e0d2378` |
| 23 | **配额门非原子 + storage 缓存被清零**:(a) audit #4:`QuotaCheck` 与 `ClaimServer` 两条语句,同一用户并发认领两台无主服可双双通过 `max_servers`(deferred-seams 曾把这条挂为“只能在真 PG 上关闭”);(b) 更隐蔽:PATCH 资源时 `UpdateServerResources(..., 0)` 把本不能改的 storage 缓存写 0,而缓存列是配额聚合的**唯一**输入 → 此后该服的 storage 维度在配额里凭空消失 | (a) 新增 pgint 并发测试:修复前两台全赢;修复后恰 1 赢 + 1 `ErrQuotaExceeded`,DB 只 1 行 owned;(b) hermetic 测试 `TestPatchServerPreservesStorageCache`(修复前 `resourceUpdates` 里 storage=0) | (a) 门槛进 `ClaimServer`:同一事务内 `pg_advisory_xact_lock(hashtext(user_id))` + 四维重查(与 `QuotaCheck` 共用 `quotaAllows` 防漂移)+ 行 `FOR UPDATE`,两个 claim handler 把 `ErrQuotaExceeded` 映射为与串行一致的 403;(b) resize 前读取现值并透传 storage。提交 `bb68fef`;deferred-seams 对应条目核销 |
| 24 | **reaper 警告信从不真正投递**:`felis reaper` 从不装配任何 Warner(`r.Warner` 恒 nil),而 `maybeWarn` 对 nil warner / 投递失败一律照样 `MarkWarned` + `warned++` → 每位有主的服都在**无人收到提醒**的情况下 15 天后被静默回收,跑批日志还谎报"已警告 N 台"。红线⑤的"best-effort 不阻塞回收"被误读成了"失败也要记成已通知" | hermetic 单测 3 例(坏 notifier / nil warner → 不盖章、成功才盖章)。真机 drill(auditfix20):13d idle 实收 1 封(收件人=owner 的验证邮箱、subject 含 3d);重跑不重发;14.5d 补发 1d(elif 档位);SMTP 端口打坏 → 日志 `warn delivery failed; will retry next run` 且**不盖章**,恢复后补发;邮箱未验证 → `has no verified email` 不盖章;owner NULL → 静默跳过;边界 11d23h 不发 / 12d1m 发;全程 `world_backups` 保持 4 条不动(纯警告零备份副作用) | `mail.SendNotice` + `mailWarner`:查 owner 的 verified email → `SendNotice`;nil/失败不盖章、下次重试;reaper pod 模板加 optional `FELIS_SMTP_PASSWORD` env;`felis setup` 复制 felis-smtp 镜像并在「configure email」刷新 minecraft ns 的 felis-smtp+felis-config 镜像;docs/troubleshooting §10 更新。提交 `8e7c7bb` |
| 25 | **idle auto-stop 从未触发(三层复合缺陷)**:条件齐备的空载服永远不停。① CRD status schema 未声明 `emptySince`,apiserver **pruning** 掉计时戳(`unknown field "status.emptySince"`),计时器每次读回都是 nil;② 即使戳幸存,Running 空载稳态**没有任何 watch 事件**(玩家进出不碰 CRD、RCON 只在 Reconcile 内探),盖章一次后 reconcile 链停摆,无人叫醒;③ auto-stop 用整对象 `Update` 写 spec 且 OperatorRole 从未有 `minecraftservers:patch/update` → 403 `cannot update resource`。三层任一都让功能永久失效,而单测(fake client 不剪 schema、手动驱动、无 RBAC)全部覆盖不到 | 真机逐层实锤:schema 修复后 `empty=2026-09-22T20:02:13Z` 首次成功持久化;静置 88s+ 无动作、operator 日志 90s 零行(②实锤);修复前日志 5 条 403(③实锤);全修后场景 1:超时戳触发即 Stopped;场景 2:起服 → 20:11:43 盖章 → **全程无干预** → 20:12:17 自动 Stopped;翻转 Running→Stopped→Running 收敛 Running;`idle=null` 后不再自停 | ① schema 补 `emptySince`(`c04a3f0`);② `reconcileRunning` 尾部返回 RequeueAfter——空载=到点精确唤醒、有人=30s 探针周期,+3 个单测(`1c89a5e`);③ auto-stop 改 merge-patch(防 status clobber,与 reaper 的 Stop 同型)+ OperatorRole 补 `patch` + rbac 测试锚点(`f650bf8`)。真机 auditfix22 全通;troubleshooting §11 增自检三连 |
| 26 | **RCON Secret 被删 → 服务器永久锁死**:删掉 per-server RCON 密码 Secret 后,(a) operator 没有任何 watch 能看到删除(Running 稳态零事件),secret 一直不重建;(b) 一旦被任意事件带到 reconcile:新密码铸造成功,但运行中的 pod 仍持有旧密码、sts 模板无变化 → pod 永不重启 → 探针用新密码连旧密码 pod,**永久 `RconNotReachable`** 直至 300s `ReadinessTimeout`;恢复只有人工删 pod。`ensureRconSecret` 的注释却声称"heals on the next pass" | 真机:删 secret 后静置 60s 无重建、无日志;annotate 触发后 96s+ 持续 RconNotReachable(新密码 `91dd62…` vs pod 旧密码);删 pod 后 17s 恢复 Running——根因=密码漂移实锤 | ① `SetupWithManager` 增 `Owns(&corev1.Secret{})`(controller-owned,删除事件映射回 CR)——真机日志见 `source="kind source: *v1.Secret"`;② pod template 新增 `RconSecretAnnotation` = 当前密码 SHA-256 指纹(64-bit):secret 重建指纹变 → sts 自动 rollout 用上新密码,未重建则恒定不抖动;测试 2 例(跨 reconcile 稳定 / 重建必变)。提交 `56f3abd`。真机复测:删 secret 后**同秒**察觉+重建+换戳,31s 全自动恢复 Running |
| 27 | **陈旧启动锚点 → 已恢复的服务器被误判 StartupTimeout + Provisioned 永久 False**:`status.startRequestedAt` 只在 Stop 时清,成功(markRunningReady)不清——一旦服务器曾经历一次长 Starting,其 300s 预算就悬在健康运行中;此后任意一次 pod 波动(rollout/崩溃)都会立即套用旧戳判 `Failed(StartupTimeout)`。且 `markFailed` 写入的 `ConditionProvisioned=False` 无人复位,恢复后仍永久挂着失败标记 | 真机:20:19 已恢复 Running 的 test-one,20:21 因一次 stamp 注入引发的 pod rollout 被标记 `Failed StartupTimeout`(锚点残留自 20:16);恢复后 status.conditions 里 `Provisioned=False … StartupTimeout` 持续存在 | `markRunningReady`:清 `StartRequestedAt`(每次启动/恢复尝试各有独立预算)+ 复位 `ConditionProvisioned=True`;测试 2 例(Ready 清锚点、Failed→Running 后 Provisioned 恢复)。提交 `82b5a60`。真机复测:清戳生效(`startRequestedAt: None`)、pod blip 删→33s 恢复全程无 Failed、锚点重新盖章后成功清除 |
| 28 | **cfsetup 把 "policy already exists" 当成功吞掉 → fail-closed 保证可被旧策略顶替**:Cloudflare Access 的 `CreateAccessPolicy` 在收到 already-exists 时直接返回 nil;若该 app 上已有一条更宽松的旧策略(改 identity 后重跑、或此前手工配置),守卫 op.console 的仍是旧策略,而 Setup 报告成功——`validateFailClosed` 只校验过"我们构建的策略",从未校验证留在线上的那条 | 代码审查(集成侧无真实 CF 账号,无法真机):httptest 3 例复刻——已存在同名策略、缺席、POST 竞态。修复前第 1 例吞错返回 nil 且不发 PUT | 改为按名 upsert:lookup → PUT 覆盖守卫体 → 缺席才 POST(POST 撞 already-exists → 重查后 PUT,绝不吞)。`apiPost/apiPut` 共用一个 `apiWrite`。提交 `30857df` |
| 29 | **"configure email" 的工作负载镜像复制从不生效**:`replicateSMTPToWorkloadNamespace` 把硬编码 `namespace: felis` 的 felis-smtp manifest 用 `-n minecraft apply` 发出——kubectl 拒绝 namespace 冲突(`the namespace from the provided object ... does not match`)→ 第一段直接 return err,**第二段 felis-config 的镜像复制根本不会执行**。于是任何"装好后再改 SMTP"的部署,reaper 的 pre-reap 警告永远拿不到新配置(这正是 8e7c7bb 添加该刷新要解决的事),且失败仅 warning 不中止 | 真机 kubectl 行为实验:`-n minecraft apply` 带 `namespace: felis` 的 manifest → `error: … does not match …`;修复后(manifest namespace=minecraft)→ `accepted-namespace=minecraft`(server dry-run);felis-config 渲染+apply 序列同为 accepted | `smtpSecretManifest(password, namespace)` 显式参数(控制面调用传 "felis",镜像传工作负载 ns);回归测试 `TestSMTPSecretManifestCarriesTargetNamespace`。提交 `ed722d5` |
## 待决策台账(未修)
前两日台账的 #7、#11–#15 已全部修复(本批核销,证据见下),不再挂在"未修"里:
| # | 主题 | 状态 |
|---|------|------|
| 7 | 异步失败不可感知 | ✅ 修复:`GET /api/v1/servers/{name}/jobs`(`ff7c57c`,含 RBAC `jobs:list` 与 OpenAPI);真机 drill 用它看到 running→succeeded |
| 11 | PG 断连报 401 | ✅ 修复:会话存储故障改为 503(`2a8f897`) |
| 12 | ready 门滞后 | ✅ 修复:operator 启动期内每 2s 重探 RCON(`a2df2f2`) |
| 13 | setup URL 截断 | ✅ 修复:TUI 折行(`abce381`) |
| 14 | controller-runtime 日志噪音 | ✅ 修复:SetLogger 接 slog(`a415246`) |
| 15 | reaper 启用未演练 | ✅ 已演练(见"第二日"节;CronJob 仍 `suspend=true` 防误删) |
## 可达性分级(#1–#79;#62 立案后剔除)——这些缺陷真实使用中到底谁能踩到
回应质疑"是不是全在测边界条件 / 只有内部 hook 才能触发":
- **hook 的边界**:演练中使用的内部手段(读服务端日志取 OTP 码、service token 打内部面代 approve、本地 SMTP sink、tmux 驱动 TUI)只是**测试仪表**——代替"真实收邮件 / 键盘敲击 / 操作台点击";触发路径本身是公开路径,真实用户做同一动作走同一段代码。真正**本环境不可复现**的只有 #28(无真实 CF 账号,仅 httptest 复刻),另有 #4/#5/#14 属工具/卫生级(无运行时触发面)。
- **分级口径**:① 日常=普通使用或默认安装下即会踩到;② 运维=升级/重启/故障恢复/磁盘满/删资源/闲置回收等真实运维动作会踩到;③ 窗口=需要并发或特定状态时序,真实但概率低;④ 工具=仅工具判定/审查;决策=非缺陷(评估/演练项)。
| # | 真实触发路径(一句话) | 分级 |
|---|------------------------|------|
| 1 | 恢复失败后 10 分钟内重试(必然的恢复动作)→ 202 但 Job 不动 | ② |
| 2 | 默认安装下第一次点「备份」→ Job FailedMount(felis-config 缺在 minecraft ns) | ① |
| 3 | #1 修复所需 RBAC(单独不产生用户可见行为) | ② |
| 4 | gofmt/staticcheck 漂移 + CI 门禁缺失 | ④ |
| 5 | 依赖漏洞(govulncheck 判定可达):需特定输入 | ④ |
| 6 | 默认安装即踩:备份/恢复恒 503、回收链无旗标不可启用 | ① |
| 7 | 任何异步操作(备份/构建)想查结果:无端点(功能缺口) | ① |
| 8 | 磁盘打满事故链(控制面被逐 → 镜像 GC → 全体 ImagePullBackOff) | ② |
| 9 | 非缺陷:升级策略评估 → runbook | 决策 |
| 10 | 非缺陷:备份语义评估 → 文档化 | 决策 |
| 11 | PG 掉线期间任何会话校验 → 401 误导 | ② |
| 12 | 每次起服:就绪感知滞后 | ① |
| 13 | 首次安装向导:长 URL 截断 | ① |
| 14 | operator 日志噪音(卫生项) | ④ |
| 15 | 流程项:reaper 启用补演练 | 决策 |
| 16 | 邮箱登录主路径:错码/重放 → 500、5 次锁定失效 | ① |
| 17 | 面板 op-login pending 列表字段缺失 | ① |
| 18 | 绑定码并发兑换(双击/双端即可)→ 部分 500 | ③ |
| 19 | 按文档启用回收 → CronJob 永远无法调度 | ② |
| 20 | 两账号验证同一邮箱 → 登录码归属不定(安全) | ③ |
| 21 | 面板改任一用户邮箱 → verified 未清、码寄旧地址 | ① |
| 22 | 全新安装 owner 层不可达(用户管理全路由) | ① |
| 23 | (a) 并发认领两服可双双过配额;(b) PATCH 任意服资源 → storage 缓存清零 | ①(a=③) |
| 24 | 有主服闲置 12–15d:警告信从不投递、静默回收 | ② |
| 25 | 启用空载停机后等超时:三层复合缺陷永不生效 | ② |
| 26 | 运维删/轮换 RCON Secret → 永久锁死至人工删 pod | ② |
| 27 | 曾长启动的服 + 任意 pod 波动 → 误判 StartupTimeout | ② |
| 28 | 重跑 cfsetup 且线上有旧策略 → fail-closed 被顶替;⚠无真实 CF 账号,仅 httptest | ④ |
| 29 | 装后改 SMTP(configure email)→ 复制链从未生效 | ② |
| 30 | 对不存在/刚被删的用户 id 操作 → 500 而非 404 | ③ |
| 31 | 真实 owner 打开 /admin/builds 被 mock 残留误判 | ① |
| 32 | 新浏览器首访 /admin/builds 播种两个假 404 | ①(轻) |
| 33 | 禁用/已删账号重走邮箱登录门 → 可复活(安全) | ① |
| 34 | 同族:死账号游戏内身份面仍存活 | ① |
| 35 | 任何真实服务器保存过一次后:所有备份/fileread 必败 | ① |
| 36 | 任何构建失败的提交者永远看不到结果 | ① |
| 37 | 管理端 fleet:login/lobby 行全是死操作 | ①(轻) |
| 38 | felis manifests 输出残留过期指导 | ②(文档) |
| 39 | 断玻璃 reset owner 用非在位名 → 铸第二 owner 永不可清 | ② |
| 40 | 断玻璃 add operator 撞名 → 裸 SQLSTATE | ② |
| 41 | 断玻璃 Sync 选系统服 → 裸内部错误 | ② |
| 42 | 对无世界盘服备份/恢复 → 202 后静默卡 30 分钟 | ② |
| 43 | #42 配套:409 一刀切文案 | ②(轻) |
| 44 | 重跑 `felis setup` 完成 s/c reconfigure → 状态框退回“Setup complete”(信息一致性) | ①(轻) |
| 45 | 评审看不到将被执行的 recipe(执行的是上传 blob 里的 Dockerfile)→ 盲批 | ① |
| 46 | 大层推送打死 registry(256Mi 模板被 OOM kill;实测 475MB 层)——用户构建/安装器入仓即触发 | ① |
| 47 | Mac 打包树构建:`._*.sql` 旁文件混入 embed → `felis migrate` 全挂 | ② |
| 48 | 安装器重跑:每镜像 start/stop docker 触发 systemd 限流 → 批量入仓中途断 | ② |
| 49 | 安装器重跑:`[registry]`/`[archive]` 运维配置静默回退(构建/S3/reaper 行为回默认) | ② |
| 50 | 配过 SMTP 的安装按文档重跑安装器 → 生成的注释块被 `[smtp]` carry 吞并并每次 +1(纯膨胀) | ② |
| 51 | 改配置/轮换 DB 凭据后:工作负载 ns 的 `felis-config` 副本永不刷新 → backup/reaper 静默读旧配置 | ② |
| 52 | 向导内 `s`/`c` 重配置后:工作负载 ns 的 `felis-config` 镜像停在上一次运行快照,直到下次 `felis setup`/安装器才收敛 | ② |
| 53 | 默认安装或 `felis update` 的 felis-api 检查:坐标仍指向已迁移的旧仓库,现仅靠 GitHub 301 兜住——redirect 一退休,安装器默认 URL 与更新检查全挂 | ② |
| 54 | 完成态安装上照 `felis update` 指引升级:`sudo felis setup` 只开配置控制台,任何组件都不会动(三处 run 指引 + trailer 全假) | ②(文档) |
| 55 | 配置的源回 200 但档案字段不可用(马虎自建/三方 Yggdrasil):坏源在前 → 经梯子的登录被静默吞掉(零日志、后续源不被询问) | ② |
| 56 | 给"还没进过服"的玩家加白名单/封禁/踢出(面板说成功、vanilla 实际拒绝)——任何新玩家第一次被管理就撞 | ① |
| 57 | 控制台发任何命令后无反馈(回复被丢弃、pod 日志也不回显) | ① |
| 58 | 对从未启动过的服务器点"文件"页(90s 卡死 → 误导性 504) | ① |
| 59 | 装了 LuckPerms 的服打开 LP 页:读不到被报成"未分配任何权限"(假事实),且历史伪造 `[RCON]` 输出 | ① |
| 60 | 新建服务器输入非法名/子域名(前端预检比后端弱)→ 对话框直出 Go 英文错误;账户验证撞已占邮箱同理 | ① |
| 61 | 给"还没进过服"的玩家预授权限:面板提示"仍然会实际生效"(过度承诺),实际 LuckPerms 静默丢弃(未进服的名字无法解析) | ① |
统计:① 26 | ② 25 | ③ 3+#23(a) | ④ 4 | 决策 3 = 61。第三批构建链另有 3 处未编号修复(Job requests 超小节点上限 → 永远 Pending、Kaniko chown、Trivy DB egress 被锁)——均属 ①/②。
第四十批追加:#63(demo-up 单起点化——旧镜像臂能产出"起了但无处路由"的演示机;dev/demo 路径)②;#62 经立案复核**不成立**(全部安装路径都构建 felis-velocity.jar,证据见该批节),已剔除、不计入。
第四十一批追加:#64(三个 vendored `gradlew` 以 100644 提交、无可执行位——README 教的构建命令在全新 clone 上直接 `Permission denied`;无任何 CI/安装路径跑过这三个模块)①;#65(三个装载器 mod 的 `license` 仍写 MIT、Forge/NeoForge 的 `issueTrackerURL` 是 `example.invalid` 占位,与仓库 AGPL-3.0-only 相悖——fabric loader 启动时会打印该字段)②(低)。
第四十二批追加:#66(英文 README 缺中文版"使用方式"里的私有仓库安装 workaround 与重跑升级说明——英文读者照文档第一步即 404、无任何指引)①(轻)。
第四十三批追加:#67(build Job 从不设 `ttlSecondsAfterFinished`——每构建一次就永久留下一个完成 Job+Pod,完成 Pod 计入节点 pod 预算(stock k3s 110),构建量上来后新构建全 Pending)②(时间维度;随构建量从②滑向①)。
第四十三批追加(二):#68(构建 context 拉取单次尝试——控制面滚动/重启/驱逐恰好撞上构建窗口时,fetch initContainer 一次 `connection refused` 直接打成终态 `Failed`;`BackoffLimit=0` 无第二次 Pod,代价 = 人工重审重提)②(运维:升级/重启/故障恢复撞上正在进行的构建;构建量大时概率上升)。
第四十五批追加:#69(README 中英"审批通过后自动构建**并部署**"过度承诺——数据模型无目标服务器、部署实为"选用该镜像";① 轻);#70(kaniko 拉内建底座默认 HTTPS——`--insecure` 只覆盖推,任何 `FROM registry.felis.svc:5000/…` 构建必败 ①);#71(kaniko 以 drop-ALL 解包底座层,chown 必败——任何非 scratch 底座必败 ①);#72(trivy 扫 jar 必拉 Java DB、被 egress 锁拒绝——含 jar 即所有真实模组包的构建必败 ①);#73(bootstrap 测试在跑 k3s 的主机上必假失败 ④ 工具);#74(console 断连测试读写竞态 flaky ④ 工具)。
第四十六批追加:#75(提交上传面无上限——pending 无个数上限、无存储预算、create/upload 无节流;一个账号可无限堆积上下文刷爆 uploads PVC ①);#76(提交无撤回/管理员删除路径——提交者无法自救、运维无法回收占用 ①);#77(未完成引导的会话触发受保护操作 → 面板显示"无权执行此操作"而非送往 /setup;tracker #8 ①);#78(create-if-absent 使新增 CR 字段在已装机永不落地——lobby 无 RCON 故控制台死、玩家数恒 0;tracker #1 ②);#79(troubleshooting [INERT] 图例指向已不存在字段 ④ 文档)。
## 已验证事实(正向清单)
- 安装→hook 发码→Owner 绑定→passkey(虚拟认证器)→面板管理员全链路 ✅
- 建服(POST /servers)→ 唤醒(operator 拉 StatefulSet pod)→ RCON `list` → SSE 控制台 → 停止 ✅
- 文件编辑:列目录/读/写(wire 为 base64)/256KiB 413/路径穿越 5 变体全拦截/运行中 409 ✅
- **备份→篡改→恢复数据演练**:v1→备份→v2→恢复→读回 v1 ✅(G2 数据可恢复)
- **毒档案 fail-closed**:穿越/绝对路径/符号链接条目 → `archive entry escapes target` 退出码 1,零写入 ✅
- 混沌:PG 掉线(healthz 仍 200、恢复后连接池自愈)、API pod 击杀(~2s 中断)、整机重启(32s 回归、会话/CRD/停止态全保留)✅
- `felis update` 报告(k3s/velocity 有更新、私有仓库 404 优雅处理)✅
- 账户全套真机 E2E:绑定码新玩家注册(幂等/并发见 #18)、邮箱验证(onboarding 门)、邮箱 OTP 登录(错码扣次/5 次锁定后正确码也 400、重放 400、staff 账号 403 拒绝)、Passkey 注册+discoverable 登录+邮箱优先登录(虚拟认证器;Chrome 要求 `Page.bringToFront` 才能过 focus 检查)、登出吊销会话(旧 cookie 401)
- op-login 全状态机(start→status→approve→finish;早 finish 不烧码、错码扣次且请求保留、重放/重复批准/非管理员批准/未知 handle 全部按契约返回)✅
- 并发:OTP 风暴 8 路 = 1×202 + 7×429 且仅铸 1 码;绑定码并发(同码/双码)见 #18 ✅
- **reaper 全链路真机演练**(修复 #19 后):绑定挂载假世界 → `felis reaper` 一 pass:归档 tar 落盘(内容含标记文件)、`world_backups` 行 `inactive_15d`+90d 过期、**PVC 删除且宿主目录回收**、servers 行保留且 activity 时钟重置(红线②)、CRD 保留;第二 pass:伪造过期归档被驱逐(`expired=1`,文件删、行转 `deleted`),合法归档未动;`evaluated=2 reaped=1` 只动到期的世界 ✅
- **S3 存储向导全链**(第三十二批):错误凭据预检零副作用、正例落地(Secret + 两 toml + API 滚 + 状态行)、上传 → MinIO 桶 → 内部面取件、UI 回滚归位、控制面/工作负载镜像对齐 ✅
- **NetworkPolicy 栅栏真机**(第三十五批):pod 源"身份×端口"矩阵全对(允许行与拒绝行同源对照)、转发型外部源白名单闭环(拒→放→拒,ipset 与 spec 双幂等)、host/ClusterIP 路径全通、Mac→NodePort 面板 200;两条归因注记见该批 ✅
### 本轮新增真机证据(第二日)
- **默认安装的备份闭环**:apply 新 bundle → 归档 PVC 自动创建(WaitForFirstConsumer,被首个 backup Job 拉绑定)→ `POST /backup` 202 → `GET /jobs` running→succeeded → 漂移 marker → `restore-backup` → 读回备份前内容 ✅
- **回收闭环(真实 k3s 布局 + 权限)**:resolvecheck(20d idle)→ CronJob 派 Job:`evaluated=2 reaped=1`;归档含 marker、PVC+宿主目录回收、`world_backups` 记 `inactive_15d`;期间先经历 `lstat /worlds/…: permission denied`(k3s storage root 0700)→ 修 ACL/root 权限后通过 ✅
- **磁盘压力三连 drill**:①1e6 自定义类:游戏 pod 先走、控制面随后仍被逐(kubelet ranked/evicted 日志)→ 镜像 GC → ErrImagePull;②改内建 `system-cluster-critical` 后:`cannot evict a critical pod` × felis-api/operator/registry,全部保持 Running;③恢复阶段实测 `DiskPressure` 条件 ~5min(eviction-pressure-transition-period)转 False,被 GC 的游戏镜像按 runbook 重导入后 25s 恢复 ✅
- **混沌**:SIGKILL k3s 主进程 → kubectl 与 felis-api ready 均在 **10s** 内恢复(api pod 本身未被重启)✅;内存压力(2.6G tmpfs 满写):无 OOM kill、无人被逐(swap 5.9G 吸收 + 内核回收),控制面不受影响 ✅
- **构建链路定位(当时未修)**:Kaniko/Trivy 默认镜像是外部的 → 已加 `[registry] kaniko_image/trivy_image/build_cpu_limit/build_mem_limit` 覆写;上下文跨 ns 不可挂载(PVC 不能跨 ns)是设计级缺口,s3 lane 也缺凭据注入 ✅(证据:`internal/submit/blobstore.go:40`、`cmd/felis/api.go` contextBase 分支、jobspec 无 volumes/env)→ **第三批已修复,见下**
### 本轮新增真机证据(第三批:探针 + 构建链路全通)
- **控制面探针(`0c8e29b`)**:felis-api / felis-operator / registry 三个 Deployment 此前**完全没有探针**。修复后真机验证:api、operator `/readyz`+`/healthz`(internal 8081),registry `/v2/`(含命名端口解析);三个 Pod `Ready=true`、`restartCount=0`、无 `Unhealthy` 事件;operator 新增 `--health-probe-bind-address`(8081,与 metrics 8080 分离,零检查也 404 → 已注册 ping)
- **构建链路全通(`f79e5eb` + `02fd2de`)**:
- 传输:context_ref 改为 internal-face URL;`felis fetch-context` initContainer 走 service token(对象来自 `felis-build` 里按 `bootstrap`/`felis setup` 同款 Secret 复制机制落地的 `felis-service-token`;netpol 只放行控制 ns:8081)+ zip-slip 安全解包到限容 emptyDir;Kaniko `--context=/context` 只读挂载
- **真机 E2E**:玩家 OTP 登录 → 建 submission → 上传 gzip context(marker `felis-e2e-build-marker`)→ Owner approve → Job:fetch ✅ → Kaniko 构建并 push `registry.felis.svc:5000/user-uploads/<sub>:latest` ✅ → Trivy 扫描(内建镜像 DB)✅ → Job `Complete`;**从 registry 拉回镜像校验 `/hello.txt` 内容 = 上传的 marker** ✅
- 演练中真机抓到并修复 3 个单测看不到的缺陷:① Job requests=limits(2C/4Gi)在 4C/5.5G starter 上**永远 Pending**(`FailedScheduling/Insufficient memory`)→ requests 改为地板值(250m/512Mi,不超上限);② Kaniko 重拷 Dockerfile 时 chown/chmod 到源文件 owner(distroless uid 65532)在 drop-ALL 能力下失败(`copying dockerfile: chown … operation not permitted`)→ fetch 容器以 root 解包(= Kaniko 自身 uid);③ Trivy DB 默认 `mirror.gcr.io` 正被 egress 锁拒绝(fail-closed)→ 新配置 `[registry] trivy_db_repository` + 文档 §8e 固化镜像配方(`docker pull/tag/push aquasec/trivy-db:2` 进内建 registry;`--insecure` 已覆盖明文 HTTP)
### 本轮新增真机证据(第四批:PG 契约测试 + 验证邮箱唯一性 + owner 角色)
- **PG 契约测试首跑抓到 #20**:`internal/pgint` 对着真实 PG(`felis_pgint`,drop schema + 重放迁移)跑会话/OTP/op-login/绑定码/submission/build 全契约;`TestOnboardingEmailOTPRejectsTakenEmail` 当场变红 → 挖出「索引从未创建 + `ErrEmailTaken` 从未返回」。
- **迁移 0020 真机应用**:`felis migrate up` → `applied 1 migration(s): [20]`;`schema_migrations` max=20;`pg_indexes` 出现 `users_verified_email_unique`。
- **邮箱唯一性真机 409 演练(auditfix17)**:owner 已验地址被玩家 onboarding 流程二次验证 → 第一次 409 `email_taken`;**同码重放仍是 409**(证明码未被消费,符合契约;若被消费会是 400)。
- **admin 改邮箱清 verified 真机演练(auditfix18)**:PATCH player.test → `[email protected]` 后 `email_verified=f`;PATCH 回原地址仍 `f`;玩家走 onboarding 重验证 → `t`。
- **owner 角色真机闭环(auditfix18)**:owner 行按 0011 文档升级路径置 `role='owner'` → `GET /api/v1/users` 200(修复前 403);owner op-login start 铸请求+发码、finish 得 `role=owner` 会话;PATCH owner role / DELETE owner / DISABLE owner 全部 403;玩家邮箱登录回归 200。
- **配额原子门(第五批,`bb68fef`,auditfix19 已上线)**:pgint 并发实证(真实 PG 上两路并发认领:恰 1 赢 + 1 `ErrQuotaExceeded`;修复前两路全赢);`docs/deferred-seams.md` 的 audit #4 条目核销。
### 本轮新增真机证据(第六批:reaper 警告链路端到端演练,auditfix20)
- **演练装置**:独立 hostNetwork Pod(`felis:auditfix20`,SA/卷结构复刻 live CronJob)+ 宿主本地零依赖 SMTP sink(捕获完整 RFC5322 原文)+ 专用 `felis-config-drill` Secret(SMTP 指向 `127.0.0.1:1025`);场景服 `warntest`(独立 CRD,desiredState=Stopped,不占资源),期间 operator 暂停防 reconcile。
- **投递**:13d idle → 恰 1 封,`To: [email protected]`(owner 的**验证**邮箱)、subject 含 `3d`、正文含 server/remaining;`warned_3d_at` 盖章、`warned_1d_at` 留空;`world_backups` 保持 4 条(纯警告零备份)。
- **去重与档位**:重跑 0 封(已盖章不重发);14.5d idle → 补发第 2 封 = `1d` 档(elif 顺序正确,不是重复 3d),`warned_1d_at` 盖章。
- **失败语义(三连)**:SMTP 端口打坏 → 日志 `warn delivery failed; will retry next run`(err=`connection refused`)、**不盖章**,恢复端口后同一 run 补发成功;邮箱 `email_verified=false` → `resolve owner email: … has no verified email`、不盖章,恢复后补发;owner NULL → 静默跳过(无日志无警告)。
- **边界**:idle=11d23h(< 12d 阈值)不发不盖章;idle=12d1m 发。`http_code` 冒烟:升级 auditfix20 后 panel 200 / op-login 200。
- **环境还原**:warntest CRD+行删除、drill Secret/Pod 删除、operator 恢复;`servers` 表回到 resolvecheck+test-one,CRD 回到 lobby/login/resolvecheck/test-one。
### 本轮新增真机证据(第七批:idle auto-stop 三层修复,auditfix21/22)
- **发现路径**:operator 深挖时先怀疑"计时器无驱动",真机实验立刻抓到 pruning(操作日志 `unknown field "status.emptySince"`)+ 触发后 88s 无动作 + RBAC 403,三层各自独立、各自足以致死。
- **场景 1(超时点唤醒 + 写权限)**:EmptySince 已超时 11 分钟 → annotate 触发一次 → 秒级内 desiredState=Stopped、sts 0/replicas、EmptySince 清、phase=Stopped。
- **场景 2(完整自驱,决定性)**:desired=Running → pod 起 → 20:11:43 盖章 → 静置无干预 → **20:12:17 自动 Stopped**(30s 到点后 ~4s 完成 stop+scaledown+markStopped)。
- **翻转混沌**:3 秒内 Running→Stopped→Running,最终收敛 Running ready(期间 409 竞争为控制器正常噪音,controller-runtime 重试自愈)。
- **收尾**:`spec.idle` 删除后再无自停;test-one 回 Stopped、空 sts、无 EmptySince;VM 镜像 auditfix22 = `f650bf8`。
### 本轮新增真机证据(第八批:operator 自愈深挖,auditfix23/24)
- **RCON Secret 删除自愈(`56f3abd`)**:修复前——删 secret 静置 60s 零察觉、触发后 96s+ 永久 `RconNotReachable`、仅人工删 pod 可恢复;修复后——删 secret **同秒**(20:22:16)察觉(Secret watch 日志)→ 新密码 + stamp 换值(`3b24164f…`→`c465b627…`)→ rollout → **20:22:47 全自动恢复 Running**(31s)。
- **陈旧锚点(`82b5a60`)**:修复前——已恢复服务器因锚点残留被 `Failed(StartupTimeout)`、`Provisioned=False` 永挂;修复后——Ready 即清锚(`startRequestedAt: None`)、`Provisioned: True`;pod blip 演练:删除→20:24:48 重锚→20:25:19 恢复→锚清空,**全程无 Failed**。
- **StatefulSet 误删自愈**:删除 sts → **秒级**重建(PVC `world-test-one-0` 保持 Bound,Retain 策略)、29s 后 Running、stamp 保持 `c465b627…`(未触发多余 rollout)。
- **收尾**:期间 7 条 operator ERROR 全为 CR/sts 409 竞争噪音(自愈);test-one 回 Stopped。
### 本轮新增真机证据(第九批:未演练端点地毯覆盖 — 迁移 / access / window / credentials)
- **account/migrate 四步状态机全链路(首个端到端演练,0 缺陷)**:造 user2 目标账户(绑定码 `FDTGKYAQ`)→ player.test claim test-one → 内部 start(未联动 UUID=404 not_linked;源=201 initiated)→ `passkey_required 409`(有 passkey 强制强因子)→ **passkey credentials list + DELETE 204**(顺带覆盖管理端点;删后 OTP 门自动解除)→ OTP start 202 / 错码 400 / 真码 confirmed → issue-code(self=400 / missing=400 / 真目标=201)→ redeem 负例×2(源持码=400;target 错码=403 setup_required——**onboarding 门正确拦截未验证账户**;验证邮箱后重试)→ **redeem 成功**(小写+空格码兼容,`servers_moved:1`)。
- **retire 断言全对**:servers.owner→user2;migration=redeemed(含时间戳);源 disabled+软删;源活跃 session=0;源旧 cookie 401;double-redeem 400;源再加迁移 409 account_retired。**环境复原**:源复活 + 新 OTP 会话、test-one 归还无主。
- **access 玩家管理全组(0 缺陷)**:players/whitelist/banlist 三个读 projector 解析正确(`There are 0 of a max of 20 players online: `→[]);kick/permission/group/luckperms-info 调用全通(demo 服无 LuckPerms → 原样透传 `Unknown or incomplete command`,API 不掩饰);输入负例 4 连 400(bad player/bad action/bad node/bad world);`whitelist add/remove` 在空档案服上得到 vanilla `That player does not exist`(fail-closed egress 无法解析 Mojang profile——**vanilla 约束非 API 缺陷**,命令已如实送达);手写 whitelist.json + `command whitelist reload` 验证成功路径解析(`players:["E2E_Tester"]`);**audit 7 条 access.\* 全记录**。
- **updates/window**:unset→null;PUT 正例+读回;半设 400;倒置 400;clear→null 全对。
- **/fleet**:4 服全量(含运行中 lobby/login 的 endpoint 地址)一次读全。
- **bootstrap-assets crd**:嵌入式资产含 `emptySince`(auditfix24 二进制核对)。
### 本轮新增真机证据(第十批:configure email 复制链修复,auditfix25)
- **发现路径**:复核 `8e7c7bb` 新增的 replicate 逻辑时怀疑 manifest/-n namespace 冲突 → kubectl 真机实验裁决(hardcoded `namespace: felis` 版直接 `error: … does not match …`,证实第一段必错、第二段被跳过)。
- **修复后行为验证**:两条命令序列(felis-smtp manifest 携带目标 ns + `felis-config` create--dry-run|apply)均被 server 接受且落点 `minecraft`。
- **live 收敛项(随下次 `felis install/setup` 重跑)**:live CronJob 仍是旧模板(无 `FELIS_SMTP_PASSWORD` env);`felis-smtp`/镜像两 ns 均未创建(SMTP 未配置,属正常);与 operator Role 手动补 patch 同批处理。
- **部署**:`felis:auditfix25` = `ed722d5`(api/operator/reaper 三处 set;api/operator rollout 完成,panel 200)。
### 本轮新增真机证据(第十一批:submission 全路由负例矩阵 + 内部面取件 + owner 会话重铸)
- **session 重铸路径(hook)**:owner 旧 cookie 过期(401)→ 临时插 `account_links`(owner↔假 UUID)→ op-login 三件套:start 202 → 日志取码 → 内部面 `op-login/{id}/approve`(hook,经 service token)→ status `approved:true` → finish 200 `role=owner` → op 域新 cookie 生效(`/fleet` 200);**演习后已删除假链接行**。顺带复证:approve 失败时 status=false、finish 400 `op_login_invalid`(码不烧,复用同码二次 finish 成功)。
- **submission 车道负例矩阵(23/23 PASS,0 缺陷)**:create 未知字段/空名/非法字符/超长 → 全 400;非 gzip 上传 400、未知 id 404;跨用户上传 → **404 不可见**(非 403);玩家打 admin 三路由(queue/approve/reject)全 403;reject 空理由/1001 字/夹带字段 → 400,成功 200(`reject_reason`/`reviewed_by`=验主邮箱/`status=rejected` 全对);重复 reject 409、reject 后 approve 409;**无 blob approve → 400 且行保持 pending_review(未 stranded)**;玩家列表严格只见自己。
- **内部面取件**:`GET /api/v1/internal/submissions/{id}/context`(service token)→ 200 + gzip magic `1f8b` + 解包内容 = 上传 marker;无 token 401;未知 id 404。
- **演习清理**:E2E-Lane-\* 8 行已删、假链接已删;队列只留历史 "E2E *"(approved)行。
- 备注:port-forward 再次因 API pod 重建悬死(空响应)→ 按速查重启即恢复;本轮无需新镜像(纯验证)。
### 本轮新增真机证据(第十二批:users 管理面全量矩阵 → 抓到并修复 #30)
- **/users 管理面矩阵(修复前,39 PASS / 2 NOTE)**:create/get/patch/list 权限与负例全对(未知字段/非法用户名/owner 角色/未知 id → 400/404;owner 保护 403 复证);sessions 单条吊销→旧 cookie 立即 401→重登恢复、revoke-all 同效(重登两次含 61s 冷却等待,全绿);passkeys 无凭证 200 no-op;links 幂等/跨用户 409/非法 auth_source 400/解链 404 全对;disable/enable/delete/double-delete 全对。
- **缺陷 #30(两处 FK 500 + 一处假 200)**:`PUT /users/{unknown}/quotas` 与 `POST /users/{unknown}/links` 触碰 user_id 外键 → **500 internal**(契约要求 404);`GET /users/{unknown}/quotas` 回**零值视图(=unlimited)**,像 id 存在。修复:`PGRepo.requireLiveUser`(live 行 + `deleted_at IS NULL`)守卫三个方法,`handleGetQuotas`/`handleLinkAccount` 补 ErrNotFound→404 映射;单测 `TestAdminSubresourcesRequireLiveUser`(fake 同步 liveUserExists 契约)+ pgint 真库契约(ghost→ErrNotFound、live 对照、软删后→ErrNotFound)。commit `1918da2`。
- **auditfix26 真机复验**:ghost 三连 → **404 not_found**;活用户对照 → PUT/GET quotas 200(读回 2)、POST link 200、delete 200;panel 200,api/operator 滚动完成(新 pod 24s Ready)。
- 契约备注(非缺陷):软删用户 `GET /users/{id}` 仍 200(带 `disabled=true` + `deleted_at`,列表已过滤);`DELETE .../sessions/deadbeef` 幂等 200;revoke-all 对未知用户 200 no-op。
### 本轮新增真机证据(第十三批:CLI / 镜像面 / 直接构建 / 面板 CDP 全路由 + #31/#32)
- **CLI 面 19/19(VM,`felis-auditfix26`)**:`version`/`-h`/无参=exit2/未知命令=exit2;`manifests` 缺 `--velocity-cidr`、缺 `--felis-image`、坏 CIDR、坏 NodePort 全 fail-loud exit2,正例渲染 25 文档且 `--worlds-host-path` 分支出 CronJob+三条前置警示——整包 `kubectl apply --dry-run=server` **全部 `configured`**;`apply` 缺 `-f`/文件不存在/坏 JSON 各自 exit2/1/1;`migrate up` 二次幂等(`database already up to date`);`bootstrap-assets crd`(含 `emptySince`)与 `game-stack`(tar)正常;`update` 正常。
- **镜像面负例 + builds 负例**:`/images` 列表 200、空体 400、合法 201、重复幂等 201、无 `?ref` 400、删除 **204**(契约;我的 200 预期作废)、再删 404、坏 ref 400;`/images/build/unknown` GET/cancel 均 404;空体提交 400;玩家打 `/images*` 全 403。
- **`/images/build` 正路径(直接构建)**:借遗留 submission blob 作 context,`FROM scratch` 内联 Dockerfile → **202 → 12s succeeded**;`/logs` SSE 实收 Kaniko 流;真机 registry `tags/list` 出现 `e2e/direct-probe:latest`(镜像确实被推入)。留档:该 drill 镜像保留在 registry。
- **`auth/options`**:未知邮箱 → `{"methods":[]}`;玩家 → `["email_otp"]`(无 passkey 时);坏邮箱 400。
- **`PATCH /servers/{name}`**:空体/storage/空 policy/未知字段/坏名 → 400;未知服 404;displayName 设置+还原 200。**契约备注**:`displayName:""` 因 `omitempty`+merge-null 语义 = **清除该字段**(本次 drill 先误读为设空串,比对 patch 前 fleet 快照确认 test-one 原无 displayName,无数据损失)。
- **面板 CDP 全路由巡检(owner 会话,14 路由)**:全部有标题/有内容/无崩溃;唯二发现 = #31/#32(见下);`luckperms` 页对 stopped 服 409 → **优雅降级**为「服务器已休眠 + 启动」态(browser 级 409 日志属正常资源日志,非缺陷)。玩家会话 4 路由零错误且**导航分级正确**(不见管理/平台组);未登录 `/login` 渲染正常(`/me` 401 为预期探测)、`/account` 未登录重定向 `/login`。
- **缺陷 #31(面板 mock 残留·owner 判定)**:`ImageBuildPage` 用 `email==="[email protected]" || startsWith("owner@")` 猜 owner——真实 owner(`[email protected]`)**不被识别**(只能看 approved 提交),而任何 `owner@x` 邮箱都冒充 owner。改为消费 TierProvider 服务端 `isOwner`。commit `6907961`。
- **缺陷 #32(面板 mock 残留·假构建种子)**:首次访问 `/admin/builds` 向 localStorage 播种 `["bld-1","bld-2"]` → 每个新浏览器打两个必然 404 的请求(页面注释自认“in production will 404”)。移除播种。同 commit。
- **auditfix27 复验**:清掉该 localStorage 键后重访 → **只发 `/me`**、零错误、空态正确;面板 111 单测 + typecheck 全绿。
- **passkey 全流程复验(CDP 虚拟认证器,面板 UI 实操)**:注册(对话框输入→仪式→列表出现 `E2E-Key`)→ API logout → 邮箱优先 passkey 登录(`/me` 回 `[email protected] user`、落回 `/`)→ `DELETE credentials/{id}` 204 → 列表清空(环境复原)。
- **环境**:api/operator/reaper 镜像 = `felis:auditfix28`(= `6907961`);面板 200。**教训**:`pkill -f "port-forward …"` 会匹配到**执行该命令的 ssh 自身 cmdline**(命令里同时含明文 `port-forward svc/…`)→ 自杀式断连;改为 `ss -ltnp` 取 PID kill,另起一条 ssh 启动转发。
### 本轮新增真机证据(第十四批:#33 死账号复活漏洞 —— 全登录门闭环 + 资产切断)
- **发现路径**:users 管理面演习收尾核对时发现软删用户仍留 `account_links`/`quotas` 孤儿行 → 顺藤摸瓜确认 `UserByEmail` 与 `SessionUser` 均不过滤 `disabled`/`deleted_at`。
- **红证据(auditfix27,真机)**:禁用用户的邮箱门重登录 `verify=200`、`/me=200`(**禁用锁死可绕过**);**已删号**用户 `start=202`(真实发码)、`verify=200`、`/me=200`(**删号可复活**)。
- **修复(`5853589`,7 文件 / +306 行)**:① `UserByEmail` 只解析 live 账号(覆盖 email/passkey 邮箱优先/op-login/auth options 全部前置门);② `SessionUser` 同样过滤(皮带:任何门铸出的死号会话都立即失效);③ `RedeemPlayerBindCode` 两个解析臂拒绝死账号(新哨兵 `ErrPlayerAccountRetired` → 403 `account_retired`,**不消费码**,可逆重试);④ `DeleteUser` 事务内 `DELETE webauthn_credentials` + `account_links`(凭证与 `UNIQUE(mc_uuid)` 占用不再外泄);⑤ discoverable passkey resolve 补 liveness 检查。单测 `TestDeadAccountsCannotLogInOrKeepSessions` + pgint `TestDeadAccountsAreLockedOutInPG`(含"删后新铸会话不验证"与资产清点)。
- **绿证据(auditfix28,真机)**:红期为死号铸出的会话 → **401**(皮带生效);死号 `start=202` 中立且 **90s 日志窗口 0 条码**、`verify=400 invalid_code`;禁用中 `start` 中立(日志码数 1→1 不变)→ `verify=400`;**恢复启用后 `verify=200` + `/me=200`(可逆)**;bind 门死账号 → **403 `account_retired`** 且码保留 `count=1`(retryable);删号后 `account_links`/`webauthn_credentials` 行数 = 0。
- **演习清理**:红演练遗留(2 个测试账号的 quotas 行、1 条旧链接、2 条死号会话)已清除;`/fleet` 与 panel 冒烟保持 200。
### 本轮新增真机证据(第十五批:#34 死账号同族收尾 —— in-game 身份解析与链接接管,auditfix29)
- **发现路径**:#33 修完 Web 登录门后做同族复查——**in-game 面是 UUID 驱动、不走 Web 会话,因此 #33 的会话皮带盖不住它**。两处:① `UserByMCUUID`(claim / menu / wake 授权 / op-login vouch / QR link-status 五处共享解析)直接读 `account_links`、不过滤 users 的 liveness;② `VerifyLinkCode` 对「holder 已死」的接管语义未定义——一刀切 409 会把迁移退役这类真实场景锁死。
- **修复(`bb9798e`,5 文件 / +182−7)**:① `UserByMCUUID` JOIN users 过滤 `disabled`+`deleted_at`——死账号的链接在游戏面**读作未链接**(claim 412、wake 落回 policy 门、vouch 403、status `linked:false`),绝不作为遗留身份存活;② `VerifyLinkCode`:**软删** holder 的链接可由新账号凭新 mint 码**接管**(账号已亡,码证明调用者仍持有该 UUID),**禁用** holder 仍 409(接管= 绕过锁死,不允许)、任何失败都不消费码。fake 单测 `TestLinkVerifyTakesOverDeletedLinkOnly`;pgint `TestVerifyLinkCodeTakesOverDeletedLink` + `TestDeadAccountsAreLockedOutInPG` 增「禁用链接无资格 / 复启用恢复」断言。
- **绿证据(auditfix29 真机,三面六点)**:
1. `link/status`:活 `{"linked":true}` → 禁用 `{"linked":false}` → 恢复 `{"linked":true}`;
2. `claim`:活链 404(身份已解析、ghost 服不存在)→ 禁用 **412 not_linked** → 恢复 404;未知 UUID 对照 412;
3. **接管正例**:临时软删 holder(`de49df52`)→ 新 mint 码 `X3G3RZZM` 由 player.test verify → **200 linked:true**;链接行迁移到 player.test(`auth_source=mojang`)、码被消费(`account_link_codes` 0 行);演练后 holder 与链接全部还原;
4. **接管负例**:临时禁用 holder(user2)→ 新码 `X5W94S42` verify → **409 already_linked**;同码重试仍 409(**不消费**);恢复启用后**同码**由 holder verify → **200**(码保留、可逆);
5. **wake 授权**:活 owner **202**(真实启动)→ 禁用 **403 forbidden** → 恢复 **202**;test-one 已停回 Stopped 且所有权释放;
6. **op-login vouch**:禁用 owner UUID → **403 not_admin**(与未链接/非 staff 同一个拒绝面);活 owner UUID → 404 op_login_not_found(vouch 已解析、只是请求不存在)。
- **环境刷新(部署态)**:热升级遗留 `FELIS_IMAGE=felis:auditfix16`(job 侧执行器 backup/restore/fileedit/forwarding-init 引用的镜像)→ 已 `set env` api/operator 双 Deployment 至 `felis:auditfix29` 并滚动完成;api/operator/reaper 三镜像一致 = auditfix29;port-forward 重启后内部面绿;panel 200、op 面 `/me` 200。
- **运维备注**:`account_link_codes` 会积累过期行("失败不消费"的另一面),巡检顺手 `DELETE FROM account_link_codes WHERE expires_at < now()`。
### 本轮新增真机证据(第十六批:内部面残面地毯补测 —— ready / join-event / status / reclaim / blacklist / backup 负例)
- **背景**:对内部面做覆盖盘点后发现五个端点此前无真机证据:`ready`、`join-event`、`status`、`player/reclaim`、`player/blacklist`,以及 `internal/servers/{name}/backup` 的负例臂。本批全部补测(auditfix29),**0 新缺陷**。
- **`ready`**:test-one → **204**(advisory 语义);坏名 → 400。
- **`join-event`**(临时 UUID `3333…`):首次 **204**、重复 **204**(幂等);DB 实锤:`test-one.last_active_at` 15:02→05:53(重置)、`server_allowlist` 恰 1 行;缺 uuid → 400;未知服 → 404。演习后 allowlist 行删除、`last_active_at` 复原。
- **`status`**:test-one → 200 全字段(phase/desiredState/endpointMode=fallback/…);未知服 → 404。
- **`player/reclaim`**:新建(临时 UUID `4444…`)→ 200 `hold_expires_at`=+30d;**幂等重试返回同一时间戳**(首窗保留,与存储行 21:53:51.726017Z 完全一致);缺字段 → 400;**protected-admin 负例**:临时插 owner↔thirdparty 链接 → **409 protected_admin**(不 bar、不 stash)。清理后 `username_blacklist`/`player_data_holds`/`server_allowlist`/临时链接 = 0/0/0/0。
- **`player/blacklist`**:bar 后命中 `{"blacklisted":true}`;陌生 UUID → `false`。
- **`internal backup` 负例**:未知服 → 404;坏名 → 400;保留名(login)→ 400 `bad_name`(在入队前拒绝,RWO 闸门单测已覆盖)。
- **探针**:`/healthz` 200、`/readyz` `{"status":"ready"}`。
### 本轮新增真机证据(第十七批:hasJoined 多源会话校验器 —— 假 Yggdrasil 全链路 + 三方身份重写/改名)
- **背景**:`hasJoined`(velocity 指向的 vanilla sessionserver 协议面,`handleHasJoined`)此前零真机覆盖。本批用「VM 主机假 Yggdrasil + 临时将 `[[auth_source]]` 换向」的方式把正/负路径全部打通。
- **装置**:主机 python 假源(`:18099`,按 username 分流 `ftok`/`ftnotch`/`ftbadname`/`ftdown`/其余 204);`felis-config` 的 littleskin 源临时改指 `http://10.42.0.1:18099/fake`(tag=`faketest`/prefix=`FT`),api 滚动后逐项打靶。**演练后配置已还原 littleskin、装置已清理**。
- **结果(全绿,0 缺陷)**:
1. **三方身份重写**:`ftok` → 200 `{"id":"74409c3bbae93acabd2176e517b4f2a0",…}`,与本地按 `uuid.NewMD5(felisAuthNS, "faketest:native-123")` 的预算值**逐位一致**;
2. **保费名冲突改名**:`ftnotch` → `47c5527a18b03fe5a49ec50c11154dcd` + `name="FT_Notch"`(Mojang 实查 Notch=200 premium);`ftok` 的 `FtPlayer` 恰也是真实 Mojang 名(`640c1672…`)→ `FT_FtPlayer`,改名按设计触发(非保费名保持原名);
3. **敌意插件名拒绝**:假源返回 `§4admin` → **204**(不落 proxy 玩家列表);
4. **源故障不静默**:假源 500 → **503**(velocity 报 auth servers down),对照未知玩家 → 204;
5. **形状负例**:缺参 / 超长参 → 204;声明 body → **400 + `Connection: close`**(防 drain 挂连接)。
- **顺带覆盖**:`isPremiumName` 的 Mojang 实查(`api.mojang.com`,超时/错误 fail-closed=改名)——404→非保费、200→保费判别实测成立。
### 本轮新增真机证据(第十八批:缺陷 #35 —— world 执行器 uid 1000 读不了游戏服写的世界 + 面板备份模块补全)
- **发现路径**:给备份页新增「立即备份 + 最近操作」后做第一次真机 drill——面板链路全对(按钮/成功消息/运行态→终态、零坏请求),**但 backup Job 真的失败了**:`Job has reached the specified backoff limit`,pod 日志 `felis backup: archive: backup: tar walk: open /world/world/level.dat: permission denied`。
- **根因(缺陷 #35,跨模块)**:世界卷的属主是**游戏镜像自己的 UID**(我们发布的 Paper 镜像都是 root),而 Paper 保存 `level.dat` 用的是 **0600**(Files.createTempFile 默认权限)→ 固定 uid 1000 的 **backup / restore / fileedit / reaper** 四类执行器:读不了(归档 `permission denied`)、也覆盖不了(restore 写不进 600-root 的 level.dat)。此前 drill 侥幸全绿,是因为当时的世界文件由 uid-1000 工具(restore/夹具)写的;**服务器真实启动保存过一次之后**,所有备份从此必死。测试全部是 shape 断言(无集群),这条只能真机抓。
- **修复(`2010961`)**:四类执行器统一改**以 root 运行**(`runAsNonRoot:false`,省略 fsGroup 防误 chgrp),容器保持除 `DAC_OVERRIDE` 外 drop-ALL——与 operator `init-forwarding` 容器的既定先例同源("只有 root 能可靠读写这些文件");DAC_OVERRIDE 兜住"游戏镜像是非 root UID"的任意镜像场景。四份 shape 测试同步改断言。troubleshooting §10「Permissions」段落重写(uid-1000 + setfacl 时代结束)。
- **绿证据(auditfix31,真机四联 drill)**:
1. **backup**(面板 CDP 实操,红→绿同场景):点击「立即备份」→ `进行中` → **`成功`**(同页面 reload 后仍在);Job pod `runAsUser:0` + DAC_OVERRIDE;新归档 `test-one-1790115428472416736.tar.gz` 实测含 `world/level.dat`(471B)与 server.properties,共 499 条。
2. **restore**:`POST restore-backup`(bk-9df0…)→ 202 → Job **Succeeded**(root 写路径过关),日志 `restored from … into /world`。
3. **fileedit**:`GET /servers/test-one/file?path=world/level.dat` → **200**(base64-gzip 内容 628B)——修复前该请求必然 permission denied。
4. **reaper**(从 live CronJob 派生一次性 Job + 复刻钻取世界:1Gi PV/PVC + `world/level.dat` 512B **0600 root** + marker + 20d idle 行):pod **root+DAC_OVERRIDE**;`world reaped server=reapdrill2 …`、`evaluated=3 reaped=1`;归档 711B 含 marker 与 level.dat;PVC 删除;`world_backups` 得 `inactive_15d` 行;servers 行/CRD 保留(红线②)。**钻取现场全部清理**(job/CRD/行/PV/宿主目录/临时归档)。
- **面板补全(`97a64c8`)**:备份页新增「立即备份」(仅 Stopped 可用,409 原文呈现)与「最近操作」卡(GET `/servers/{name}/jobs`,running 每 5s 自刷新,failed 显示 Job 失败文本——正是它把上面这次失败暴露出来的)。面板 113 单测 + typecheck 全绿。
- **部署注意**:backup/restore/fileedit 的 Job spec 是 api **运行时渲染**,随镜像即生效;**reaper CronJob 的 pod 模板是安装期静态渲染**——本轮已按新形状热补丁 live 对象(root + DAC_OVERRIDE),`felis install/setup` 重渲染时收敛。
### 本轮新增真机证据(第十九批:面板补齐 owner 解绑通行密钥入口)
- **背景**:`DELETE /users/{id}/passkeys`(owner-tier 凭据补救:密钥丢失/被盗时切断登录脚架,且不锁死账号——邮箱码/游戏内审批仍可用)后端早已实现并审计,但面板无入口,owner 只能靠 API。属"功能缺口"而非缺陷。
- **修复(`11ac4f5`)**:用户详情页危险操作区新增「解绑通行密钥」行 + 确认对话框(`api.unbindUserPasskeys` 客户端方法 + 双语 i18n + wire-shape 测试)。
- **绿证据(auditfix32,CDP 真机全链)**:player.test 注册虚拟认证器密钥 `E2E-Unbind` → `/auth/options` 从 `["email_otp"]` 变为 `["passkey","email_otp"]` → owner 在 `/admin/users/<id>` 点「解绑通行密钥」→ 对话框 → 确认 → **凭据列表清空、options 回落 `["email_otp"]`**(passkey 门确实关闭);审计 `user.unbind_passkeys` 落账(同批还可见 `account.passkey.registered`/`backup.create`/`backup.restore`/`reap_world` 各审计行)。面板 114 单测全绿。
### 本轮新增真机证据(第二十批:缺陷 #36 —— 提交者看不到构建结果)
- **缺陷 #36(提交/构建模块·结果不可见)**:`/me/submissions` 只回审核状态;构建的成功/失败(含失败原因)只有 admin-tier `/images/build/{id}` 看得到 → **提交者永远不知道自己的包构建死了**。修复 `72c4aa3`:两个列表路由(玩家 `/me/submissions` + 管理 `/submissions`)对 `build_id` 非空的行附 `build_status`/`build_error`,来源是只读 `Builder.Get`(**绝不调 Sync**——状态推进归 15s reconcile 循环,列表渲染不碰集群);构建行已消失(ErrNotFound)→ 字段省略;其他存储错误照常 500,绝不静默吞。OpenAPI 的 Submission schema 同步。
- **绿证据(auditfix33,真机双例)**:① 失败例(上下文 Dockerfile `COPY does-not-exist`)→ 提交 → owner approve → `/me/submissions` 返回 `build_status:"failed"` + `build_error:"build job failed or scan found a CRITICAL CVE"`;② 成功例(`FROM scratch`+LABEL)→ `build_status:"succeeded"`。面板 CDP(玩家会话):展开行显示「构建状态」徽章(构建失败/构建成功)+ 失败原因 + build_id,全程零 4xx。**演习残留已清**(2 行 submission+build、2 个 context blob、1 条 whitelist 条目;`sub-bf7dc1…` 那条是更早 E2E 遗留,未动)。
### 本轮新增真机证据(第二十一批:面板文件编辑器补齐 —— 缺口而非缺陷)
- **背景**:`GET /files`、`GET/PUT /file` 后端早已全绿(可读写 `level.dat`),但面板无入口——「一行 server.properties 写错导致起不来」的修复路径只有 API。属功能缺口。
- **修复(`0a36b3f`)**:新增 `/servers/:name/files` 页:面包屑目录浏览、编辑器对话框([]byte ↔ base64 编解码)、二进制文件打开即只读(NUL/非 UTF-8 拒绝 round-trip)、>256KiB 禁用保存;**停服门前置**(世界卷 RWO,未停服时整页显示「服务器正在运行」+ 停止动作,而不是让每个调用 409);控制台右侧新增门口卡。i18n `files` 命名空间(en/zh)+ 3 条 wire-shape 测试。
- **绿证据(auditfix34,CDP owner 全链)**:根目录 → `world/` 导航;`felis-e2e-marker.txt`(原 `v1\n`)打开 → 追加 → 保存「已保存 …」→ **API 读回一致** → 重开一致 → 还原原始字节 → 落盘复核一致(零残留);`level.dat` 打开为只读 + 二进制提示;控制台门口卡存在;`audit_logs` 两条 `file.write`(edit/restore 各一);全程零 API 4xx/5xx。面板 117 单测 + typecheck 全绿。
### 本轮新增真机证据(第二十二批:缺陷 #37 —— 系统服务在面板里全是死操作)
- **缺陷 #37(面板·fleet 死操作)**:`login`/`lobby` 是平台自建系统服务、名字在保留名单里,于是**每个 per-server 路由都用 `ValidateServerName` 拒绝**(400 `bad_name`)——但管理端 fleet 表格给这两行渲染 认领/停止/唤醒/控制台,全是死操作(控制台链接点进去也是一页 `invalid server name: reserved`)。修复 `2f90851`:`naming.IsSystemServer` 作为唯一事实源;fleet 行附 `system:true`;面板把这两行渲染为「系统服务」纯标签(owner 列 + 操作列),不再给任何动作。玩家侧不受影响(`/me/servers` 本就不含系统服务)。
- **绿证据(auditfix35,真机)**:`GET /fleet`(owner)→ `lobby system:true`、`login system:true`、`test-one/resolvecheck` 无 flag;面板 CDP:两行「系统服务」、**0 按钮 0 链接**;`test-one` 行照常 认领/控制台、无系统标记;零 4xx。Go 侧新增 `TestIsSystemServer` + `TestFleetAdminRead` 的 system 子测试。
### 本轮新增真机证据(第二十三批:缺陷 #38 + 多节点回收缺口 —— reaper 提示过期 / 钉节点)
- **缺陷 #38(CLI·提示过期)**:`felis manifests` 渲染 reaper 时的 stderr 提示还在教“uid 1000 需要 `setfacl -m u:1000:x` 才能遍历存储根”——#35 之后 reaper 已改为 **root + DAC_OVERRIDE**,这条指导已失效且会误导运维(照做无害但白做,真问题被掩盖)。同批落 **多节点回收缺口**(结论第 5 条):新增 `--reaper-node`,reaper CronJob 的 pod 渲染 `nodeSelector kubernetes.io/hostname=<node>`;多节点集群必须钉在存世界的节点,否则可能调度到 hostPath 为空的节点。`--reaper-node` 无 `--worlds-host-path` 时 fail-loud exit 2。修复 `daf7602`(双测:`TestReaperCronJob_NodePin`、`TestManifestsReaperNodePin`)。
- **绿证据(auditfix36,真机 render + dry-run + 收敛 diff)**:
1. **固定渲染**:`felis manifests --felis-image felis:auditfix36 --velocity-cidr 10.211.55.6/32 --panel-node-port 30443 --worlds-host-path /var/lib/rancher/k3s/storage --archive-local-path /var/lib/felis/archives --reaper-node localhost.localdomain` → exit 0;bundle 内 `kubernetes.io/hostname: localhost.localdomain`;stderr **0 处** uid 1000 / setfacl,改为「已钉到节点 …」;
2. **负例**:`--reaper-node` 无 `--worlds-host-path` → exit 2 + 原文;不传 node 的渲染 stderr 仍完整保留「NO nodeSelector … 多节点必须传 --reaper-node」警示,bundle 内 0 个 selector;
3. **`kubectl apply --dry-run=server -f -`**:整包 **全部 configured**(含新 nodeSelector 的 CronJob);
4. **再安装收敛性 diff**(`kubectl diff`,本批新增的收敛证据):与 live 对比只剩 **一个对象**(reaper CronJob)两处实质增量——`+env FELIS_SMTP_PASSWORD`(热升级期未补的模板字段)与本次显式传入的 `+nodeSelector`;其余 24 份文档(api/operator Deployment、RBAC、NetworkPolicy、PVC、registry)**零差异**——即“热补丁过的 live 对象”与“当前代码重渲染”已收敛(reaper 安全上下文 root+DAC_OVERRIDE 两侧一致,无 diff)。
### 本轮新增真机证据(第二十四批:S3 上传通道真机演练 —— 此前只有单测的暗路径)
- **背景**:`internal/submit/s3store.go`(S3ContextStore)此前只有单测,本装是 local 路径(`user_uploads_context = "/var/lib/felis/uploads"`)从未激活。本批用「VM 宿主 MinIO(quay.io 镜像,:9000)+ 临时把 felis-config 切到 `s3://felis-user-uploads` + `[registry.s3] endpoint=http://10.211.55.6:9000` + 创建 `felis-uploads-s3` Secret」把整条通道打通,**全程 0 缺陷**、演练后还原并逐项复核。
- **绿证据(auditfix36,真机全链)**:
1. 切 S3 后 api 滚动启动 **无** “S3 user-uploads store not configured” 告警(凭据解析成功);pod 内 busybox 探针实测可达 `http://10.211.55.6:9000/minio/health/live`(rc=0);
2. 玩家提交 + 上传 → **200**,对象实测落桶:`sub-12c953e12b0cd144/context.tar.gz`(197B,mc ls 实见);
3. owner approve → 构建 Job `build-bld-1790118222183394193` **status.succeeded=1**(fetch-context 从内部面流式取件 = api 自 S3 读回成功);`/me/submissions` 的 `build_status` 收敛为 `succeeded`;
4. **回滚**:felis-config 还原(sha256 与演练前备份**逐字节一致**)、删除 `felis-uploads-s3`、api 滚动;再演练一次本地路径:新提交上传 **200** 且 blob 实测落在 uploads PVC(context.tar.gz 197B);启动日志仅剩 smtp/jwks 两条既有提示;
5. **清理**:2 行 submission + 1 行 build 删除、回滚演练 blob 删除、S3 构建产物从 whitelist 摘除(204)、MinIO 容器 + 两个镜像移除;`/fleet` 200。
- **遗留观察(非缺陷)**:registry 里保留本次推送的 `user-uploads/sub-12c953e12b0cd144:latest` 层数据(与早前 direct-probe 同类,filesystem registry 无删除接口);`felis setup` 的「S3 存储」向导屏本身未演练(列入 TUI 逐屏待办)。
### 本轮新增真机证据(第二十五批:breakGlass 控制台逐屏全量 + 备份/恢复门禁 —— 缺陷 #39–#43)
- **背景**:breakGlass 此前只验过 Owner 首装与 Add Operator happy path。本批把菜单四操作(Owner reset / Add Operator / Halt / Sync)+ 首装(bootstrap)分支全部逐屏真机走完,并打穿 Sync/恢复背后的 API 门禁;共抓 5 个缺陷、全部修复复验。驱动方式:VM tmux(`remain-on-exit on` 才能读回 alt-screen 撕掉后的 durable summary)。
- **先落的正向证据(无缺陷)**:Halt——`test-one` Running → 选中 → 卡「is stopping」→ CRD `desiredState=Stopped`、pod 收敛消失、审计 `break_glass.halt`;面板 `wake`/`stop` 两个恢复杠杆均 202(复验后还原)。Sync 正路径——test-one(Stopped)→ 卡 backup started → Job 6s Complete → 归档落 `felis-backups` PVC(167MB)→ `world_backups` 行 `present` → `/api/v1/backups` 可见。Sync 负路径——对 Running 选 → 友好 409 卡(不误烧冷却)。
- **缺陷 #39(owner 席位可被静默复制,且不可清理)**:恢复模式下用非在位席位名做「reset」→ `UpsertOwner` insert 臂**铸出第二个 owner 行**、原席位继续存活;面板对任何 owner 行都删/降/禁 403 → 永久无法收敛回单席。修复 `55d515d`:`provisionOwner` 先查在位席位(新 `PGRepo.OwnerUsername`),非席位名 → `ownerSeatTakenError`(`Is api.ErrConflict` → TUI 路由回表单并**指名**应输入的用户名);bootstrap(无席位)与同席位名复位原样。真机双验:新名被拒(表单原位显示 `an Owner already exists as "08595879-…" — enter that username to reset the Owner`);改席位名复位成功(owner 恒 1 行、id/邮箱不变);面板删除保护同步实证(对新 owner 行 DELETE → 403)。
- **缺陷 #40(operator 撞名 = 裸 SQLSTATE,「换名重试」分支在真机从未生效)**:`InsertOperator` 冲突返回原始驱动错误(23505),TUI 却按 `api.ErrConflict` 判定「可恢复、换名重试」——fake 与 PG 漂移。真机复现:输入已存在用户名 → 控制台 exit 1 + 裸错误。修复 `55d515d`:`isUniqueViolation → ErrConflict` + pgint 契约断言。复验:撞名**回到表单**提示换名 → `drill-op-2` 成功(审计 `break_glass.operator_create` 落账,演练行已清)。
- **缺陷 #41(Sync picker 死选项)**:picker 列出 login/lobby,而备份 API 对它们**永远失败**(保留名 + 无 servers 行);真机选 lobby → 裸内部错误 + exit 1。修复 `ac3a557`:`backupPickable` 过滤系统服(halt picker 保留它们,断玻璃完整权力)。
- **缺陷 #42(缺失世界盘 → 202 后静默卡死 30 分钟)**:对「从未启动/已被回收」的服备份或恢复:202 → Job → Pod `persistentvolumeclaim "world-<name>-0" not found` **Pending 至 deadline**,全程零失败记录。真机用已回收的 `resolvecheck` 复现(留证后删除)。修复 `508a1c0`:`Cluster.WorldVolumeExists`(直接 Get 与 Job 挂载**同名**的 PVC)+ 两 handler 409 `no_world_volume`("start it once to create it, then retry")。**修复首跑翻出配套 RBAC 洞**:felis-api SA 无 `persistentvolumeclaims:get`(403 被吞成 500「internal error」)→ `APIMinecraftRole` 补 get-only 规则 + rbac 测试锚点。复验(auditfix38):resolvecheck 两面 409 + 友好文案;test-one 照常 202 → Job 10s Complete → 新行落库。
- **缺陷 #43(409 一刀切文案)**:TUI 把所有 409 当停服门 → 世界盘拒绝会展示错误原因。修复 `ac3a557`:按 body 的 `error.code` 分流(无 code 的旧体仍按停服门)。复验:无盘 pick 显示 API 原文;停服门文案不变。
- **同步完成**:① bootstrap 分支 scratch 库演练(`migrate up` 20 迁移 → 无菜单/无认证直接铸 owner;`local_auth_enabled=true`;审计 `break_glass.bootstrap`;exit 0;库/hba 规则/临时配置即测即清);② 台面收敛:reaper CronJob `suspend=true/auditfix36/无 pin` → `suspend=false / felis:auditfix38 / nodeSelector=localhost.localdomain`;③ 镜像升级 `auditfix37→38`(api/operator + `FELIS_IMAGE`)。
- **流程修正(教训)**:pgint 一度误用**本机 Docker Desktop**(启动 daemon + 临时 PG 容器)——已完全清理(容器/镜像删除、daemon 退出),并改为**经 ssh 隧道用 VM 的 postgres** 运行(`ssh -L 15433:127.0.0.1:5432` → `postgres://felis:***@localhost:15433/felis_pgint?sslmode=disable`)。勿再在本机跑容器。
### 本轮新增真机证据(第二十六批:`felis setup` 重跑向导逐屏 + 构建链 pin 回验 —— 缺陷 #44)
- **范围**:本机已装机,故覆盖"重跑状态屏 + c/s/e 三条 reconfigure 流";首装屏(postgres/owner/connect/edge/storage/smtp/mc-bind/migration/preflight/summary)此前各批已有定点真机证据(安装闭环 / 第二~三批 / 第九批 / 第二十五批),本轮不重复。
- **逐屏走查(auditfix38→39,VM tmux)**:
1. 重跑 → 直落状态屏「✓ Felis is already set up.」(owner/connect 不触碰;host bootstrap 已就绪跳过)✓
2. `c` → 三选一 chooser(Local / Cloudflare+Access / Reverse proxy + 警示语)渲染 ✓,esc 无损返回。
3. `s` → chooser 预选当前后端;S3 分支表单(Endpoint/Bucket/Region/AK/SK + 提示)渲染 ✓。**观察:reconfigure 非只读**——选「Local disk」即 apply(写 /etc 两文件 + 重渲染 felis-config + 滚 API);从 S3 表单 esc 退回会把选择重置为 Local 预选。
4. `e` → SMTP 表单渲染 ✓;esc 直接回状态屏、零副作用(felis-smtp 未创建)✓。
- **缺陷 #44(CLI·重跑框脱落)**:重跑后完成 `s`/`c` reconfigure,落回首装 summary「✓ Setup complete.」——丢了 alreadySetUp 框(smtp 路径有专门分支,storage/connect 漏)。修复 `abb5910`(`showSummary` 透传 `m.result.alreadySetUp`)+ 回归测试 `TestRootReconfigureStorageKeepsStatusFraming`。真机复验(auditfix39):storage reconfigure 完成 → 「✓ Felis is already set up.」+ storage recap ✓。
- **演练事故(自曝;环境 drift,非产品缺陷)**:首次 storage 走查意外触发 Local apply——它从 `/etc/felis/felis.pod.toml` 重渲染 Secret,而构建链 pin 值当初**只热补在 live Secret、不在 /etc 文件** → 重渲染清空 pin(放任则下次构建死在 `:latest` 拉取 + trivy DB egress)。当场修复:pin 值写回 `/etc/felis/felis.host.toml` + `/etc/felis/felis.pod.toml` → 重渲染 Secret → 滚 API;再做第二次 storage apply,重渲染后 pin 仍在(drift 修复耐久)。
- **构建链回验(direct build,真机)**:Job 规格实证 `kaniko=gcr.io/kaniko-project/executor:v1.24.0`、`trivy=aquasec/trivy:0.74.0`、`--db-repository registry.felis.svc:5000/mirror/trivy-db:2`(registry 仍有 `mirror/trivy-db`);`POST /images/build`(context=遗留 `sub-bf7dc18…` blob,内部面取件)→ 202 → **succeeded**;registry `e2e/pins-check` 落位;whitelist 条目已摘除(204)、Job 已清。
- **文档修正(`ae6e925`)**:§8e 原「编辑 felis.toml 后重启 felis-api」不完整(API 挂的是 Secret)→ 改为「写进 /etc 两文件 → 重渲染 Secret → roll」,并写明三种无效/易损做法(只 restart / 只改 host 文件 / 只热补 live Secret——后者会在下次 reconfigure 被冲掉)。
- 收尾:api 1/1、panel 200、tmux 全清。
### 本轮新增真机证据(第二十七批:告警模块落地 —— 内部面 /metrics + 规则集 + 真实构建失败实弹演练)
- **范围**:把"指标 → 规则 → 告警"链路从零补到可交付:API 内部面 `/metrics`(`94f71ee`)、`deploy/alerts/` 规则与 promtool 单测(`43df08b`)、真机实弹演练(本批)。
- **/metrics(`94f71ee`)**:`felis_image_build_failures_total` 此前只在进程内存里、无任何 scrape 出口。修复:internal face(8081)新增 `GET /metrics`(服务面 Public 路由,语义同 /healthz);单测断言外部面 404。真机:port-forward `svc/felis-api-internal 18081:8081` → 200 且含 `felis_image_build_failures_total`;operator `:8080` 提供 `felis_servers_total` / `felis_start_duration_seconds_*`。
- **规则集(`43df08b`)**:`deploy/alerts/felis-alerts.yaml`(plain Prometheus)5 条——构建失败 increase>0 / 起服 p90>300s / 磁盘可用<15% / DiskPressure / 内存可用<10%;`felis-prometheusrule.yaml` 为 prometheus-operator twin(脚本比对两文件 groups 一致);`felis-alerts_test.yml` 为 promtool 单测。VM 上 promtool 3.14.0 实跑:`check rules` + `test rules` 双 SUCCESS。
- **实弹演练(真实构建失败 → pending → firing)**:
1. 打包含 `COPY does-not-exist` 的 Dockerfile 上传到 uploads PVC `sub-alertdrill` → `POST /images/build`(`e2e/alert-drill2:latest`);
2. Kaniko `failed to get fileinfo for /context/does-not-exist` → Job Failed、build 行 `failed`;
3. 真实 Prometheus(宿主 `:19090`)scrape `127.0.0.1:18081`(api) 与 `:18080`(operator) 双 target up;`felis_image_build_failures_total{job="felis-api"}=1`;
4. `FelisImageBuildFailures` pending(activeAt 08:28:14Z)→ **08:33:14Z 准时 firing**(`for: 5m` 精确到期),labels/annotations 完整。
- **清尾(残留全清)**:whitelist `e2e/alert-drill` 摘除(204);`sub-alertdrill` 目录、`/tmp/drillctx`、两个演练 Job、tmux `prom`/`fwd`/`alertpoll`、`/root/prom-drill`(promtool+prometheus 二进制)全删;DB `%alert-drill%` 行删除(whitelist/build 复核 count=0);宿主无残留监听/进程。演练期间 live api 进程内计数器=1(重启归零,属演练事实)。
- **可达性追加**:#44 定级 ①(轻)——已装机环境重跑 `felis setup` 完成 storage/connect reconfigure 即触发。
### 本轮新增真机证据(第二十八批:缺陷 #45 —— 审核门"盲批":评审看不到将被执行的 recipe)
- **缺陷 #45(构建 lane·审核语义)**:被执行的 Dockerfile 永远来自**上传上下文压缩包根目录的 `Dockerfile`**(`build/jobspec.go` 钉死 `--dockerfile=Dockerfile`),API 的 `dockerfile` 字段**仅审计存档**(`submit.auditDockerfile`、`build.Request` 注释均已声明)——但审核者没有任何路径能看到它:`GET /api/v1/submissions` 不含 blob 内容、面板只显示 `context_ref` 文本、内部面取件路由是 service-token(评审用不了)→ "人工审核是门禁"事实上是**盲批**。同批口径缺口:`POST /api/v1/images/build` 的 `dockerfile` 字段在 OpenAPI 里无任何说明(易被当成"将被执行"),面板表单也把该框呈现为"Dockerfile 内容 *"。
- **修复(`168a375`)**:
- 新增 admin-tier `GET /api/v1/submissions/{id}/context`:评审下载与构建 Pod 同源同字节的 `context.tar.gz`;`Content-Disposition: attachment` + `nosniff`(攻击者提供的归档只下载、不渲染);审计 `submission.context.download`(actor=评审者、target=submission id)。
- 内部面取件 handler 共享 `openSubmissionContext`/`streamSubmissionContext`(行为不变,原测试锁定)。
- OpenAPI:新路由 + `/api/v1/images/build` 字段描述补全(明说"执行的是 context 根目录的 Dockerfile;`dockerfile` 仅审计")。
- 面板:SubmissionsPage 展开区新增「下载上下文」按钮(spinner/错误呈现,i18n en/zh);ImageBuildPage 表单补审计说明行。
- 测试:`TestAdminSubmissionContextRoute`(流式/404/503)+ admin-only 矩阵加该路由 + OpenAPI parity 强制文档;面板 117 单测 + 构建、go vet/go test 全绿。
- **真机验证(auditfix41 已部署;owner 会话经 op-login + 内部面代 approve 重铸)**:
- admin 下载 `sub-bf7dc18e9dd97dc2` → **200**,`attachment; filename="context.tar.gz"`、`application/gzip`、`nosniff`;sha256 `205496f2…` **与 uploads PVC blob 逐字节一致**;
- 内部面(Bearer=felis-service-token)同 blob → 200 + 同 sha256(重构未破坏构建取件路径);无 token → 401;
- 有效玩家会话(console host 隔离,仅 adminOnly 生效)→ **403**;匿名 → 401;不存在 id(admin)→ 404;
- 审计落账:`audit_logs` = `[email protected] | submission.context.download | sub-bf7dc18e9dd97dc2`;
- 面板产物:服务端 index.html 引用新构建 `index-CwFSKpgW.js`,bundle 内含新按钮逻辑(grep 命中 3 处)。
- **可达性追加**:#45 定级 ①——每一次真实的"用户提交 → 管理员审核"都会踩到(审核者此前无法查看将被执行的内容)。
- **hook 链补齐(可复用)**:staff 账号走邮件登录门会被设计拒绝(refuse staff)→ owner 会话铸法:`op-login/start`(email=felis-owner@example.com)→ VM 日志 grep `email-otp` 取码(no-Mailer fallback)→ 内部面 `op-login/{id}/approve`(Bearer=felis-service-token;body `approver_uuid`=owner 的 mc_uuid)→ `op-login/finish`(curl -c 存 cookie)。
### 本轮新增真机证据(第二十九批:镜像耐久落地 —— registry 托管 + 回环拉取路径;#46–#49)
- **背景(结论清单第 2 条)**:磁盘压力演练证明 kubelet 会 GC 掉"当前无人使用"的镜像 → ImagePullBackOff,恢复依赖人工重导入。本批让镜像自愈:自建镜像全部托管进内建 registry,节点侧 pull 经回环 hostPort(节点 containerd 到 Service VIP 是死路,实测 "Empty reply")。
- **平台侧 `a9b275a`**:registry 容器端口加 `hostPort 127.0.0.1:5000`;同提交修 **#46** —— registry 独立资源模板(1 CPU / 2Gi):旧模板 256Mi 在实测推 475MB 层时被 OOM kill(dmesg `oom-kill … registry, oom_score_adj=989`,上传中断),2Gi 下同一推送 2 秒完成。
- **安装器侧 `13d64e0` + 文档 `fa0e8d7`**:自建镜像规范 ref = `registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo`;import 进 containerd 就用该名(首启命中本地,免 registry round-trip),`deploy_bundle` 之后统一 `push_images_to_registry` 入仓(推 `127.0.0.1:5000`;registry 只认主机名之后的路径 —— 推/拉落点一致)。`configure_registry_mirror` 写 `registries.yaml`(`registry.felis.svc:5000 → http://127.0.0.1:5000`),内容不变不重启 k3s;`import_registry_image` 预缓存 registry:2(重跑走跳过分支)。迁移 **0021** 把 recommended 白名单重指到 registry ref(线上实查两行已落)。
- **真机三次重跑**(`/opt/felis/src` = 765a892 快照;`FELIS_SKIP_FETCH=1 FELIS_INSTALL_MODE=full FELIS_IMAGE=registry.felis.svc:5000/felis/felis:auditfix42`):
- run1 失败 = **#47**:Mac tar 的 `._*` 旁文件混入构建上下文,`._0004_*.sql` 被 //go:embed → 新二进制的 `felis migrate` 报 `non-numeric version "."`,安装器停在 run_migrations。修复 `5fa8b74`(.dockerignore 排除 `._*`/`.DS_Store`)。复验方式:故意在暂存树留 `._zz_probe_junk.sql` → 重建后 `migrations applied`(过滤生效)。
- run2 失败 = **#48**:每个镜像一对 `systemctl start/stop docker` 触发 systemd 限流(`Start request repeated too quickly / start-limit-hit`),第 4 个镜像(paper)未入仓。修复 `c7e585e`(整批一次 start/stop,单测锚定)。run3 零 `[fail]`:4 镜像全部入仓(push digest ×4 实收)。
- run3 收敛实查:api/operator/reaper = `registry.felis.svc:5000/felis/felis:auditfix42`(registry ref 首滚命中本地 import);registry 模板 1 CPU/2Gi 生效;reaper CronJob 同步换 ref 且 nodeSelector 保留;login/lobby Running 于 registry ref;迁移 applied。
- **GC 演练(本批验收本体)**:
- A 控制面:`ctr images rm …/felis/felis:auditfix42` + `ctr content prune references` → 本地 ref 消失 → `rollout restart felis-api` → 事件 `Pulling` → `Pulled … Successfully pulled image … in 25ms`,ref 恢复、pod Running。
- B 游戏:rm `…/felis/lobby:demo` → 删 `lobby-0` → `Successfully pulled … in 10ms … Image size: 182357358 bytes`,Running。
- #46 复验:整轮重建 + 4 推送期间 dmesg `oom-kill` 计数不变(仍 2,历史)。
- **构建 lane 复验**:`POST /images/build`(context=遗留 `sub-bf7dc18…`;ref `registry.felis.svc:5000/e2e/durable-check2:latest`)→ 202 → succeeded;registry `e2e/durable-check2` tags 落位;Job 事件见 felis/kaniko/trivy 三镜像就绪;白名单条目摘除(204)、Job 清。
- **#49(同一条升级路径的第二类静默回退)**:`write_felis_toml` 重写整表([registry] 仅 url+build_namespace、[archive] 仅 store+local_path)→ §8e 的 kaniko/trivy pin、构建上限、uploads 后端、[registry.s3]、reaper 的 retention/warn_before/max_local_bytes 在重跑时全部丢失(S3 安装切回 local、构建回退被 egress 拒绝的上游 executor)。修复 `765a892`+`b8e554d`(沿用 [smtp]/[[auth_source]] 的 carry 模式;url/build_namespace/store/local_path 保持安装器所有)。真机复验:预置 `retention = "30d"`,run3 后 host toml / pod toml / felis-config Secret 三层都在,且构建 lane 直接用 carry 的配置跑通(上条)。
- **可达性**:#46 ①(用户构建大层或安装器入仓即触发;实测 475MB 层);#47 ②(Mac 打包树构建路径,真机踩中);#48 ②(安装器重跑,真机踩中);#49 ②(升级=重跑安装器,真机踩中)。
- **口径/遗留**:`demo-up.sh` 未改(dev/demo 路径,本地 tag 导入维持原状);kaniko/trivy 编译默认值未动(§8e 改为"镜像进 registry"配方);registry 2Gi 为渲染常量(暂未开 flag);drill 残留:registry 里 `e2e/*` 小镜像留档,`durable-check2` 白名单条目已摘。
### 本轮新增真机证据(第三十批:缺陷 #50 —— `[smtp]` carry 吞掉下一节的注释块,每次重跑 +1)
- **发现路径**:为 `felis setup` 首装连续走查做前置盘点时读 live 配置——`/etc/felis/felis.host.toml` 已经堆了 **3 份**、`felis.pod.toml` **4 份**重复的 Yggdrasil 注释块(同一段文案逐次叠加);用脚本自带的提取器实测:一次重跑会把 3 份全部当作 `[smtp]` 内容带走,再叠一份模板注释 → **每次重跑 +1、无上界**(host/pod 增速不同步,现场 3 vs 4 即历史残迹)。
- **根因**:`persisted_smtp_block` 打印"`[smtp]` 到下一个 section header 之间"的**所有行**;generated 注释块正好落在这段区间里 → 被吞并。同文件的自称"cached on first call"缓存因写方是命令替换(子 shell)从未生效,pod 写实际上重复抽取刚被重写的 host,加剧了两文件的不同步。纯注释膨胀、无功能损失,但属 #49 同族 carry 语义缺陷(把不属于自己的内容也带走了)。
- **修复 `4d3c85f`**:carry 改为白名单(section header + 键行),与 #49 的 `[registry]`/`[archive]` 提取同型;删掉失效缓存说明。`deploy/bootstrap_test.sh` 新增用例:配置值被携带 / 不吞注释行 / 不越节 / 写回后二次抽取**字节稳定**(幂等)。
- **真机验证(auditfix43,两次连续全量重跑)**:
- run1(`bootstrap-auditfix43.log`):零 `[fail]`;host 注释块 **3→1**、pod **4→1**;与 run 前快照 diff 恰为 21/31 行(全部是被删掉的重复注释)——值零漂移(kaniko/trivy pin、`trivy_db_repository`、uploads local、`[registry.s3]` 空、`retention="30d"`、auth_source 原样、空 `[smtp]` 节保留;两文件 db host 仍分别为 127.0.0.1 / 10.211.55.6)。
- run2(`bootstrap-auditfix43b.log`,紧接再跑):零 `[fail]`,仍 **1/1** —— 收敛证明(旧代码此处会 1→2 继续增长)。
- 部署同步:api/operator = `registry.felis.svc:5000/felis/felis:auditfix43`;reaper CronJob 同 ref 且 nodeSelector 保留;registry `felis/felis` tags = auditfix41/42/43;panel 200;host 二进制已从新镜像提取。
- **可达性**:#50 ②——配置过 email(存在 `[smtp]` 节)的安装,按文档升级=重跑安装器即触发;危害=配置注释无限膨胀(每次 +1),无功能损失。
### 本轮新增真机证据(第三十一批:首装连续走查 + 缺陷 #51 —— 工作负载 `felis-config` 副本永不刷新)
**一、`felis setup` 首装单次连续走查(队列第 1 项,完成;0 缺陷)**
- **装置**:scratch 库 `felis_scratch`(新建 + `felis migrate up` 21 条迁移)+ scratch 配置(真实 pod toml 副本,仅换库名;root 0600);hook 直插 `account_link_codes` 一枚绑定码(等价 `/link` 内网端点写入,代替"进服拿码");tmux 驱动 `felis setup -config`;k8s 只读复用(登录门 Ready 等待通过)。
- **连续走查(一条会话走完)**:Preflight(自动:PG✓/迁移 21 applied/面板✓)→ **MC 绑定**(输码 → working →「✓ Owner account is ready.」+ 一次性 setup URL)→ **连接 chooser**(Local)→ **存储 chooser**(Local → working → ~20s 后「✓ Local storage configured.」,含 Secret 重渲染 + API rollout)→ **首装 Summary**(「✓ Setup complete.」+ owner/setup URL/access/storage/panel 卡片 + c/s/e 提示)→ **轨道回顾**(← 依次只读 recap Storage→Connection→Owner→Preflight,→/esc 回到前台)→ Enter 退出 → stdout 汇总(`Owner account … provisioned (passwordless)`、`Recorded as "root"`、setup URL、Admin console)→ **EXIT=0**。
- **结果**:移动端/文案/切换全部符合设计,**0 新缺陷**;scratch 库侧复核:owner(role=owner) 1 行、绑定码已消费(0)、setup_tokens 1、`local_auth_enabled=true`。
- **副作用与还原(如实记录)**:存储 apply 会把 `/etc/felis` 两文件重写为 Go encoder 形态(无注释、含空值键如 `[archive.s3]`,**值零漂移**:kaniko/trivy pin、uploads local、retention 30d、auth_source 原样),并重渲染控制面 Secret + 滚 API——这是该向导的既定行为;演练后按 pre 快照整文件还原(sha256 逐字节一致),两 ns Secret 重渲染复核一致,scratch 库/配置/hba 行/tmux 全部清理。
**二、缺陷 #51(工作负载 `felis-config` 副本永不刷新)**
- **发现路径**:上条还原核对时发现 minecraft ns 的 `felis-config` 是**旧形态**(encoder 式)而 felis ns 已是模板式 → 挖出 `ensureSecretReplica` 的「绝不覆盖既有副本」(凭据语义:防冲掉手工轮换值)把 **felis-config 也纳入只建不更**,而 bootstrap 只 apply 控制 ns。后果:改配置后(DB 凭据轮换、[archive] 保留策略调整、root domain 等)backup/restore/fileedit Job 与 reaper 永远读旧副本 → 静默失效(如备份认证失败)。
- **红证据(auditfix43,真机)**:向 minecraft 副本注入 `# drill-51-stale-marker` → 运行 `felis setup` → 输出 `- config (minecraft ns): skipped (already exists)`;副本 sha `5ec2ff6e…` 保持,控制面 `e4791fe1…` 不同(陈旧坐实)。
- **修复 `328e570`**:① setup 侧:`ensureSecretReplica` 增 `refreshExisting`(仅 felis-config 传 true)——源缺失降级 skip、内容一致 skip(`already current`)、不同则原地 Update;凭据类保持 create-if-absent;新增 `updated` 结果与「refreshed from the control namespace」文案。② bootstrap 侧:新增 `apply_felis_config_secrets()`,每次运行同时 apply 控制 ns + 工作负载 ns 两份(同渲染自最新 pod toml)。单测:Go 新增 4 例(陈旧刷新/一致跳过/空键补写/源缺失降级);`bootstrap_test.sh` 新增 4 断言(两 ns apply、同一 pod toml 渲染、恰 2 次 apply)。
- **绿证据(auditfix44,真机)**:A) 安装器重跑(标记仍在副本中)→ 日志出现 `secret/felis-config configured`(工作负载 ns 被刷新)→ 副本 sha 与控制面一致、标记 0;B) 再注入标记 → 运行**新** `felis setup` → 输出 `- config (minecraft ns): refreshed from the control namespace`、凭据仍 `skipped (already exists)`、副本 sha 一致、`EXIT=0`。部署=auditfix44(api/operator/reaper),panel 200,host 二进制随镜像 `docker cp` 刷新。
- **可达性**:#51 ②——升级=重跑安装器、或改完配置跑 setup 即触发;旧行为下 backup/reaper 静默使用旧配置(DB 轮换后备份全挂)。
**三、环境修复(非产品)**:现场 pg_hba 缺 `felis_pgint` 规则(按文档走 ssh 隧道跑 pgint 会 ident 失败)——补回 `host felis_pgint felis 127.0.0.1/32 scram-sha-256` 并复测连接成功。
### 本轮新增真机证据(第三十二批:S3 存储向导逐屏走查收尾 + 缺陷 #52 —— 向导 apply 不刷新工作负载 `felis-config` 镜像)
**一、S3 存储向导屏走查(队列第 2 项,完成;0 功能缺陷,走查自身暴露镜像缺口 → #52)**
- **装置**:VM 宿主 MinIO 容器(`quay.io/minio/minio`,`felis`/`felis-drill-9000`,:9000)+ bucket `felis-wizard-uploads`;tmux 驱动 `felis setup` 重跑向导;行动前先留 preS3 快照(两 toml + 两 ns Secret + sha256)。
- **负例**:错误凭据 → `✗ Could not save storage settings.` + `submit: s3 credentials rejected: The Access Key Id you provided does not exist`(`CheckS3Access` 预检先于一切写入);**零副作用**(无 Secret、两 toml sha 不变);`esc` 返回编辑时已填值保留(密钥掩码)✓
- **正例**:修正凭据 → ~10s working → `✓ Object storage configured.` → Enter → 状态屏 `storage s3://felis-wizard-uploads · http://10.211.55.6:9000` ✓
- **落地核对**:`felis-uploads-s3` Secret(access_key_id/secret_access_key);两 toml `user_uploads_context` + `[registry.s3]` endpoint/refs;API 滚动;启动日志仅既有的 smtp/jwks 警告 ✓
- **功能链(batch24 同款)**:player.test 邮箱 OTP 登录(hook 取码)→ `POST /api/v1/me/submissions` 201 → context 上传 200 → MinIO 桶实见 `sub-853a4e2ba4ba6443/context.tar.gz` → 内部面取回 200 + tar 内容正确 ✓
- **UI 回滚(`s` → Local)**:两 toml 归位(`/var/lib/felis/uploads`、`[registry.s3]` 归空)、控制面 Secret 更新、API 滚毕;与 preS3 快照的差异仅「s3 字段回环 + encoder 形态」(注释丢失属该向导既定行为)——**值零漂移** ✓
- **观察(不计缺陷)**:切回 Local 后 `felis-uploads-s3` Secret 残留(演练按清理流程删除;是否自动清理属产品取舍)。
**二、缺陷 #52(向导内 apply 只刷新控制面,工作负载镜像滞后到下一次运行)**
- **发现路径**:回滚后按计划复核「两 ns 重渲染」——minecraft 副本仍为 S3 内容(`4fb80ed8…`),而控制面与两 toml 已 local(`df206074…`)。
- **红证据(auditfix44)**:① S3 方向:S3 apply(20:18:44)后副本停在 local 内容(20:22:53 快照 = `e4791fe1…`)达 4 分钟;② Local 方向:回滚 apply(20:23:41)后副本停在 S3 内容。副本 managedFields 两笔写入(12:13:58Z `kubectl-client-side-apply` = 安装器 rerun 的双 ns apply;12:23:06Z manager `felis` = 下一次 setup 启动的 refresh)都不是 apply 时刻——**apply 本身不碰副本**。
- **根因**:`applyFelisConfigSecret`(storage/connection/edge 三条 apply 的公共出口)只 apply 控制 ns;镜像刷新只存在于 `felis setup` 启动(#51)与安装器。email 路径早有显式镜像刷新,storage/connection 是漏网的两条。
- **修复 `de7fb2c`**:镜像刷新移入 `applyFelisConfigSecret`(best-effort + stderr 警告;控制面-only 安装无工作负载 ns 时降级不阻塞);smtp helper 去掉重复块。门禁:`gofmt`/`go vet`/`go test ./...`/`bootstrap_test.sh` 全绿。
- **绿证据(`v0.0.0+fix52`,宿主二进制 sha `0bd49467…` 已装 `/usr/local/bin/felis`,旧版留 `/root/felis-auditfix44.bin`;真机双向)**:Run1 从 Local `s`→S3:apply 后 `control = mirror = podtoml = 4fb80ed8…`(S3 渲染;旧代码此刻镜像会停在 local);Run2 `s`→Local:`control = mirror = podtoml = df206074…`;副本 managedFields 写者 = `kubectl-client-side-apply` @ 12:33:40Z / 12:34:59Z(正是 apply 时刻)。Run1 启动块另见 `config (minecraft ns): refreshed from the control namespace`(#51 机制照常先收敛一次旧账)。
- **可达性**:#52 ②——任何 `s`/`c` 重配置即触发;危害等级低(镜像消费者 backup/restore/fileedit/reaper 当前不读被改动字段——`UserUploadsContext` 仅 `api.go` 消费——但「配置动了、镜像没动」正是 #51 要消灭的静默滞后类,且与 email 路径的既定行为不一致)。
**三、清理与还原**:MinIO 容器/卷/两镜像、`felis-uploads-s3` Secret、DB 行(`image_submissions` `sub-853a4e2ba4ba6443`)、/tmp 残留(player-cookies/drill-ctx/svc-tok 等)全清;docker 停;收尾 `control = mirror = df206074`(ALIGNED)、API 滚毕 Running、panel/healthz 200 ✓。
### 本轮新增真机证据(第三十三批:`/updates` 维护窗口 API+UI 走查全绿;#53 跟随仓库迁移;#54 update 升级指引全假)
**零、仓库迁移(背景)**:remote 已改 `[email protected]:FelisMC/Felis.git`(本批两个修复随 `2c6739a`、`397a400` 直推 main)。带 token 实测:旧 `MliroLirrorsIngenuity/Felis` 路径 301(GitHub rename redirect)、新路径 200——旧坐标当前仍能工作但全靠 redirect。新仓库**尚无 stable release**(`releases/latest` 带 token 也 404):felis-api 更新检查的 404 属发布流程事实,非代码缺陷。
**一、`/updates` 维护窗口 API 走查(完成;全绿)**
- 读:GET 未设置 → `{"not_before":null,"not_after":null}`。
- 负例全按预期拒绝:半设 / 倒序 / 相等 / 坏 JSON / 未知字段 → 400;缺 `Content-Type` → 415。
- 正例:写入 → 回读一致 → **API pod 重启后仍在**(已落 DB,非内存态)。
- 鉴权:无 cookie → 401;player 会话 → 403(新铸 player 会话,留 `/tmp/player-cookies.txt`)。
- 收尾:DB 回到 `{null,null}`。
**二、`/updates` 面板 UI 走查(CDP,完成;全绿零 console 错误)**
- 状态流转逐屏:生效中 → 过期 → 计划 → 未设置(截图 `/tmp/updates-{1..5}-*.png`)。
- 倒序提交 → 校验文案正确;清除后表单清空 + 「未设置」提示。
- 核账:DB 收尾 `{null,null}`;审计 `updates.window_set` 恰 3 笔(20:46:18 / :20 / :21)。
- 驱动:`/tmp/cdp-updates2.js`(bun + 原生 WebSocket;须 `Network.setCookie` 注入 owner 会话,否则新标签页 401 跳登录)。
- 观察(不计缺陷):窗口自身存取已验证;「窗口被 runner 消费」的端到端链路(Applier/Notifier)仍属 INTEGRATION-ONLY 设计(deferred-seams),不要当缺陷重复修。
**三、#53(旧仓库坐标残留)—— 提交 `2c6739a`**
- 红(`v0.0.0+fix52` 实机):`felis update` → `github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 — …`。
- 绿(`v0.0.0+fix53` 实机):同一命令 → `github: FelisMC/Felis releases/latest returned HTTP 404 — …`。
- 范围:`internal/updater/topology.go`、`deploy/bootstrap.sh`(默认 `FELIS_REPO_URL` + 两处 UA)、README ×2、两个测试夹具(6 文件 9 处);门禁四件套全绿。
- 可达性:②——默认安装 / 每次 `felis update` 都读该坐标;危害在 redirect 退休时兑现(安装与更新一起挂)。
**四、#54(`felis update` 的 apply 指引在既有安装上全是错的)—— 提交 `397a400`**
- 红(`v0.0.0+fix52` 实机文本):`--panel --force` → `run: sudo felis setup` + “felis setup is idempotent and re-runs the installer…” + felis-api「例外」块。
- 决定性证据(完成态安装):`felis setup` 跑 8s 退出,`/opt/felis/velocity/velocity.jar` mtime/hash 不变、零 bootstrap 输出;同刻安装器重跑日志有 `resolving the newest Velocity 3.5.1 build`。代码侧 `shouldRunHostBootstrapBeforeConfig` 仅在 4 个 marker 不全时进 bootstrap——既有安装上 `felis setup` 只开配置 TUI。
- 绿(`v0.0.0+fix54` 实机;宿主二进制已换 `/usr/local/bin/felis`,fix52 留档 `/root/felis-fix52.bin` sha `0bd49467…`):
- `--panel --force` / `--velocity --force` → `run: curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash` + 单条 trailer(channel 未持久化 caveat、私有仓库 token'd form、`felis setup is not this path`);
- `--mc` 无命令无 trailer;`--all` trailer 恰一次;`--force`/`LatestKnown` 语义不变。
- 文案同步:`docs/troubleshooting.md §15` 删掉 “or `sudo felis setup`”、补 channel caveat;测试改为 `TestApplyGuidancePointsEveryComponentAtTheInstaller`(钉安装器路径 + 禁 `run: sudo felis setup`)。
- 可达性:②(文档/指引;同 #38 型)。
### 本轮新增真机证据(第三十四批:`felis nano` 全链走查 60/60;缺陷 #55 —— 坏源静默挡住验证梯子)
**零、装置(全部在 VM,不进产品环境)**:`/srv/nanotest/nano-stub.py`=可控 Yggdrasil 源,路径即行为——`/ok/<name>`、`/okid/<name>/<id>`、`/props/<name>`、`/badname/<case>`(tooshort/toolong/badchar/section/noname)、`/none`=204、`/boom`=500、`/redir`=302、`/garbage`、`/emptyid`、`/slow/<secs>`;跑在 127.0.0.1:9900,临时单元 `nano-stub`,请求日志 `/srv/nanotest/stub.log`。驱动脚本:`nano-matrix.sh`(HTTP 矩阵 S1–S7)、`nano-service.sh`(systemd 层)、`nano-config.sh`(配置校验)、`nano-fw.sh`(firewalld);`nano-harness.sh` 从 HEAD 版 `bootstrap.sh` 提取产品函数在 scratch 路径跑(`felis-nano-test` 改名件,不碰产品单元)。
**一、HTTP 矩阵 S1–S7(60 项断言)**
- 红(修复前 `v0.0.0+audit-nano1`):`SUMMARY pass=47 fail=13`——13 红全在 #55 语义内:5 个坏名用例「应 503+日志、实得静默 204」(状态+日志 ×2=10),failover 3(坏源在前登录不落第二源 ×2、无日志 ×1)。留档 `matrix-red.r2.log`。
- 绿(fix55,同装置重跑):`SUMMARY pass=60 fail=0`。留档 `matrix-fix55.r2.log` 与 `matrix-fix55/`(更早原跑在 `matrix/`)。
- 覆盖面:S1 输入校验(缺参/超长→204、POST→405、未知路径→404、带体 GET→400、HEAD→204、400 带 `Connection: close`)、S2 三方登录全链(前缀 + felisAuthNS UUID 确定性、ip/serverId 转发保真)、S3 premium 撞名改名(`LS_`)与免费名不动、S4 失败模式(不可达/500/302/garbage/204 → 跳过、记日志、503)、S5 多源优先级 + 坏源后 failover、S6 日志纪律(长 URI 截断、控制字节转义)、S7 properties 转发。
**二、systemd 层(`nano-service.sh`,真 unit + 真 DynamicUser;重跑落盘 `nano-service.r2.log`,0 FAIL)**:装服务→起服务;重跑语义(配置字节不动、`resolve_nano_listen` 从 unit 读回端点);drain(stop 期间 in-flight 登录照样答完,实测 stop 等待 5335ms);SIGKILL 后自动重启;端点迁移 8081→8099 后新址应答、旧址关闭;三种启动文案(loopback「Bound to loopback」/ CIDR「admits」/ 无 CIDR「WARNING」)。
**三、配置层(`nano-config.sh`,LoadNano 16/16,落盘 `nano-config.r2.log`)**:未知键、空/重复/保留 tag、colon、空白、坏 prefix、重复 prefix(大小写不敏感)、非 http scheme、URL 带 query、明文公网 http、缺文件、`[server] listen` 忽略警告、Mojang-only 可服务。
**四、firewalld(`nano-fw.sh`,3/3,落盘 `nano-fw.r2.log`)**:关旧「开全源」8081/tcp、仅对 CIDR 开 rich rule、loopback 不开洞、无 CIDR 警告、快照清理还原。
**五、缺陷 #55(提交 `9dad61f`)**
- 红(fix54 实机):配置源回 200 但档案字段不可用(三方源名字进不了 MC 字符集;身份源 UUID 不解析)→ 静默 204、零日志;且筛查在 `handleHasJoined` 直接 `return` → [坏源, 好源] 顺序下 204,好源根本没被询问。
- 根因:两个 200-后筛查在 handler 层(200 已赢下梯子之后),而同类坏答案(200 无档案/非 200/不可达)在 `resolveHasJoined` 里是「记日志 + skip + failed」——一致性缺口。
- 修复:筛查移入 `resolveHasJoined`(身份源 `uuid.Parse`、三方源 `mcUsernameRe`)→ 命中记日志(`unusable profile name` / `unparseable profile id`)+ `failed=true` + `continue`;handler 守卫保留为最后防线(注释更新)。单测 +3(含 `log.Writer()` 捕获断言)。
- 可达性:#55 ②——需要配置了一个「回 200 但字段不可用」的源(马虎的自建/三方 Yggdrasil);坏源在前时经梯子的登录全部被吞(静默、零日志),安全侧(坏名不落玩家列表)保留。
**六、观察(不计缺陷)**:premium 冷查询 fail-closed 抖动——api.mojang.com 从 VM 首查偶发接近 2s 超时 → 偶按「疑似 premium」给免费名加前缀(`FelisNanoStub1` → `LS_FelisNanoStub`);符合 `isPremiumName` 既定取舍(误加前缀=外观代价,误放行=抢名),记观察不修。
### 本轮新增真机证据(第三十五批:NetworkPolicy 真机强制矩阵全绿 —— 端到端白名单闭环 + felis-velocity 刷新存活验证)
**零、装置与现场**:三张策略 live 于 `minecraft` ns(`felis-default-deny-ingress` / `felis-allow-rcon-from-control-plane` / `felis-allow-game-from-velocity`;apply 于 09-22T05:37:24Z,与渲染收敛 diff 零漂移);k3s 参数无 `--disable-network-policy`,**强制执行实测生效**;kube-router 机制实证:per-pod `KUBE-POD-FW-*` 链、未标记流量 `REJECT --reject-with icmp-port-unreachable`、每链首条 `--src-type LOCAL -j ACCEPT`(本节点豁免)、ipBlock 落 ipset `KUBE-SRC-*`(白名单成员可直接查)。探测法:**nsenter 进真实 pod 网络命名空间(源 IP=真实 pod IP)+ netns/veth 合成「转发型外部源」10.99.0.2(模拟 velocity 在另一台机器)**;全程不改产品代码。
**一、pod 源矩阵(源=真实 pod netns)**
- api(felis,命中 RCON 白名单标签)→ lobby:25575 = **OPEN**;同源 → lobby:25565 = **REFUSED**(端口级区分 ✓)
- registry(felis,非匹配)→ lobby:25575/25565 = **REFUSED**;同 pod → api:8081 = OPEN(对照:无策略命名空间不受限)
- coredns(kube-system)→ login:25565 = **REFUSED**
- 收尾复跑全矩阵与首轮逐行一致(`np-matrix-run1/2.log`)。
**二、host(=velocity 同机侧)**:→ login/lobby 的 ClusterIP 与 podIP :25565 **全 OPEN**(velocity 注册的后端路径实际可用);→ api-internal:8081 OPEN;→ lobby:25575 亦 OPEN——归因:kube-router 每 pod 链的 `--src-type LOCAL -j ACCEPT`(**本节点流量豁免,kube-router 设计行为**,kubelet 探针等依赖它;netpol 语义无法对节点自身收口,节点 root 本在 TCB 内)。**注记①:node-local 豁免。**
**三、转发型外部源(最严苛模拟)**
- 基线:netns(10.99.0.2) → 全部服务器端口(podIP 与 ClusterIP、25565/25575)= REFUSED,包级可见 netpol 的 icmp-port-unreachable(`td3.log`)。
- **白名单闭环**:临时把 10.99.0.0/24 加入 allow-game → ipset 即时生效(members:`10.99.0.0/24` + `10.211.55.6`)→ **lobby-svc:25565 与 login-svc:25565 = OPEN**(ClusterIP=velocity 实际拨号形态);25575 仍 REFUSED(端口维度不破)→ 回滚 → spec 哈希逐字节一致(md5 `e6b07440…`)、ipset 复原、复测全部 REFUSED。
- 层间归因:firewalld 规则含 `ct status dnat accept`(**DNAT 后的服务流量放行**;非 DNAT 转发走 forward policy 的 `admin-prohibited` 拒绝)——受支持拓扑(velocity 同机=LOCAL、ClusterIP、Mac→NodePort 面板 200)全部实测可用;「外部未经服务直连 podIP」不属于任何产品流。**注记②:firewalld 只放 DNAT 服务流。**
- 装置保养:veth 未归区时 firewalld 会拒其转发属装置噪声(已归因);public 区临时挂载已摘、netns/veth 已删、策略零残留(spec diff 为空)。
**四、felis-velocity 刷新循环存活(顺带验证)**:20:18–20:42 的 `server list refresh failed` 全部落在 #52 演练的 API rollout 窗口(成功不打日志属设计);tcpdump 450s 窗口抓到 52 条 `GET /api/v1/servers` 载荷行(成对出现=双抓包点看到同一请求,折算约 26 次 ≈ 每 15–17s 一次,与 `REGISTRATION_REFRESH=15s` 常量吻合)及对应 200 响应;"keeping current registrations" 为设计降级。非缺陷。
**五、留档(VM `/srv/npdrill/`)**:`np-matrix.sh`+`np-matrix-run1/2.log`、`np-netns.sh`、`netns-probe.py`、`np-cidr-test.sh/.log`、`spec-before/after.json`、`td3.log`(包级归因)、`tcpdump-8081.log`(刷新存活)、`iptables-save.txt` 与 `nft-rules.txt`(现场快照)。
### 本轮新增真机证据(第三十六批:面板错误文案全量本地化(#60)+ Run4c 管理面写操作复核 0 缺陷)
**零、现场**:三个缺陷批次(#56–#59)已按「一缺陷一 commit」推 main 并部署 `auditfix59`;本批完成 #60 后部署 `auditfix60`(api/operator/reaper 三处 + 宿主 CLI,`felis version` = v0.0.0+fix60)。面板门禁:typecheck 0 错、vitest 117/117、vite build 通过(注意项目测试是 `bun run test`=vitest run;裸 `bun test` 是 Bun 内置 runner,解析不了 `@/` 别名,41 个 fail 系假象,勿误报)。
**一、Run4c 管理面写操作复核(4 段全真机,最终 0 缺陷;3 处为测试脚本自身误报,均已自纠)**
1. **镜像添加+删除**:添加 `registry.felis.svc:5000/e2e/probe:1` → 列表出现、API 落库(`added_at=14:50:53Z`)✅。删除初测"未生效"系脚本用 `tr` 找行——该表是 `div` 网格,`tr` 选择器命中 0 → 从未点到按钮。换 `div[class*=grid-cols-12] → button[title]` 重测:confirm("确定删除此镜像吗?")自动接受 → 页面 1ms 内刷新、**API 侧同 ref 从 5 条降至 4 条**——删除通道无缺陷。
2. **构建负例**:外部 registry(`docker.io/...`)+ 不存在 context → 对话框红字拒绝 `build: invalid request: image reference … must target the internal registry "registry.felis.svc:5000"`(截图 `run4c-3`)。脚本"未捕获"系其正则先命中了侧边栏"镜像"二字——误报。
3. **创建重复用户**:对话框内正确显示 **"该用户名已被使用。"**、对话框保持打开(截图 `run6-1`)。此前 run4c 的"dialog-closed"系脚本点错页面(进到了用户详情页);且该轮实际是一次**正例**:真管理员的 username 是 UUID 串(`08595879-…`,role=`owner` 是角色名),字面用户名 `owner` 当时空闲、创建确实成功——测试件已删(DELETE→200,列表 7→6)。
4. **创建重复子域名**:正确拒绝 **"该子域名已被占用。"**(截图 `run4c-5`)✅
**二、缺陷 #60(提交 `c9af548`):37 个用户可达错误码显示生英文**
- 发现路径:run4c 复核 `email_taken` 时做了一次系统性对照——后端 `newError` 共 **71 个错误码**,面板 `humanizeError` 只映射 **23 个**;其余走 `default` 分支直接透传 `err.message`(多为 Go 包装嵌套的英文,如 `invalid server name: naming: invalid server name "BadName": must match ^[a-z0-9-]{3,32}$`)。
- 三个代表真机验证(修复前→修复后):
- `bad_name`:新建服务器"名称"填 `BadName`(前端只查非空)→ 前:对话框直出上述嵌套英文;后:**"服务器名称不合法:需为 3–32 位小写字母、数字或连字符,且不能使用保留名。"**(`run8-1`)
- `bad_subdomain`:子域名填 `ab`(前端 RE 允许 1–2 字符,后端要求 ≥3)→ 后:**"子域名不合法:…"**(`run8-2`)
- `email_taken`:DB hook 造"他人已验证邮箱" → 账户页真发码(no-mailer 日志取码)→ verify 409 → 后:**"该邮箱已在其他账户上完成验证;请直接用该邮箱登录,或换一个地址。"**(`run9-1`)
- 修复:api.ts +37 case、zh/en errors.json 各 +37 键(纯映射,无逻辑变更)。**剩 11 个故意不补**:`bad_request`/`conflict`/`internal`/`panic`/`not_found`(通用兜底)、`forbidden`/`not_admin`/`unauthorized`(状态码分支已在 default 覆盖)、`not_ready`(内部面)、`setup_required`(bootstrap 期 SPA 流控信号,403 文案可用)、`unsupported_media_type`(CSRF 门,面板永不可达)。
- 遗留观察(不计缺陷):构建负例的英文前缀 `build: invalid request:` 仍会透传——该错误无独立码、message 即最终文案;对管理员可读,记观察。
**三、观察(不计缺陷)**
- UserDetailPage 对 owner/对自己都显示删除按钮;真的点了会得 403 + 具体 message("the owner account cannot be deleted from the panel"),但前端 403 分支统一显示"你无权执行此操作"——笼统但不算错,记观察。
- 字面用户名 `owner` 可注册(用户名无保留名单;角色由服务端管理、无提权路径),记观察。
**四、留档(Mac)**:脚本 `/tmp/cdp-run5.js`(镜像删除复测)、`run6.js`(重复用户复测)、`run7/run8.js`(bad_name/bad_subdomain 前后对照)、`run9b.js`(email_taken,内置 ssh 取码);截图 `run5-1`/`run6-1`/`run7-1/2`/`run8-1/2`/`run9-1`。VM `/opt/felis/src` = `c9af548` 快照。
### 本轮新增真机证据(第三十七批:管理面交互收尾 0 缺陷 + 缺陷 #61 —— LuckPerms 写操作的过度承诺)
**零、现场**:`auditfix61`(api/operator/reaper 三处 + 宿主 CLI)。面板门禁 typecheck/vitest 117/build 全绿后部署。
**一、管理面交互收尾(3 项全真机,0 缺陷)**
1. **submissions approve/reject**:player 会话(邮箱 OTP 铸造,no-mailer 日志取码)现场造两条 pending(`sub-489eda47fe4bc12d` / `sub-b1dcc4b4a2cc5f16`,各上传 tar.gz context)→ 面板「通过」→ 状态变「审核通过」且**自动构建 `bld-1790176562566802097` 到 succeeded**(点击到构建完成全链闭环);「驳回」→ 对话框必填原因 → 状态变「已拒绝」。全程 UI 零报错。
2. **用户会话撤销**:player.test 铸 2 条新会话(共 4 条)→ UI「单独撤销」首条 → **精确生效**(cookie5 → 401、cookie4 → 200、列表 4→3);「全部撤销」→ 提示「所有会话已撤销。」、列表清空、cookie4 也 401。
3. **ServerCard 启动交互**:面板点 test-one「启动」→ `Starting` → `Running/ready`。
**二、缺陷 #61(提交 `b4ef42d`):LuckPerms 页对写操作"过度承诺"**
- 发现路径:LP 页写操作交互测试(给 `E2E_Tester` 写权限)→ 历史卡显示绿色 success +「(服务器未返回输出)」→ **落盘取证发现根本没写入**。
- 真机实验矩阵(test-one,LP 5.5.85 / H2 存储):
| 操作 | RCON 回包 | 落盘(`lp export` 实测) |
|---|---|---|
| 面板 UI:`E2E_Tester permission set e2e.ui.write.test` | 空 | ❌ |
| 面板控制台:`E2E_Tester … set probe.test true` | 空 | ❌ |
| 面板控制台:`<UUID> … set probe2.test true` | 空 | ✅ |
| **独立 Python RCON 客户端**(绕开 felis):`E2E_Tester … set probe3.test` | 空 | ❌ |
| 独立客户端:`<UUID> … set probe4.test` | "Another command…"(异步提示) | ✅ |
- 结论:**felis 只如实转发空回包;名字写入的静默失败是 LuckPerms 自身行为**(独立客户端 1:1 复现)。
- 根因(LP config.yml 官方注释背书):`use-server-uuid-cache: false`(LP 默认)→ "commands using a player's username will not work **unless the player has joined since LuckPerms was first installed**"——**未进过服的玩家名永远无法解析**(与有无外网无关)。
- 面板问题:旧文案承诺"授予与撤销操作仍然会实际生效"(#59 遗留半句)→ 对"给还没来过的玩家预授权"场景是不成立的承诺。
- 修复(纯文案,双语):`luckperms_no_reply` → "……按玩家名的操作只对「自 LuckPerms 安装以来进过本服」的玩家可靠,对没进过服的玩家名可能静默不生效";`luckperms_no_output` → "(服务器未返回输出,无法确认结果)"。真机复验(`run16b`):读提示与写占位均按新文案显示。
- 可达性:#61 ①——给未进服的玩家预授权是日常动作,此前面板显示绿色成功构成误导。
- 局限(记观察):读提示块只在"读取为空"时渲染;已有部分玩家数据的服上不会露出这段说明(后续增强候选:常驻说明或写前对未解析名字的提示)。
**三、观察(不计缺陷)**
- LP 命令是**串行异步**执行:快速连发第二条会得到 `§7[§b§lL§3§lP§7]§r §7Another command is being executed, waiting for it to finish...`(带颜色代码原文,面板如实显示);`lp export` 回包时有时无(响应与执行解耦)。felis 层无责。
- submissions 徽章(`submissions.json`:"审核通过/已拒绝")与过滤标签(`admin.json`:"已通过/已驳回")两套词,语义均可,记观察。
- 控制台对空回包命令只显示 echo、无占位提示(终端风格);LP 页有占位文案。记观察。
**四、留档(Mac)**:`/tmp/cdp-run10.js`(approve/reject)、`run11a/11b.js`(会话撤销)、`run12a.js`(启动交互)、`run13/14/15/15b/16/16b.js`(LP 全链)。VM 独立探针:`/tmp/rcon_probe.py`、`/tmp/rcon_one.py`(+ `/tmp/rcon_pw.txt`);LP 导出样本在 `/data/plugins/LuckPerms/luckperms-2026-09-23-15-*.json.gz`。测试数据:两条 pending 提交(一 approved 一 rejected,构建 succeeded/failed 各一)。VM `/opt/felis/src` = `b4ef42d` 快照。
### 本轮新增真机证据(第三十八批:#56–#59 证据回填 + 服务器详情/运维面收尾 0 缺陷)
**零、现场**:`auditfix61`(api/operator/reaper 三处 + 宿主 CLI,`felis version` = v0.0.0+fix61)。本批两部分:把 #56–#59 四项修复的真机证据落节(此前只存在于 commit message),并把队列里剩余的服务器详情页/运维面交互复跑做完(0 缺陷)。
**一、缺陷 #56–#59 证据回填(部署 `auditfix59`;四项均 = 修复前真机触发 + 修复后复验)**
1. **#56(`0790f8d`)玩家管理操作把 RCON 回包扔掉、只报 canned 成功。** 触发路径:给"还没进过服"的玩家加白名单/封禁——vanilla 对没见过的名字回 `That player does not exist` 并拒绝,而旧面板无论如何都显示成功。修复后回包逐字上屏:`whitelist add E2E_Bad` → `Added E2E_Bad to the whitelist`(磁盘同步真写入);负例 `NoSuchPlayerXYZ` → `That player does not exist`(不再伪装成功);静默服务器回退本地化文案。
2. **#57(`70c988e`)控制台命令的回复无处显示。** 触发路径:控制台发任何命令——`sendCommand` 拿得到 RCON 回包但被丢弃、pod 日志也不回显命令输出,等于零反馈。修复后 echo + 回包以终端样式渲染在提示符上方(实测 `list`)。
3. **#58(`4d4cdd6`)无世界盘服务器的文件页 90s 卡死。** 触发路径:对从未启动/已回收的服点"文件"页——旧行为建 Job → Pod `FailedScheduling (pvc not found)` Pending 到 90s 超时 → 误导性 504 `files_timeout`。修复:file 路由补上与 backup/restore 同款 `WorldVolumeExists` 门 → 快速 409 `no_world_volume` + 面板双语文案。本批现场复查:`resolvecheck`(无 world PVC)→ 409 `no_world_volume`,0.03s;`test-one`(有盘)→ 200(~2s)。
4. **#59(`a2ff2a1`)LuckPerms 页把"读不到"报成"没有"。** 触发路径:打开装了 LP 的服的管理页——LP 5.5.85 的 `lp` 命令 RCON 回包全为空(独立 RCON 客户端 1:1 复现),读投影永远为空,旧页面却断言"没有父组/没有显式节点"(假事实)、写历史伪造 `[RCON]` 行。修复:原始回包随 rosters 同款披露渲染;空回包显式提示、不再假断言;历史占位不再伪造输出。
**二、服务器详情页交互复跑(5 项全真机,0 缺陷)**
1. **ServerCard 停止**:test-one `Stopping` → `Stopped`(启动在第三十七批;收尾复查 `desiredState: Stopped`)。
2. **ServerFiles 写流程**:编辑 motd → 保存「已保存」→ 重开读回一致(`motd=Felis E2E files drill`)→ 还原默认 `A Minecraft Server`(收尾复查确认)。
3. **备份/恢复**:立即备份 → Job `succeeded`;恢复 → 确认对话框(破坏性警告原文)→ Job `restore-test-one` `Complete`(5s)→ UI「恢复 27秒钟前 成功」→ 唤醒 `Running/ready`(恢复后的世界可加载)。
4. **白名单**:加/负例/移除全链复跑(证据见 #56)——终态磁盘只剩 `E2E_Tester`(88 字节)。
5. **封禁/解封**:封禁(内联确认)→ `Banned E2E_Bad: Banned by an operator.` + 磁盘写入;解封 → `Unbanned E2E_Bad`;`banned-players.json` 终态 `[]`。
**三、运维面复查(3 项,0 缺陷)**
1. **`/admin/updates` 复跑**(队列"深挖"项):设置窗口 → 「维护窗口更新成功。」+ 状态卡更新;负例 end<start → 前端「结束时间必须在开始时间之后。」、后端 400 `end must be after start`;清除 → 「当前未设置维护窗口」。
2. **metrics 端点**:内部面 `/metrics` 正常——`felis_*` 样本 18 条 / 3 个指标族(build 失败计数、回收计数、起服时长直方图)。
3. **CLI 覆盖核销**:`run` 分发表 17 项(15 个用户面命令 + `bootstrap-assets`/`init-forwarding` 两个容器内部入口)在历批演练中均已有真机记录,本批逐项核销无遗漏。
**四、接口语义注记与收尾**
- 文件 API 的 `path` 是**相对路径**:空串 / `.` = 根(正常列出)、`config` 下钻正常;字面 `/` 被路径约束拒绝(400 `bad_path: path escapes from parent`)——面板从不发绝对路径(`joinPath` 只拼相对段),无用户面影响;本批复查脚本初次误用 `/` 时曾见 13s 延迟,属首次 fileedit Job 冷启动,非卡死。
- 收尾静止态:test-one `Stopped`;三个名单 = whitelist `E2E_Tester` / banned `[]` / ops `[]`;`E2E_Bad` 仅余 latest.log 与 usercache.json(日志与 Mojang 缓存)。
- 留档(Mac):脚本 `/tmp/cdp-run17a/17b/17c`(停止、写文件、还原)、`run18/18b/19`(备份、恢复对话框、完成等待)、`run20a/20b`(白名单加/减)、`run21/21b/22`(封禁重试、封禁、解封)、`run23`(维护窗口)+同名 `-out.json`;截图 `run17b-1`、`run18-1..4`、`run18b-1,2`、`run19-1`、`run20a-1,2`、`run20b-1`、`run21b-1`、`run23-1..3`。
### 本轮新增真机证据(第三十九批:reaper 多节点实机 —— VM 克隆双节点验证)
**零、装置**:`prlctl clone "CentOS Linux 9 Stream" --linked --name felis-node2`(linked 克隆,初始 1.4M)→ node2 改 host 名 `felis-node2`,停用/禁用 `felis-velocity`、`k3s.service`(server 形态)、docker,以 k3s-agent 加入同一集群(`https://10.211.55.6:6443`,复用 node1 node-token)。克隆副作用 = node2 自带 node1 storage 副本(6 项 / 3.1G)——已移开,模拟真实新节点。两节点均 Ready(v1.36.4+k3s1);为克隆,node1 经历一次正常重启,重启后组件/服务全回归(docker 随自启后又停回 `inactive`)。
**一、pin 正向**:`kubectl create job reaper-b39-ok --from=cronjob/felis-reaper`(live 模板原样,nodeSelector=`localhost.localdomain`;node2 无 taint、是合法调度候选)→ pod 落在 `localhost.localdomain`(nodeSelector 命中;backups PVC 的 affinity 亦指向同节点——两者本就该同节点),4s 跑通:`evaluated=2 reaped=0 warned=0 skipped=0 evicted=0 expired=0`、Job `Complete`。即:渲染出的 pin 在真实双节点集群把 reaper 钉在持盘节点。
**二、错位 pin 反向对照**:同模板把 nodeSelector 改成 `felis-node2` → pod 停在 **Pending**,scheduler 事件原文:`0/2 nodes are available: 1 node(s) didn't match PersistentVolume's node affinity, 1 node(s) didn't match Pod's node affinity/selector`——node1 被错位 selector 拒绝、node2 被 `felis-backups` PVC 的 volume node affinity 拒绝(PV `pvc-0b4fbda6-…`,nodeAffinity=`localhost.localdomain`、hostPath=`/var/lib/rancher/k3s/storage/pvc-0b4fbda6-…_minecraft_felis-backups`)。**结论:本部署形态下 pin 写错是 fail-closed(卡住 + 明确调度事件),不会在无世界节点上静默执行**;真正要防的是首跑顺序(全新多节点安装时,第一次 reaper 运行会把 backups PV 落在其所在节点)——这正是渲染默认要求 `--reaper-node` 并 stderr 警告的原因,维持现状不修。
**三、装置回收**:drill job ×2 删除(minecraft ns 无残留);node2 关机 → `kubectl delete node felis-node2` → `prlctl delete felis-node2`(VM+克隆件删除)。收尾 `get nodes` = 单节点 `localhost.localdomain`,felis/游戏 pod 全 Running,CronJob `suspend=false` + nodeSelector 原样,docker `inactive`。重启副作用按既有说明处理:Mac 侧 443 入面板依赖的手工 `socat` 中继(本台账开头注记「重启 VM 后需重开」)随重启消失——已按原样重开(`TCP6-LISTEN:443,ipv6only=0,reuseaddr,fork TCP:127.0.0.1:30443`)并复核 op.console 面板 / api-me / player 面板皆 200。
**四、留档**:node2 agent 上线日志(`k3s agent is up and running`、VXLAN subnet event 来自 10.211.55.6);两向 job 的 pod/调度事件原文(见上);`/root/node2-storage-copy/`(3.1G,随 VM 删除)。
### 本轮新增真机证据(第四十批:Java 插件层收口 —— #63 demo-up 单起点、CI Java 门禁、插件面真机 E2E)
**零、现场**:本批不改 Go/面板(控制面维持 `auditfix61`);产出 = `deploy/demo-up.sh` 重写(`f5a76cf`)+ `plugins/test.sh` 与 CI `plugins` 作业(`c59b387`)。三组真机验证(T1/T2/T3)在 VM 上针对该两文件跑完;CI 首跑即绿。
**一、#62 复核:不成立(立案后剔除,未计为缺陷)**
- 主张:"fresh 非 demo 安装跳过 Velocity 插件构建 → proxy 静默不路由"。复核三条路径,**每条都构建 felis-velocity.jar**:
1. 源码臂:`felis-install.log`(首装)L1530 `building felis-velocity.jar`、L1553 `staged`;`bootstrap-auditfix{42,42b,42c,43,43b,44}.log` 每次重跑同两行。
2. 嵌入 tar 臂(fresh release/TUI 路径;本 VM 从未走过):`felis bootstrap-assets game-stack | tar -x` → 37 文件(含 `plugins/velocity`、`plugins/shared`,构建输入齐全)→ `docker run --rm -v /tmp/gs62:/src:z -w /src/plugins/velocity gradle:8.14-jdk21 gradle --no-daemon clean build` → **BUILD SUCCESSFUL in 35s**,产出唯一 `felis-velocity-0.1.0.jar`(71117B,与现装同尺寸)。
3. 调用图:`build_velocity_plugin` 唯一调用点在 `build_game_stack` 末尾;`build_game_stack` 在 full 模式主链路(L2868)必经,nano 早返回,不存在可绕开的"demo 分支"。
- 结论:代码阅读误判,剔除。真正会跳过插件构建的是 demo-up.sh 的旧镜像臂 → 即 #63(本批修复)。
**二、#63(`f5a76cf`):demo-up.sh 双起源 → 单起点(129 → 63 行)**
- 红证据(读取 + 真机口径核对):① 版本分叉——demo-up 硬编码 `PAPER_MC_VERSION:=1.21.8`,bootstrap `resolve_game_jars` 从 Limbo CI 产物名推导(T3 实测当日 Limbo 2026.0.3-ALPHA / MC 26.3;同一登录两跳必须同协议);② 其镜像臂从不构建 felis-velocity.jar,导入的是本地 tag(`felis-limbo:demo` 等),与 bootstrap 写入的 registry refs 不一致 → 死件、旧基座上可致"起了但无处路由";③ 起 docker 后从不停(违反 13d64e0 起"构建方以 docker 停收尾"的约定);wiring 段在现代基座上只是 no-op。
- 修复后 = bootstrap(未 SKIP 时)+ 三项硬校验(`felis.host.toml`、`[velocity]` 段、`felis-velocity.jar`)+ `felis setup` 交棒;镜像逻辑全删。
- 真机验证:**T1** 健康基座(`SKIP_BOOTSTRAP=1 SKIP_SETUP=1`)→ 通过、exit 0;**T2** 移走 jar → `ERROR: /opt/felis/velocity/plugins/felis-velocity.jar missing — … silently routes nothing …`、exit 1(随即还原 71117B root:root 0644);**T3** 默认臂全量重跑(`FELIS_SKIP_FETCH=1 FELIS_INSTALL_MODE=full FELIS_IMAGE=…:auditfix61 FELIS_WORLDS_HOST_PATH=… SKIP_SETUP=1`,日志 `/root/demo-up-b40.log`):L733 插件构建、L752 `staged`、L941 交棒行、`[fail]`=0;收尾态 `felis-velocity active / k3s active / docker inactive`。
- 附带实录(同一日志窗):重跑期间 felis-api 滚动,velocity 00:22:25/00:22:40 两次 `refresh failed … keeping current registrations.`(失败保旧——规格 §11 承诺的真机上演)→ 00:22:55 恢复注册 → 00:22:58 服务随 `install_velocity` 重启并重载插件(新 pid)→ 00:22:59 `routing ready` → 00:23:29 收敛到 pod IP。
**三、Java 层 CI 门禁(`c59b387`):三个手工测试 + 三个装机 jar 首进 CI**
- 缺口:`plugins/*/test` 三个 main 测试从未被任何 build/CI 运行;velocity/paper/limbo 三个"装机即用"jar 只在 bootstrap/Dockerfile 编译(Go CI 从不碰 Java)。
- 产出:`plugins/test.sh`(Maven Central 取 adventure 三 jar,pinned+sha256;三测试 javac+java;三生产编译,limbo 按 bootstrap 同源解析当日版本)+ `ci.yml` 新增 `plugins` 作业(temurin 21 + Gradle 8.14)。
- 首跑两次真跑抓出两处"从未被跑过"的假设错误:① `InviteCardTest` 一行式缺 examination-api(adventure-api 4.26.1 的 `Component` 签名引用 `Examinable`,javac 编译期即需);② limbo 的 `com.loohp:Limbo:+` **永远不可解析**(LOOHP 仓无 maven-metadata,404 实查)→ 裸 `gradle -p plugins/limbo build`(plugins/README 原文)从来不可行;二者均已修(脚本 + 该测试 javadoc + README)。
- 验证:VM 容器三测试 `OK (32/36/48 checks)`;velocity 21s / paper 29s / limbo 13s(2026.0.3-ALPHA)全 `BUILD SUCCESSFUL`。GitHub Actions run `35888009965`(c59b387)= go/shell/panel/plugins **4/4 success**(`plugins` 作业首跑即绿)。
**四、插件层真机 E2E(不依赖真实客户端的面)**
- 自建纯 socket 探针 `/tmp/mcprobe.py`(status = 服务器列表 ping;login = 登录首包),重装前与重装后各跑一遍,结果一致:
- 子域 MOTD(§11 只读缓存 + 相位):`test-one` → `« test-one » 休眠中,加入即唤醒 / sleeping — join to wake`;`lobby`/`login` → `« … » 在线 / online`;`resolvecheck` → 休眠中;`nosuchxyz.<root>` → 回落 `A Felis server`(非 felis 子域不被劫持)。
- 登录边界:`LoginStart`(含现代协议 UUID 字段)→ 服务端首包 `0x01 EncryptionRequest` ⇒ 边缘 online-mode 强制成立。
- 加载面:`Loaded plugin felis-link 0.2.0`(共 5 插件)。
- 口径:真实账号进服(limbo 门 → 菜单 → 转服)仍按"进服跳过"决定不演练;以上为不依赖客户端的最大真机面。(首次探测因探针漏发 UUID 字段被静默关闭——探针缺陷,非服务端问题;补齐后一次通过。)
**五、观察(不计缺陷)**
- ViaVersion 自报有 5.12.0(当前 5.11.0):bootstrap 的 pin 是 FL-007 实测版本,属刻意,不随提示升级。
- limbo 插件编译期一条 deprecated API 提示(`FelisLimboPlugin` 用/覆写已弃用 API):门禁下可见、不阻塞,留观察。
### 本轮新增真机证据(第四十一批:装载器 mod 层收口 —— #64 exec 位、#65 元数据、编译+起服双门禁)
**零、现场**:控制面与面板不动;产出三个 commit:`f6048f2`(#64 修复)、`c2fe6a6`(#65 元数据)、`aa5abfa`(`plugins/test-mods.sh` + CI `mods` 作业 + release 双门禁 + README 记录)——这是最后一块"从未被任何自动化碰过"的模块面(fabric / forge / neoforge 三个装载器 mod)。
**一、#64(`f6048f2`):三个 vendored `gradlew` 缺可执行位——文档正路第一步即失败**
- 红证据:`git ls-tree origin/main` 三个 `gradlew` 全为 `100644`(blob `b9bb139f…`);`plugins/README.md` 第 212–214 行原文教 `plugins/fabric/gradlew -p plugins/fabric build`(forge/neoforge 同式);VM 全新 clone 实跑 → `bash: line 1: ./gradlew: Permission denied`,exit 126。
- 为何从未暴露:无 CI、无安装路径,README 是唯一入口——直到本批第一次真跑才现形。
- 修复:`chmod +x` + `git update-index --chmod=+x`(mode 100644→100755 ×3)。
- 可达性:①(全新用户照文档走的第一步)。
**二、#65(`c2fe6a6`):mod 元数据与仓库 LICENSE 相悖**
- 三处 `license = "MIT"`(`fabric.mod.json`、forge/neoforge 的 `mods.toml`)vs 仓库 `README.md:73` 的 `AGPL-3.0-only` + LICENSE 全文。时间线:mods 2026-06-26 加入(`93f143f`)、LICENSE 2026-07-12 才落地(`037eb24`)——陈旧残留。
- 另两处 `issueTrackerURL = "https://example.invalid/felis"`(占位域名)→ `https://github.com/FelisMC/Felis/issues`。
- 用户面:fabric loader 启动打印 license 字段;`mods.toml` 被两个 loader 解析。
- 可达性:②(低;元数据/合规面,非功能)。
**三、双门禁落地(`aa5abfa`)**
- `plugins/test-mods.sh`:JDK 17 下依次跑三个模块的 vendored wrapper(`./gradlew --no-daemon build`);java 大版本非 17 直接 fail-loud(三模块目标是 Java-17 的 Minecraft 线)。
- `ci.yml` 新增 `mods` 作业(temurin 17 + `gradle/actions/setup-gradle@v4`,版本由各模块 wrapper 自管);`release.yml` 发布前加两道 Java 门禁(JDK21 `plugins/test.sh` + JDK17 `plugins/test-mods.sh`)。
- README:Status 更新为 compile + boot 双验证;Building 段补 limbo `-PlimboVersion=<release>` 要求与两个门禁说明。
- CI:run `35892544563`(`aa5abfa`)——`mods` 作业首跑即绿,五作业全过(shell / mods / panel / go / plugins 5/5)。
**四、三个模块的真机编译 + 起服 E2E(本批核心证据)**
- 编译(VM 容器 `eclipse-temurin:17-jdk` + 持久 gradle 缓存):三模块 `BUILD SUCCESSFUL`(首跑 4–5 分钟级/模块),产物 `felis-{fabric,forge,neoforge}-0.1.0.jar` = 29276 / 28855 / 28562 B;留档 `/root/mods-build-b41.log`。
- 起服(三个真实专用服,控制台直驱 `/link`,随后 `linkx` 未知命令对照、`stop` 正常停服):
- **fabric** 1.20.1 + loader 0.19.5 + fabric-api:`Felis link ready; /link is registered.` → `Done (10.670s)!` → `/link 只能由玩家执行 / /link can only be run by a player.` → `Unknown or incomplete command`(linkx)→ `Stopping server`;`EXIT_fabric_boot=0`。
- **forge** 1.20.1-47.3.0:`Felis link ready; /link will be registered.`(modloading-worker)→ `Done (11.751s)!` → 同四段证据;`EXIT_forge_boot=0`。
- **neoforge** 20.4.251(1.20.4):`Done (16.094s)!` → 同四段证据;`EXIT_neoforge_boot=0`。
- 口径:三服均为"装好 mod 后真启动"的场景,同时验证 loader 真实解析修改后的元数据(license=AGPL、tracker URL)后正常装载;完整取码链(玩家在游戏内 `/link`)按"进服跳过"决定不演练,装载/注册/拒绝面已全部真机成立。
- 留档:`/root/mods-e2e-b41.log`、`/root/mods-e2e-b41-resume{2}.log`、`/opt/felis/mods-e2e/`(496M,三服目录保留);收尾态:docker `inactive`、k3s/felis-velocity `active`、磁盘 17G free。
**五、演习装置自身两次修正(非产品缺陷,如实记录)**
- 控制台驱动脚本 `/root/modserver-drive.sh` 初版把 FIFO 写端先开(`exec 3> console`)→ 自我死锁(内核 `wait_for_partner`、`State: S`),java 从未启动(无 `server.out` 实证);改 `exec 3<> console`(RDWR 打开不阻塞)后三服全通。
- forge 安装器首跑:`libraries.minecraft.net` 连接失败 → `com.google.code.findbugs:jsr305:3.0.2` 下载失败、`There was an error during installation`、无 `run.sh`;原样重试即 `The server installed successfully`(VM 网络抖动,非 forge/产品问题)。
**六、观察(不计缺陷)**
- 无新增。
### 本轮新增真机证据(第四十二批:覆盖面对账 + 文档/工具收尾 —— #66 README_EN、§28 图对齐、sync.sh 本地修复)
**零、现场**:本批不碰产品代码;对"所有模块均已演练"的结论做**独立对账**(枚举仓库全部模块/交付物 × 台账覆盖),只挖到文档与本地工具层面的残留。commits:`62c5a8a`(#66 README_EN)、`62e8c87`(§28 序列图对齐)。
**一、覆盖面对账(本批核心,逐项)**
- 仓库卫生:`git ls-files` 全量扫描无 `.DS_Store`/`*.log`/`node_modules`/构建产物误入库;顶层 `felis` 二进制(83MB)被 `.gitignore` 正确忽略;旧仓库名 `MliroLirrorsIngenuity` 残留仅存在于台账(历史记录)与本地未跟踪二进制。
- 未竟工作标记:`TODO|FIXME|XXX|HACK` 全仓(go/ts/sh/md/yml)仅 1 处命中——`internal/api/api_test.go:263` 的 `context.TODO()`(合法测试写法),无未完成实现标记。
- CLI 派发面:`cmd/felis/run.go` 的 usage ↔ dispatch 表有双向测试强约束(`run_test.go`,含显式 `undocumentedCommands` 允许表:`bootstrap-assets`/`init-forwarding` 等 in-Pod 入口),机制已核实成立。
- CI 覆盖:`shell` 作业按 shebang 对**全部 tracked `*.sh`** 做语法检查(`git ls-files '*.sh'` 逐文件 `bash -n`/`sh -n`)并运行 `deploy/bootstrap_test.sh`;`plugins`/`mods` 各自作业。无"存在但从无入口运行"的脚本。
- `internal/*` 22 个包逐一对账:0 提及者(`internal/apis`)为 CRD 类型包(被 40+ 站点引用、随 operator/manifests 演练间接全覆盖);`deploy/{limbo,lobby,paper}` 为三张游戏镜像定义,随每次起服演练执行。真盲区=无。
- `panel` 全部 17 条页面路由 × 既有批次(CDP 巡检、深挖、写流程)覆盖;`scripts/` 三脚本为**本地私有**(`.git/info/exclude`,不入库)。
- `docs/` 四件:troubleshooting(多批引用)、openapi(含 parity 强制测试)、deferred-seams(内容为现行状态,含 CLOSED/verified-live 注记,抽查与代码一致)、sequence-diagrams(发现漂移 → 修复,见第三节)。
**二、#66(`62c5a8a`):README_EN 与中文 README 漂移**
- 红证据:`README.md`(中文)"使用方式"含**必需**的私有仓库 workaround(仓库私有 → 裸 raw 命令 404;token 经 `curl --config -` 由 stdin 下发、不进 argv;`sudo -E` 传给安装器)与"重跑=升级 felis-api、通道不继承(跟 main 需 `FELIS_VERSION_BOOTSTRAP=dev`)"两段;`README_EN.md` 是 `README.md:6` 链接的英文入口,这两段**完全没有** → 英文读者照文档第一步就 404 且无指引。
- 修复:两段译入 EN(保留原命令与语义;`grep` 校验 token/config/dev 三关键串在位)。可达性:①(轻)——英文首装路径第一步。
**三、§28 序列图对齐(`62e8c87`)**
- 漂移点(两处,与修复后代码实读对照):① Claim Transaction 图注仍是 audit #4 前形态(`SELECT EXISTS` + 事务外 pre-check)→ 更新为现行实现:`pg_advisory_xact_lock(user_id)` + 行 `FOR UPDATE` + **事务内**四维配额门(图上 `QuotaAvailable` 只是 fast path);② Link Flow 更新为:取码 `FOR UPDATE`、同用户重验幂等、退役(soft-deleted)账户链接被当场接管——409 仅对"其他活跃用户"。均以 `internal/api/pgrepo.go` 实读 + handler 映射核对。
- 惯例依据:该文件有维护史(`676407d docs(diagrams): align §28 sequence diagrams with implemented routes`)。
**四、本地工具修复(私有脚本,不入库):`scripts/sync.sh` 排除 `plugins/` 与 `go:embed` 冲突**
- 红证据(本机复现):以 sync.sh 原排除清单 rsync 一个 fresh 目标 → `go build` 立即失败:`bootstrap_asset.go:30:12: pattern plugins/limbo/build.gradle: no matching files found`(`bootstrap_asset.go` embed 了 plugins/{limbo,paper,velocity,shared} 源码;排除清单成文于 embed 引入之前)。存量目标靠 rsync 对排除路径的"保护"而看似正常,fresh 目标必炸。
- 修复:排除项从 `plugins/` 改为 `plugins/*/build|.gradle|bin/`(与 `.gitignore` 同义);`watch.sh` 同步去掉 plugins 的忽略。验证:新清单 rsync fresh 目标 → `go build` 绿(83.7MB 二进制产出);两脚本 `bash -n` 通过。因脚本被 `.git/info/exclude` 私有,修复留在工作机、无 commit。
**五、release.yml 静态复核(该流水线从未执行过——仓库当前无任何 tag/release)**
- 三处输入/输出契约实核:① 发布资产名 `felis-linux-amd64/arm64` ↔ `deploy/bootstrap.sh` 的 `asset="felis-linux-${arch}"` 及收敛检查 `felis ${FELIS_REF}`(版本首行契约两侧一致);② `out/linux_amd64/usr/local/bin/felis` ↔ Dockerfile 最终层 `COPY /out/felis /usr/local/bin/felis`;③ 版本断言 `felis ${GITHUB_REF_NAME}` ↔ `cmdVersion` 的 `felis %s` 首行。
- 唯一无法在本机完全复刻的 `file(1)` 机器类型断言:用交叉编译的 linux/arm64 二进制实测输出 `ELF 64-bit LSB executable, ARM aarch64, …`,断言串 `ARM aarch64` 匹配(BSD/Ubuntu file 同源 magic)。
- 结论:整条流水线未跑属"尚未切 tag"的发布流程事实(第四十批已注记);切首个 tag 前无待修项。
**六、观察(不计缺陷)**
- `internal/api/handlers_account.go:169` 的 reclaim 选择仍为 Java/Velocity 侧 CODE-ONLY(deferred-seams 已记 accepting)——维持。
### 本轮新增真机证据(第四十三批:构建 Job 收尸 #67 + #68 上下文拉取重试 + 生命周期/故障注入压测)
**零、现场**:控制面从 auditfix61 升到 **auditfix62**(含 #67;自建镜像推入内建 registry,api/operator/reaper 三处滚动,`felis version` = `v0.0.0+fix62`)。本批证据 = 6 轮起停循环 + 3 处故障注入 + PG 断连语义 + 并发风暴 + 持续轮询。
**一、#67(`2755e41`):build Job 从不回收——每构建一个、完成 Pod 永久堆积**
- 红证据(真机):`felis-build` 里 5 个完成 Job/Pod 最长 26h 无人回收;全仓 TTL 对照——fileedit 2m / backup 10m / restore 10m / reaper CronJob 3+3 历史,**唯独 build 没有 `ttlSecondsAfterFinished`**;且 `build.go:69` 的注释写着 "(e.g. GC'd); treated as failed"——预期的 GC 从未存在。危害:完成 Pod 计入节点 pod 预算(stock k3s 110),构建量一上来先把节点塞满,新构建全部 Pending;etcd/磁盘同步膨胀。
- 修复:`buildJobTTL = 7 * 24h`(终态后计时;日志路由是失败分诊面故取长窗;`Sync` 对终态幂等、`JobUnknown` 只影响非终态 → 删除后的唯一代价是日志 404)。单测 `TestBuildJobIsReapedAfterCompletion`。
- 真机验证(auditfix62):① 真实路径触发构建(owner 会话 → `POST /api/v1/images/build`,借遗留 blob 作 context)→ 新 Job `ttlSecondsAfterFinished=604800`、构建 25s `succeeded`;② 对旧 Job 打 `ttl=30s` → 40s 内 Job+Pod 被 TTL 控制器收走(机制实证);③ 旧堆积 5 个里 1 个已收、4 个留存对照。
**二、生命周期 ×6 + 三处故障注入(operator / api / postgresql)**
- 每轮:CR `desiredState` Running → 等 Running → Stopped → 等 pod 消失。结果:**6/6 全收敛**——Running 23–29s(注入轮与无注入轮无差)、Stopped 3s、pod 每轮如期消失;收尾 conditions 无 Failed 残留(Ready=False / RconReached=False = Stopped 正常;Provisioned=True)。
- 注入 1(cycle 2,删 operator pod @t+2s):Running 仍 23s 达成(新 operator 立即接管 reconcile)。
- 注入 2(cycle 4,`systemctl restart postgresql`):CR 路径无感;另做定点验证:PG 停机窗口 `/me` = **503**(非 401,#11 语义保持)、`/healthz` = 200(存活探针独立)、内部 status = 200(纯集群读);PG 恢复后 `/me` = 200。
- 注入 3(cycle 5,删 api pod @t+2s):内部轮询出现 2 次 `http=000`(约 6s 窗口)后自愈;Running 28s 达成。
- 观察(不计缺陷):高频翻杆期间 operator 报 9 条 `Operation cannot be fulfilled ... object has been modified`(乐观锁冲突 → 重排队自愈),目标时间无差;controller-runtime 正常重试语义,生产低频操作下更罕见。
**三、并发风暴 + 拒绝面 + 持续轮询**
- 20 路并发 internal status + 10 路 `/me`:30/30 全 200。
- 无主 + ownerOnly 服 wake:403 `forbidden`(正确拒绝面,非 5xx)。
- 持续轮询(status + healthz + me,5s 一轮 × 300 轮 ≈ 25 分钟,900 样本):收盘 `LONG POLL DONE fails=27`——27 个失败样本**全部**落在三个自导演练窗口内(02:24:15/20 红证据杀 api〔6〕、02:29:48/53 auditfix63 滚动升级〔6〕、02:30:53–02:31:13 #68 确定性复现 scale api→0〔15〕),窗口外 873 样本全 200、零自发失败;每个窗口在动作结束后 ≤1 个探测周期(~5s)内恢复。
**四、留档与收尾**
- 留档:`/root/soak43.sh` + `soak43.log`、`/root/poll43-long.sh` + `poll43-long.log`、`/root/mint-owner-43.sh`(owner 会话重铸)、`/root/felis-image-build43.log`(镜像构建)、`/tmp/owner-jar43.txt`(owner cookie);镜像 `10.43.182.43:5000/felis/felis:auditfix62` 已入 registry(`e2e/ttl-probe:latest` = 本批 drill 镜像,保留)。
- VM 状态:docker 用毕已停(inactive)、k3s/velocity active;控制面三处 = auditfix62。
**五、#68(`b76d0ac`):构建 context 拉取对控制面重启零容忍——一次拒连即终态失败**
- 红证据(真机):删 api pod 后触发构建 `bld-1790187851749257160` → Job `Failed`;pod 时间线 `18:24:12Z` 起、context-fetch `18:24:13Z` exit 1,日志原文 `dial tcp 10.43.237.249:8081: connect: connection refused`;Kaniko 从未启动(PodInitializing);新 api pod 11s 后就绪、同 blob 前后各一次构建均 25s succeeded——纯粹"单次尝试"造成的无谓失败。
- 修复(`b76d0ac`):`fetchContextWithRetry`——传输错误/5xx 重试至 45s 窗口(3s 间隔),4xx 快速失败(是答案不是抖动);URL 校验提前为 usage error(2);保留原错误文案前缀。窗口/间隔为 vars(测试可缩窗);4 例单测:5xx 后恢复、拒连后恢复、404 不重试、窗口耗尽(全绿,含 `-race`)。
- 真机验证(auditfix63,确定性优先,不复刻竞态):`scale api→0` → 以复刻 Job 直建(由红证据 Job 派生:换名 `build-retry68`、剥 controller 标签与 selector、镜像改 auditfix63,复用 `felis-service-token` Secret)→ fetch 连续 8 次 `connection refused; retrying`(实捕日志)→ `scale api→1` → fetch 恢复、Kaniko 构建推送、trivy 干净 → **Job Complete**(`18:30:52Z→18:31:24Z`,32s);同形场景对照旧版 = 终态 Failed。正常路径回归:api 在线直构 `bld-1790188306964162960` → `succeeded`(10s),fetch 日志**零重试行**(重试不引入正常路径开销)。
- 附带发现(运维):本批升级沿用 `set image` 后 `FELIS_IMAGE` env 仍停 `auditfix61`(两代落后)——构建 Job 的 fetch 容器实际一直在用旧镜像;已把 api/operator 的 `FELIS_IMAGE` 与 api/operator/reaper 三处镜像全部对齐 `auditfix63`。真实安装器路径随 manifests 重渲注入该 env,手工升级须成对改(升级清单事项)。
### 本轮新增真机证据(第四十四批:modpack 规模上下文全链 + 并发构建)
**零、现场**:控制面 = auditfix63(含 #67/#68;api/operator `FELIS_IMAGE` 同版本),VM docker 停。本批目标 = 把"真实模组包大小"这条链压到生产尺度:此前所有提交/构建演练的上下文都是 KB 级。
**一、200 MiB 上下文全链(零缺陷)**
- 构造:5×40 MiB `/dev/urandom` 载荷 + Dockerfile(随机数据不可压缩 = 最坏情形),tar.gz = 209,750,275 B。
- 上传(真实 app 路由):`POST /api/v1/me/submissions/{id}/context`(owner 会话)→ 200,**0.44s / ~481 MB/s**;宿主核对 PVC 落盘 `sub-cc160b3ad19080f9/context.tar.gz`。服务端读写超时设计(ReadTimeout/WriteTimeout 有意不设、仅 ReadHeaderTimeout)与流式落盘(io.Copy → cappedReader)均按预期,无内存尖峰。
- 批准 → 构建(`bld-1790189227767686191`):fetch init **1s**(200MB 内部面拉取 + 解压,无重试行)、kaniko **4s**(unpack/COPY/snapshot/push gzip 层 **209,738,853 B**)、trivy **4s**(0 findings);Job Complete,全链 13s(18:47:07Z→18:47:20Z)。registry manifest 实查层大小一致(推送非虚)。
- 覆盖点随验:提交构建 Job `ttlSecondsAfterFinished=604800`(#67 对真实提交路同样生效);`/me/submissions` `build_status=succeeded`(#36 面在 200MB 规模下仍成立)。
- 资源采样(2s 粒度,构建太快可能漏尖峰,标注局限):无 OOM——api 17.9MB / kaniko 33.4MB / registry 34.1MB / trivy 35.9MB;emptyDir+快照为临时占用,随 Job TTL 回收。
- 1 GiB 帽的单测已有覆盖(`internal/submit/submit_test.go` oversize→ErrInvalid);帽下真机 = 本条(帽上真机不划算,不做)。
**二、并发构建 ×3(零缺陷)**
- 三路同时 `POST /api/v1/images/build`(各自 ref)→ 202×3;三个 Job 均 **11s Complete**,无串扰:每个 Job destination 与 tag 一一对应(conc-b44-1/2/3 各落位)、`ttl=604800` 各自在。
**三、磁盘运维注记(非缺陷)**
- 连续镜像构建(docker build)把 VM 根盘从 11G free 压到 4.3G;`docker builder prune -af` 一键回收 **6.5G**(回 11G free)。即每轮全量重建约需 5–7G 构建缓存空间;生产主机根盘建议 ≥40G 并定期 prune(如需可在部署文档补一行,本批未改码)。
- 留档:`/root/bigctx.tar.gz`(200MiB 上下文原件)、`/root/bigctx/`、`/root/stats44.log`、`/root/stats44-build.log`;提交 `sub-cc160b3ad19080f9`(blob 与 `user-uploads` 镜像保留作尺寸样本);镜像 `e2e/conc-b44-{1,2,3}:latest`。
### 本轮新增真机证据(第四十五批:模组构建链三连修 #70/#71/#72 + 提交→装服全链闭环 + 文档/测试修 #69/#73/#74)
**零、现场**:控制面 auditfix63→66(每修一版真机重演);本批目标 = 真实用户会写的那种模组包 Dockerfile(`FROM <平台底座>` + 内容层)。此前所有构建演练都是 `FROM scratch`——用户真实形态的第一步从未跑过,本批连撞三个缺陷。
**一、#70(`ac403b9`):kaniko 拉底座缺 pull 侧 insecure 标志**
- 红证据(真机):提交 `FROM registry.felis.svc:5000/felis/paper:demo`(`bld-1790189685537480076`)→ kaniko `Retrieving image manifest …` 后即败:`Get "https://registry.felis.svc:5000/v2/": http: server gave HTTP response to HTTPS client`。根因:Job 只给了 `--insecure --skip-tls-verify`——kaniko v1.24 `--help` 实测二者均只覆盖 **push**,拉取默认 HTTPS,撞上明文 registry。
- 修复:补 `--insecure-pull` / `--skip-tls-verify-pull`(对称补齐;构建命名空间 egress 本就只许 DNS/内建 registry,不扩大可达面)。单测断言两 flag 在参。
- 真机复验:下一次构建 `Retrieving image manifest` 不再报错,进入解包阶段(随即暴露 #71)。
**二、#71(`6e47730`):kaniko 解包底座层被 drop-ALL 卡死**
- 红证据(真机):`error building image: error building stage: failed to get filesystem from image: chown /etc/gshadow: operation not permitted`——kaniko 以 root 解包 tar 层要把文件 chown 到层里记录的属主(root:shadow 等),drop-ALL 后 CAP_CHOWN/CAP_FOWNER/DAC_OVERRIDE 全无。`FROM scratch` 全用 COPY(属主= kaniko 自己)所以从未暴露;**任何真实底座**必触。
- 修复:仅 kaniko 容器补回最小能力集 `CHOWN + DAC_OVERRIDE + FOWNER`(fetch/trivy 保持 drop-ALL 基线;单测断言"恰好三枚"防漂移)。
- 真机复验:解包通过、COPY、推送成功(进入 trivy 阶段,撞上 #72)。
**三、#72(`b14bacb`):trivy 扫 jar 必拉 Java DB——egress 锁下必败**
- 红证据(真机):kaniko 全绿后 trivy `FATAL … Unable to initialize the Java DB … failed to download artifact from mirror.gcr.io/aquasec/trivy-java-db:1: connection refused`。Java DB 按需下载:镜像一有 jar 就触发——模组包 = jar 集合,故所有真实用户构建都会倒在扫描门;旧构建全 scratch(无 jar)从未触发。
- 修复:新增 `[registry] trivy_java_db_repository`(→ `--java-db-repository`,与 vuln-DB 旋钮对称);installer carry 白名单收录;§8e 配方补 Java DB 镜像步骤;deferred-seams 更新。
- 真机复验(第 4 次尝试,auditfix66):Job spec 实测含 `--db-repository registry.felis.svc:5000/mirror/trivy-db:2 --java-db-repository registry.felis.svc:5000/mirror/trivy-java-db:1`;扫描 `ubuntu 26.04` + `paper/paper.jar` + `pebble` 全 0 漏洞;**Job Complete**(全链 27s,`bld-1790190788310114143`)。
**四、capstone:提交 → 装服 → 起服,全链闭环(零缺陷)**
- 第 4 次尝试产物 `user-uploads/sub-c006cbd633317ccb:latest`(= `felis/paper:demo` + 用户 marker,经完整提交管道构建):
- 白名单自动收录(`source=built`)→ 产品路由 `PATCH /api/v1/servers/test-one {"image": …}` **200**(`image_not_whitelisted` 门通过);
- `POST /servers/test-one/wake` 202 → pod `test-one-0` Running 1/1(25s 内);
- `kubectl -n minecraft exec test-one-0 -- cat /felis-probe-marker.txt` → **`felis-probe44-ok`**(用户构建上下文的内容确在运行中的服务器里);
- 服务器日志 `Done (4.678s)!`(paper 完整启动);RCON 线程应答平台就绪探针;
- 收尾:stop 202、pod 消失、镜像回滚 `felis/paper:demo`、desiredState=Stopped。
- 口径:这是"玩家上传模组包 → 服主审批 → 自动构建 → 选用为该服镜像 → 起服"的机制全链(真实客户端进服仍按决定跳过)。
**五、#69(`5cbfa89`):README 过度承诺"自动部署"**
- 复核(读全):提交数据模型无目标服务器字段;approve→build→白名单是唯一自动化;部署 = 服务器编辑框选镜像(产品路由与 UI 均在);面板文案本身只承诺"自动触发安全构建"。
- 修复:README 中英两行改为"自动构建;产物进入镜像白名单,可直接选用为服务器镜像完成部署"。可达性 ①(轻)。
**六、#73(`2961beb`):bootstrap 测试在跑 k3s 的主机上必假失败(工具级)**
- 红证据:VM(真实 k3s 主机)跑 `bootstrap_test.sh`(**原版同现**)→ `FAIL a missing worlds root is warned about`:用例拿真实路径 `/var/lib/rancher/k3s/storage` 期望 WARN,而该目录在"安装器真正要服务的机器"上必然存在 → 假红;CI 从未见到(runner 无此路径)。
- 修复:warn 情形改用保证不存在的路径;并把 `trivy_db_repository` / `trivy_java_db_repository` carry 断言补齐。VM 复跑 **140 PASS / 0 FAIL**。
**七、#74(`96aa817`):CI flaky——console 断连测试读写竞态(工具级)**
- 红证据:run `35907662213`(纯 README 提交)go 作业红:`TestServerConsoleDisconnectTeardown: expected the first event before disconnect, got ""`;`gh run rerun --failed` 即绿 → 竞态确认。
- 根因:测试只等"源读到第一行"就 cancel,"中继把事件写入响应"尚未发生;cancel 落进缝里时 body 为空。
- 修复:加 `firstDataWriter` 信号(写完成后再 cancel,写→读有 happens-before);本地 `-count=60` 与 `-race ×5` 全绿;CI 复跑 `5cbfa89` = success。
**八、运维注记(非缺陷)**
- 升级滚动撞 kubelet **ephemeral-storage 压力**:auditfix64 滚动时新 api/operator pod `Pending 4m46s`(事件 `untolerated taint(s)`),02:58:12 kubelet eviction manager 回收后自动调度成功。诱因 = 构建把根盘压到 89%(docker 构建缓存 + kaniko/emptyDir 临时层);`docker builder prune -af` 两清共回收 ~11G,余 11G free。生产清单:根盘 ≥40G + 升级前清构建缓存(与批 44 注记合并)。
- 中间失败构建留下的 registry 镜像(`user-uploads/sub-83f6…`、`sub-171e…` = 已推未过门;`sub-c006…` = capstone 产物)留档;测试服 `test-one` 已回 `felis/paper:demo` + Stopped;docker 停。
### 本轮新增真机证据(第四十六批:上传面 #75/#76 真机闭环 + tracker #8/#1 收口 #77/#78 + 文案 #79 + 发布链 rc smoke)
**零、现场**:代码三条 commit 先落(`43699b4` #77 面板跳转、`c57daaf` #78 converge、`b9ebc87` #79 INERT 文案),随后按 `b9ebc87` 重建镜像 = `registry.felis.svc:5000/felis/felis:auditfix77`(`v0.0.0+fix77`;宿主源 `10.43.182.43:5000`,docker build → push),控制面 api/operator + minecraft ns `felis-reaper` CronJob + api/operator 的 `FELIS_IMAGE` 四处成对齐;`/opt/felis/src` = b9ebc87 快照(上一版 `src.bak46`);宿主 drill 二进制 `/root/felis-fix77.bin`(Mac 侧 `GOOS=linux GOARCH=arm64` 交叉编译,`-X main.version=v0.0.0+fix77`)。
**一、#75 真机闭环(配额/限流,owner 会话,auditfix77)**——留档 `/root/probe77/create-lane.log`、`lane2.log`:
- **create 失败不烧窗口**:坏 JSON → 400,**同一秒**合法 create → 201(`sub-811427aa7dcf6807`)——`release` 路径生效;
- **create 冷却**:同用户 30s 内第二条 → 429 `submission_cooldown`;
- **pending 上限**:攒到 5 条 pending → 第 6 条 → 403 `submission_quota_exceeded`;
- **upload 失败不烧窗口**:对已审核行上传 → 409 `already_reviewed`,**同一秒**对 pending 行上传 → 200;
- **upload 冷却**:15s 内第二次 → 429 `submission_cooldown`;等 16s → 200;
- **存储预算**:把 S2 的 blob 稀疏 `truncate` 到让 owner 已存字节 = 2GiB−100B(`blocks=8`,不占盘)→ 上传 → 403 `submission_quota_exceeded`;管理员删除 S2(行 + 目录双清)后 → 同一上传 → 200(预算即时释放)。
- 口径:预算按 blob 的实际占用聚合(`Blobs.Size` 遍历汇总,正是"用久了就超"的同一条读取路径);稀疏垫付只是把"已经存了 2GiB 的用户"这一状态合成出来。
**二、#76 真机闭环(撤回 + 管理员删除)**——留档 `/root/probe77/lane3.log`:
- 撤回自己 pending → 200;DB 行 = 0、uploads 目录 = gone;重复撤回 → 404;对已审核 → 409 `already_reviewed`;对他人 id → 404 `not_found`(owner 判定先于状态判定,不泄露他人行状态);
- 管理员删除(pending)→ 200;重复 → 404;行与 blob 双清(lane2 的 S2 同证:`s2 rows=0 dir=gone`);
- 面板侧 CDP:**撤回两步确认**(展开行 → 「撤回提交」→「确认撤回」→ 行消失)与 **admin 队列每行两步删除**(trash →「确认删除」→ 行消失、计数回落)双双走过;截图 `/tmp/withdraw-{1,2}-*.png`、`/tmp/admdelete-{1,2}-*.png`(Mac),脚本 `/tmp/cdp-withdraw76.js`、`/tmp/cdp-admdelete76.js`。
- 真机注记(非缺陷,分级 ZT 的对照实证):admin 面只在 `op.console.<root>` 主机可用——同一 owner 会话在玩家面板主机上 `is_admin=false`(admin 路由 403),在 `op.console.<root>` 上 `is_admin=true`(200);面板按要求显示「无权访问」而不是假装能点。
**三、#77(tracker #8,`43699b4`)真机闭环:锁定会话 → /setup**
- 铸一个未完成引导会话(SQL:`sha256('drill-locked-77')` → `sessions`,用户 `3f2c1b0a-…`,`email_verified=f`、无 passkey)→ `GET /me` 200、`GET /me/submissions` → 403 `setup_required`(后端本就正确);
- CDP 用该 cookie 打开 `/submissions`:页面请求 `/me/submissions` 收 403(code=`setup_required`)→ **自动跳转 `/setup`**(`location.href` 实测)→ 向导渲染「初始化你的账户 / 第一步 · 填写邮箱」(截图 `/tmp/setup8-redirect.png`;网络事件 = 200/403(/setup/status 200) 序列在脚本输出里)。
- 顺带确认:Dashboard 首屏三请求(`/me`、`/me/servers`、`account/link/start`)都是 `SetupAllowed`——所以补丁前用户是在"点受保护操作"时才撞到那句误报「无权执行此操作」。
**四、#78(tracker #1,`c57daaf`)真机演练:`felis converge`**
- 现场先剥后补:`kubectl patch` 清空 lobby `spec.rcon` 与 login `spec.startup.healthHTTPPort`(复刻老装机 CR 缺后加字段的状态)→ `/root/felis-fix77.bin converge` → 两条全部填回(`rcon{enabled:true,secretRef:lobby-rcon/password}`、`healthHTTPPort:8080`);**第二次运行 → 双双 `already converged`**(幂等);operator 侧滚动收尾,login/lobby 回 Running/Ready(日志 `converge-{1,2}.log`、`crs-before.yaml`)。
- 语义边界(单测):非零值一律不碰(运维自换的 secretRef 生还)、缺席 CR 只提示(创建仍归 setup)、占名非系统角色 CR 拒绝收敛、未配镜像跳过。
**五、#79(`b9ebc87`,文档)**:troubleshooting 的 `[INERT]` 图例仍写「§12 列着唯一仍适用的字段」,而唯一候选 `spec.storage.retainOnDelete` 早已移除、§12 自述「每个字段都有 controller 读」。改为「今天没有字段处于该状态」并指向 §13 的移除记录。
**六、发布链 rc smoke(tag `v0.1.0-rc1` → release run `35949621233`)**:
**六、发布链 rc smoke(tag `v0.1.0-rc1` → release run `35949621233`):首次全绿**
- run 步骤实况:go vet/test ✅ → `plugins/test.sh` ✅ → `plugins/test-mods.sh` ✅(**release.yml 史上第一次真正跑这三关**)→ buildx 双架构构建 ✅ → **stamp 校验** ✅(`felis v0.1.0-rc1` + arm64 ELF 断言)→ publish ✅;
- 产物释出:`felis-linux-amd64`(61,378,722 B)与 `felis-linux-arm64`(57,344,162 B)双资产,release 标记 `prerelease=true`;
- **prerelease 语义复核**:`GET /repos/FelisMC/Felis/releases/latest` → **404**(RC 没有顶掉 latest——正是 release.yml 注释里防的那件事);bootstrap 默认通道在无 stable 时按设计给出显式指引后 die、`felis update` 的 404 报错可读——两个行为本批均实测;
- 产物级复验(目标架构实机):arm64 资产 scp 到 VM 执行 → `felis v0.1.0-rc1`(`go1.26.8 linux/arm64`),且直接可用:`/root/felis-rc1.bin converge` 打现集群 → 双服 `already converged`;
- 留档:`/root/felis-rc1.bin`;asset 副本 Mac `/tmp/rc1b/felis-linux-arm64`。
- 剩余(产物决定,未代拍板):仓库仍无 **stable** release → 新装走默认 release 通道会以指引性报错 die(提示改 dev 或等 stable);切首个稳定 tag(如 `v0.1.0`)即让默认通道与 `felis update` 真正上线,建议维护者择时执行。
**七、运维注记(非缺陷)**
- pg_hba 的 `host felis_pgint felis 127.0.0.1/32 scram-sha-256` 行**再次丢失**(批 31 补过一次)——补回并 reload 后 pgint 才能连。重装/动过 PG 后先查这条(已写进速查)。
- 升级域注记:#75/#76 的 pgint 新断言(`CountPendingSubmissionsBy`、`DeletePendingSubmission`/`DeleteSubmission` CAS)首次上真 PG:**17/17 全绿**。
- 根盘:镜像构建后 83% → `docker builder prune -af` 回收 3.5G,docker 停回 inactive。
- CI:`43699b4` success;`c57daaf` 被后一 commit 的并发策略取消(同分支 cancel-in-progress),其树被 `b9ebc87` 的 success 完整覆盖;`b9ebc87` success。
- tracker 收编:**关** #10/#20/#21/#22(#20→`d9246dd`、#21→`f5a76cf`、#22→`3f2b28d`、#10→`23792d6`,均附证据评论)+ **关** #1/#8(本轮实现并真机验证);**注记**(保持打开作老装机待办)#2/#3/#4;**留存** #9/#12/#13/#15(真增强,超出生产可用主干,未动)。
## 结论:离"生产可用"还差什么(按优先级)
1. ~~构建链路的上下文通道~~ ✅ **已修**(`f79e5eb`/`02fd2de`,真机全链路含拉回校验;Trivy DB 需按 §8e 镜像一次)。
2. ~~镜像耐久~~ ✅ **已完成**(第二十九批:`a9b275a`/`13d64e0`/`fa0e8d7` + `c7e585e`;三次真机重跑 + GC 两演练——控制面与游戏镜像被 GC 后自动回拉;同批修复并复验 #46/#47/#48/#49)。
3. ~~PG 级契约测试~~ ✅ **已落地**(`2a55a0d`):`internal/pgint`(`-tags pgint`,需 `FELIS_TEST_PG_URL` 指向名字含 `pgint` 的库,harness 会 drop schema + 重放真实迁移)已覆盖会话/OTP/op-login/绑定码/submission/build/owner 角色;**首跑即抓到 #20**(索引与 ErrEmailTaken 从未存在)。运行方式见 CONTRIBUTING.md。
4. ~~面板把 /jobs 显示出来~~ ✅ **已完成**(`97a64c8`,备份页「最近操作」卡,正是它把 #35 暴露出来的)。
5. ~~多节点回收~~ ✅ **已修**(`daf7602`:`--reaper-node` → CronJob pod `kubernetes.io/hostname` nodeSelector,真机 render/dry-run/收敛 diff 三连;单节点部署不传即维持原状)。附带核销缺陷 #38(渲染提示里过期的 uid-1000/setfacl 指导)。
6. ~~告警~~ ✅ **已完成**(`94f71ee` + `43df08b` + 第二十七批实弹演练:真实构建失败 → pending → 08:33:14Z firing;规则随 `deploy/alerts/` 交付)。
7. ~~玩家可见的构建结果~~ ✅ **已修**(`72c4aa3`,缺陷 #36:列表路由附 `build_status`/`build_error`,面板「我的提交」展开行呈现,真机双例验证)。
8. ~~面板文件编辑器入口~~ ✅ **已补**(`0a36b3f`,缺口补齐,真机 CDP 全链)。
9. ~~fleet 系统服务死操作~~ ✅ **已修**(`2f90851`,缺陷 #37)。
10. ~~S3 上传通道演练 + 存储向导逐屏~~ ✅ **已演练**(第二十四批:上传通道全链、0 缺陷;第三十二批:向导屏逐屏 + 上传取件 + UI 回滚,0 功能缺陷并修 #52)。
11. ~~breakGlass 控制台逐屏~~ ✅ **已演练**(第二十五批:菜单 4 操作 + bootstrap 分支,缺陷 #39–#43 全修全验)。~~`felis setup` 全屏向导逐屏~~ ✅ **重跑侧已逐屏**(第二十六批),**首装侧连续走查也已完成**(第三十一批:scratch 库单次连续走完全部屏幕 + 轨道回顾,0 缺陷)。
12. ~~`felis nano` 全链~~ ✅ **已走查**(第三十四批:HTTP 矩阵红 13→绿 60/60;systemd/配置/firewalld 三层重跑全绿;同批修复 #55——坏源静默挡住验证梯子)。
13. ~~NetworkPolicy 真机强制矩阵~~ ✅ **已演练**(第三十五批:pod 源端口级矩阵全绿、转发型外部源「拒→放→拒」白名单闭环、host/ClusterIP/NodePort 路径全通;两条已归因注记——node-local 流量豁免、firewalld 仅放 DNAT 服务流)。
14. ~~面板错误文案全量本地化~~ ✅ **已修**(第三十六批:#60,37 个用户可达码补齐双语映射,三类真机验证;管理面写操作 Run4c 复核 0 缺陷)。
15. ~~管理面交互级收尾(submissions / 会话撤销 / 启动)~~ ✅ **已演练**(第三十七批:点击→构建→succeeded 全链、单条/全部撤销精确验证;同批修 #61——LP 写承诺按 LP 官方语义收窄,真机复验)。
16. ~~缺陷 #56–#59(玩家管理回包 / 控制台回显 / 文件页 fail-fast / LP 诚实化)~~ ✅ **已修并归档**(第三十八批回填真机证据:修复前触发 + 修复后复验,四项全绿)。
17. ~~服务器详情与运维面收尾~~ ✅ **已演练**(第三十八批:ServerFiles 写流程、备份/恢复全链、白名单/封禁闭环、ServerCard 停止、`/admin/updates` 复跑、metrics、CLI 核销——0 缺陷)。
18. ~~reaper 多节点实机~~ ✅ **已演练**(第三十九批:VM 克隆双节点 k3s——pin 命中持盘节点并跑通、错位 pin fail-closed Pending、装置全回收;`felis-backups` PVC 的 volume node affinity 构成第二层保险)。
19. ~~Java 插件层自动化缺口~~ ✅ **已补**(第四十批:`plugins/test.sh` + CI `plugins` 作业——三个手工测试与三个装机 jar 全进门禁,CI 4/4 绿;首跑抓出两处"没跑过"的错误:InviteCardTest 编译类路径缺 examination-api、limbo `+` 版本不可解析。同批关闭 demo-up 单起点化 #63、并复核剔除 #62)。
20. ~~装载器 mod 层(fabric/forge/neoforge)~~ ✅ **已收口**(第四十一批:#64 三个 `gradlew` 补可执行位、#65 元数据对齐 AGPL-3.0-only;`plugins/test-mods.sh` + CI `mods` 作业 + release 双门禁;三真实专用服起服 E2E——mod 装载、`/link` 注册、控制台拒绝全绿。CI run `35892544563` = 5/5,`mods` 作业首跑即绿)。
21. ~~全仓覆盖面对账 + 文档收尾~~ ✅ **已完成**(第四十二批:22 个 internal 包 × 面板 17 条页面路由 × 全部 tracked 脚本 × 仓库卫生逐项对账,无盲区;#66 README_EN 修复;§28 两张序列图对齐现行实现;本地 `scripts/sync.sh` 排除清单与 `go:embed` 冲突——修复+复现验证〔脚本私有,不入库〕)。
22. ~~构建层资源回收~~ ✅ **已修**(第四十三批:#67 build Job TTL——完成 Job/Pod 不再无限堆积;真机:新构建 Job `ttl=604800` + 旧 Job 打 30s TTL 实测被 TTL 控制器收走。同批:6 轮生命周期循环 + operator/api/pg 三处注入全绿、PG 503 语义复验、30 路并发全 200)。
23. ~~构建上下文拉取的短暂断连~~ ✅ **已修**(第四十三批追加:#68——fetch-context 有界重试〔45s 窗口 / 3s 间隔;4xx 不重试〕;真机:api 停机期连续 8 次拒连全部重试、恢复后构建 32s Complete;正常路径无重试开销。附:演练升级暴露的 `FELIS_IMAGE` 两代错位已对齐 auditfix63)。
24. ~~modpack 规模上下文全链 + 并发构建~~ ✅ **已演练**(第四十四批:200MiB 上下文上传 0.44s → 批准 → 构建全链 13s、三路并发构建全绿;TTL 与玩家可见面在规模下复验,零缺陷。见该批节)。
25. ~~内建底座构建(`FROM registry.felis.svc:5000/…`)~~ ✅ **已修并真机闭环**(第四十五批三连修:#70 拉取缺 `--insecure-pull`、#71 drop-ALL 卡解包、#72 trivy Java DB 未镜像——"真实模组包"形态的三个必踩点;修复后 paper 底座构建 27s Complete、`paper.jar` 扫描 0 漏洞)。
26. ~~提交 → 装服 → 起服 闭环~~ ✅ **已演练**(第四十五批 capstone:构建产物 PATCH 为服务器镜像〔白名单门通过〕→ wake → Running 1/1 → 用户 marker 可读 + `Done (4.678s)!`;test-one 已复原为 paper:demo/Stopped)。
27. ~~文档与测试工具面~~ ✅ **已修**(第四十五批:#69 README 部署承诺对齐实现;#73 bootstrap 测试密闭化〔VM 140 PASS〕;#74 console 断连测试消抖;CI 全绿)。
28. ~~提交上传面的上限与生命周期~~ ✅ **已修并真机闭环**(第四十六批:#75 per-user create/upload 冷却〔429〕、pending ≤5〔403〕、2GiB 存储预算〔403〕;#76 撤回〔owner+pending CAS〕与管理员删除〔行+blob 双清〕;面板两步确认 CDP 全绿)。
29. ~~未完成引导的导航(tracker #8)~~ ✅ **已修并真机闭环**(第四十六批 #77:`403 setup_required` → 自动 `/setup`;CDP 复验)。
30. ~~已装机系统服的新增 CR 字段(tracker #1)~~ ✅ **已修并真机闭环**(第四十六批 #78:`sudo felis converge`——零值才填、非零不覆写;剥字段→填回→幂等三连真机过)。
31. ~~troubleshooting `[INERT]` 图例~~ ✅ **已修**(第四十六批 #79,文档级)。
## 剩余待演练队列(截至第四十三批)
- ~~reaper 多节点实机~~ ✅ 已完结(第三十九批:临时克隆第二节点组双节点 k3s 实机——pin 命中持盘节点并跑通、错位 pin fail-closed Pending;装置已回收)。
- ~~面板交互级收尾~~ ✅ 已完结(第三十六–三十八批:管理面写操作、submissions/会话、服务器详情页全链,0 缺陷)。
- ~~`demo-up.sh`~~ ✅ 已完结(第四十批:单起点化 #63——bootstrap 装机 + 三项硬校验 + `felis setup` 交棒;T1/T2/T3 真机全绿)。
- ~~Java 插件层自动化~~ ✅ 已补(第四十批:`plugins/test.sh` + CI `plugins` 作业;三个手工测试与三个装机 jar 首进 CI)。
- ~~装载器 mod 层(fabric/forge/neoforge)~~ ✅ 已收口(第四十一批:#64/#65 修复 + `plugins/test-mods.sh`/CI `mods`/release 双门禁 + 三真实专用服起服 E2E)。
- ~~全仓覆盖面对账~~ ✅ 已复核(第四十二批:模块×台账逐项对账无盲区;批次产出为文档/本地工具层面的修复,见该批节)。
- ~~长稳/故障注入(首轮)~~ ✅ 已完成(第四十三批:6 轮起停循环 + 3 处注入 + PG 断连 503 + 30 路并发 + 25 分钟持续轮询;见该批节)。
- ~~modpack 规模上下文 + 并发构建~~ ✅ 已完成(第四十四批:200MiB 上下文全链 13s + 并发 ×3 全绿;见该批节)。
- ~~提交 → 装服 全链(含平台底座 FROM)~~ ✅ 已完成(第四十五批:三连修 #70/#71/#72 + capstone 提交→装服→起服;见该批节)。
**队列现已清空。**(真实游戏客户端进服已按决定跳过,用内部 mint/approve hook 链替代;不依赖客户端的最大真机面见第四十批节 + 第四十一批节〔装载器三服起服〕;第四十二批为覆盖面对账复核。)
## 系统性观察
- **PGRepo 与接口契约/fake 漂移**(#16/#17/#18):`fakeRepo` 与接口注释是对的、PG 实现在细节上落后,单测全绿也发现不了。建议后续引入 PG 级契约测试(testcontainers 或针对关键写路径的集成测试),重点覆盖「接口注释承诺了字段/错误码/生命周期」的方法。
- 部署知识:交叉编译必须先把 `panel/dist` 拷进 `internal/panel/static` 再 `go build`,否则镜像内 SPA 缺失(页面显示 “assets were not built”)。本批镜像为 `felis:auditfix7`。
## 复现入口速查
- 面板会话 cookie:`/tmp/owner-cookies2.txt`(2026-09-23 铸;auditfix42 部署后仍有效;staff 邮件门被设计拒绝 → 走第二十八批的 op-login + 内部面 approve hook 链);API 助手:`/tmp/fcurl.sh`
- 测试服:`test-one`(minecraft ns,stopped);合法备份 `bk-47ee2e7e96e5a4ca9d0e51b805518bac`
- CDP 调试口:Mac `127.0.0.1:9333`(独立 Chrome,profile `/tmp/felis-chrome2`);WebAuthn 虚拟认证器需在**同一 CDP 会话**内完成仪式,且先 `Page.bringToFront`(否则 NotAllowedError: page does not have focus)
- 已部署到 VM:控制面(api/operator/reaper + `FELIS_IMAGE`)= `registry.felis.svc:5000/felis/felis:auditfix61`(含至 #61 的全部修复;镜像托管于内建 registry);系统服/推荐 ref = `registry.felis.svc:5000/felis/{limbo,lobby,paper}:demo`;源码快照 `/opt/felis/src`(`b4ef42d`;上一版在 `/opt/felis/src.bak`,更早 `src.old43`);host 二进制 = `/usr/local/bin/felis` = `v0.0.0+fix61`(含 #53–#61;旧件留档 `/root/felis-fix54.bin`(sha `a0b29153…`)、`/root/felis-fix52.bin`、`/root/felis-auditfix44.bin`);**docker 守护进程在第三十七批构建后已停(用前 `sudo systemctl start docker`)**;**world 执行器以 root+DAC_OVERRIDE 运行**;reaper CronJob(minecraft ns)`suspend=false / …auditfix61 / nodeSelector=localhost.localdomain`;迁移 `schema_migrations` max=21;SMTP 密码 env 待「configure email」刷新时落地
- 安装器重跑配方(第三十批复用;先同步 `/opt/felis/src` 再跑):`cd /root && FELIS_SKIP_FETCH=1 FELIS_INSTALL_MODE=full FELIS_IMAGE=registry.felis.svc:5000/felis/felis:<tag> FELIS_WORLDS_HOST_PATH=/var/lib/rancher/k3s/storage nohup bash /opt/felis/src/deploy/bootstrap.sh > /root/bootstrap-<tag>.log 2>&1 &`;完成后 `grep -c '\[fail\]'` 应为 0
- 第三十一批 drill 留档:`/root/preTUI43/`(首装走查前快照:两 toml + 两 ns Secret + sha256、走查后 Secret 快照)、`/root/pre51-replica.toml`/`post51-replica.toml`/`post51b-replica.toml`(#51 红/绿副本证据);走查残留(scratch 库、scratch 配置、hba 行、tmux 会话)均已清理
- 第三十二批 drill 留档:`/root/preS3/`(S3 走查前快照:两 toml + 两 ns Secret + sha256 + deploys.txt)、`/root/s3wiz/`(posts3 快照、postroll-sha256、fix52-red.txt、fix52-green.txt)、`/root/felis-fix52`(含 #52 的宿主二进制);MinIO 容器+卷+两镜像、`felis-uploads-s3` Secret、DB 行(`image_submissions` `sub-853a4e2ba4ba6443`)、/tmp 残留均已清理,docker 已停
- 第三十三批 drill 留档(VM):`/tmp/felis-fix53`、`/tmp/felis-fix54`(宿主二进制;sha `9afd141d…` / `a0b29153…`)、`/root/felis-fix52.bin`;`/tmp/player-cookies.txt`(player 会话);面板 CDP 驱动脚本 `/tmp/cdp-updates2.js`(Mac 侧)
- 第三十四批 drill 留档(VM):`/srv/nanotest/`(nano 装置:`nano-stub.py` + 5 驱动脚本 + `matrix/` 红原跑 / `matrix-fix55/` 绿原跑)与四层重跑落盘 `matrix-red.r2.log`(`pass=47 fail=13`)/ `matrix-fix55.r2.log`(`pass=60 fail=0`)/ `nano-service.r2.log` / `nano-config.r2.log` / `nano-fw.r2.log`;`nano-stub` 临时单元仍在跑(收尾 `systemctl stop nano-stub`);宿主二进制 `/usr/local/bin/felis-nano-test` = fix55 改名件(sha `2c8c34fd…`,与 `/usr/local/bin/felis` 同物)
- 第三十五批 drill 留档(VM):`/srv/npdrill/`(netpol 装置全套:`np-matrix.sh` + `np-matrix-run1/2.log`、`np-netns.sh`、`netns-probe.py`、`np-cidr-test.sh/.log`、`spec-before/after.json`、`td3.log`(包级归因)、`tcpdump-8081.log`(刷新存活)、`iptables-save.txt` 与 `nft-rules.txt`);netns/veth、firewalld 临时挂载、策略改动全部已清理/复原(spec md5 `e6b07440…` 前后一致)
- pgint 正确跑法(走 VM 的 PG,勿在本机起容器):Mac 侧短隧道 `ssh -6 -i ~/.ssh/id_ed25519 -N -L 15433:127.0.0.1:5432 root@…` → `FELIS_TEST_PG_URL='postgres://felis:<pw>@localhost:15433/felis_pgint?sslmode=disable' go test -tags pgint ./internal/pgint/ -count=1`(pg_hba 需 `host felis_pgint felis 127.0.0.1/32 scram-sha-256`;第三十一批把现场丢失的这条补回;隧道用完即 kill)
- breakGlass TUI 驱动法:VM tmux `new-session -d -s bg -x 160 -y 45 "/root/felis-auditfixNN breakGlass"` + `set-window-option -t bg remain-on-exit on`(否则退出摘要读不到);send-keys/capture-pane 驱动;退出码 1 = 操作失败(卡上会显示原因)
- 构建链路 drill 现成条件:`felis-build` 里有 `felis-service-token`(Secret 复制);`felis-config` 的 kaniko/trivy pin 指向内建 registry 的 mirror 路径(`registry.felis.svc:5000/mirror/...`,见 §8e)——本批后安装器重跑会 **carry** 这些 pin(不再被写盘冲掉;#49 复验点);直构一行:`POST /api/v1/images/build`(context=`http://felis-api-internal.felis.svc.cluster.local:8081/api/v1/internal/submissions/sub-bf7dc18e9dd97dc2/context`)
- registry 工具:`curl -s http://127.0.0.1:5000/v2/_catalog`、`/v2/<repo>/tags/list`;**取 manifest 必须带 `Accept:`(OCI index/manifest list;缺 Accept 的 GET 会 404,别误判没推上)**;GC 演练配方:`k3s ctr images rm <ref>`(要清到 blob 级再加 `k3s ctr content prune references`)→ 删 pod / `rollout restart` → `kubectl describe` 看 `Pulled … in …ms`
- setup 重跑向导(第二十六批):状态屏 `c/s/e` 三流已走查;**reconfigure 非只读**——storage 选 Local / 提交 S3 即 apply + 滚 API(幂等),email 表单 esc 无副作用;从 S3 表单 esc 退回会重置为 Local 预选;存储屏(第三十二批补充):chooser 预选当前值(↑↓ 切换),Local 直接 enter 提交、S3 五字段 `enter next`/末字段 `enter submit`,提交后 10–40s(含 API rollout status 180s 上限)
- VM 内部面:`k3s kubectl -n felis port-forward svc/felis-api-internal 18081:8081`(Pod 重建后转发会悬死,需重启);reaper CronJob(minecraft ns)已应用且已收敛(`suspend=false`)
- 第三十八批收尾复查:docker `inactive`;控制面三处 = `auditfix61`;`/opt/felis/src` = `b4ef42d` tarball 展开(无 `.git`);reaper `suspend=false` + `nodeSelector=localhost.localdomain`;test-one `Stopped`;`resolvecheck` 文件接口 409 `no_world_volume`(0.03s)
- 第三十九批 drill(已回收):临时 VM `felis-node2`(`prlctl clone --linked` 克隆 node1;node2 停用 felis-velocity/k3s-server/docker 后以 agent 加入)——流程与两向对照见该批节;node1 为克隆经历过一次正常重启(组件全回归、docker 停回 `inactive`);收尾 job/节点对象/VM 全部删除,集群回到单节点
- 第四十批工具与留档:VM `/tmp/mcprobe.py`(纯 socket MC 探针:`python3 /tmp/mcprobe.py status <sub>.<root>`;`login <sub>.<root> <name>` 打登录首包);插件门禁跑法 = `docker run --rm -v /opt/felis/src:/src:z -w /src gradle:8.14-jdk21 bash plugins/test.sh`(或宿主 `bash plugins/test.sh`,需 JDK 21 + Gradle 8.14;日志 `/root/plugins-test-b40.log`);T3(demo-up 默认臂全量重跑)日志 `/root/demo-up-b40.log`;`/opt/felis/src` = `c59b387` 快照(上一版 `src.bak40`;此后全量重跑用 `sudo bash /opt/felis/src/deploy/demo-up.sh` 即可,SKIP 开关见脚本头);CI run `35888009965`(含新 `plugins` 作业首跑绿)
- 第四十三批工具与留档:控制面三处 = `registry.felis.svc:5000/felis/felis:auditfix62`(含 #67;镜像源 `10.43.182.43:5000/felis/felis:auditfix62`,宿主 docker build 后 `docker push 10.43.182.43:5000/...`);`/opt/felis/src` = `2755e41` 快照(上一版 `src.bak42`);owner 会话重铸 = `bash /root/mint-owner-43.sh`(cookie 落 `/tmp/owner-jar43.txt`);直构触发(含 TTL 检查)= `curl -sk -b /tmp/owner-jar43.txt -X POST https://127.0.0.1:30443/api/v1/images/build -H 'Content-Type: application/json' -d '{"image_ref":"registry.felis.svc:5000/e2e/ttl-probe:latest","dockerfile":"FROM scratch\nLABEL felis=x\n","context_ref":"http://felis-api-internal.felis.svc.cluster.local:8081/api/v1/internal/submissions/sub-bf7dc18e9dd97dc2/context"}'` → `kubectl -n felis-build get job build-bld-<id> -o jsonpath='{.spec.ttlSecondsAfterFinished}'` 应为 `604800`;压测编排 `/root/soak43.sh`(日志 `soak43.log`,含 6 轮循环 + 3 注入)、长轮询 `/root/poll43-long.sh`(日志 `poll43-long.log`);镜像构建日志 `/root/felis-image-build43.log`。注意:`lobby`/`login` 是保留名,internal per-server 路由(status/wake 等)对它们**按设计**返回 400,勿作缺陷误报。
- 第四十三批追加(#68)留档:控制面 api/operator/reaper 镜像 + api/operator `FELIS_IMAGE` = `registry.felis.svc:5000/felis/felis:auditfix63`(`felis version` = `v0.0.0+fix63`;升级时 `set image` 与 `FELIS_IMAGE` env 必须成对改,否则构建 Job 的 fetch 容器会悄悄用旧镜像);`/opt/felis/src` = `b76d0ac` 快照(上一版 `src.bak43` = `2755e41`);复刻 Job `/root/job68-new.json`(由红证据 Job `build-bld-1790187851749257160` 派生:换名 `build-retry68`、剥 controller uid 标签与 selector、镜像改 auditfix63、复用 `felis-service-token` Secret);重演验证法:`kubectl -n felis scale deploy/felis-api --replicas=0` → `kubectl -n felis-build create -f /root/job68-new.json` → `kubectl -n felis-build logs <pod> -c context-fetch`(应见 `connection refused; retrying`)→ `scale --replicas=1` → Job 收敛 `Complete`;对照:api 在线直构成功且 fetch 日志无重试行。
- 第四十四批工具与留档:200MiB 上下文构造 = 5×40MiB `head -c 41943040 /dev/urandom` + `Dockerfile`(`FROM scratch` + `COPY payload /payload`),`tar -czf /root/bigctx.tar.gz Dockerfile payload`;上传 = `curl -sk -b /tmp/owner-jar43.txt -X POST -H "Content-Type: application/gzip" --data-binary @/root/bigctx.tar.gz https://127.0.0.1:30443/api/v1/me/submissions/<id>/context`;采样 = `k3s crictl stats`(此版 **无 `--no-trunc`**,列解析取 `$2=NAME $4=MEM`);docker 缓存清理 = `systemctl start docker && docker builder prune -af && systemctl stop docker`;留档 `/root/stats44.log`、`/root/stats44-build.log`、提交 `sub-cc160b3ad19080f9`、镜像 `e2e/conc-b44-{1,2,3}`。
- 第四十五批工具与留档:从平台底座构建的配方 = 提交上下文含 `Dockerfile`(`FROM registry.felis.svc:5000/felis/paper:demo` + `COPY marker.txt /felis-probe-marker.txt`);装服验证 = `PATCH /api/v1/servers/test-one {"image":"registry.felis.svc:5000/user-uploads/<sub>:latest"}` → `POST /servers/test-one/wake` → `kubectl -n minecraft exec test-one-0 -- cat /felis-probe-marker.txt`;trivy 双 DB 镜像 = `mirror/trivy-db:2` + `mirror/trivy-java-db:1`(Java DB digest `5766dfbb…`;两键在 `/etc/felis/felis.{host,pod}.toml` 与 felis-config Secret 双副本);`bootstrap_test.sh` 需在 Linux 跑(VM 留档 `/root/btest72/`,应 140 PASS);本批提交样本 `sub-{b45d416f4af811ef(无镜像),83f677e41acd80c3,171e8f967f0de3bc,c006cbd633317ccb,cc160b3ad19080f9}`;留档 `/root/probe44/`、`/root/probe44.tar.gz`、`/root/probe44*.sid`。
- 第四十六批工具与留档:控制面三处 = `registry.felis.svc:5000/felis/felis:auditfix77`(宿主 docker 源 `10.43.182.43:5000/felis/felis:auditfix77`,`v0.0.0+fix77`;api/operator + minecraft ns `felis-reaper` CronJob + api/operator `FELIS_IMAGE` 四处成对改);`/opt/felis/src` = `b9ebc87` 快照(上一版 `src.bak46`);宿主 drill 二进制 `/root/felis-fix77.bin`(Mac 侧 `CGO_ENABLED=0 GOOS=linux GOARCH=arm64 go build -ldflags="-s -w -X main.version=v0.0.0+fix77" ./cmd/felis`;scp 到 IPv6 主机时主机位必须写 `root@[fdb2:…]` 方括号);#75/#76/#77/#78 留档 `/root/probe77/`(`create-lane.log`、`lane2.log`、`lane3.log`〔含玩家会话矩阵 + 分级 ZT 对照 + curl 7.76 在 `-H Host:` 下**不发 jar cookie**的坑——host 路由 drill 用显式 `-H "Cookie: felis_session=…"`〕、`converge-1.log`/`converge-2.log`、`crs-before.yaml`、各步请求/响应 json);面板 CDP 截图(Mac):`/tmp/setup8-redirect.png`、`/tmp/withdraw-{1,2}-*.png`、`/tmp/admdelete-{1,2}-*.png`,脚本 `/tmp/cdp-setup8.js`、`/tmp/cdp-withdraw76.js`、`/tmp/cdp-admdelete76.js`;锁定会话铸法 = `sha256('drill-locked-77')` 直插 `sessions`(用户 `3f2c1b0a-…`);玩家会话铸法 = `[email protected]` 邮件 OTP(码在 felis-api 日志 `no Mailer configured` 行;`edge-dis2` 是软删死账号,登录门按设计排除);pgint 前置:`pg_hba` 需 `host felis_pgint felis 127.0.0.1/32 scram-sha-256`(本批第二次丢失并补回);发布链 smoke = tag `v0.1.0-rc1`(release run `35949621233`),prerelease 不移动 `/releases/latest`——无 stable 时 bootstrap 默认通道给出显式指引后 die、`felis update` 404 报错可读;首个稳定 tag 是产物决定(未代拍板)。
+14
View File
@@ -67,6 +67,20 @@ go test ./internal/api
go test ./cmd/felis go test ./cmd/felis
``` ```
The hermetic suites run against in-memory fakes; the business stores' SQL is
verified separately against a real Postgres, on a throwaway database whose name
must contain `pgint` (the harness drops and recreates its schema and replays the
embedded migrations):
```bash
FELIS_TEST_PG_URL='postgres://felis:***@127.0.0.1:5432/felis_pgint?sslmode=disable' \
go test -tags pgint ./internal/pgint/ -v
```
Run it after touching anything under `internal/api/pgrepo.go`, `internal/submit`,
or `internal/build` that speaks SQL: the fakes encode the contract, and this
suite exists to catch the drift between the fakes and the real queries.
Build the CLI: Build the CLI:
```bash ```bash
+5 -5
View File
@@ -17,10 +17,10 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
- **即开即玩**:玩家尝试连接时自动唤醒服务器,空闲后自动休眠,像游戏主机一样省资源。 - **即开即玩**:玩家尝试连接时自动唤醒服务器,空闲后自动休眠,像游戏主机一样省资源。
- **Web 控制面板**:浏览器中查看服务器状态、在线玩家与资源用量,管理备份与恢复。 - **Web 控制面板**:浏览器中查看服务器状态、在线玩家与资源用量,管理备份与恢复。
- **自动备份与恢复**:定时将世界打包存档,支持从任意备份点一键回滚。 - **备份与恢复**:一键把整服数据(世界、配置、插件/模组,即整个 /data 卷)打包进集群内的归档库,支持从任意备份点回滚;默认安装就已启用(归档 PVC 与路径由安装器一并生成)。
- **智慧回收**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间。 - **智慧回收(可选开启)**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间;安装时设置 `FELIS_WORLDS_HOST_PATH`(k3s 默认 `/var/lib/rancher/k3s/storage`)即启用每日回收,不设置则不删任何世界。
- **多核心支持**:兼容 Paper、Fabric、Forge、NeoForge,经由 Velocity 代理统一入口。 - **多核心支持**:兼容 Paper、Fabric、Forge、NeoForge,经由 Velocity 代理统一入口。
- **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建并部署。 - **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建;构建产物进入镜像白名单,可直接选用为服务器镜像完成部署。
- **Passkey 登录**:支持指纹、面容、硬件密钥等无密码认证方式。 - **Passkey 登录**:支持指纹、面容、硬件密钥等无密码认证方式。
- **零信任安全**:面板流量由 Cloudflare Access 保护,集群内 API 不暴露到公网。 - **零信任安全**:面板流量由 Cloudflare Access 保护,集群内 API 不暴露到公网。
@@ -29,7 +29,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
在准备好的 Linux 主机上执行: 在准备好的 Linux 主机上执行:
```bash ```bash
curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/deploy/bootstrap.sh | sudo bash curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
``` ```
脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。 脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。
@@ -41,7 +41,7 @@ curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/de
> export FELIS_GITHUB_TOKEN=<对本仓库有读权限的 token> > export FELIS_GITHUB_TOKEN=<对本仓库有读权限的 token>
> printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \ > printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \
> | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \ > | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \
> https://api.github.com/repos/MliroLirrorsIngenuity/Felis/contents/deploy/bootstrap.sh \ > https://api.github.com/repos/FelisMC/Felis/contents/deploy/bootstrap.sh \
> | sudo -E bash > | sudo -E bash
> ``` > ```
> >
+25 -4
View File
@@ -17,10 +17,10 @@ Table of Contents
- **Wake on Join**: Servers start automatically when a player connects, and stop when idle — like hibernate for your server. - **Wake on Join**: Servers start automatically when a player connects, and stop when idle — like hibernate for your server.
- **Web Dashboard**: Monitor server status, online players, and resource usage from your browser, with backup and restore management. - **Web Dashboard**: Monitor server status, online players, and resource usage from your browser, with backup and restore management.
- **Auto Backup & Restore**: Scheduled world backups with one-click rollback from any backup point. - **Backup & Restore**: One-click snapshots of a server's whole data volume (worlds, config, plugins/mods — the entire /data volume) into the cluster's archive store, with rollback from any backup point — enabled by default (the installer renders the archive PVC and its path).
- **World Reaper**: Worlds idle for more than 15 days are automatically backed up and removed to free disk space. - **World Reaper** (opt in): Worlds idle for more than 15 days are automatically backed up and removed to free disk space. Enable it by setting `FELIS_WORLDS_HOST_PATH` at install time (on k3s: `/var/lib/rancher/k3s/storage`); without it, no world is ever deleted.
- **Multi-core Support**: Compatible with Paper, Fabric, Forge, and NeoForge, federated behind a Velocity proxy. - **Multi-core Support**: Compatible with Paper, Fabric, Forge, and NeoForge, federated behind a Velocity proxy.
- **Modpack Submission**: Players submit custom modpacks; admin approval triggers automatic build and deployment. - **Modpack Submission**: Players submit custom modpacks; admin approval triggers an automatic build, and the result is whitelisted as a server image you can select to deploy.
- **Passkey Login**: Passwordless authentication via fingerprint, face recognition, or hardware security keys. - **Passkey Login**: Passwordless authentication via fingerprint, face recognition, or hardware security keys.
- **Zero Trust Security**: Panel traffic protected by Cloudflare Access; the internal API is never exposed to the internet. - **Zero Trust Security**: Panel traffic protected by Cloudflare Access; the internal API is never exposed to the internet.
@@ -29,11 +29,32 @@ Table of Contents
On a prepared Linux host, run: On a prepared Linux host, run:
```bash ```bash
curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/deploy/bootstrap.sh | sudo bash curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
``` ```
The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain. The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain.
> **This repository is currently private**, so the command above returns 404. Use the
> credentialed form instead; the installer itself needs the same token to resolve and
> download the release, so pass it through with `sudo -E`:
>
> ```bash
> export FELIS_GITHUB_TOKEN=<a token with read access to this repository>
> printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \
> | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \
> https://api.github.com/repos/FelisMC/Felis/contents/deploy/bootstrap.sh \
> | sudo -E bash
> ```
>
> The token reaches `curl --config -` over stdin instead of the command line: argv is
> readable by any local user via `/proc`, and that is exactly why the installer's
> internal `github_api` uses the same form.
Rerunning this command is also how you upgrade felis-api to a newer version (`felis setup`
cannot — it uses the binary already installed on the host). The rerun keeps the installed
root domain but **not** the channel: if this host follows main, also
`export FELIS_VERSION_BOOTSTRAP=dev`.
## Build from Source ## Build from Source
Felis is built with Go and Node.js: Felis is built with Go and Node.js:
+6 -5
View File
@@ -13,8 +13,8 @@ var bootstrapScript string
//go:embed deploy/crd/*.yaml //go:embed deploy/crd/*.yaml
var bootstrapAssets embed.FS var bootstrapAssets embed.FS
// gameStackAssets carries everything deploy/bootstrap.sh needs to build the two // gameStackAssets carries everything deploy/bootstrap.sh needs to build the three
// always-on game images (login limbo + lobby) and the Velocity plugin, for the TUI // game images (login limbo, lobby, plain Paper) and the Velocity plugin, for the TUI
// install path — which pipes the embedded bootstrap.sh into bash and therefore has // install path — which pipes the embedded bootstrap.sh into bash and therefore has
// NO source checkout on disk to build from. // NO source checkout on disk to build from.
// //
@@ -26,6 +26,7 @@ var bootstrapAssets embed.FS
// //
//go:embed deploy/limbo/Dockerfile deploy/limbo/entrypoint.sh //go:embed deploy/limbo/Dockerfile deploy/limbo/entrypoint.sh
//go:embed deploy/lobby/Dockerfile deploy/lobby/entrypoint.sh //go:embed deploy/lobby/Dockerfile deploy/lobby/entrypoint.sh
//go:embed deploy/paper/Dockerfile deploy/paper/entrypoint.sh
//go:embed plugins/limbo/build.gradle plugins/limbo/settings.gradle plugins/limbo/src //go:embed plugins/limbo/build.gradle plugins/limbo/settings.gradle plugins/limbo/src
//go:embed plugins/paper/build.gradle plugins/paper/settings.gradle plugins/paper/src //go:embed plugins/paper/build.gradle plugins/paper/settings.gradle plugins/paper/src
//go:embed plugins/velocity/build.gradle plugins/velocity/settings.gradle plugins/velocity/src //go:embed plugins/velocity/build.gradle plugins/velocity/settings.gradle plugins/velocity/src
@@ -46,9 +47,9 @@ func GameStackTar(w io.Writer) error {
if err != nil { if err != nil {
return err return err
} }
// Mode 0644 for everything: entrypoint.sh is invoked as `sh <file>` by both // Mode 0644 for everything: entrypoint.sh is invoked as `sh <file>` by all
// Dockerfiles precisely because the +x bit does not survive a Windows checkout, // three Dockerfiles precisely because the +x bit does not survive a Windows
// so nothing here needs to be executable. // checkout, so nothing here needs to be executable.
if err := tw.WriteHeader(&tar.Header{ if err := tw.WriteHeader(&tar.Header{
Name: path, Name: path,
Mode: 0o644, Mode: 0o644,
+101
View File
@@ -1,6 +1,8 @@
package felis package felis
import ( import (
"io/fs"
"regexp"
"strings" "strings"
"testing" "testing"
) )
@@ -64,6 +66,37 @@ func TestLobbyLuckPermsWiringIsConsistent(t *testing.T) {
} }
} }
// The Paper jar digest rides the same cross-file contract as LuckPerms above: bootstrap.sh
// resolves "url sha256" out of Fill's content-addressed download URL and passes the digest
// as a build-arg the Dockerfile must require and verify. docker only WARNS about an unknown
// --build-arg, so a renamed arg would surface as a required-arg failure on a real host
// mid-install — this test is the only compile step the pairing gets.
//
// Both images pull the same jar from the same URL, so both have to check it: a gate on one
// of them leaves the other booting on whatever bytes happened to arrive.
func TestPaperJarDigestWiringIsConsistent(t *testing.T) {
const arg = "PAPER_JAR_SHA256"
if n := strings.Count(BootstrapScript(), "--build-arg "+arg+"="); n < 2 {
t.Errorf("bootstrap.sh passes --build-arg %s %d time(s); the lobby and the "+
"plain-Paper build each need it", arg, n)
}
for _, name := range []string{"deploy/lobby/Dockerfile", "deploy/paper/Dockerfile"} {
dockerfile := readGameStackFile(t, name)
if !strings.Contains(dockerfile, "ARG "+arg) {
t.Errorf("%s declares no ARG %s", name, arg)
}
if !strings.Contains(dockerfile, `if [ -z "${PAPER_JAR_SHA256:-}" ]`) {
t.Errorf("%s does not fail the build when %s is unset", name, arg)
}
// Requiring the arg is not the same as spending it, and which file gets hashed
// matters as much as the command: a `sha256sum -c` over some other download
// would satisfy a bare substring check while paper.jar still arrives unchecked.
if !strings.Contains(dockerfile, `echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c`) {
t.Errorf("%s never verifies /paper/paper.jar against %s", name, arg)
}
}
}
// A 1.8 client joining a protocol-47 backend dies on the first chunk unless ViaVersion's // A 1.8 client joining a protocol-47 backend dies on the first chunk unless ViaVersion's
// serverside block-connection tracking is off: under modern forwarding the Velocity injector // serverside block-connection tracking is off: under modern forwarding the Velocity injector
// reports 1.13 as the lowest supported protocol, ConnectionData.init() returns early on that, // reports 1.13 as the lowest supported protocol, ConnectionData.init() returns early on that,
@@ -104,6 +137,74 @@ func TestBootstrapPinsViaBlockConnectionsOff(t *testing.T) {
} }
} }
// The embed list and the images bootstrap.sh builds are two lists nobody reconciles.
// deploy/paper shipped an image build without ever being added to gameStackAssets, and
// nothing said so: a checkout on disk satisfies the build either way, and the tar is
// only the build context on the path that has no checkout — `curl | bash`, where the
// third `docker build -f` then names a file that was never unpacked. So derive the
// inputs from the script and from each Dockerfile's own COPY lines instead of restating
// them here; a fourth image inherits the check for free.
func TestGameStackTarCarriesEveryBuildInput(t *testing.T) {
// Matches the path only when GAME_STACK_DIR is followed by one, which skips the
// build-context arguments (`"$GAME_STACK_DIR"`, `"${GAME_STACK_DIR}:/src:z"`) and
// the glob for gradle's output, none of which are inputs this tar has to carry.
found := regexp.MustCompile(`\$\{GAME_STACK_DIR\}/(\S+?)"`).FindAllStringSubmatch(BootstrapScript(), -1)
var paths []string
seen := map[string]bool{}
for _, m := range found {
if !seen[m[1]] {
seen[m[1]] = true
paths = append(paths, m[1])
}
}
// Guards the regex itself: a rewrite of how bootstrap.sh spells the build context
// would otherwise turn this test into an unconditional pass. It has to come before
// the loop — a missing file in there is fatal, and a floor placed after it would
// never be reached to say that the regex, not the tar, is what went wrong.
if len(paths) < 3 {
t.Fatalf("only %d game-stack path(s) resolved out of bootstrap.sh; the limbo, "+
"lobby and paper Dockerfiles are all built from ${GAME_STACK_DIR}", len(paths))
}
for _, path := range paths {
requireEmbedded(t, path)
// A Dockerfile that arrives without the files it COPYs fails just as late and
// just as far from here; the deploy/paper gap was missing its entrypoint too.
for _, src := range copySources(t, path) {
requireEmbedded(t, src)
}
}
}
// copySources lists the build-context paths a Dockerfile COPYs in, skipping the
// --from=<stage> copies, whose sources are produced by an earlier stage rather than
// unpacked from the tar.
func copySources(t *testing.T, dockerfile string) []string {
t.Helper()
var out []string
// Continuations are joined first: a COPY split across lines would otherwise be two
// fragments, neither of them starting with COPY followed by a source, and its
// source would slip past unchecked.
body := strings.ReplaceAll(readGameStackFile(t, dockerfile), "\\\n", " ")
for line := range strings.SplitSeq(body, "\n") {
f := strings.Fields(line)
if len(f) < 2 || f[0] != "COPY" || strings.HasPrefix(f[1], "--") {
continue
}
out = append(out, strings.TrimSuffix(f[1], "/"))
}
return out
}
// fs.Stat rather than ReadFile: half of these are directories (`COPY plugins/shared/`),
// and embed.FS answers for those too.
func requireEmbedded(t *testing.T, path string) {
t.Helper()
if _, err := fs.Stat(gameStackAssets, path); err != nil {
t.Errorf("%s is a game-stack build input but is not in gameStackAssets; an "+
"install with no source checkout dies on it: %v", path, err)
}
}
func readGameStackFile(t *testing.T, name string) string { func readGameStackFile(t *testing.T, name string) string {
t.Helper() t.Helper()
b, err := gameStackAssets.ReadFile(name) b, err := gameStackAssets.ReadFile(name)
+62 -23
View File
@@ -18,6 +18,7 @@ import (
"felis.lolicon.best/internal/config" "felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/fileedit" "felis.lolicon.best/internal/fileedit"
"felis.lolicon.best/internal/mail" "felis.lolicon.best/internal/mail"
"felis.lolicon.best/internal/naming"
"felis.lolicon.best/internal/panel" "felis.lolicon.best/internal/panel"
"felis.lolicon.best/internal/passkey" "felis.lolicon.best/internal/passkey"
"felis.lolicon.best/internal/platform" "felis.lolicon.best/internal/platform"
@@ -142,10 +143,14 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
// build namespace and pushes to the internal registry. The build Pod never // build namespace and pushes to the internal registry. The build Pod never
// holds DB credentials — felis-api owns the PG store and admits scanned // holds DB credentials — felis-api owns the PG store and admits scanned
// images, so the Builder is constructed here with both bindings. // images, so the Builder is constructed here with both bindings.
buildCfg := buildConfig(cfg)
// The fetch initContainer runs THIS image's fetch-context entrypoint, so the
// build config carries the api's own image (the platform sets FELIS_IMAGE).
buildCfg.FelisImage = os.Getenv("FELIS_IMAGE")
builder := &build.Builder{ builder := &build.Builder{
Store: build.NewPGStore(drv.DB()), Store: build.NewPGStore(drv.DB()),
Jobs: build.NewK8sJobs(cl, buildConfig(cfg)), Jobs: build.NewK8sJobs(cl, buildCfg),
Config: buildConfig(cfg), Config: buildCfg,
} }
// User-modpack approval lane (user-directed extension over §16; see // User-modpack approval lane (user-directed extension over §16; see
@@ -159,14 +164,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
// The blob upload transport is selected by the shape of user_uploads_context — // The blob upload transport is selected by the shape of user_uploads_context —
// the two backends the setup wizard chooses between. A local path wires // the two backends the setup wizard chooses between. A local path wires
// LocalContextStore (the mounted uploads PVC); an s3:// base wires // LocalContextStore (the mounted uploads PVC); an s3:// base wires
// S3ContextStore when its credentials resolve. Either way the store's target is // S3ContextStore when its credentials resolve. Anything else — or an s3:// base
// derived from the SAME config field the context ref uses, so the blob lands // with no credentials configured — leaves Blobs nil so POST
// exactly where Kaniko's --context points. Anything else — or an s3:// base with
// no credentials configured — leaves Blobs nil so POST
// /me/submissions/{id}/context returns 503, honest like the restore executor // /me/submissions/{id}/context returns 503, honest like the restore executor
// when its PVC is not supplied. (Letting the sandboxed Kaniko build Pod READ the // when its PVC is not supplied.
// context — PVC mount for local, creds+egress for S3 — is a separate deployment //
// integration.) // Reading the blob back is the API's job, not Kaniko's: the build Pod runs in
// another namespace and can neither mount the uploads PVC (a PVC does not cross
// namespaces) nor hold object-store credentials, so ContextBaseURL makes the
// derived context ref an internal-face URL that the build Job's fetch
// initContainer streams (cmd/felis fetch-context). The platform renders this
// address into the api Deployment (felis API base URL env); the fallback keeps
// a hand-rolled deployment working under the platform's default control
// namespace.
contextBase := cfg.Registry.UserUploadsContext contextBase := cfg.Registry.UserUploadsContext
var blobs submit.Blobs var blobs submit.Blobs
switch { switch {
@@ -186,11 +196,12 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
fmt.Fprintf(stderr, "felis api: user-uploads context %q is neither a local path nor an s3:// base — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", contextBase) fmt.Fprintf(stderr, "felis api: user-uploads context %q is neither a local path nor an s3:// base — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", contextBase)
} }
submissions := &submit.Manager{ submissions := &submit.Manager{
Store: submit.NewPGStore(drv.DB()), Store: submit.NewPGStore(drv.DB()),
Builds: builder, Builds: builder,
Registry: cfg.Registry.URL, Registry: cfg.Registry.URL,
ContextStore: contextBase, ContextStore: contextBase,
Blobs: blobs, ContextBaseURL: internalAPIBaseURL(),
Blobs: blobs,
} }
// Restore subsystem (spec §7): the weak-SA restore Job mounts the target // Restore subsystem (spec §7): the weak-SA restore Job mounts the target
@@ -259,6 +270,7 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
Builder: builder, Builder: builder,
Restorer: restorer, Restorer: restorer,
Backuper: backuper, Backuper: backuper,
JobStatus: api.NewK8sJobStatus(cl, cfg.K8s.Namespace),
Files: files, Files: files,
Submissions: submissions, Submissions: submissions,
Mailer: mailer, Mailer: mailer,
@@ -279,6 +291,11 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
AdminHostname: cfg.Auth.AdminHostname, AdminHostname: cfg.Auth.AdminHostname,
PanelHostname: cfg.Auth.PanelHostname, PanelHostname: cfg.Auth.PanelHostname,
WakeCooldown: 30 * time.Second, WakeCooldown: 30 * time.Second,
// The user-modpack lane's per-user throttles: a create spaces out
// review-queue rows, an upload spaces out (up to 1 GiB) context streams.
// Separate keys, so the normal create→upload sequence stays immediate.
SubmitCreateCooldown: 30 * time.Second,
SubmitUploadCooldown: 15 * time.Second,
// Bound concurrent console/build-log SSE streams per principal. Generous enough // Bound concurrent console/build-log SSE streams per principal. Generous enough
// for legitimate multi-tab / multi-server watching, while capping how many // for legitimate multi-tab / multi-server watching, while capping how many
// upstream follow connections a single caller can tie up if their streams stall. // upstream follow connections a single caller can tie up if their streams stall.
@@ -286,15 +303,14 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
} }
fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)") fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)")
// Felis-nano: wire the multi-source hasJoined multiplexer only when third-party auth // Felis-nano: the multi-source hasJoined multiplexer. Mojang leads as the code-owned
// sources are configured. Mojang leads as the code-owned identity anchor (正版优先); // identity anchor (正版优先); config can only append namespace-rewritten third-party
// config can only append namespace-rewritten third-party sources, never a trusted one, // sources, never a trusted one, so a misconfig cannot reopen the impersonation hole.
// so a misconfig cannot reopen the impersonation hole. No sources = a.AuthSources stays // Wired unconditionally: the installer points Velocity at this route whether or not any
// nil = the endpoint 204s every login (ships off). // [[auth_source]] is configured, so an empty list has to mean a Mojang-only relay, the
if len(cfg.AuthSources) > 0 { // same as under `felis nano`. A nil list would 204 every login, premium ones included.
a.AuthSources = authSourcesFromConfig(cfg.AuthSources) a.AuthSources = authSourcesFromConfig(cfg.AuthSources)
fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources)) fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources))
}
// Passkey (WebAuthn) verifier (spec §14, Phase 6). One relying party spans BOTH // Passkey (WebAuthn) verifier (spec §14, Phase 6). One relying party spans BOTH
// web faces: the RP id is the panel hostname (console.<root>), and because that is // web faces: the RP id is the panel hostname (console.<root>), and because that is
@@ -402,9 +418,32 @@ func buildConfig(cfg *config.Config) build.Config {
return build.Config{ return build.Config{
Namespace: cfg.Registry.BuildNamespace, Namespace: cfg.Registry.BuildNamespace,
RegistryURL: cfg.Registry.URL, RegistryURL: cfg.Registry.URL,
// Empty overrides fall back to the build package's defaults, so an
// install that has not imported kaniko/trivy keeps the compiled-in refs
// (and fails loudly on pull rather than silently building with the wrong
// image).
KanikoImage: cfg.Registry.KanikoImage,
TrivyImage: cfg.Registry.TrivyImage,
CPULimit: cfg.Registry.BuildCPULimit,
MemLimit: cfg.Registry.BuildMemLimit,
// Empty keeps Trivy's own default; an install with builds points this at
// the internal DB mirror (see config.RegistryConfig.TrivyDBRepository).
TrivyDBRepository: cfg.Registry.TrivyDBRepository,
TrivyJavaDBRepository: cfg.Registry.TrivyJavaDBRepository,
} }
} }
// internalAPIBaseURL resolves the platform's internal-face base URL: the address
// the platform rendered into this pod (felis API base URL env), or — for a
// hand-rolled deployment that set none — the platform default control namespace,
// the same fallback setup.go uses to hand the login gate its address.
func internalAPIBaseURL() string {
if base := os.Getenv(naming.EnvAPIBaseURL); base != "" {
return base
}
return platform.InternalAPIBaseURL(platform.DefaultControlNamespace)
}
// uploadsSchemeRE matches a leading URL scheme like "s3://" or "gs://". // uploadsSchemeRE matches a leading URL scheme like "s3://" or "gs://".
var uploadsSchemeRE = regexp.MustCompile(`^[a-zA-Z][a-zA-Z0-9+.-]*://`) var uploadsSchemeRE = regexp.MustCompile(`^[a-zA-Z][a-zA-Z0-9+.-]*://`)
+65
View File
@@ -3,8 +3,73 @@ package main
import ( import (
"net/http" "net/http"
"testing" "testing"
"felis.lolicon.best/internal/config"
) )
// TestAuthSourcesFromConfig pins the one place the hasJoined identity anchor is decided:
// Mojang is prepended in code, first, and is the only source whose UUIDs are trusted as-is.
// The empty case matters on its own — both `felis api` and `felis nano` call this with a
// config that has no [[auth_source]] at all, and that has to be a Mojang-only relay rather
// than an empty list that rejects every login.
func TestAuthSourcesFromConfig(t *testing.T) {
for _, tc := range []struct {
name string
configured []config.AuthSourceConfig
}{
{"no configured sources", nil},
{"configured sources", []config.AuthSourceConfig{
{Tag: "littleskin", Prefix: "LS", URL: "https://littleskin.example/hasJoined"},
{Tag: "guild", Prefix: "GD", URL: "https://guild.example/hasJoined"},
}},
} {
t.Run(tc.name, func(t *testing.T) {
got := authSourcesFromConfig(tc.configured)
if len(got) != len(tc.configured)+1 {
t.Fatalf("got %d sources, want Mojang + %d configured", len(got), len(tc.configured))
}
if got[0].Tag != "mojang" || got[0].URL != mojangSessionServer || !got[0].Identity {
t.Errorf("first source = %+v, want the Mojang identity anchor", got[0])
}
for i, c := range tc.configured {
s := got[i+1]
if s.Identity {
t.Errorf("configured source %q is marked Identity; only Mojang may be", c.Tag)
}
if s.Tag != c.Tag || s.Prefix != c.Prefix || s.URL != c.URL {
t.Errorf("source %d = %+v, want %+v in config order", i+1, s, c)
}
}
})
}
}
// TestBuildConfig_ProjectsOverrides pins the [registry] overrides reaching the
// build subsystem: unset fields must stay EMPTY (the build package's compiled-in
// defaults apply there, not here), and set fields must pass through verbatim —
// an air-gapped install points these at its imported mirrors.
func TestBuildConfig_ProjectsOverrides(t *testing.T) {
empty := buildConfig(&config.Config{})
if empty.KanikoImage != "" || empty.TrivyImage != "" || empty.CPULimit != "" || empty.MemLimit != "" {
t.Errorf("empty registry config must project empty overrides (defaults live in internal/build), got %+v", empty)
}
full := buildConfig(&config.Config{Registry: config.RegistryConfig{
URL: "registry.felis.svc:5000",
BuildNamespace: "felis-build",
KanikoImage: "reg/kaniko:v1",
TrivyImage: "reg/trivy:v1",
BuildCPULimit: "1",
BuildMemLimit: "2Gi",
}})
if full.KanikoImage != "reg/kaniko:v1" || full.TrivyImage != "reg/trivy:v1" ||
full.CPULimit != "1" || full.MemLimit != "2Gi" {
t.Errorf("registry overrides did not reach build.Config: %+v", full)
}
if full.Namespace != "felis-build" || full.RegistryURL != "registry.felis.svc:5000" {
t.Errorf("namespace/registry url must keep projecting: %+v", full)
}
}
// TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every // TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every
// felis-api listener: the shared factory must bound the header and idle phases // felis-api listener: the shared factory must bound the header and idle phases
// (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because // (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because
+1 -1
View File
@@ -220,7 +220,7 @@ func buildMinecraftServerFromApplyRequest(req applyRequest, namespace string) (*
} }
memLim, ok := limits[corev1.ResourceMemory] memLim, ok := limits[corev1.ResourceMemory]
if !ok || memLim.IsZero() { if !ok || memLim.IsZero() {
return nil, fmt.Errorf("internal error: refusing to create a server without a memory ceiling (§22)") return nil, fmt.Errorf("internal error: refusing to create a server without a memory ceiling")
} }
// ---- storage ---- // ---- storage ----
+1 -1
View File
@@ -116,7 +116,7 @@ func newBackupID() string {
var b [16]byte var b [16]byte
if _, err := rand.Read(b[:]); err != nil { if _, err := rand.Read(b[:]); err != nil {
// crypto/rand failure is fatal and unrecoverable; a time-based fallback would // crypto/rand failure is fatal and unrecoverable; a time-based fallback would
// be a weaker ID for no benefit. ponytail: panic is the honest failure here. // be a weaker ID for no benefit. A panic is the honest failure here.
panic("felis backup: crypto/rand: " + err.Error()) panic("felis backup: crypto/rand: " + err.Error())
} }
return "bk-" + hex.EncodeToString(b[:]) return "bk-" + hex.EncodeToString(b[:])
+34 -10
View File
@@ -51,7 +51,7 @@ func resolveInternalAPI(ctx context.Context, cl client.Client, controlNamespace
} }
token = string(sec.Data[naming.ServiceTokenSecretKey]) token = string(sec.Data[naming.ServiceTokenSecretKey])
if token == "" { if token == "" {
return "", "", fmt.Errorf("Secret %s has no %s key", naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey) return "", "", fmt.Errorf("secret %s has no %s key", naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey)
} }
return fmt.Sprintf("http://%s:%d", ip, platform.APIInternalPort), token, nil return fmt.Sprintf("http://%s:%d", ip, platform.APIInternalPort), token, nil
@@ -86,23 +86,31 @@ func requestBackup(ctx context.Context, hc *http.Client, baseURL, token, name, o
// backupErrorFromResponse turns a non-202 into a human message. The well-known codes get // backupErrorFromResponse turns a non-202 into a human message. The well-known codes get
// an operator-facing explanation; anything else falls back to the API's // an operator-facing explanation; anything else falls back to the API's
// {"error":{message}} body, then the bare status code. // {"error":{code,message}} body, then the bare status code.
func backupErrorFromResponse(resp *http.Response) error { func backupErrorFromResponse(resp *http.Response) error {
switch resp.StatusCode {
case http.StatusConflict: // not_stopped
return fmt.Errorf("the server must be stopped before its world can be backed up — halt it first")
case http.StatusServiceUnavailable: // backup_unavailable
return fmt.Errorf("the backup subsystem is not configured on felis-api (FELIS_IMAGE / FELIS_BACKUP_PVC unset)")
case http.StatusNotFound:
return fmt.Errorf("no such server")
}
var e struct { var e struct {
Error struct { Error struct {
Code string `json:"code"`
Message string `json:"message"` Message string `json:"message"`
} `json:"error"` } `json:"error"`
} }
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<16)) raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<16))
_ = json.Unmarshal(raw, &e) _ = json.Unmarshal(raw, &e)
switch resp.StatusCode {
case http.StatusConflict:
// Two refusals share 409: the stopped gate and the missing-world-volume
// gate. The body's code distinguishes them; a code-less body reads as the
// stopped gate (the only 409 before the volume gate existed), and any other
// coded 409 falls through to the API's own operator text.
if e.Error.Code == "" || e.Error.Code == "not_stopped" {
return fmt.Errorf("the server must be stopped before its world can be backed up — halt it first")
}
case http.StatusServiceUnavailable: // backup_unavailable
return fmt.Errorf("the backup subsystem is not configured on felis-api (FELIS_IMAGE / FELIS_BACKUP_PVC unset)")
case http.StatusNotFound:
return fmt.Errorf("no such server")
}
if e.Error.Message != "" { if e.Error.Message != "" {
return fmt.Errorf("felis-api: %s", e.Error.Message) return fmt.Errorf("felis-api: %s", e.Error.Message)
} }
@@ -120,3 +128,19 @@ func performBackupNow(ctx context.Context, cl client.Client, controlNamespace, n
hc := &http.Client{Timeout: 10 * time.Second} hc := &http.Client{Timeout: 10 * time.Second}
return requestBackup(ctx, hc, baseURL, token, name, osUser) return requestBackup(ctx, hc, baseURL, token, name, osUser)
} }
// backupPickable narrows the backup picker to servers the backup API can accept.
// System servers (login/lobby) are excluded: they have no row in the servers
// table and carry reserved names, so every attempt dies in name validation —
// offering them would be a dead pick. The halt picker keeps them on purpose
// (break-glass retains full power over system servers); only the API-backed
// backup op cannot reach them.
func backupPickable(servers []haltableServer) []haltableServer {
out := make([]haltableServer, 0, len(servers))
for _, s := range servers {
if !s.system {
out = append(out, s)
}
}
return out
}
+32 -3
View File
@@ -114,16 +114,23 @@ func TestRequestBackup(t *testing.T) {
cases := []struct { cases := []struct {
name string name string
code int code int
body string // optional JSON error body
expect string expect string
}{ }{
{"409 not_stopped", http.StatusConflict, "must be stopped"}, {"409 not_stopped", http.StatusConflict, "", "must be stopped"},
{"503 backup_unavailable", http.StatusServiceUnavailable, "not configured"}, {"409 no_world_volume surfaces the API's own text", http.StatusConflict,
{"404 not found", http.StatusNotFound, "no such server"}, `{"error":{"code":"no_world_volume","message":"this server has no world volume yet — start it once to create it, then retry"}}`,
"no world volume yet"},
{"503 backup_unavailable", http.StatusServiceUnavailable, "", "not configured"},
{"404 not found", http.StatusNotFound, "", "no such server"},
} }
for _, tc := range cases { for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) { t.Run(tc.name, func(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(tc.code) w.WriteHeader(tc.code)
if tc.body != "" {
_, _ = io.WriteString(w, tc.body)
}
})) }))
defer srv.Close() defer srv.Close()
_, err := requestBackup(context.Background(), hc, srv.URL, "tok", "survival", "alice") _, err := requestBackup(context.Background(), hc, srv.URL, "tok", "survival", "alice")
@@ -142,3 +149,25 @@ func TestRequestBackup(t *testing.T) {
} }
}) })
} }
// The Sync picker must not offer system servers: the backup API validates names
// and resolves a servers-table row, so a lobby/login pick can only die in
// validation — a dead choice in an emergency console.
func TestBackupPickable(t *testing.T) {
got := backupPickable([]haltableServer{
{name: "lobby", phase: "Running", system: true},
{name: "login", phase: "Running", system: true},
{name: "test-one", phase: "Stopped"},
})
if len(got) != 1 || got[0].name != "test-one" || got[0].system {
t.Fatalf("backupPickable = %+v, want only the user server", got)
}
// Survivors keep their input order (the picker's cursor math depends on it).
got = backupPickable([]haltableServer{
{name: "alpha"}, {name: "login", system: true}, {name: "beta"},
})
if len(got) != 2 || got[0].name != "alpha" || got[1].name != "beta" {
t.Fatalf("backupPickable order = %+v, want [alpha beta]", got)
}
}
+47 -18
View File
@@ -76,12 +76,19 @@ type ownerStore interface {
AdminExists(ctx context.Context) (bool, error) AdminExists(ctx context.Context) (bool, error)
// UserByUsername loads a staff login projection. // UserByUsername loads a staff login projection.
UserByUsername(ctx context.Context, username string) (*api.StaffUser, error) UserByUsername(ctx context.Context, username string) (*api.StaffUser, error)
// OwnerUsername names the single active Owner seat, or "" when none exists.
// provisionOwner refuses to re-target anything but this username: with the
// seat occupied, a fresh name would mint a second owner row (UpsertOwner's
// insert arm) while the existing — possibly compromised — seat stays live,
// and no supported path can delete an owner row afterwards.
OwnerUsername(ctx context.Context) (string, error)
UpsertOwner(ctx context.Context, id, username, email string) error UpsertOwner(ctx context.Context, id, username, email string) error
// InsertOperator mints a NEW Operator staff account. Unlike UpsertOwner it is // InsertOperator mints a NEW Operator staff account. Unlike UpsertOwner it is
// insert-only: a username already taken is a conflict (api.ErrConflict), never a // insert-only: a username already taken is a conflict (api.ErrConflict), never a
// silent reset, so adding an Operator can never clobber the Owner or an existing // silent reset, so adding an Operator can never clobber the Owner or an existing
// Operator. The row is role=admin, identical in shape to the Owner — Felis has no // Operator. The row is role=admin — an Operator is staff BELOW the single
// separate operator DB role (migration 0003: staff = role=admin). // role=owner identity (migration 0011 adds that role); the two are the only
// staff roles.
InsertOperator(ctx context.Context, id, username, email string) error InsertOperator(ctx context.Context, id, username, email string) error
// CompleteOwnerSetup atomically consumes the in-game link code, creates or // CompleteOwnerSetup atomically consumes the in-game link code, creates or
// promotes the bound Owner, enables local auth, and stores the one-time setup // promotes the bound Owner, enables local auth, and stores the one-time setup
@@ -266,20 +273,31 @@ func authenticateAdmin(ctx context.Context, s ownerStore, username string) (matc
if err != nil { if err != nil {
return "", false, err return "", false, err
} }
if u.Role != "admin" { // Staff means admin OR owner: recovery attribution must accept the Owner (the
// primary break-glass identity), not just plain admins.
if u.Role != "admin" && u.Role != "owner" {
return "", false, nil return "", false, nil
} }
return u.Username, true, nil return u.Username, true, nil
} }
// provisionOwner mints or resets the single Owner account direct-to-Postgres, // provisionOwner mints or resets the single Owner account direct-to-Postgres,
// passwordless. The account is role=admin with no password — the Owner completes // passwordless. The account is role=owner with no password — the Owner completes
// passwordless login setup via the web setup-token flow after `felis setup`. // passwordless login setup via the web setup-token flow after `felis setup`.
// With a seat already occupied the reset must name that seat (ownerSeatTakenError
// otherwise): the upsert's insert arm would silently mint a SECOND owner, and
// every owner row is undeletable through the panel, so the tier could never
// converge back to one.
func provisionOwner(ctx context.Context, s ownerStore, username, email string) error { func provisionOwner(ctx context.Context, s ownerStore, username, email string) error {
username = strings.TrimSpace(username) username = strings.TrimSpace(username)
if username == "" { if username == "" {
return errors.New("owner username is required") return errors.New("owner username is required")
} }
if seat, err := s.OwnerUsername(ctx); err != nil {
return fmt.Errorf("check the owner seat: %w", err)
} else if seat != "" && seat != username {
return &ownerSeatTakenError{seat: seat}
}
id := newOwnerID() id := newOwnerID()
if id == "" { if id == "" {
return errors.New("generate owner id: entropy source failed") return errors.New("generate owner id: entropy source failed")
@@ -290,13 +308,26 @@ func provisionOwner(ctx context.Context, s ownerStore, username, email string) e
return nil return nil
} }
// provisionOperator mints a NEW Operator staff account direct-to-Postgres. Like the // ownerSeatTakenError refuses an Owner reset that names anything but the
// Owner it is role=admin and passwordless — Felis has no separate operator DB role, // occupied seat, naming it so the operator can retype. Is reports
// so an Operator is simply an additional staff admin (migration 0003). UNLIKE // api.ErrConflict so the TUI's recoverable-error branch (shared with the
// provisionOwner, which upserts the single Owner and resets it on a username // operator path's taken-name clash) routes back to the form instead of ending
// conflict, this is insert-only: a username already taken returns api.ErrConflict // the console.
// rather than overwriting a live account, so adding an Operator can never silently type ownerSeatTakenError struct{ seat string }
// clobber the Owner's or another Operator's account.
func (e *ownerSeatTakenError) Error() string {
return fmt.Sprintf("an Owner already exists as %q — enter that username to reset the Owner", e.seat)
}
func (e *ownerSeatTakenError) Is(target error) bool { return target == api.ErrConflict }
// provisionOperator mints a NEW Operator staff account direct-to-Postgres. It is
// role=admin and passwordless — an additional staff admin below the single
// role=owner identity (migrations 0003 + 0011). UNLIKE provisionOwner, which
// upserts the single Owner and resets it on a username conflict, this is
// insert-only: a username already taken returns api.ErrConflict rather than
// overwriting a live account, so adding an Operator can never silently clobber
// the Owner's or another Operator's account.
func provisionOperator(ctx context.Context, s ownerStore, username, email string) error { func provisionOperator(ctx context.Context, s ownerStore, username, email string) error {
username = strings.TrimSpace(username) username = strings.TrimSpace(username)
if username == "" { if username == "" {
@@ -387,7 +418,7 @@ func newSetupToken() (raw, hash string, err error) {
// performSetupMCBind is the `felis setup` Owner-establishment path: the operator // performSetupMCBind is the `felis setup` Owner-establishment path: the operator
// binds their Minecraft account via a one-time link code the login gate handed // binds their Minecraft account via a one-time link code the login gate handed
// them in-game, the bound user is promoted to role='admin' (passwordless Owner), // them in-game, the bound user is promoted to role='owner' (passwordless Owner),
// local auth is enabled, and a one-time setup URL is minted for the first web // local auth is enabled, and a one-time setup URL is minted for the first web
// login where the Owner verifies email / enrolls a passkey. adminHostname is the // login where the Owner verifies email / enrolls a passkey. adminHostname is the
// operator-console host the URL points at (op.console.<root>): the Owner is staff, // operator-console host the URL points at (op.console.<root>): the Owner is staff,
@@ -571,12 +602,10 @@ type breakGlassResult struct {
backupStatus string backupStatus string
// Cloudflare-specific edge detail (set only when connectMethod is Cloudflare) // Cloudflare-specific edge detail (set only when connectMethod is Cloudflare)
edgeConfigured bool edgeConfigured bool
edgeAud string edgeAud string
edgeRoutedHosts []string edgeRoutedHosts []string
edgeConfigPath string edgeConfigPath string
edgePanelHostname string
edgeAdminHostname string
} }
type consoleMode string type consoleMode string
+57 -8
View File
@@ -19,14 +19,15 @@ import (
// terminal. The design is passwordless: accounts carry no credential, and the // terminal. The design is passwordless: accounts carry no credential, and the
// Owner completes first-login through the setup-token web flow. // Owner completes first-login through the setup-token web flow.
type fakeOwnerStore struct { type fakeOwnerStore struct {
upserts []upsertCall upserts []upsertCall
inserts []upsertCall inserts []upsertCall
settings map[string][]byte settings map[string][]byte
audits []api.AuditEntry audits []api.AuditEntry
tokens []setupTokenCall tokens []setupTokenCall
redeems []redeemCall redeems []redeemCall
users map[string]*api.StaffUser // keyed by username users map[string]*api.StaffUser // keyed by username
admins bool // AdminExists answer admins bool // AdminExists answer
ownerSeat string // OwnerUsername answer: the occupied seat, "" when none
// CompleteOwnerSetup's success result. redeemUserID defaults to the fresh id // CompleteOwnerSetup's success result. redeemUserID defaults to the fresh id
// the caller passes (the unlinked-UUID case) when left empty. // the caller passes (the unlinked-UUID case) when left empty.
@@ -40,6 +41,7 @@ type fakeOwnerStore struct {
auditErr error auditErr error
userErr error // non-not-found error from UserByUsername userErr error // non-not-found error from UserByUsername
adminErr error adminErr error
seatErr error
redeemErr error redeemErr error
createTokenErr error createTokenErr error
} }
@@ -80,6 +82,15 @@ func (f *fakeOwnerStore) UserByUsername(_ context.Context, username string) (*ap
return nil, api.ErrNotFound return nil, api.ErrNotFound
} }
// OwnerUsername reports the single active Owner seat. Tests set ownerSeat; the
// zero value models a fresh install where bootstrap is free to mint.
func (f *fakeOwnerStore) OwnerUsername(_ context.Context) (string, error) {
if f.seatErr != nil {
return "", f.seatErr
}
return f.ownerSeat, nil
}
func (f *fakeOwnerStore) UpsertOwner(_ context.Context, id, username, email string) error { func (f *fakeOwnerStore) UpsertOwner(_ context.Context, id, username, email string) error {
if f.upsertErr != nil { if f.upsertErr != nil {
return f.upsertErr return f.upsertErr
@@ -201,6 +212,34 @@ func TestProvisionOwner(t *testing.T) {
} }
}) })
t.Run("an occupied seat refuses any other username", func(t *testing.T) {
// The seat is the single owner row: upserting a fresh name would take the
// insert arm and mint a SECOND owner, while the existing seat — possibly the
// compromised account this reset was meant to replace — stays live, and no
// supported path can delete an owner row.
f := &fakeOwnerStore{ownerSeat: "seat-holder"}
err := provisionOwner(ctx, f, "someone-else", "")
if !errors.Is(err, api.ErrConflict) {
t.Fatalf("error = %v, want it to wrap api.ErrConflict so the TUI routes back to the form", err)
}
if !strings.Contains(err.Error(), `"seat-holder"`) {
t.Errorf("error = %q, want it to name the occupied seat", err)
}
if len(f.upserts) != 0 {
t.Errorf("want no write against an occupied seat, got %d", len(f.upserts))
}
})
t.Run("the occupied seat's own username still resets", func(t *testing.T) {
f := &fakeOwnerStore{ownerSeat: "seat-holder"}
if err := provisionOwner(ctx, f, "seat-holder", "[email protected]"); err != nil {
t.Fatalf("provisionOwner(reset): %v", err)
}
if len(f.upserts) != 1 || f.upserts[0].username != "seat-holder" || f.upserts[0].email != "[email protected]" {
t.Fatalf("want 1 reset upsert for the seat, got %+v", f.upserts)
}
})
t.Run("propagates a store error", func(t *testing.T) { t.Run("propagates a store error", func(t *testing.T) {
f := &fakeOwnerStore{upsertErr: errors.New("boom")} f := &fakeOwnerStore{upsertErr: errors.New("boom")}
if err := provisionOwner(ctx, f, "owner", ""); err == nil { if err := provisionOwner(ctx, f, "owner", ""); err == nil {
@@ -264,6 +303,16 @@ func TestAuthenticateAdmin(t *testing.T) {
} }
}) })
t.Run("the owner role attributes like an admin", func(t *testing.T) {
owner := mkAdmin("root")
owner.Role = "owner" // the platform owner is staff too (migration 0011)
f := &fakeOwnerStore{users: map[string]*api.StaffUser{"root": owner}}
matched, ok, err := authenticateAdmin(ctx, f, "root")
if err != nil || !ok || matched != "root" {
t.Fatalf("authenticateAdmin(owner) = (%q, %v, %v), want (root, true, nil)", matched, ok, err)
}
})
t.Run("an unknown user is a non-match, not an error", func(t *testing.T) { t.Run("an unknown user is a non-match, not an error", func(t *testing.T) {
f := &fakeOwnerStore{} f := &fakeOwnerStore{}
_, ok, err := authenticateAdmin(ctx, f, "nobody") _, ok, err := authenticateAdmin(ctx, f, "nobody")
+73
View File
@@ -0,0 +1,73 @@
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"os"
"strings"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/platform"
)
// cmdConverge is the explicit convergence pass over already-installed system
// servers (#1). Provisioning is create-if-absent, so a field the desired spec
// gained after an install (spec.rcon, spec.startup.healthHTTPPort, a derived env
// key) never reaches the existing CR — and nothing says so. This command fills
// exactly those zero-value fields; see convergeSystemServers for the full contract
// and why it is a separate, operator-timed step rather than part of setup.
//
// It reads the same host config as setup (the control plane's felis.toml) and
// talks to the cluster with the local kubeconfig, so it must run as root on the
// control-plane host.
func cmdConverge(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("converge", flag.ContinueOnError)
fs.SetOutput(stderr)
cfgPath := fs.String("config", defaultSetupConfigPath, "path to felis.toml")
if err := fs.Parse(args); err != nil {
if errors.Is(err, flag.ErrHelp) {
return 0
}
return 2
}
if os.Geteuid() != 0 {
fmt.Fprintln(stderr, "felis converge: refused — converging needs the cluster credentials, so it must run as root (try: sudo felis converge)")
return 1
}
cfg, err := config.Load(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis converge: %v\n", err)
fmt.Fprintln(stderr, "If this host was never installed, run `sudo felis setup` first.")
return 1
}
cl, err := buildSystemServerClient()
if err != nil {
fmt.Fprintf(stderr, "felis converge: %v\n", err)
return 1
}
controlNS := platform.DefaultControlNamespace
outcomes := convergeSystemServers(context.Background(), cl, cfg.K8s.Namespace,
cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage,
platform.InternalAPIBaseURL(controlNS), cfg.Server.RootDomain,
defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname))
fmt.Fprintln(stdout, "felis converge: filling fields an installed system server predates (operator-set values are never overwritten):")
exit := 0
for _, o := range outcomes {
switch {
case o.err != nil:
fmt.Fprintf(stdout, " - %s: ERROR %v\n", o.name, o.err)
exit = 1
case len(o.changes) > 0:
fmt.Fprintf(stdout, " - %s: updated (%s)\n", o.name, strings.Join(o.changes, ", "))
default:
fmt.Fprintf(stdout, " - %s: %s\n", o.name, o.skipped)
}
}
return exit
}
+185
View File
@@ -0,0 +1,185 @@
package main
import (
"context"
"slices"
"strings"
"testing"
"felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/naming"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/client/fake"
)
// converge is the explicit pass over an installed system server whose CR predates
// a field the desired spec has since gained (#1). It must fill exactly the
// zero-valued whitelist fields and the derived env, and must not touch anything a
// non-zero value already occupies — that is the operator's.
func TestConvergeSystemServersFillsPredatedFields(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
// An old install: the lobby CR was created before the desired spec began
// rendering spec.rcon, and the login CR before the HTTP readiness gate existed.
// One derived env key is absent entirely (as if it were added later), and one
// hand-added env var plus a non-whitelisted spec field must survive.
lobby, err := lobbySystemServer("reg/lobby:1", "minecraft")
if err != nil {
t.Fatalf("build lobby: %v", err)
}
lobby.Spec.Rcon = v1alpha1.RconSpec{}
lobby.Spec.JavaMemory = "999Mi"
login, err := loginSystemServer("reg/limbo:1", "minecraft",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
if err != nil {
t.Fatalf("build login: %v", err)
}
login.Spec.Startup.HealthHTTPPort = 0
kept := login.Spec.Env
login.Spec.Env = nil
for _, e := range kept {
if e.Name != envPanelHostname {
login.Spec.Env = append(login.Spec.Env, e)
}
}
login.Spec.Env = append(login.Spec.Env, v1alpha1.EnvVar{Name: "OPERATOR_TUNING", Value: "keep-me"})
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(lobby, login).Build()
outcomes := convergeSystemServers(ctx, cl, "minecraft", "reg/limbo:1", "reg/lobby:1",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
byName := map[string]systemServerOutcome{}
for _, o := range outcomes {
if o.err != nil {
t.Fatalf("%s: unexpected error: %v", o.name, o.err)
}
byName[o.name] = o
}
lobbyOut := byName[naming.SystemLobbyServer]
if len(lobbyOut.changes) != 1 || lobbyOut.changes[0] != "spec.rcon" {
t.Errorf("lobby changes = %v, want [spec.rcon] (only the zero-valued field)", lobbyOut.changes)
}
loginOut := byName[naming.SystemLoginServer]
if !slices.Contains(loginOut.changes, "spec.startup.healthHTTPPort") || !slices.Contains(loginOut.changes, "env "+envPanelHostname) {
t.Errorf("login changes = %v, want the health port plus the missing derived env key", loginOut.changes)
}
var gotLobby v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLobbyServer}, &gotLobby); err != nil {
t.Fatalf("get lobby: %v", err)
}
if !gotLobby.Spec.Rcon.Enabled ||
gotLobby.Spec.Rcon.SecretRef.Name != naming.RconSecretName(naming.SystemLobbyServer) ||
gotLobby.Spec.Rcon.SecretRef.Key != naming.RconSecretKey {
t.Errorf("lobby rcon = %+v, want the desired block with the %s secret",
gotLobby.Spec.Rcon, naming.RconSecretName(naming.SystemLobbyServer))
}
if gotLobby.Spec.JavaMemory != "999Mi" {
t.Errorf("lobby javaMemory = %q, want 999Mi — converge fills new fields, it does not rewrite the spec", gotLobby.Spec.JavaMemory)
}
var gotLogin v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLoginServer}, &gotLogin); err != nil {
t.Fatalf("get login: %v", err)
}
if gotLogin.Spec.Startup.HealthHTTPPort != felisLimboHealthPort {
t.Errorf("login healthHTTPPort = %d, want %d", gotLogin.Spec.Startup.HealthHTTPPort, felisLimboHealthPort)
}
env := map[string]string{}
for _, e := range gotLogin.Spec.Env {
env[e.Name] = e.Value
}
if env[envPanelHostname] != "console.mc.example.net" {
t.Errorf("%s was not added back: %q", envPanelHostname, env[envPanelHostname])
}
if env["OPERATOR_TUNING"] != "keep-me" {
t.Error("a hand-added env var was dropped; converge only touches config-derived names")
}
}
// A field already holding a non-zero value belongs to the operator: converge must
// report "already converged" and write nothing.
func TestConvergeSystemServersLeavesNonZeroFieldsAlone(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
lobby, err := lobbySystemServer("reg/lobby:1", "minecraft")
if err != nil {
t.Fatalf("build lobby: %v", err)
}
lobby.Spec.Rcon = v1alpha1.RconSpec{
Enabled: true,
SecretRef: v1alpha1.SecretKeyRef{Name: "operator-rotated", Key: "password"},
}
login, err := loginSystemServer("reg/limbo:1", "minecraft",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
if err != nil {
t.Fatalf("build login: %v", err)
}
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(lobby, login).Build()
for _, o := range convergeSystemServers(ctx, cl, "minecraft", "reg/limbo:1", "reg/lobby:1",
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net") {
if o.err != nil {
t.Fatalf("%s: unexpected error: %v", o.name, o.err)
}
if len(o.changes) != 0 || o.skipped != "already converged" {
t.Errorf("%s outcome = %+v, want already converged with no writes", o.name, o)
}
}
var got v1alpha1.MinecraftServer
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLobbyServer}, &got); err != nil {
t.Fatalf("get lobby: %v", err)
}
if got.Spec.Rcon.SecretRef.Name != "operator-rotated" {
t.Errorf("lobby rcon secretRef = %q — converge overwrote a field the operator had already set",
got.Spec.Rcon.SecretRef.Name)
}
}
// Guards: an absent CR is reported (creation is setup's job), a foreign CR is
// refused rather than adopted, and an unset image skips like the provisioner does.
func TestConvergeSystemServersGuards(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
run := func(cl client.Client, loginImage, lobbyImage string) []systemServerOutcome {
return convergeSystemServers(ctx, cl, "minecraft", loginImage, lobbyImage,
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
}
t.Run("absent CRs are reported, not created", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).Build()
for _, o := range run(cl, "reg/limbo:1", "reg/lobby:1") {
if o.err != nil {
t.Fatalf("%s: %v", o.name, o.err)
}
if o.created || !strings.Contains(o.skipped, "not present") {
t.Errorf("%s outcome = %+v, want a not-present skip", o.name, o)
}
}
})
t.Run("foreign CR is refused", func(t *testing.T) {
foreign := &v1alpha1.MinecraftServer{}
foreign.Name = naming.SystemLoginServer
foreign.Namespace = "minecraft"
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(foreign).Build()
out := run(cl, "reg/limbo:1", "")
if len(out) != 2 {
t.Fatalf("outcomes = %d, want 2", len(out))
}
if out[0].err == nil || !strings.Contains(out[0].err.Error(), "not marked") {
t.Fatalf("login error = %v, want an unmarked-name refusal", out[0].err)
}
})
t.Run("unset image skips", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).Build()
out := run(cl, "", "reg/lobby:1")
if out[0].skipped != "image not configured" {
t.Errorf("login skipped = %q, want %q", out[0].skipped, "image not configured")
}
})
}
+202
View File
@@ -0,0 +1,202 @@
package main
import (
"archive/tar"
"compress/gzip"
"context"
"errors"
"flag"
"fmt"
"io"
"net/http"
"os"
"os/signal"
"path/filepath"
"strings"
"syscall"
"time"
)
// cmdFetchContext is the in-Pod entrypoint the build Job's context-fetch
// initContainer runs. It reads the blob the platform stored for a submission
// from the felis-api INTERNAL face (with a bounded retry — see
// fetchContextWithRetry) and extracts it into the shared emptyDir the Kaniko
// container then builds from.
//
// Why this exists: the build Pod runs in the build namespace, where it can neither
// mount the control-plane uploads PVC (a PVC does not cross namespaces) nor hold
// object-store credentials, so the API that WROTE the blob is the transport. The
// route is service-token-gated; the token arrives through a namespace-local Secret
// mounted only into this initContainer, never into Kaniko's — so the untrusted
// Dockerfile's build steps have no credential to read (their containers share no
// environment, no PID namespace, and Kaniko itself mounts the context read-only).
//
// The extraction is deliberately paranoid: the tarball is attacker-controlled
// input, so absolute paths, ".." escapes, links, and special files are refused
// rather than sanitized. Kaniko treats the extracted tree as hostile regardless
// (spec §16), but the pod's own filesystem still must not be written outside the
// context directory it was given.
func cmdFetchContext(args []string, _, stderr io.Writer) int {
fs := flag.NewFlagSet("fetch-context", flag.ContinueOnError)
fs.SetOutput(stderr)
url := fs.String("url", "", "internal-face URL of the submission's build-context tarball")
out := fs.String("out", "/context", "directory to extract the build context into")
if err := fs.Parse(args); err != nil {
return 2
}
if *url == "" {
fmt.Fprintln(stderr, "felis fetch-context: --url is required")
return 2
}
token := os.Getenv("FELIS_SERVICE_TOKEN")
if token == "" {
fmt.Fprintln(stderr, "felis fetch-context: FELIS_SERVICE_TOKEN is empty — the internal face rejects anonymous reads")
return 2
}
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
// Validate the URL once up front: a bad one is a usage error (2), not
// something to sit in the retry loop.
if _, err := http.NewRequest(http.MethodGet, *url, nil); err != nil {
fmt.Fprintf(stderr, "felis fetch-context: bad --url: %v\n", err)
return 2
}
// No overall client timeout: a legitimate modpack context can be large and the
// Job's activeDeadlineSeconds is the real bound. The header timeout catches a
// wedged endpoint without capping a healthy download.
client := &http.Client{Transport: &http.Transport{ResponseHeaderTimeout: time.Minute}}
resp, err := fetchContextWithRetry(ctx, client, *url, token, stderr)
if err != nil {
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
return 1
}
defer resp.Body.Close()
if err := extractTarGz(resp.Body, *out); err != nil {
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
return 1
}
return 0
}
// fetchRetryInterval/fetchRetryWindow bound how long the fetch waits out a
// control-plane blip before giving up. The api pod being replaced is a normal
// event (rollout, eviction, a chaos drill), and without a retry one refused
// dial turns it into a failed build: BackoffLimit=0 gives the Job no second
// Pod, so the terminal verdict costs a manual re-approval — the live drill hit
// exactly this (context-fetch exit 1 on `connect: connection refused` while
// the api pod rolled; the new pod was serving 11 seconds later and the same
// 198-byte blob). The window is tiny next to the Job's 30-minute
// activeDeadline; a 4xx (missing blob, rejected token) still fails fast.
//
// Vars, not consts, so tests can shrink the window.
var (
fetchRetryInterval = 3 * time.Second
fetchRetryWindow = 45 * time.Second
)
// fetchContextWithRetry GETs the context tarball, retrying transport failures
// and 5xx responses until fetchRetryWindow runs out. A 4xx is an answer, not a
// blip — retrying it only delays the honest error.
func fetchContextWithRetry(ctx context.Context, client *http.Client, url, token string, stderr io.Writer) (*http.Response, error) {
deadline := time.Now().Add(fetchRetryWindow)
for {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, fmt.Errorf("bad --url: %w", err)
}
req.Header.Set("Authorization", "Bearer "+token)
resp, err := client.Do(req)
if err == nil && resp.StatusCode == http.StatusOK {
return resp, nil
}
if err == nil {
status := resp.Status
_ = resp.Body.Close()
err = fmt.Errorf("GET returned %s", status)
if resp.StatusCode < 500 {
return nil, err
}
}
if ctx.Err() != nil {
return nil, fmt.Errorf("GET failed: %w", err)
}
if time.Now().After(deadline) {
return nil, fmt.Errorf("GET failed (retried for %s): %w", fetchRetryWindow, err)
}
fmt.Fprintf(stderr, "felis fetch-context: %v; retrying (the internal face may be restarting)\n", err)
select {
case <-ctx.Done():
return nil, fmt.Errorf("GET failed: %w", err)
case <-time.After(fetchRetryInterval):
}
}
}
// extractTarGz streams a gzip'd tarball into root, creating directories as
// needed. Every entry is vetted BEFORE anything is written: a path that is
// absolute or escapes root (via ".."), a link (symlink or hardlink), or any
// special file kind aborts the whole extraction. Refusing rather than skipping is
// deliberate — a context that needs one of those constructs is not a context this
// transport carries, and silently dropping entries would build from a corpus the
// submitter did not upload.
func extractTarGz(r io.Reader, root string) error {
if err := os.MkdirAll(root, 0o755); err != nil {
return fmt.Errorf("create context dir: %w", err)
}
zr, err := gzip.NewReader(r)
if err != nil {
return fmt.Errorf("context is not a valid gzip tarball: %w", err)
}
defer zr.Close()
tr := tar.NewReader(zr)
for {
hdr, err := tr.Next()
if errors.Is(err, io.EOF) {
return nil
}
if err != nil {
return fmt.Errorf("read context tarball: %w", err)
}
name := filepath.Clean(hdr.Name)
if name == "." {
continue
}
// The zip-slip guard: reject, never rewrite. filepath.Clean collapses any
// "a/../../b", so these two checks are sufficient once Clean has run.
if filepath.IsAbs(name) || name == ".." || strings.HasPrefix(name, ".."+string(filepath.Separator)) {
return fmt.Errorf("context entry %q escapes the context directory", hdr.Name)
}
target := filepath.Join(root, name)
switch hdr.Typeflag {
case tar.TypeDir:
if err := os.MkdirAll(target, 0o755); err != nil {
return fmt.Errorf("create %q: %w", name, err)
}
case tar.TypeReg, tar.TypeRegA:
if err := os.MkdirAll(filepath.Dir(target), 0o755); err != nil {
return fmt.Errorf("create parent of %q: %w", name, err)
}
mode := os.FileMode(0o644)
if hdr.FileInfo().Mode()&0o111 != 0 {
mode = 0o755 // preserve executability (entrypoint scripts), nothing else
}
f, err := os.OpenFile(target, os.O_CREATE|os.O_WRONLY|os.O_TRUNC, mode)
if err != nil {
return fmt.Errorf("create %q: %w", name, err)
}
if _, err := io.Copy(f, tr); err != nil {
_ = f.Close()
return fmt.Errorf("write %q: %w", name, err)
}
if err := f.Close(); err != nil {
return fmt.Errorf("close %q: %w", name, err)
}
default:
return fmt.Errorf("context entry %q has unsupported type %q (links and special files are refused)", hdr.Name, string(hdr.Typeflag))
}
}
}
+315
View File
@@ -0,0 +1,315 @@
package main
import (
"archive/tar"
"bytes"
"compress/gzip"
"io"
"net"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strings"
"sync/atomic"
"testing"
"time"
)
type tarEntry struct {
name string
body string
mode int64
typ byte
linkname string
}
// tgzBody builds an in-memory .tar.gz from entries, preserving each entry's type
// and mode so the tests can exercise the guards with exactly the bytes an
// attacker could upload.
func tgzBody(t *testing.T, entries ...tarEntry) []byte {
t.Helper()
var buf bytes.Buffer
zw := gzip.NewWriter(&buf)
tw := tar.NewWriter(zw)
for _, e := range entries {
typ := e.typ
if typ == 0 {
typ = tar.TypeReg
}
mode := e.mode
if mode == 0 {
mode = 0o644
}
hdr := &tar.Header{Name: e.name, Typeflag: typ, Mode: mode, Size: int64(len(e.body))}
if typ == tar.TypeSymlink {
hdr.Linkname = e.linkname
hdr.Size = 0
}
if err := tw.WriteHeader(hdr); err != nil {
t.Fatalf("write header %q: %v", e.name, err)
}
if hdr.Size > 0 {
if _, err := tw.Write([]byte(e.body)); err != nil {
t.Fatalf("write body %q: %v", e.name, err)
}
}
}
if err := tw.Close(); err != nil {
t.Fatalf("close tar: %v", err)
}
if err := zw.Close(); err != nil {
t.Fatalf("close gzip: %v", err)
}
return buf.Bytes()
}
// A normal context extracts with its tree intact, and the executable bit that
// modpack entrypoints rely on survives.
func TestExtractTarGzRoundTrip(t *testing.T) {
dir := t.TempDir()
body := tgzBody(t,
tarEntry{name: "Dockerfile", body: "FROM scratch\n"},
tarEntry{name: "mods/example.jar", body: "jar-bytes"},
tarEntry{name: "start.sh", body: "#!/bin/sh\n", mode: 0o755},
tarEntry{name: "mods/", typ: tar.TypeDir, mode: 0o755},
)
if err := extractTarGz(bytes.NewReader(body), dir); err != nil {
t.Fatalf("extract: %v", err)
}
for name, want := range map[string]string{
"Dockerfile": "FROM scratch\n",
"mods/example.jar": "jar-bytes",
} {
got, err := os.ReadFile(filepath.Join(dir, name))
if err != nil || string(got) != want {
t.Fatalf("%s = (%q, %v), want %q", name, got, err, want)
}
}
fi, err := os.Stat(filepath.Join(dir, "start.sh"))
if err != nil || fi.Mode()&0o111 == 0 {
t.Fatalf("entrypoint script lost its exec bit: %v (%v)", fi, err)
}
}
// The guards: "..", absolute paths, symlinks, and special files are refused whole
// — nothing escapes, and nothing is silently skipped.
func TestExtractTarGzRefusesEscapes(t *testing.T) {
cases := []struct {
name string
entries []tarEntry
}{
{"dotdot", []tarEntry{{name: "../outside", body: "x"}}},
{"nested dotdot", []tarEntry{{name: "a/../../outside", body: "x"}}},
{"absolute", []tarEntry{{name: "/etc/outside", body: "x"}}},
{"symlink", []tarEntry{{name: "link", typ: tar.TypeSymlink, linkname: "/etc"}}},
{"hardlink", []tarEntry{{name: "hard", typ: tar.TypeLink, linkname: "somewhere"}}},
{"device", []tarEntry{{name: "dev", typ: tar.TypeChar}}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
dir := t.TempDir()
if err := extractTarGz(bytes.NewReader(tgzBody(t, tc.entries...)), dir); err == nil {
t.Fatal("extract accepted a hostile entry, want an error")
}
// Nothing may have been written outside the target (or at all).
entries, _ := os.ReadDir(dir)
if len(entries) != 0 {
t.Fatalf("hostile archive left %d entries behind", len(entries))
}
})
}
}
// The command end to end: it dials the URL with the bearer token from the
// environment, and refuses to run without it (the internal face would 401
// anyway; failing at parse time is the honest earlier error).
func TestCmdFetchContextFetchAndExtract(t *testing.T) {
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Header.Get("Authorization") != "Bearer test-token" {
w.WriteHeader(http.StatusUnauthorized)
return
}
w.Header().Set("Content-Type", "application/gzip")
_, _ = w.Write(body)
}))
defer srv.Close()
dir := t.TempDir()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + dir}, io.Discard, io.Discard); code != 0 {
t.Fatalf("cmdFetchContext exit = %d, want 0", code)
}
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
}
// No token: refuse before dialing.
t.Setenv("FELIS_SERVICE_TOKEN", "")
var stderr bytes.Buffer
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 2 {
t.Fatalf("missing token exit = %d, want 2 (stderr %q)", code, stderr.String())
}
// A non-200 answer (e.g. the route's 404 for a never-uploaded context) fails.
srv404 := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusNotFound)
}))
defer srv404.Close()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
if code := cmdFetchContext([]string{"--url=" + srv404.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, io.Discard); code != 1 {
t.Fatalf("404 exit = %d, want 1", code)
}
}
// A body that is not a gzip tarball must fail the extraction rather than produce
// an empty (or partial) context Kaniko would then try to build.
func TestExtractTarGzRejectsNonGzip(t *testing.T) {
dir := t.TempDir()
err := extractTarGz(strings.NewReader("not a tarball"), dir)
if err == nil || !strings.Contains(err.Error(), "gzip") {
t.Fatalf("err = %v, want a gzip complaint", err)
}
}
// shrinkFetchWindow swaps the retry knobs for a faster test and restores them
// afterwards, so no test leaks a tiny window into another.
func shrinkFetchWindow(t *testing.T, interval, window time.Duration) {
t.Helper()
oldInterval, oldWindow := fetchRetryInterval, fetchRetryWindow
fetchRetryInterval, fetchRetryWindow = interval, window
t.Cleanup(func() { fetchRetryInterval, fetchRetryWindow = oldInterval, oldWindow })
}
// A control-plane blip mid-fetch is survived: a 5xx on the first attempt is
// retried and the second attempt's tarball extracts. This walks back the live
// drill's failure, where the api pod rolled mid-fetch and the single attempt
// died, failing the build Job.
func TestFetchContextRetriesThroughBlip(t *testing.T) {
shrinkFetchWindow(t, 10*time.Millisecond, time.Second)
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
var calls int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
if atomic.AddInt32(&calls, 1) == 1 {
w.WriteHeader(http.StatusBadGateway) // the port is up, the API is not
return
}
_, _ = w.Write(body)
}))
defer srv.Close()
dir := t.TempDir()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + dir}, io.Discard, &stderr); code != 0 {
t.Fatalf("exit = %d, want 0 (stderr %q)", code, stderr.String())
}
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
}
if !strings.Contains(stderr.String(), "retrying") {
t.Fatalf("stderr %q does not mention the retry", stderr.String())
}
}
// The live drill's exact shape: the dial itself is refused (the api pod is
// gone and no endpoint answers). A refused dial is retried like any other
// transport failure, and once the face is back the fetch completes.
func TestFetchContextRetriesRefusedDial(t *testing.T) {
shrinkFetchWindow(t, 10*time.Millisecond, 5*time.Second)
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
// Borrow a listen address, then close it: the first attempts dial into a
// refused connection, exactly like a restarting control plane.
probe := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
addr := strings.TrimPrefix(probe.URL, "http://")
probe.Close()
dir := t.TempDir()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
// Start the fetch; while the retry loop burns refused dials, bring the same
// address back.
result := make(chan int, 1)
go func() {
result <- cmdFetchContext([]string{"--url=http://" + addr + "/sub-1/context", "--out=" + dir}, io.Discard, &stderr)
}()
time.Sleep(100 * time.Millisecond) // let a handful of dials be refused
ln, err := net.Listen("tcp", addr)
if err != nil {
t.Fatalf("rebind %s: %v", addr, err)
}
back := &http.Server{Handler: http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Header.Get("Authorization") != "Bearer test-token" {
w.WriteHeader(http.StatusUnauthorized)
return
}
_, _ = w.Write(body)
})}
defer back.Close()
go func() { _ = back.Serve(ln) }()
code := <-result
if code != 0 {
t.Fatalf("exit = %d, want 0 (stderr %q)", code, stderr.String())
}
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
}
if !strings.Contains(stderr.String(), "retrying") {
t.Fatalf("stderr %q does not mention the retry", stderr.String())
}
}
// A 4xx is an answer, not a blip: a missing/never-uploaded context fails
// immediately — no retry loop burns the build's deadline on a terminal error.
func TestFetchContextDoesNotRetry4xx(t *testing.T) {
shrinkFetchWindow(t, 5*time.Millisecond, 200*time.Millisecond)
var calls int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
atomic.AddInt32(&calls, 1)
w.WriteHeader(http.StatusNotFound)
}))
defer srv.Close()
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
t.Fatalf("exit = %d, want 1 (stderr %q)", code, stderr.String())
}
if got := atomic.LoadInt32(&calls); got != 1 {
t.Fatalf("server saw %d attempts, want exactly 1", got)
}
if strings.Contains(stderr.String(), "retrying") {
t.Fatalf("stderr %q mentions a retry for a terminal 4xx", stderr.String())
}
}
// The retry is bounded: an internal face that stays down does not hang the
// build pod; the window runs out and the fetch reports the exhausted retries.
func TestFetchContextGivesUpAfterWindow(t *testing.T) {
shrinkFetchWindow(t, 5*time.Millisecond, 60*time.Millisecond)
var calls int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
atomic.AddInt32(&calls, 1)
w.WriteHeader(http.StatusServiceUnavailable)
}))
defer srv.Close() // the face is up but never healthy: 503 forever
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
var stderr bytes.Buffer
start := time.Now()
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
t.Fatalf("exit = %d, want 1 (stderr %q)", code, stderr.String())
}
if elapsed := time.Since(start); elapsed > 5*time.Second {
t.Fatalf("gave up after %v; the window is supposed to bound it", elapsed)
}
if got := atomic.LoadInt32(&calls); got < 2 {
t.Fatalf("server saw %d attempts, want at least one retry", got)
}
if !strings.Contains(stderr.String(), "retried for") {
t.Fatalf("stderr %q does not report the exhausted retry window", stderr.String())
}
}
+1 -1
View File
@@ -26,7 +26,7 @@ const defaultForwardingDataDir = "/data"
// volume of unknown ownership; 0666/0777 then let a non-root Paper rewrite the // volume of unknown ownership; 0666/0777 then let a non-root Paper rewrite the
// same files on boot. // same files on boot.
// //
// ponytail: relies on the initContainer running as root to write into a volume of // This relies on the initContainer running as root to write into a volume of
// unknown ownership; that is how the operator schedules it. If that ever changes, // unknown ownership; that is how the operator schedules it. If that ever changes,
// give the server pod an fsGroup so the shared volume is group-writable instead. // give the server pod an fsGroup so the shared volume is group-writable instead.
const ( const (
+45 -24
View File
@@ -26,8 +26,10 @@ func (m *multiFlag) Set(v string) error {
// felis-reaper identity only when the retention reaper is enabled, gated with // felis-reaper identity only when the retention reaper is enabled, gated with
// its CronJob), the weak build/restore Job SAs, the build/minecraft // its CronJob), the weak build/restore Job SAs, the build/minecraft
// NetworkPolicies, and the running control-plane workloads (felis-api/operator // NetworkPolicies, and the running control-plane workloads (felis-api/operator
// Deployments + the in-cluster registry Deployment/Service/PVC) — as a single // Deployments + the in-cluster registry Deployment/Service/PVC + the
// multi-document YAML stream on stdout, ready for `kubectl apply -f -`. // world-archive PVC that backs backup/restore, unless --backup-pvc is emptied)
// — as a single multi-document YAML stream on stdout, ready for
// `kubectl apply -f -`.
// //
// It is a pure renderer: it never contacts a cluster and holds no credentials. // It is a pure renderer: it never contacts a cluster and holds no credentials.
// --velocity-cidr records the proxy host addresses allowed by the game NetworkPolicy. // --velocity-cidr records the proxy host addresses allowed by the game NetworkPolicy.
@@ -44,9 +46,10 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
panelNodePort := fs.Int("panel-node-port", int(platform.DefaultPanelNodePort), "NodePort that exposes the built-in HTTPS panel/API origin") panelNodePort := fs.Int("panel-node-port", int(platform.DefaultPanelNodePort), "NodePort that exposes the built-in HTTPS panel/API origin")
felisImage := fs.String("felis-image", "", "container image the felis-api/operator Deployments run, also passed through as FELIS_IMAGE (REQUIRED)") felisImage := fs.String("felis-image", "", "container image the felis-api/operator Deployments run, also passed through as FELIS_IMAGE (REQUIRED)")
registryImage := fs.String("registry-image", "", "in-cluster registry image (default: registry:2)") registryImage := fs.String("registry-image", "", "in-cluster registry image (default: registry:2)")
backupPVC := fs.String("backup-pvc", "", "name of the backup PVC advertised to the restore executor via FELIS_BACKUP_PVC (default none = restore endpoint returns 503)") backupPVC := fs.String("backup-pvc", "felis-backups", "name of the world-archive PVC this bundle renders in the Minecraft namespace and advertises to the backup/restore executors via FELIS_BACKUP_PVC (default: felis-backups; pass an empty value to render none, leaving backup/restore answering 503)")
worldsHostPath := fs.String("worlds-host-path", "", "node directory under which each world PVC is visible as <path>/<pvc>; enables the reaper CronJob (requires --backup-pvc and --archive-local-path)") worldsHostPath := fs.String("worlds-host-path", "", "node directory the reaper reads worlds from: each world PVC resolves as <path>/<pvc>, or as the stock local-path directory <path>/<pv-name>_<ns>_<pvc-name> (k3s storage root: /var/lib/rancher/k3s/storage); enables the reaper CronJob (requires --archive-local-path and a non-empty --backup-pvc)")
archiveLocalPath := fs.String("archive-local-path", "", "path the backup PVC is mounted at in the reaper CronJob; MUST equal felis.toml [archive] local_path") archiveLocalPath := fs.String("archive-local-path", "", "path the backup PVC is mounted at in the reaper CronJob; MUST equal felis.toml [archive] local_path")
reaperNode := fs.String("reaper-node", "", "node that holds --worlds-host-path: pins the reaper CronJob's pod there via nodeSelector kubernetes.io/hostname (multi-node clusters need this, or the reaper may schedule where the hostPath is empty)")
var velocityCIDRs multiFlag var velocityCIDRs multiFlag
fs.Var(&velocityCIDRs, "velocity-cidr", "CIDR of a Velocity proxy host allowed to reach game port 25565 (repeatable, REQUIRED)") fs.Var(&velocityCIDRs, "velocity-cidr", "CIDR of a Velocity proxy host allowed to reach game port 25565 (repeatable, REQUIRED)")
var packageCIDRs multiFlag var packageCIDRs multiFlag
@@ -81,36 +84,53 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
fmt.Fprintf(stderr, "felis manifests: --panel-node-port must be in Kubernetes NodePort range 30000-32767 (got %d)\n", *panelNodePort) fmt.Fprintf(stderr, "felis manifests: --panel-node-port must be in Kubernetes NodePort range 30000-32767 (got %d)\n", *panelNodePort)
return 2 return 2
} }
// The node pin exists only for the reaper's hostPath: naming a node without the
// worlds root would be silently dropped (no CronJob renders), so fail loud like
// the storage-trio check below.
if *reaperNode != "" && *worldsHostPath == "" {
fmt.Fprintln(stderr, "felis manifests: --reaper-node requires --worlds-host-path "+
"(it pins the reaper CronJob, which renders only with the retention storage trio)")
return 2
}
// Retention/reaper rendering is opt-in and needs all three storage coordinates // Retention/reaper rendering is opt-in and needs a storage topology together:
// together: where worlds live (to read+archive them), the backup PVC (to write // where worlds live (to read+archive them), a backup PVC (to write archives
// archives into), and the path it is mounted at (which MUST equal felis.toml // into — rendered from --backup-pvc), and the path it is mounted at (which MUST
// [archive] local_path so tarLocal's absolute archive refs resolve). A partial // equal felis.toml [archive] local_path so tarLocal's absolute archive refs
// configuration is almost certainly an operator mistake, so fail loud rather than // resolve). A partial configuration is almost certainly an operator mistake, so
// silently drop retention. Asking for it without the other two is rejected; an // fail loud rather than silently drop retention or render a reaper with nowhere
// empty trio renders the bundle WITHOUT the reaper and says so. // to write. The backup PVC itself defaults to felis-backups (it is what makes a
// default install's backup endpoint work at all); retention additionally needs
// --worlds-host-path.
if *worldsHostPath != "" { if *worldsHostPath != "" {
if *backupPVC == "" || *archiveLocalPath == "" { if *backupPVC == "" || *archiveLocalPath == "" {
fmt.Fprintln(stderr, "felis manifests: --worlds-host-path enables the reaper CronJob and requires "+ fmt.Fprintln(stderr, "felis manifests: --worlds-host-path enables the reaper CronJob and requires "+
"--backup-pvc and --archive-local-path too (--archive-local-path must equal felis.toml [archive] local_path)") "--archive-local-path (must equal felis.toml [archive] local_path) and a non-empty --backup-pvc "+
"(the archive store; default felis-backups)")
return 2 return 2
} }
// The reaper WILL render. Two deployment preconditions this generator cannot // The reaper WILL render. Two deployment facts this generator cannot check
// check would SILENTLY turn retention into a no-op if unmet — surface them as // would silently turn retention into a no-op if unmet — surface them as
// loudly as the fail-closed cases above, so an operator is never left with a // loudly as the fail-closed cases above, so an operator is never left with a
// reaper that reaps an empty directory. (Both are also in the WorldsHostPath // reaper that reaps nothing. (Both are also in the WorldsHostPath flag/field
// flag/field docs, but nobody deploying from stdout reads those.) // docs, but nobody deploying from stdout reads those.)
pin := "the CronJob sets NO nodeSelector: a single-node starter pins it to the worlds implicitly, but on a " +
"multi-node cluster you MUST pass --reaper-node <name> (or add a nodeSelector) for the node holding the " +
"worlds, or the reaper may schedule where the hostPath is empty"
if *reaperNode != "" {
pin = fmt.Sprintf("the CronJob is pinned to node %q via kubernetes.io/hostname — keep this pointed at the "+
"node that actually holds the world volumes", *reaperNode)
}
fmt.Fprintf(stderr, "felis manifests: note: rendering the retention reaper CronJob (worlds hostPath %q). "+ fmt.Fprintf(stderr, "felis manifests: note: rendering the retention reaper CronJob (worlds hostPath %q). "+
"Two preconditions are NOT verified here:\n"+ "These points are NOT verified here:\n"+
" - each world PVC must be visible at %s/<pvc> on the node: a stock local-path-provisioner lays "+ " - the node's world volumes must actually live below %s: the reaper resolves a world as "+
"volumes under PV-name paths (.../pvc-<uuid>_<ns>_<pvc>/), so unless the worlds StorageClass is "+ "%s/<pvc>, then as the stock local-path directory <path>/<pv-name>_<ns>_<pvc-name> (what k3s "+
"arranged to expose <path>/<pvc>, the reaper tars an empty directory;\n"+ "writes under /var/lib/rancher/k3s/storage). Any other provisioner needs its volumes exposed as "+
" - the CronJob sets NO nodeSelector: a single-node starter pins it to the worlds implicitly, but "+ "<path>/<pvc>, or each candidate's archive fails and the world is preserved;\n"+
"on a multi-node cluster you MUST add a nodeSelector for the node holding the worlds, or the reaper "+ " - %s.\n", *worldsHostPath, *worldsHostPath, *worldsHostPath, pin)
"may schedule where the hostPath is empty.\n", *worldsHostPath, *worldsHostPath)
} else { } else {
fmt.Fprintln(stderr, "felis manifests: note: retention reaper CronJob not rendered "+ fmt.Fprintln(stderr, "felis manifests: note: retention reaper CronJob not rendered "+
"(pass --worlds-host-path, --backup-pvc and --archive-local-path to enable it)") "(pass --worlds-host-path and --archive-local-path — the archive PVC defaults to felis-backups — to enable it)")
} }
out, err := platform.RenderYAML(platform.Params{ out, err := platform.RenderYAML(platform.Params{
@@ -124,6 +144,7 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
RegistryImage: *registryImage, RegistryImage: *registryImage,
BackupPVC: *backupPVC, BackupPVC: *backupPVC,
WorldsHostPath: *worldsHostPath, WorldsHostPath: *worldsHostPath,
ReaperNode: *reaperNode,
ArchiveLocalPath: *archiveLocalPath, ArchiveLocalPath: *archiveLocalPath,
VelocityCIDRs: []string(velocityCIDRs), VelocityCIDRs: []string(velocityCIDRs),
PackageSourceCIDRs: []string(packageCIDRs), PackageSourceCIDRs: []string(packageCIDRs),
+71 -8
View File
@@ -77,6 +77,10 @@ func TestManifestsRendersBundle(t *testing.T) {
"10.0.0.5/32", "10.0.0.5/32",
// The felis image flows through to the Deployments. // The felis image flows through to the Deployments.
"registry.felis.svc:5000/felis:v1", "registry.felis.svc:5000/felis:v1",
// Backup works out of the box: the archive PVC renders and the api gets
// the env that wires the backup/restore executors to it.
"name: felis-backups",
"name: FELIS_BACKUP_PVC",
} { } {
if !strings.Contains(text, want) { if !strings.Contains(text, want) {
t.Errorf("rendered bundle missing %q", want) t.Errorf("rendered bundle missing %q", want)
@@ -95,15 +99,19 @@ func TestManifestsRendersBundle(t *testing.T) {
} }
} }
// TestManifestsReaperRequiresTrio proves --worlds-host-path is a fail-loud opt-in: // TestManifestsReaperRequiresStorage proves --worlds-host-path is a fail-loud
// asking for the reaper without the backup PVC and its mount path (which must equal // opt-in: asking for the reaper without a writable archive store (the backup PVC,
// [archive] local_path) is rejected rather than silently dropping retention. // which defaults to felis-backups but can be emptied) and its mount path (which
func TestManifestsReaperRequiresTrio(t *testing.T) { // must equal [archive] local_path) is rejected rather than silently dropping
// retention or deleting worlds it could not archive first.
func TestManifestsReaperRequiresStorage(t *testing.T) {
base := []string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32", "--worlds-host-path", "/var/lib/felis/worlds"} base := []string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32", "--worlds-host-path", "/var/lib/felis/worlds"}
for _, extra := range [][]string{ for _, extra := range [][]string{
{}, // neither backup-pvc nor archive-local-path {}, // missing archive-local-path (backup-pvc defaults)
{"--backup-pvc", "felis-backups"}, // missing archive-local-path {"--backup-pvc", "other"}, // still missing archive-local-path
{"--archive-local-path", "/backups"}, // missing backup-pvc // A reaper with no archive store would have nowhere to write the archive
// it must verify before deleting a world; emptying the PVC is rejected.
{"--archive-local-path", "/backups", "--backup-pvc="},
} { } {
var out, errBuf bytes.Buffer var out, errBuf bytes.Buffer
code := run(append(append([]string{}, base...), extra...), &out, &errBuf) code := run(append(append([]string{}, base...), extra...), &out, &errBuf)
@@ -119,6 +127,23 @@ func TestManifestsReaperRequiresTrio(t *testing.T) {
} }
} }
// TestManifestsBackupPVCOptOut proves --backup-pvc= renders a bundle with no
// archive store at all: no PVC and no FELIS_BACKUP_PVC env, so backup/restore
// answer 503 instead of pointing Jobs at a claim nobody provisions.
func TestManifestsBackupPVCOptOut(t *testing.T) {
var out, errBuf bytes.Buffer
code := run([]string{"manifests", "--felis-image", "reg/felis:test",
"--velocity-cidr", "10.0.0.5/32", "--backup-pvc="}, &out, &errBuf)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, errBuf.String())
}
for _, absent := range []string{"felis-backups", "FELIS_BACKUP_PVC"} {
if strings.Contains(out.String(), absent) {
t.Errorf("--backup-pvc= bundle must not contain %q", absent)
}
}
}
// TestManifestsRendersReaper proves the happy path with the full retention trio: // TestManifestsRendersReaper proves the happy path with the full retention trio:
// a batch/v1 CronJob is emitted, named felis-reaper, mounting the backup PVC at the // a batch/v1 CronJob is emitted, named felis-reaper, mounting the backup PVC at the
// supplied archive path. // supplied archive path.
@@ -150,9 +175,47 @@ func TestManifestsRendersReaper(t *testing.T) {
// this generator cannot verify (else a misarranged hostPath silently no-ops // this generator cannot verify (else a misarranged hostPath silently no-ops
// retention): the <path>/<pvc> arrangement-dependency and the multi-node // retention): the <path>/<pvc> arrangement-dependency and the multi-node
// nodeSelector hazard. // nodeSelector hazard.
for _, want := range []string{"local-path-provisioner", "nodeSelector"} { for _, want := range []string{"local-path", "nodeSelector"} {
if !strings.Contains(errBuf.String(), want) { if !strings.Contains(errBuf.String(), want) {
t.Errorf("reaper render must warn operators about %q on stderr, got %q", want, errBuf.String()) t.Errorf("reaper render must warn operators about %q on stderr, got %q", want, errBuf.String())
} }
} }
} }
// TestManifestsReaperNodePin: --reaper-node pins the rendered CronJob's pod via
// kubernetes.io/hostname and replaces the "no nodeSelector" hazard note with the
// pin confirmation; using it without the worlds root is a fail-loud 2.
func TestManifestsReaperNodePin(t *testing.T) {
var out, errBuf bytes.Buffer
code := run([]string{
"manifests",
"--felis-image", "registry.felis.svc:5000/felis:v1",
"--velocity-cidr", "10.0.0.5/32",
"--worlds-host-path", "/var/lib/felis/worlds",
"--archive-local-path", "/backups",
"--reaper-node", "node-a",
}, &out, &errBuf)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, errBuf.String())
}
for _, want := range []string{
"kubernetes.io/hostname: node-a",
} {
if !strings.Contains(out.String(), want) {
t.Errorf("pinned render missing %q", want)
}
}
if !strings.Contains(errBuf.String(), "node-a") {
t.Errorf("stderr must confirm the pin, got %q", errBuf.String())
}
var out2, err2 bytes.Buffer
if code := run([]string{
"manifests",
"--felis-image", "registry.felis.svc:5000/felis:v1",
"--velocity-cidr", "10.0.0.5/32",
"--reaper-node", "node-a",
}, &out2, &err2); code != 2 {
t.Errorf("--reaper-node without --worlds-host-path: exit = %d, want 2", code)
}
}
+71 -17
View File
@@ -29,7 +29,12 @@ import (
"flag" "flag"
"fmt" "fmt"
"io" "io"
"net"
"net/http" "net/http"
"os"
"os/signal"
"syscall"
"time"
"felis.lolicon.best/internal/api" "felis.lolicon.best/internal/api"
"felis.lolicon.best/internal/config" "felis.lolicon.best/internal/config"
@@ -37,21 +42,27 @@ import (
// nanoStubRepo satisfies api.Repo but implements only the one method handleHasJoined calls. // nanoStubRepo satisfies api.Repo but implements only the one method handleHasJoined calls.
// The reclaim username blacklist is a felis-api/DB concern; a nano host has no Postgres, so // The reclaim username blacklist is a felis-api/DB concern; a nano host has no Postgres, so
// nothing is barred here. ponytail: a real blacklist would need the very DB nano exists to // nothing is barred here. A real blacklist would need the very DB nano exists to
// avoid — YAGNI until a nano host grows a reclaim store. // avoid — YAGNI until a nano host grows a reclaim store.
type nanoStubRepo struct{ api.Repo } type nanoStubRepo struct{ api.Repo }
func (nanoStubRepo) IsUsernameBlacklisted(context.Context, string) (bool, error) { return false, nil } func (nanoStubRepo) IsUsernameBlacklisted(context.Context, string) (bool, error) { return false, nil }
// nanoDefaultListen is loopback because hasJoined carries no auth token (Velocity speaks the
// vanilla sessionserver protocol), so a public bind is an open auth relay: anyone can point
// their proxy at it and spend this host's egress IP on Mojang. A same-host Velocity reaches
// 127.0.0.1; serving an off-host proxy is an explicit -listen opt-in.
const nanoDefaultListen = "127.0.0.1:8081"
// nanoLogURIMax is room for a real hasJoined query (a 16-character name, a 41-character
// serverId, an address) several times over.
const nanoLogURIMax = 256
func cmdNano(args []string, stdout, stderr io.Writer) int { func cmdNano(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("nano", flag.ContinueOnError) fs := flag.NewFlagSet("nano", flag.ContinueOnError)
fs.SetOutput(stderr) fs.SetOutput(stderr)
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (reads [[auth_source]])") cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (reads [[auth_source]])")
// Loopback default: hasJoined carries no auth token (authlib speaks the vanilla listen := fs.String("listen", nanoDefaultListen, "listen address for the hasJoined endpoint")
// sessionserver protocol), so a public bind is an open auth relay — anyone can point
// their proxy at it and spend this host's egress IP on Mojang. A same-host Velocity
// reaches 127.0.0.1; serving an off-host proxy is an explicit -listen opt-in.
listen := fs.String("listen", "127.0.0.1:8081", "listen address for the hasJoined endpoint")
if err := fs.Parse(args); err != nil { if err := fs.Parse(args); err != nil {
return 2 return 2
} }
@@ -61,24 +72,67 @@ func cmdNano(args []string, stdout, stderr io.Writer) int {
fmt.Fprintln(stderr, "felis nano:", err) fmt.Fprintln(stderr, "felis nano:", err)
return 1 return 1
} }
// [server] listen belongs to felis api. Someone moving nano off loopback naturally reaches
// for it, and without this line would get connection refused with no hint why.
if cfg.Server.Listen != "" {
fmt.Fprintf(stderr, "felis nano: [server] listen = %q is ignored; nano binds -listen (%s), which the installer sets from FELIS_NANO_LISTEN\n", cfg.Server.Listen, *listen)
}
handler := api.HasJoinedHandler(authSourcesFromConfig(cfg.AuthSources), nanoStubRepo{})
fmt.Fprintf(stderr, "felis nano: hasJoined multiplexer on %s — Mojang + %d third-party source(s)\n", *listen, len(cfg.AuthSources)) fmt.Fprintf(stderr, "felis nano: hasJoined multiplexer on %s — Mojang + %d third-party source(s)\n", *listen, len(cfg.AuthSources))
for i, s := range cfg.AuthSources { for i, s := range cfg.AuthSources {
fmt.Fprintf(stderr, " [%d] %s -> %s\n", i+1, s.Tag, s.URL) fmt.Fprintf(stderr, " [%d] %s -> %s\n", i+1, s.Tag, s.URL)
} }
// Log each request so a live login attempt is visible while testing against a real ln, err := net.Listen("tcp", *listen)
// Velocity — "is authlib even reaching me?" is the first question during verification. if err != nil {
logged := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
fmt.Fprintf(stderr, "felis nano: %s %s\n", r.Method, r.RequestURI)
handler.ServeHTTP(w, r)
})
srv := newAPIServer(*listen, logged)
if err := srv.ListenAndServe(); err != nil {
fmt.Fprintln(stderr, "felis nano:", err) fmt.Fprintln(stderr, "felis nano:", err)
return 1 return 1
} }
return 0 ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
return serveNano(ctx, newAPIServer(*listen, nanoHandler(cfg.AuthSources, stderr)), ln, stderr)
}
// nanoDrainTimeout outlasts the source scan of any realistic list (each source is given
// five seconds) and stays well inside systemd's default 90-second stop timeout.
const nanoDrainTimeout = 30 * time.Second
// serveNano serves until ctx ends, then drains. A restart, the documented way to pick up a
// config edit, sends SIGTERM; without the drain a login already waiting on an upstream has
// its connection reset, and Velocity tells that player the auth servers are down.
func serveNano(ctx context.Context, srv *http.Server, ln net.Listener, stderr io.Writer) int {
errc := make(chan error, 1)
go func() { errc <- srv.Serve(ln) }()
select {
case err := <-errc:
fmt.Fprintln(stderr, "felis nano:", err)
return 1
case <-ctx.Done():
shutdownCtx, cancel := context.WithTimeout(context.Background(), nanoDrainTimeout)
defer cancel()
if err := srv.Shutdown(shutdownCtx); err != nil {
fmt.Fprintln(stderr, "felis nano: shutdown:", err)
return 1
}
return 0
}
}
// nanoHandler is what felis nano serves: the shared hasJoined handler, Mojang first, behind
// a request log.
func nanoHandler(sources []config.AuthSourceConfig, stderr io.Writer) http.Handler {
handler := api.HasJoinedHandler(authSourcesFromConfig(sources), nanoStubRepo{})
// Log each request so a live login attempt is visible while testing against a real
// Velocity — "is Velocity even reaching me?" is the first question during verification.
// The URI is the caller's text: quoted so a control or bidi character cannot rewrite the
// line and invalid UTF-8 cannot turn the journal entry into a blob, and capped so one
// request cannot write a megabyte of log.
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
uri := r.RequestURI
if len(uri) > nanoLogURIMax {
uri = uri[:nanoLogURIMax] + "..."
}
fmt.Fprintf(stderr, "felis nano: %s %q\n", r.Method, uri)
handler.ServeHTTP(w, r)
})
} }
+139
View File
@@ -0,0 +1,139 @@
package main
import (
"bytes"
"context"
"io"
"net"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
"time"
"unicode/utf8"
"felis.lolicon.best/internal/api"
)
// [server] listen in a nano config reads like the bind address but is not one; nano must
// say so. The -listen value cannot be bound, so cmdNano returns right after loading.
func TestNanoWarnsThatServerListenIsIgnored(t *testing.T) {
cfg := filepath.Join(t.TempDir(), "felis.toml")
if err := os.WriteFile(cfg, []byte("[server]\nlisten = \"0.0.0.0:9999\"\n"), 0o600); err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
if rc := cmdNano([]string{"-config", cfg, "-listen", "127.0.0.1:-1"}, io.Discard, &stderr); rc != 1 {
t.Fatalf("cmdNano = %d, want 1 from the unbindable -listen", rc)
}
if !strings.Contains(stderr.String(), `listen = "0.0.0.0:9999" is ignored`) {
t.Fatalf("stderr %q should say the configured listen is ignored", stderr.String())
}
// With no [server] table at all there is nothing to warn about.
if err := os.WriteFile(cfg, nil, 0o600); err != nil {
t.Fatal(err)
}
stderr.Reset()
_ = cmdNano([]string{"-config", cfg, "-listen", "127.0.0.1:-1"}, io.Discard, &stderr)
if strings.Contains(stderr.String(), "is ignored") {
t.Fatalf("stderr %q warns about a listen the operator never set", stderr.String())
}
}
// A stop signal that lands while a login is waiting on an upstream must let that login
// finish: the request is answered, and serveNano returns only afterwards.
func TestNanoDrainsInFlightLoginOnShutdown(t *testing.T) {
entered, release := make(chan struct{}), make(chan struct{})
srv := newAPIServer("", http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
close(entered)
<-release
w.WriteHeader(http.StatusNoContent)
}))
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
ctx, stop := context.WithCancel(context.Background())
done := make(chan int, 1)
go func() { done <- serveNano(ctx, srv, ln, io.Discard) }()
got := make(chan int, 1)
go func() {
resp, err := http.Get("http://" + ln.Addr().String() + "/session/minecraft/hasJoined")
if err != nil {
got <- -1
return
}
resp.Body.Close()
got <- resp.StatusCode
}()
<-entered
stop()
select {
case <-done:
t.Fatal("serveNano returned while a login was still in flight")
case <-time.After(200 * time.Millisecond):
}
close(release)
if code := <-got; code != http.StatusNoContent {
t.Fatalf("in-flight login got %d, want its answer (204)", code)
}
if rc := <-done; rc != 0 {
t.Fatalf("serveNano = %d after a clean drain, want 0", rc)
}
}
// The nano delivery path: the shared handler behind nano's stub store must admit a login its
// source validated. nanoStubRepo implements only the bar-list lookup, so a new store call in
// handleHasJoined would reach its nil embedded Repo and panic here, while the full-api tests,
// which use a complete fake store, stay green.
func TestNanoAdmitsAValidatedLogin(t *testing.T) {
const id = "069a79f444e94726a5befca90e38aaf5"
ygg := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = io.WriteString(w, `{"id":"`+id+`","name":"Notch"}`)
}))
defer ygg.Close()
h := api.HasJoinedHandler([]api.AuthSource{{Tag: "mojang", URL: ygg.URL, Identity: true}}, nanoStubRepo{})
w := httptest.NewRecorder()
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, "/session/minecraft/hasJoined?username=Notch&serverId=abc", nil))
if w.Code != http.StatusOK || !strings.Contains(w.Body.String(), id) {
t.Fatalf("code = %d body = %q, want the validated profile", w.Code, w.Body.String())
}
}
// An unauthenticated relay on a public address spends this host's Mojang rate limit for
// anyone who finds it, so the default bind has to stay loopback.
func TestNanoListensOnLoopbackByDefault(t *testing.T) {
host, _, err := net.SplitHostPort(nanoDefaultListen)
if ip := net.ParseIP(host); err != nil || ip == nil || !ip.IsLoopback() {
t.Fatalf("default -listen %q is not a loopback address", nanoDefaultListen)
}
}
// The request log prints text the caller chose. A bidi override must not reorder the line,
// an invalid byte must not make journald store the entry as a blob, and a huge query must
// not become a huge log line. serverId is left out so the handler answers without asking
// any source.
func TestNanoRequestLogIsQuotedAndCapped(t *testing.T) {
const rlo = rune(0x202e) // RIGHT-TO-LEFT OVERRIDE
var log bytes.Buffer
h := nanoHandler(nil, &log)
target := "/session/minecraft/hasJoined?username=" + string(rlo) + "evil" + string([]byte{0x9b}) + "31m" + strings.Repeat("a", 4096)
w := httptest.NewRecorder()
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, target, nil))
line := log.String()
if strings.ContainsRune(line, rlo) || !utf8.ValidString(line) {
t.Fatalf("raw caller bytes reached the log: %q", line)
}
if escaped := strings.Trim(strconv.QuoteRune(rlo), "'"); !strings.Contains(line, escaped) {
t.Fatalf("log line %q should show the override escaped as %s", line, escaped)
}
if len(line) > 2*nanoLogURIMax {
t.Fatalf("log line is %d bytes for a %d-byte URI; want it capped", len(line), len(target))
}
}
+31 -2
View File
@@ -4,16 +4,19 @@ import (
"flag" "flag"
"fmt" "fmt"
"io" "io"
"log/slog"
"os" "os"
"felis.lolicon.best/internal/apis/felis/v1alpha1" "felis.lolicon.best/internal/apis/felis/v1alpha1"
felismetrics "felis.lolicon.best/internal/metrics" felismetrics "felis.lolicon.best/internal/metrics"
"felis.lolicon.best/internal/operator" "felis.lolicon.best/internal/operator"
"github.com/go-logr/logr"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
utilruntime "k8s.io/apimachinery/pkg/util/runtime" utilruntime "k8s.io/apimachinery/pkg/util/runtime"
clientgoscheme "k8s.io/client-go/kubernetes/scheme" clientgoscheme "k8s.io/client-go/kubernetes/scheme"
ctrl "sigs.k8s.io/controller-runtime" ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/cache" "sigs.k8s.io/controller-runtime/pkg/cache"
"sigs.k8s.io/controller-runtime/pkg/healthz"
ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics"
metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server" metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server"
) )
@@ -25,6 +28,11 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
fs := flag.NewFlagSet("operator", flag.ContinueOnError) fs := flag.NewFlagSet("operator", flag.ContinueOnError)
fs.SetOutput(stderr) fs.SetOutput(stderr)
metricsAddr := fs.String("metrics-bind-address", ":8080", "address the metric endpoint binds to") metricsAddr := fs.String("metrics-bind-address", ":8080", "address the metric endpoint binds to")
// healthAddr serves the manager's health endpoints (/healthz, /readyz) that the
// Deployment's probes dial. Without it the operator pod would carry no probe at
// all, and a wedged manager would keep its endpoint forever. It must differ from
// metricsAddr: the metrics server owns :8080.
healthAddr := fs.String("health-probe-bind-address", ":8081", "address the health probe endpoint binds to")
// namespace MUST equal the [k8s] namespace felis-api is configured with, and // namespace MUST equal the [k8s] namespace felis-api is configured with, and
// the deployment manifests (felis manifests) render both from one value. It // the deployment manifests (felis manifests) render both from one value. It
// scopes the manager's cache (informers) to a single namespace so the operator // scopes the manager's cache (informers) to a single namespace so the operator
@@ -42,9 +50,16 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
utilruntime.Must(clientgoscheme.AddToScheme(scheme)) utilruntime.Must(clientgoscheme.AddToScheme(scheme))
utilruntime.Must(v1alpha1.AddToScheme(scheme)) utilruntime.Must(v1alpha1.AddToScheme(scheme))
// controller-runtime logs through its own logr sink; without one, its first
// reconcile prints "log.SetLogger(...) was never called" ATTACHED TO A FULL
// GOROUTINE STACK — pure noise, not signal. Route it to slog's default handler
// so its messages appear as ordinary stderr lines.
ctrl.SetLogger(logr.FromSlogHandler(slog.Default().Handler()))
mgr, err := ctrl.NewManager(ctrl.GetConfigOrDie(), ctrl.Options{ mgr, err := ctrl.NewManager(ctrl.GetConfigOrDie(), ctrl.Options{
Scheme: scheme, Scheme: scheme,
Metrics: metricsserver.Options{BindAddress: *metricsAddr}, Metrics: metricsserver.Options{BindAddress: *metricsAddr},
HealthProbeBindAddress: *healthAddr,
// Scope every informer to the single watched namespace. Without this the // Scope every informer to the single watched namespace. Without this the
// cached client (mgr.GetClient) would LIST/WATCH cluster-wide, which a // cached client (mgr.GetClient) would LIST/WATCH cluster-wide, which a
// namespaced Role cannot grant — the operator would fail closed at runtime // namespaced Role cannot grant — the operator would fail closed at runtime
@@ -60,6 +75,20 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
} }
fmt.Fprintf(stderr, "felis operator: watching namespace %q\n", *namespace) fmt.Fprintf(stderr, "felis operator: watching namespace %q\n", *namespace)
// Register the two probe endpoints. controller-runtime only mounts /healthz and
// /readyz once at least one check is registered, so a bare listener would 404.
// The checks are the canonical always-pass ping: the probes' contract is "the
// manager process is up and serving", and a dependency hiccup (e.g. an API blip)
// must not restart the operator.
if err := mgr.AddHealthzCheck("ping", healthz.Ping); err != nil {
fmt.Fprintf(stderr, "felis operator: register healthz check: %v\n", err)
return 1
}
if err := mgr.AddReadyzCheck("ping", healthz.Ping); err != nil {
fmt.Fprintf(stderr, "felis operator: register readyz check: %v\n", err)
return 1
}
// Publish the named felis_* metrics (spec §23) on the endpoint the manager // Publish the named felis_* metrics (spec §23) on the endpoint the manager
// already serves (metricsAddr). controller-runtime's metrics server exposes // already serves (metricsAddr). controller-runtime's metrics server exposes
// its global Registry, so registering into it is all that is needed for // its global Registry, so registering into it is all that is needed for
+135 -20
View File
@@ -1,9 +1,13 @@
package main package main
import ( import (
"context"
"database/sql"
"errors"
"flag" "flag"
"fmt" "fmt"
"io" "io"
"os"
"path/filepath" "path/filepath"
"strconv" "strconv"
"strings" "strings"
@@ -12,8 +16,11 @@ import (
"felis.lolicon.best/internal/apis/felis/v1alpha1" "felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/backup" "felis.lolicon.best/internal/backup"
"felis.lolicon.best/internal/config" "felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/mail"
"felis.lolicon.best/internal/platform"
"felis.lolicon.best/internal/reaper" "felis.lolicon.best/internal/reaper"
"felis.lolicon.best/internal/store" "felis.lolicon.best/internal/store"
corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
utilruntime "k8s.io/apimachinery/pkg/util/runtime" utilruntime "k8s.io/apimachinery/pkg/util/runtime"
clientgoscheme "k8s.io/client-go/kubernetes/scheme" clientgoscheme "k8s.io/client-go/kubernetes/scheme"
@@ -30,7 +37,7 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("reaper", flag.ContinueOnError) fs := flag.NewFlagSet("reaper", flag.ContinueOnError)
fs.SetOutput(stderr) fs.SetOutput(stderr)
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml") cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
worldsRoot := fs.String("worlds-root", "/worlds", "mount root under which world PVCs are visible (tarLocal: <root>/<pvc>)") worldsRoot := fs.String("worlds-root", "/worlds", "mount root under which world PVCs are visible (tarLocal: <root>/<pvc>, else the stock local-path <root>/<pv-name>_<ns>_<pvc-name>)")
if err := fs.Parse(args); err != nil { if err := fs.Parse(args); err != nil {
return 2 return 2
} }
@@ -47,21 +54,8 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
return 1 return 1
} }
archiver, err := buildArchiver(cfg, *worldsRoot)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: %v\n", err)
return 1
}
ctx := ctrl.SetupSignalHandler() ctx := ctrl.SetupSignalHandler()
drv, err := store.Open(ctx, cfg.Database.URL)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: open database: %v\n", err)
return 1
}
defer drv.Close()
scheme := runtime.NewScheme() scheme := runtime.NewScheme()
utilruntime.Must(clientgoscheme.AddToScheme(scheme)) utilruntime.Must(clientgoscheme.AddToScheme(scheme))
utilruntime.Must(v1alpha1.AddToScheme(scheme)) utilruntime.Must(v1alpha1.AddToScheme(scheme))
@@ -71,6 +65,19 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
return 1 return 1
} }
archiver, err := buildArchiver(ctx, cfg, *worldsRoot, cl)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: %v\n", err)
return 1
}
drv, err := store.Open(ctx, cfg.Database.URL)
if err != nil {
fmt.Fprintf(stderr, "felis reaper: open database: %v\n", err)
return 1
}
defer drv.Close()
r := &reaper.Reaper{ r := &reaper.Reaper{
Cfg: rcfg, Cfg: rcfg,
Store: reaper.NewPGStore(drv.DB()), Store: reaper.NewPGStore(drv.DB()),
@@ -78,6 +85,47 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
Archiver: archiver, Archiver: archiver,
} }
// Pre-reap warnings go out by email when [smtp] is configured (the same
// relay and password_ref convention felis-api uses); without it the channel
// stays nil and the reaper logs each suppressed warning instead of stamping
// it, so a later SMTP setup still gets to warn. The owner must have a
// VERIFIED address — that flag is what proves the mailbox.
if cfg.SMTP.Host != "" {
passRef := cfg.SMTP.PasswordRef
if passRef == "" {
passRef = platform.SMTPPasswordEnv
}
password := os.Getenv(passRef)
if cfg.SMTP.Username != "" && password == "" {
fmt.Fprintf(stderr, "felis reaper: warning: [smtp] username is set but credentials env %s is empty — warning emails will fail AUTH\n", passRef)
}
db := drv.DB()
r.Warner = &mailWarner{
lookupEmail: func(ctx context.Context, ownerID string) (string, error) {
var email string
switch err := db.QueryRowContext(ctx,
`SELECT email FROM users
WHERE id = $1 AND email_verified = true AND COALESCE(email, '') <> ''`,
ownerID).Scan(&email); {
case errors.Is(err, sql.ErrNoRows):
return "", fmt.Errorf("owner %s has no verified email", ownerID)
case err != nil:
return "", err
}
return email, nil
},
notifier: &mail.SMTP{
Host: cfg.SMTP.Host,
Port: cfg.SMTP.Port,
From: cfg.SMTP.From,
Username: cfg.SMTP.Username,
Password: password,
},
}
} else {
fmt.Fprintln(stderr, "felis reaper: [smtp] not configured — pre-reap warnings are logged and NOT marked sent")
}
sum, err := r.RunOnce(ctx) sum, err := r.RunOnce(ctx)
if err != nil { if err != nil {
fmt.Fprintf(stderr, "felis reaper: %v\n", err) fmt.Fprintf(stderr, "felis reaper: %v\n", err)
@@ -88,6 +136,38 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
return 0 return 0
} }
// mailWarner delivers a pre-reap notice to the owner's verified email — the
// only channel this build can reach. Unowned owners and owners who never proved
// a mailbox yield an error; the reaper retries such notices on its next run and
// never lets them block the reap (red line ⑤).
type mailWarner struct {
lookupEmail func(ctx context.Context, ownerID string) (string, error)
notifier noticeNotifier
}
// noticeNotifier is the slice of mail.SMTP the warner needs (injected in tests).
type noticeNotifier interface {
SendNotice(ctx context.Context, email, subject, body string) error
}
func (w *mailWarner) Warn(ctx context.Context, ownerID, server, remaining string) error {
email, err := w.lookupEmail(ctx, ownerID)
if err != nil {
return fmt.Errorf("resolve owner email: %w", err)
}
subject := fmt.Sprintf("Felis: 服务器 %s 将在 %s 后回收 · server reaped in %s", server, remaining, remaining)
body := fmt.Sprintf(
"Felis 世界回收提醒 / world-reaper notice\r\n"+
"\r\n"+
"服务器 / Server: %s\r\n"+
"距回收 / Time left: %s\r\n"+
"\r\n"+
"闲置的服务器会先自动备份,再释放世界;有人加入游戏即可重置倒计时。\r\n"+
"Idle servers are backed up and then released; any join resets the countdown.\r\n",
server, remaining)
return w.notifier.SendNotice(ctx, email, subject, body)
}
// reaperConfig derives the reaper's retention windows from felis.toml. The 15d // reaperConfig derives the reaper's retention windows from felis.toml. The 15d
// idle deadline is fixed by §18; only the warning offsets, retention, and the // idle deadline is fixed by §18; only the warning offsets, retention, and the
// store soft-cap are configurable (§24). // store soft-cap are configurable (§24).
@@ -122,22 +202,57 @@ func reaperConfig(cfg *config.Config) (reaper.Config, error) {
} }
// buildArchiver constructs the WorldArchiver. Only tarLocal is implemented in // buildArchiver constructs the WorldArchiver. Only tarLocal is implemented in
// this build; the resolver maps each world PVC to <worldsRoot>/<pvc>, the mount // this build; the resolver maps each world PVC to its directory under worldsRoot
// convention the reaper Job is deployed with. // (resolveWorldDir).
func buildArchiver(cfg *config.Config, worldsRoot string) (backup.WorldArchiver, error) { func buildArchiver(ctx context.Context, cfg *config.Config, worldsRoot string, cl client.Client) (backup.WorldArchiver, error) {
switch cfg.Archive.Store { switch cfg.Archive.Store {
case "tarLocal": case "tarLocal":
return &backup.TarLocal{ return &backup.TarLocal{
BackupRoot: cfg.Archive.LocalPath, BackupRoot: cfg.Archive.LocalPath,
Resolve: func(pvc string) (string, error) { Resolve: resolveWorldDir(ctx, cl, cfg.K8s.Namespace, worldsRoot),
return filepath.Join(worldsRoot, pvc), nil
},
}, nil }, nil
default: default:
return nil, fmt.Errorf("[archive] store %q is not implemented in this build (only tarLocal)", cfg.Archive.Store) return nil, fmt.Errorf("[archive] store %q is not implemented in this build (only tarLocal)", cfg.Archive.Store)
} }
} }
// resolveWorldDir maps a world PVC to its directory under worldsRoot, supporting
// the two layouts a Felis host actually has:
//
// 1. <root>/<pvc> — the reaper's documented arrangement (worlds exposed by PVC
// name, e.g. via mounting each volume or a crafted storage class).
// 2. <root>/<pv-name>_<namespace>_<pvc-name> — what a stock k3s install gets:
// local-path-provisioner stores every volume under its storage root as that
// exact directory name. Without this arm, retention on a default install could
// only ever fail to find a world (a no-op reaper, or worse an operator
// arranging paths by hand).
//
// The second path is derived EXACTLY from the live PVC's spec.volumeName, never
// from a glob: a leftover directory of an old, deleted PV must never be mistaken
// for the world the PVC currently binds, because the reaper archives the resolved
// directory and then deletes that PVC — archiving stale bytes and deleting the
// real world would be data loss. When neither path exists the first is returned,
// so the archive walk fails loudly against the documented path.
func resolveWorldDir(ctx context.Context, cl client.Client, namespace, worldsRoot string) backup.PVCResolver {
return func(pvc string) (string, error) {
direct := filepath.Join(worldsRoot, pvc)
if _, err := os.Stat(direct); err == nil {
return direct, nil
}
var claim corev1.PersistentVolumeClaim
if err := cl.Get(ctx, client.ObjectKey{Namespace: namespace, Name: pvc}, &claim); err != nil {
return "", fmt.Errorf("resolve world PVC %s: %w", pvc, err)
}
if pv := claim.Spec.VolumeName; pv != "" {
volDir := filepath.Join(worldsRoot, fmt.Sprintf("%s_%s_%s", pv, claim.Namespace, claim.Name))
if _, err := os.Stat(volDir); err == nil {
return volDir, nil
}
}
return direct, nil
}
}
// parseSpanDuration parses the human spans used in felis.toml's [archive] table: // parseSpanDuration parses the human spans used in felis.toml's [archive] table:
// "3mo" (months≈30d), "15d" (days), or any time.ParseDuration unit ("12h"). // "3mo" (months≈30d), "15d" (days), or any time.ParseDuration unit ("12h").
func parseSpanDuration(s string) (time.Duration, error) { func parseSpanDuration(s string) (time.Duration, error) {
+123
View File
@@ -0,0 +1,123 @@
package main
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client/fake"
)
// TestResolveWorldDir pins the two world layouts the reaper must find, and the
// fail-closed miss. The stock local-path arm is derived from the live PVC's
// volumeName — a name-based guess (glob) could tar a stale deleted PV's bytes and
// then delete the current world, which is why it is read from the API instead.
func TestResolveWorldDir(t *testing.T) {
ctx := context.Background()
root := t.TempDir()
// Arrange a world under the documented <root>/<pvc> layout.
named := filepath.Join(root, "world-named-0")
if err := os.MkdirAll(named, 0o750); err != nil {
t.Fatal(err)
}
// Arrange a second world the way k3s local-path stores it.
pvDir := filepath.Join(root, "pvc-11111111-2222-3333-4444-555555555555_minecraft_world-live-0")
if err := os.MkdirAll(pvDir, 0o750); err != nil {
t.Fatal(err)
}
claim := &corev1.PersistentVolumeClaim{
ObjectMeta: metav1.ObjectMeta{Name: "world-live-0", Namespace: "minecraft"},
Spec: corev1.PersistentVolumeClaimSpec{
VolumeName: "pvc-11111111-2222-3333-4444-555555555555",
},
}
cl := fake.NewClientBuilder().WithScheme(haltScheme(t)).WithObjects(claim).Build()
resolve := resolveWorldDir(ctx, cl, "minecraft", root)
t.Run("documented name layout wins", func(t *testing.T) {
got, err := resolve("world-named-0")
if err != nil || got != named {
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, named)
}
})
t.Run("stock local-path layout resolves exactly", func(t *testing.T) {
got, err := resolve("world-live-0")
if err != nil || got != pvDir {
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, pvDir)
}
})
t.Run("neither layout present falls back to the documented path", func(t *testing.T) {
// The claim exists but its directory does not: return the documented path so
// the archive walk fails there, and the reaper preserves the world.
missing := &corev1.PersistentVolumeClaim{
ObjectMeta: metav1.ObjectMeta{Name: "world-gone-0", Namespace: "minecraft"},
Spec: corev1.PersistentVolumeClaimSpec{VolumeName: "pvc-99999999-0000-0000-0000-000000000000"},
}
cl := fake.NewClientBuilder().WithScheme(haltScheme(t)).WithObjects(missing).Build()
got, err := resolveWorldDir(ctx, cl, "minecraft", root)("world-gone-0")
if err != nil || got != filepath.Join(root, "world-gone-0") {
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, filepath.Join(root, "world-gone-0"))
}
})
t.Run("unknown pvc is an error, not a guess", func(t *testing.T) {
_, err := resolve("world-unknown-0")
if err == nil || !strings.Contains(err.Error(), "resolve world PVC world-unknown-0") {
t.Fatalf("err = %v, want a resolve-world-PVC error", err)
}
})
}
// The pre-reap warner resolves the owner's VERIFIED email and hands the notice
// to the mailer. Every failure (no verified address, relay refusal) returns an
// error so the reaper retries on its next run instead of stamping a notice
// nobody received.
func TestMailWarner(t *testing.T) {
lookup := func(email string, err error) func(context.Context, string) (string, error) {
return func(context.Context, string) (string, error) { return email, err }
}
n := &captureNotifier{}
w := &mailWarner{lookupEmail: lookup("[email protected]", nil), notifier: n}
if err := w.Warn(context.Background(), "u1", "survival", "3d"); err != nil {
t.Fatalf("Warn: %v", err)
}
if n.email != "[email protected]" || !strings.Contains(n.subject, "survival") || !strings.Contains(n.subject, "3d") {
t.Fatalf("notice envelope = (%q, %q)", n.email, n.subject)
}
if !strings.Contains(n.body, "survival") || !strings.Contains(n.body, "3d") {
t.Fatalf("body missing server/remaining:\n%s", n.body)
}
w = &mailWarner{lookupEmail: lookup("", errors.New("owner u2 has no verified email")), notifier: n}
if err := w.Warn(context.Background(), "u2", "survival", "3d"); err == nil || !strings.Contains(err.Error(), "verified email") {
t.Fatalf("unverified owner = %v, want the lookup error surfaced", err)
}
w = &mailWarner{lookupEmail: lookup("[email protected]", nil), notifier: &captureNotifier{err: errors.New("relay down")}}
if err := w.Warn(context.Background(), "u1", "survival", "3d"); err == nil || !strings.Contains(err.Error(), "relay down") {
t.Fatalf("relay failure = %v, want it surfaced", err)
}
}
type captureNotifier struct {
email, subject, body string
err error
}
func (n *captureNotifier) SendNotice(_ context.Context, email, subject, body string) error {
if n.err != nil {
return n.err
}
n.email, n.subject, n.body = email, subject, body
return nil
}
+4
View File
@@ -19,9 +19,11 @@ Commands:
restore Extract a world archive into a world volume (internal Job entrypoint) restore Extract a world archive into a world volume (internal Job entrypoint)
backup Archive a world into the backup store and record it (internal Job entrypoint) backup Archive a world into the backup store and record it (internal Job entrypoint)
files List/read/write one file in a stopped server's world (internal Job entrypoint) files List/read/write one file in a stopped server's world (internal Job entrypoint)
fetch-context Fetch and extract a submission's build context (internal Job entrypoint)
manifests Render the control-plane RBAC + NetworkPolicy install bundle as YAML manifests Render the control-plane RBAC + NetworkPolicy install bundle as YAML
apply Create a MinecraftServer CRD (direct K8s write; use -f server.json) apply Create a MinecraftServer CRD (direct K8s write; use -f server.json)
setup Run host bootstrap + first-run setup console (TUI; requires root/sudo) setup Run host bootstrap + first-run setup console (TUI; requires root/sudo)
converge Fill in fields a newer desired spec added to already-installed system servers
version Print the build stamp of this binary version Print the build stamp of this binary
update Report which platform components have updates available update Report which platform components have updates available
breakGlass Open the local break-glass emergency console (TUI; requires root/sudo) breakGlass Open the local break-glass emergency console (TUI; requires root/sudo)
@@ -47,9 +49,11 @@ var commands = map[string]func(args []string, stdout, stderr io.Writer) int{
"restore": cmdRestore, "restore": cmdRestore,
"backup": cmdBackup, "backup": cmdBackup,
"files": cmdFiles, "files": cmdFiles,
"fetch-context": cmdFetchContext,
"manifests": cmdManifests, "manifests": cmdManifests,
"apply": cmdApply, "apply": cmdApply,
"setup": cmdSetup, "setup": cmdSetup,
"converge": cmdConverge,
"breakGlass": cmdBreakGlass, "breakGlass": cmdBreakGlass,
"bootstrap-assets": cmdBootstrapAssets, "bootstrap-assets": cmdBootstrapAssets,
"init-forwarding": cmdInitForwarding, "init-forwarding": cmdInitForwarding,
+37 -7
View File
@@ -225,16 +225,44 @@ func provisionSystemServers(ctx context.Context, cfg *config.Config, out io.Writ
// renamed it must replicate the Secret by hand. // renamed it must replicate the Secret by hand.
controlNS := platform.DefaultControlNamespace controlNS := platform.DefaultControlNamespace
apiBaseURL := platform.InternalAPIBaseURL(controlNS) apiBaseURL := platform.InternalAPIBaseURL(controlNS)
// Both Secrets must land in the minecraft namespace before the pods that mount // These Secrets must land in the minecraft namespace before the pods that
// them are created: the service token (login authenticates to felis-api with it) // mount them are created: the service token (login authenticates to felis-api
// and the Velocity forwarding secret (every backend verifies the proxy's signed // with it), the Velocity forwarding secret (every backend verifies the proxy's
// handshake with it — without it the login gate would derive an OFFLINE UUID and // signed handshake with it — without it the login gate would derive an OFFLINE
// the Owner would bind the wrong Minecraft identity). // UUID and the Owner would bind the wrong Minecraft identity), and felis-config
// (the on-demand BACKUP Job runs in the minecraft namespace and mounts it to
// self-record its world_backups row; without the replica the Job's volume
// mount fails and every backup request strands in the cluster).
// An empty build_namespace means the build system's compiled-in default; the
// replica must target the namespace the Jobs actually run in.
buildNS := cfg.Registry.BuildNamespace
if buildNS == "" {
buildNS = platform.DefaultBuildNamespace
}
secretOutcomes := []systemServerOutcome{ secretOutcomes := []systemServerOutcome{
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace, ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token"), naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "minecraft ns", false),
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace, ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
naming.ForwardingSecretName, naming.ForwardingSecretKey, "forwarding-secret"), naming.ForwardingSecretName, naming.ForwardingSecretKey, "forwarding-secret", "minecraft ns", false),
// refresh=true: felis-config is the rendered config, not a credential. The
// backup/restore/fileedit Jobs and the reaper mount this copy, so a re-run
// must update it when the control plane's render has moved on (a stale copy
// e.g. keeps an old database URL after a credential rotation).
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
"felis-config", "felis.toml", "config", "minecraft ns", true),
// The reaper's pre-reap warning emails authenticate with the same relay
// password felis-api uses; the reaper pod runs in the minecraft namespace,
// where a secretKeyRef resolves only against a local mirror. Skipped while
// the relay is not configured yet — the "configure email" screen refreshes
// both mirrors when it applies.
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
"felis-smtp", "password", "smtp", "minecraft ns", false),
// The build namespace needs the same token: the build Job's fetch
// initContainer reads the submission context from the internal face. Best
// effort — a deployment that only installs the control plane simply never
// builds a user submission.
ensureSecretReplica(ctx, cl, controlNS, buildNS,
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "felis-build ns", false),
} }
outcomes := ensureSystemServers(ctx, cl, cfg.K8s.Namespace, cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage, apiBaseURL, cfg.Server.RootDomain, defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname)) outcomes := ensureSystemServers(ctx, cl, cfg.K8s.Namespace, cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage, apiBaseURL, cfg.Server.RootDomain, defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname))
outcomes = append(secretOutcomes, outcomes...) outcomes = append(secretOutcomes, outcomes...)
@@ -245,6 +273,8 @@ func provisionSystemServers(ctx context.Context, cfg *config.Config, out io.Writ
fmt.Fprintf(out, " - %s: ERROR %v\n", o.name, o.err) fmt.Fprintf(out, " - %s: ERROR %v\n", o.name, o.err)
case o.created: case o.created:
fmt.Fprintf(out, " - %s: created (DesiredState=Running)\n", o.name) fmt.Fprintf(out, " - %s: created (DesiredState=Running)\n", o.name)
case o.updated:
fmt.Fprintf(out, " - %s: refreshed from the control namespace\n", o.name)
default: default:
fmt.Fprintf(out, " - %s: skipped (%s)\n", o.name, o.skipped) fmt.Fprintf(out, " - %s: skipped (%s)\n", o.name, o.skipped)
} }
+194 -35
View File
@@ -1,6 +1,7 @@
package main package main
import ( import (
"bytes"
"context" "context"
"fmt" "fmt"
"time" "time"
@@ -95,7 +96,7 @@ const felisLimboHealthPort int32 = 8080
// fail-safes to readiness-only, so a login pod that has the URL/domain but not yet // fail-safes to readiness-only, so a login pod that has the URL/domain but not yet
// the token is safe (it simply does not authenticate) rather than broken. // the token is safe (it simply does not authenticate) rather than broken.
const ( const (
envAPIBaseURL = "FELIS_API_BASE_URL" envAPIBaseURL = naming.EnvAPIBaseURL
envRootDomain = "FELIS_ROOT_DOMAIN" envRootDomain = "FELIS_ROOT_DOMAIN"
envPanelHostname = "FELIS_PANEL_HOSTNAME" envPanelHostname = "FELIS_PANEL_HOSTNAME"
envLobbyServer = "FELIS_LOBBY_SERVER" envLobbyServer = "FELIS_LOBBY_SERVER"
@@ -255,10 +256,31 @@ func buildSystemServerClient() (client.Client, error) {
// setup can report it without the provisioner deciding on the output format. // setup can report it without the provisioner deciding on the output format.
type systemServerOutcome struct { type systemServerOutcome struct {
name string name string
created bool // true = we created it this run created bool // true = we created it this run
available bool // true = the required object now exists updated bool // true = we refreshed an existing replica from the source
skipped string // non-empty = why it was skipped (image unset / already exists) available bool // true = the required object now exists
err error // non-nil = create failed skipped string // non-empty = why it was skipped (image unset / already exists)
err error // non-nil = create failed
changes []string // converge only: the fields this pass filled
}
// systemServerPlan is one system service in the provisioner's table: its name,
// the image config gives it, and the pure builder for its desired CR.
type systemServerPlan struct {
name string
image string
build func(image, namespace string) (*v1alpha1.MinecraftServer, error)
}
// systemServerPlans is the single description of the login+lobby pair, shared by
// ensureSystemServers (create-if-absent) and convergeSystemServers (field fill).
func systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerPlan {
return []systemServerPlan{
{name: naming.SystemLoginServer, image: loginImage, build: func(image, ns string) (*v1alpha1.MinecraftServer, error) {
return loginSystemServer(image, ns, apiBaseURL, rootDomain, panelHostname)
}},
{name: naming.SystemLobbyServer, image: lobbyImage, build: lobbySystemServer},
}
} }
// ensureSystemServers idempotently creates the login and lobby system services. // ensureSystemServers idempotently creates the login and lobby system services.
@@ -269,17 +291,7 @@ type systemServerOutcome struct {
// K8s client and namespace; this function performs no signal-handler or client // K8s client and namespace; this function performs no signal-handler or client
// setup of its own. // setup of its own.
func ensureSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome { func ensureSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome {
type plan struct { plans := systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname)
name string
image string
build func(image, namespace string) (*v1alpha1.MinecraftServer, error)
}
plans := []plan{
{name: naming.SystemLoginServer, image: loginImage, build: func(image, ns string) (*v1alpha1.MinecraftServer, error) {
return loginSystemServer(image, ns, apiBaseURL, rootDomain, panelHostname)
}},
{name: naming.SystemLobbyServer, image: lobbyImage, build: lobbySystemServer},
}
outcomes := make([]systemServerOutcome, 0, len(plans)) outcomes := make([]systemServerOutcome, 0, len(plans))
for _, p := range plans { for _, p := range plans {
@@ -379,12 +391,7 @@ var derivedSystemEnv = map[string]bool{
// deliberate removal is indistinguishable from drift and re-adding it would fight the // deliberate removal is indistinguishable from drift and re-adding it would fight the
// operator every run. // operator every run.
func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired *v1alpha1.MinecraftServer) (bool, error) { func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired *v1alpha1.MinecraftServer) (bool, error) {
want := make(map[string]string, len(derivedSystemEnv)) want := derivedEnvWanted(desired)
for _, e := range desired.Spec.Env {
if derivedSystemEnv[e.Name] {
want[e.Name] = e.Value
}
}
changed := false changed := false
for i, e := range existing.Spec.Env { for i, e := range existing.Spec.Env {
@@ -402,6 +409,116 @@ func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired
return true, nil return true, nil
} }
// derivedEnvWanted maps the derived env keys of desired onto their values.
func derivedEnvWanted(desired *v1alpha1.MinecraftServer) map[string]string {
want := make(map[string]string, len(derivedSystemEnv))
for _, e := range desired.Spec.Env {
if derivedSystemEnv[e.Name] {
want[e.Name] = e.Value
}
}
return want
}
// convergeSystemServers is the explicit convergence pass over already-installed
// system servers (#1). ensureSystemServers is create-if-absent by design — an
// existing CR is left alone so a re-run cannot clobber an operator's edits — and
// that leaves no path for a field the DESIRED spec gained after the install:
// spec.rcon (the lobby's write channel), spec.startup.healthHTTPPort (the login
// gate's readiness probe), or a config-derived env key that did not exist yet.
// Such fields sit at their zero value forever while re-running setup reports
// success, which is exactly the reported "configuration updates never reach an
// installed deployment" symptom.
//
// This pass fills exactly those zero-value fields and the config-derived env keys,
// and nothing else: a field already holding a non-zero value is the operator's and
// is never overwritten. It is an explicit command rather than an implicit step of
// setup because some fills need an ordering only the operator knows — enabling
// RCON or the HTTP readiness gate on a server whose image predates the listener
// would hold that server in Starting until it was marked Failed. Rebuild (or
// upgrade) the images first, then run this.
func convergeSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome {
outcomes := make([]systemServerOutcome, 0, 2)
for _, p := range systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname) {
if p.image == "" {
outcomes = append(outcomes, systemServerOutcome{name: p.name, skipped: "image not configured"})
continue
}
desired, err := p.build(p.image, namespace)
if err != nil {
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: err})
continue
}
var existing v1alpha1.MinecraftServer
switch err := cl.Get(ctx, client.ObjectKeyFromObject(desired), &existing); {
case apierrors.IsNotFound(err):
outcomes = append(outcomes, systemServerOutcome{name: p.name,
skipped: "not present — run `sudo felis setup` first"})
continue
case err != nil:
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: err})
continue
}
if existing.Labels[v1alpha1.LabelSystemRole] != p.name {
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: fmt.Errorf(
"existing MinecraftServer %s/%s is not marked as the Felis %q system role; refusing to converge it",
namespace, p.name, p.name,
)})
continue
}
var changes []string
if existing.Spec.Rcon == (v1alpha1.RconSpec{}) && desired.Spec.Rcon != (v1alpha1.RconSpec{}) {
existing.Spec.Rcon = desired.Spec.Rcon
changes = append(changes, "spec.rcon")
}
if existing.Spec.Startup.HealthHTTPPort == 0 && desired.Spec.Startup.HealthHTTPPort != 0 {
existing.Spec.Startup.HealthHTTPPort = desired.Spec.Startup.HealthHTTPPort
changes = append(changes, "spec.startup.healthHTTPPort")
}
changes = append(changes, convergeDerivedEnv(&existing, desired)...)
if len(changes) == 0 {
outcomes = append(outcomes, systemServerOutcome{name: p.name, available: true, skipped: "already converged"})
continue
}
if err := cl.Update(ctx, &existing); err != nil {
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: fmt.Errorf("converge %s: %w", p.name, err)})
continue
}
outcomes = append(outcomes, systemServerOutcome{name: p.name, available: true, updated: true, changes: changes})
}
return outcomes
}
// convergeDerivedEnv makes the config-derived env match the desired values: a key
// whose value drifted is overwritten, and a key missing entirely is added. This is
// the wider half of the same explicit pass — refreshDerivedEnv's present-only loop
// can never introduce a NEW key, which is how a derived key added after an install
// never reached it at all.
func convergeDerivedEnv(existing, desired *v1alpha1.MinecraftServer) []string {
want := derivedEnvWanted(desired)
var changes []string
present := make(map[string]bool, len(existing.Spec.Env))
for i := range existing.Spec.Env {
e := &existing.Spec.Env[i]
present[e.Name] = true
if v, ok := want[e.Name]; ok && v != e.Value {
e.Value = v
changes = append(changes, "env "+e.Name)
}
}
for _, e := range desired.Spec.Env {
if !derivedSystemEnv[e.Name] || present[e.Name] {
continue
}
existing.Spec.Env = append(existing.Spec.Env, e)
changes = append(changes, "env "+e.Name)
}
return changes
}
// The login gate is a hard prerequisite of the Owner bind, so setup waits for it // The login gate is a hard prerequisite of the Owner bind, so setup waits for it
// rather than racing it. The ceiling covers a cold image pull on a fresh node; // rather than racing it. The ceiling covers a cold image pull on a fresh node;
// the poll is fast enough that a warm start feels immediate. // the poll is fast enough that a warm start feels immediate.
@@ -471,26 +588,35 @@ func phaseOrPending(p v1alpha1.Phase) string {
return string(p) return string(p)
} }
// ensureSecretReplica copies one Secret from the control namespace into the minecraft // ensureSecretReplica copies one Secret from the control namespace into a workload
// namespace so a backend pod can mount it via secretKeyRef. A secretKeyRef is // namespace (minecraft — or the build namespace, whose fetch initContainer reads the
// namespace-local, but the backends run in the minecraft namespace while the sources // context from the felis-api internal face with the same token) so a pod can mount it
// of truth live beside the control plane — so without this replica the operator's // via secretKeyRef. A secretKeyRef is namespace-local, but those workloads do not run
// injected secretKeyRef would dangle and wedge the pod in CreateContainerConfigError. // beside the control plane — so without this replica the secretKeyRef would dangle and
// wedge the pod in CreateContainerConfigError.
// //
// Two Secrets need it, for different reasons: the service token (login only — it // Three Secrets need it, for different reasons: the service token (the login limbo and
// authenticates the limbo plugin to the felis-api internal face) and the Velocity // the build Pod's context fetch — both authenticate to the felis-api internal face),
// modern-forwarding secret (every backend — it is how a backend knows a login really // the Velocity modern-forwarding secret (every backend — it is how a backend knows
// came from the proxy, and so that the player's UUID is Mojang-verified rather than // a login really came from the proxy, and so that the player's UUID is Mojang-verified
// offline-derived). // rather than offline-derived), and the SMTP relay password (the reaper's pre-reap
// warning emails; the felis-config mirror is what carries [smtp] into its pod).
// //
// It is create-if-absent: an existing replica is left untouched so a hand-rotated // It is create-if-absent: an existing replica is left untouched so a hand-rotated
// value in the minecraft namespace is never clobbered (to rotate, delete the replica // value in the workload namespace is never clobbered (to rotate, delete the replica
// and re-run setup). Best-effort like the rest of the provisioner: a missing source or // and re-run setup). Best-effort like the rest of the provisioner: a missing source or
// a create failure degrades to a reported outcome, never a hard setup failure. It // a create failure degrades to a reported outcome, never a hard setup failure. It
// copies only Type and Data — never labels/annotations/ownerRefs — so the replica // copies only Type and Data — never labels/annotations/ownerRefs — so the replica
// carries no accidental GC owner or managed-by lineage. // carries no accidental GC owner or managed-by lineage.
func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace, minecraftNamespace, secretName, secretKey, label string) systemServerOutcome { //
name := label + " (minecraft ns)" // refreshExisting switches the felis-config mirror to refresh-in-place: that Secret is
// a rendered config, never a hand-rotated credential, and the workload Jobs that mount
// it (backup/restore/fileedit) plus the reaper silently misbehave on a stale copy —
// e.g. after a database credential rotation the control plane moves on while every
// backup Job keeps failing auth. Credential Secrets keep the never-overwrite rule so a
// rotated value survives; to rotate those, delete the replica and re-run setup.
func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace, minecraftNamespace, secretName, secretKey, label, where string, refreshExisting bool) systemServerOutcome {
name := label + " (" + where + ")"
validate := func(secret *corev1.Secret, location, skipped string) systemServerOutcome { validate := func(secret *corev1.Secret, location, skipped string) systemServerOutcome {
if len(secret.Data[secretKey]) == 0 { if len(secret.Data[secretKey]) == 0 {
return systemServerOutcome{name: name, skipped: fmt.Sprintf( return systemServerOutcome{name: name, skipped: fmt.Sprintf(
@@ -498,6 +624,33 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
} }
return systemServerOutcome{name: name, available: true, skipped: skipped} return systemServerOutcome{name: name, available: true, skipped: skipped}
} }
// refreshFromControl updates an existing replica from the control-namespace source
// when the rendered key differs. Only the felis-config mirror opts in.
refreshFromControl := func(existing *corev1.Secret) systemServerOutcome {
var src corev1.Secret
if err := cl.Get(ctx, client.ObjectKey{Namespace: controlNamespace, Name: secretName}, &src); err != nil {
if apierrors.IsNotFound(err) {
return systemServerOutcome{name: name, skipped: fmt.Sprintf(
"source Secret %s/%s not found — provision it (deploy/bootstrap.sh), then re-run setup",
controlNamespace, secretName)}
}
return systemServerOutcome{name: name, err: err}
}
if out := validate(&src, controlNamespace, ""); !out.available {
return out
}
if bytes.Equal(existing.Data[secretKey], src.Data[secretKey]) {
return validate(existing, minecraftNamespace, "already current")
}
if existing.Data == nil {
existing.Data = map[string][]byte{}
}
existing.Data[secretKey] = src.Data[secretKey]
if err := cl.Update(ctx, existing); err != nil {
return systemServerOutcome{name: name, err: err}
}
return systemServerOutcome{name: name, updated: true, available: true}
}
if controlNamespace == minecraftNamespace { if controlNamespace == minecraftNamespace {
// Same namespace needs no replica, but the source still has to exist. // Same namespace needs no replica, but the source still has to exist.
var existing corev1.Secret var existing corev1.Secret
@@ -516,6 +669,9 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
var existing corev1.Secret var existing corev1.Secret
getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing) getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing)
if getErr == nil { if getErr == nil {
if refreshExisting {
return refreshFromControl(&existing)
}
return validate(&existing, minecraftNamespace, "already exists") return validate(&existing, minecraftNamespace, "already exists")
} }
if !apierrors.IsNotFound(getErr) { if !apierrors.IsNotFound(getErr) {
@@ -544,6 +700,9 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
if getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing); getErr != nil { if getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing); getErr != nil {
return systemServerOutcome{name: name, err: getErr} return systemServerOutcome{name: name, err: getErr}
} }
if refreshExisting {
return refreshFromControl(&existing)
}
return validate(&existing, minecraftNamespace, "already exists") return validate(&existing, minecraftNamespace, "already exists")
} }
return systemServerOutcome{name: name, err: err} return systemServerOutcome{name: name, err: err}
+87 -6
View File
@@ -156,7 +156,7 @@ func TestEnsureSecretReplica(t *testing.T) {
} }
replicate := func(cl client.Client, controlNS, mcNS string) systemServerOutcome { replicate := func(cl client.Client, controlNS, mcNS string) systemServerOutcome {
return ensureSecretReplica(ctx, cl, controlNS, mcNS, return ensureSecretReplica(ctx, cl, controlNS, mcNS,
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token") naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "minecraft ns", false)
} }
t.Run("replicates when absent", func(t *testing.T) { t.Run("replicates when absent", func(t *testing.T) {
@@ -251,6 +251,87 @@ func TestEnsureSecretReplica(t *testing.T) {
}) })
} }
// The felis-config mirror is the one replica that must refresh: it is a rendered
// config, and a stale workload-side copy (backup/restore/fileedit Jobs, the reaper)
// misbehaves silently — a rotated database credential keeps the control plane moving
// while every backup Job keeps failing auth. Credential Secrets keep create-if-absent.
func TestEnsureSecretReplicaRefresh(t *testing.T) {
scheme := newSystemServerScheme(t)
ctx := context.Background()
configSecret := func(ns, body string) *corev1.Secret {
return &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: "felis-config", Namespace: ns},
Type: corev1.SecretTypeOpaque,
Data: map[string][]byte{"felis.toml": []byte(body)},
}
}
refresh := func(cl client.Client) systemServerOutcome {
return ensureSecretReplica(ctx, cl, "felis", "minecraft",
"felis-config", "felis.toml", "config", "minecraft ns", true)
}
replicaBody := func(t *testing.T, cl client.Client) string {
t.Helper()
var got corev1.Secret
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: "felis-config"}, &got); err != nil {
t.Fatalf("get replica: %v", err)
}
return string(got.Data["felis.toml"])
}
t.Run("refreshes a stale config replica", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("felis", "current"),
configSecret("minecraft", "stale"),
).Build()
out := refresh(cl)
if out.err != nil || !out.updated || !out.available {
t.Fatalf("outcome = %+v, want refreshed", out)
}
if got := replicaBody(t, cl); got != "current" {
t.Errorf("replica = %q, want current", got)
}
})
t.Run("leaves a current config replica alone", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("felis", "same"),
configSecret("minecraft", "same"),
).Build()
out := refresh(cl)
if out.err != nil || out.updated || !out.available || out.skipped != "already current" {
t.Fatalf("outcome = %+v, want already current", out)
}
})
t.Run("fills an empty-key replica", func(t *testing.T) {
empty := &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: "felis-config", Namespace: "minecraft"},
Data: map[string][]byte{},
}
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("felis", "current"), empty).Build()
out := refresh(cl)
if out.err != nil || !out.updated {
t.Fatalf("outcome = %+v, want refreshed", out)
}
if got := replicaBody(t, cl); got != "current" {
t.Errorf("replica = %q, want current", got)
}
})
t.Run("missing source degrades to a skip", func(t *testing.T) {
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
configSecret("minecraft", "stale")).Build()
out := refresh(cl)
if out.err != nil || out.updated || out.available || out.skipped == "" {
t.Fatalf("outcome = %+v, want skipped (source missing)", out)
}
if got := replicaBody(t, cl); got != "stale" {
t.Errorf("replica = %q, want untouched stale", got)
}
})
}
func TestRequiredProvisioningError(t *testing.T) { func TestRequiredProvisioningError(t *testing.T) {
ready := []systemServerOutcome{ ready := []systemServerOutcome{
{name: "service-token (minecraft ns)", available: true}, {name: "service-token (minecraft ns)", available: true},
@@ -497,7 +578,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
// human would have added and a memory bump off the built-in default. // human would have added and a memory bump off the built-in default.
stale := func() *v1alpha1.MinecraftServer { stale := func() *v1alpha1.MinecraftServer {
ms, err := loginSystemServer("felis-limbo:demo", "minecraft", ms, err := loginSystemServer("felis-limbo:demo", "minecraft",
"http://old.internal:8081", "159.223.32.51.nip.io", "console.159.223.32.51.nip.io") "http://old.internal:8081", "203.0.113.10.nip.io", "console.203.0.113.10.nip.io")
if err != nil { if err != nil {
t.Fatalf("build stale login server: %v", err) t.Fatalf("build stale login server: %v", err)
} }
@@ -509,7 +590,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
run := func(cl client.Client) []systemServerOutcome { run := func(cl client.Client) []systemServerOutcome {
return ensureSystemServers(ctx, cl, "minecraft", "felis-limbo:demo", "felis-lobby:demo", return ensureSystemServers(ctx, cl, "minecraft", "felis-limbo:demo", "felis-lobby:demo",
"http://felis-api-internal.felis.svc.cluster.local:8081", "http://felis-api-internal.felis.svc.cluster.local:8081",
"mc.flyemoji.network", "console.mc.flyemoji.network") "mc.example.net", "console.mc.example.net")
} }
envOf := func(t *testing.T, cl client.Client) map[string]string { envOf := func(t *testing.T, cl client.Client) map[string]string {
@@ -537,11 +618,11 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
} }
} }
env := envOf(t, cl) env := envOf(t, cl)
if env[envPanelHostname] != "console.mc.flyemoji.network" { if env[envPanelHostname] != "console.mc.example.net" {
t.Errorf("%s = %q — players are still being sent to the old console", t.Errorf("%s = %q — players are still being sent to the old console",
envPanelHostname, env[envPanelHostname]) envPanelHostname, env[envPanelHostname])
} }
if env[envRootDomain] != "mc.flyemoji.network" { if env[envRootDomain] != "mc.example.net" {
t.Errorf("%s = %q, want the new root domain", envRootDomain, env[envRootDomain]) t.Errorf("%s = %q, want the new root domain", envRootDomain, env[envRootDomain])
} }
}) })
@@ -568,7 +649,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
t.Run("reports no refresh when config already matches", func(t *testing.T) { t.Run("reports no refresh when config already matches", func(t *testing.T) {
fresh, err := loginSystemServer("felis-limbo:demo", "minecraft", fresh, err := loginSystemServer("felis-limbo:demo", "minecraft",
"http://felis-api-internal.felis.svc.cluster.local:8081", "http://felis-api-internal.felis.svc.cluster.local:8081",
"mc.flyemoji.network", "console.mc.flyemoji.network") "mc.example.net", "console.mc.example.net")
if err != nil { if err != nil {
t.Fatalf("build fresh login server: %v", err) t.Fatalf("build fresh login server: %v", err)
} }
+1 -1
View File
@@ -86,7 +86,7 @@ func (m *backupModel) loadCmd() tea.Cmd {
if err != nil { if err != nil {
return backupListMsg{err: fmt.Errorf("list servers: %w", err)} return backupListMsg{err: fmt.Errorf("list servers: %w", err)}
} }
return backupListMsg{cl: cl, servers: servers} return backupListMsg{cl: cl, servers: backupPickable(servers)}
} }
} }
+40 -1
View File
@@ -133,6 +133,11 @@ func writeConfig(path string, cfg *config.Config) error {
return os.Rename(tmpPath, path) return os.Rename(tmpPath, path)
} }
// applyFelisConfigSecret applies the rendered config to the control namespace and
// then converges the workload-namespace mirror best-effort. The mirror feeds the
// backup/restore/fileedit Jobs and the reaper; without this refresh a reconfigure
// here would leave those readers on the previous render until the next `felis
// setup` run (startup pass) or installer re-run.
func applyFelisConfigSecret(ctx context.Context) error { func applyFelisConfigSecret(ctx context.Context) error {
out, err := kubectlOutput(ctx, out, err := kubectlOutput(ctx,
"-n", "felis", "create", "secret", "generic", "felis-config", "-n", "felis", "create", "secret", "generic", "felis-config",
@@ -142,7 +147,41 @@ func applyFelisConfigSecret(ctx context.Context) error {
if err != nil { if err != nil {
return err return err
} }
return kubectlWithInput(ctx, out, "apply", "-f", "-") if err := kubectlWithInput(ctx, out, "apply", "-f", "-"); err != nil {
return err
}
if err := replicateFelisConfigToWorkloadNamespace(ctx); err != nil {
fmt.Fprintf(os.Stderr, "felis setup: warning: the control-plane config is applied, but the workload-namespace mirror could not be refreshed (%v); re-run felis setup once that is fixed\n", err)
}
return nil
}
// replicateFelisConfigToWorkloadNamespace overwrites the workload-namespace
// felis-config mirror with the freshly rendered pod config. Deliberately a full
// replace, not create-if-absent: a stale mirror is exactly what silently hands
// the Jobs that mount it old settings after a reconfigure. No-op when the
// workload namespace is unset or is the control namespace itself.
func replicateFelisConfigToWorkloadNamespace(ctx context.Context) error {
cfg, err := config.Load(hostSetupConfigPath)
if err != nil {
return err
}
ns := cfg.K8s.Namespace
if ns == "" || ns == "felis" {
return nil
}
manifest, err := kubectlOutput(ctx,
"-n", ns, "create", "secret", "generic", "felis-config",
"--from-file=felis.toml="+podSetupConfigPath,
"--dry-run=client", "-o", "yaml",
)
if err != nil {
return fmt.Errorf("render felis-config for %s: %w", ns, err)
}
if err := kubectlWithInput(ctx, manifest, "-n", ns, "apply", "-f", "-"); err != nil {
return fmt.Errorf("replicate felis-config to %s: %w", ns, err)
}
return nil
} }
func installCloudflaredService(ctx context.Context, cloudflaredBin, configPath string) error { func installCloudflaredService(ctx context.Context, cloudflaredBin, configPath string) error {
+5 -1
View File
@@ -207,7 +207,11 @@ func (m *mcBindModel) doneView() string {
if box.Len() > 0 { if box.Len() > 0 {
box.WriteString("\n") box.WriteString("\n")
} }
box.WriteString(tuiLabel.Render("setup URL ") + "\n" + tuiPassword.Render(m.setupTokenURL) + "\n\n") box.WriteString(tuiLabel.Render("setup URL ") + "\n")
for _, line := range wrapDisplayURL(m.setupTokenURL, 70) {
box.WriteString(tuiPassword.Render(line) + "\n")
}
box.WriteString("\n")
box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup.\nIt is shown only once.")) box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup.\nIt is shown only once."))
} }
if m.auditWarning != "" { if m.auditWarning != "" {
+34 -7
View File
@@ -151,19 +151,46 @@ func TestOwnerModelProvisionErrorRouting(t *testing.T) {
} }
}) })
t.Run("a conflict on the Owner path is not a retry", func(t *testing.T) { t.Run("the Owner seat refusal returns to the form naming the seat", func(t *testing.T) {
// Defensive: the Owner upserts and so never conflicts, but were one ever to // Upserting a fresh username while a seat is occupied would mint a second
// surface it must end the session rather than loop the form — only the // owner, so provisionOwner refuses with ownerSeatTakenError (Is
// insert-only operator path is retryable. // api.ErrConflict) and the console must route back for a retype — the same
// recoverable contract as the operator clash, and the only Owner-path
// conflict there is.
m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true) m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true)
seatErr := &ownerSeatTakenError{seat: "seat-holder"}
next, cmd := m.Update(owProvisionMsg{err: conflict}) next, cmd := m.Update(owProvisionMsg{err: seatErr})
om := next.(*ownerModel)
if om.step != owProvision {
t.Fatalf("step = %v, want owProvision — the seat refusal is recoverable", om.step)
}
if om.provisionErr == nil || !errors.Is(om.provisionErr, api.ErrConflict) || !strings.Contains(om.provisionErr.Error(), "seat-holder") {
t.Errorf("provisionErr = %v, want the seat refusal naming the seat", om.provisionErr)
}
// Feed the rebuilt form's init message back through so its view renders;
// then the note must carry the seat name (the operator's retype cue).
if cmd != nil {
if msg := cmd(); msg != nil {
if n2, _ := om.Update(msg); n2 != nil {
om = n2.(*ownerModel)
}
}
}
if view := om.form.View(); !strings.Contains(view, "seat-holder") {
t.Errorf("the provision form must surface the seat refusal:\n%s", view)
}
})
t.Run("a generic Owner-path fault still tears the console down", func(t *testing.T) {
m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true)
next, cmd := m.Update(owProvisionMsg{err: errors.New("boom")})
om := next.(*ownerModel) om := next.(*ownerModel)
if om.provisionErr != nil { if om.provisionErr != nil {
t.Error("the Owner path recorded a retryable conflict; only the operator path retries") t.Error("a generic fault must not be treated as a retryable refusal")
} }
if res, ok := cmd().(ownerResultMsg); !ok || res.err == nil { if res, ok := cmd().(ownerResultMsg); !ok || res.err == nil {
t.Error("an Owner-path conflict should tear down via an error result") t.Error("a generic Owner-path fault should tear down via an error result")
} }
}) })
} }
+21 -10
View File
@@ -174,12 +174,13 @@ func (m *ownerModel) Update(msg tea.Msg) (tea.Model, tea.Cmd) {
case owProvisionMsg: case owProvisionMsg:
if msg.err != nil { if msg.err != nil {
// A taken Operator username is the expected, recoverable outcome of the // api.ErrConflict marks the two recoverable refusals: a taken Operator
// insert-only operator path (refusing the clash is the whole reason it is // username (insert-only clash) and an Owner reset naming anything but the
// insert-only, not an upsert). Route back to the form with a note so the // occupied seat (ownerSeatTakenError Is ErrConflict). Route back to the
// operator can pick another name, rather than tearing down the console — // form with a note so the operator can retype, rather than tearing down
// any other error is a genuine fault and still ends the session. // the console — any other error is a genuine fault and still ends the
if m.operation == bgAddOperator && errors.Is(msg.err, api.ErrConflict) { // session.
if errors.Is(msg.err, api.ErrConflict) {
m.provisionErr = msg.err m.provisionErr = msg.err
m.step = owProvision m.step = owProvision
m.form = m.sized(m.buildProvisionForm()) m.form = m.sized(m.buildProvisionForm())
@@ -364,9 +365,15 @@ func (m *ownerModel) buildProvisionForm() *huh.Form {
} }
} }
if m.provisionErr != nil { if m.provisionErr != nil {
// The only error routed back to this form is a username clash on the insert-only // Recoverable refusals routed back here: the seat refusal already names the
// operator path; show a concrete prompt to choose another name. // username to enter, so show it verbatim; the operator-name clash gets the
desc = "That username is already taken — choose a different one.\n\n" + desc // generic retry prompt.
note := "That username is already taken — choose a different one."
var seatErr *ownerSeatTakenError
if errors.As(m.provisionErr, &seatErr) {
note = seatErr.Error()
}
desc = note + "\n\n" + desc
} }
fields := []huh.Field{ fields := []huh.Field{
@@ -420,7 +427,11 @@ func (m *ownerModel) doneView() string {
var box strings.Builder var box strings.Builder
box.WriteString(tuiLabel.Render("username ") + m.username + "\n") box.WriteString(tuiLabel.Render("username ") + m.username + "\n")
if m.setupTokenURL != "" { if m.setupTokenURL != "" {
box.WriteString("\n" + tuiLabel.Render("setup URL ") + "\n" + tuiPassword.Render(m.setupTokenURL) + "\n\n") box.WriteString("\n" + tuiLabel.Render("setup URL ") + "\n")
for _, line := range wrapDisplayURL(m.setupTokenURL, 70) {
box.WriteString(tuiPassword.Render(line) + "\n")
}
box.WriteString("\n")
box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup. It is shown only once.")) box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup. It is shown only once."))
} }
if m.auditWarning != "" { if m.auditWarning != "" {
+1
View File
@@ -553,6 +553,7 @@ func (m *rootModel) showSummary() (tea.Model, tea.Cmd) {
storageLabel: m.result.storageDetail, storageLabel: m.result.storageDetail,
routedHosts: routed, routedHosts: routed,
localHint: m.result.connectMethod == connectLocal, localHint: m.result.connectMethod == connectLocal,
alreadySetUp: m.result.alreadySetUp,
}) })
} }
+28
View File
@@ -249,6 +249,34 @@ func TestRootReconfigureSMTP(t *testing.T) {
} }
} }
// TestRootReconfigureStorageKeepsStatusFraming locks the same rule for the
// "change storage" path: on a re-run, completing it must land back on the
// alreadySetUp status framing (with the updated recap), not "Setup complete."
func TestRootReconfigureStorageKeepsStatusFraming(t *testing.T) {
m := newTestRoot(true, consoleModeSetup, "")
m = drive(t, m, preflightDoneMsg{})
if _, ok := m.screen.(*summaryModel); !ok {
t.Fatalf("re-run after preflight, screen = %T, want *summaryModel", m.screen)
}
m = drive(t, m, reconfigureStorageMsg{})
if _, ok := m.screen.(*storageChooserModel); !ok {
t.Fatalf("reconfigure-storage screen = %T, want *storageChooserModel", m.screen)
}
m = drive(t, m, storageResultMsg{method: storageLocal, detail: "local disk · /var/lib/felis/uploads"})
sum, ok := m.screen.(*summaryModel)
if !ok {
t.Fatalf("after reconfigure-storage, screen = %T, want *summaryModel", m.screen)
}
if !sum.alreadySetUp {
t.Fatalf("after reconfigure-storage, summary should keep the alreadySetUp framing")
}
if sum.storageLabel != "local disk · /var/lib/felis/uploads" {
t.Fatalf("storageLabel = %q, want the updated recap", sum.storageLabel)
}
}
func TestRootRerunLandsOnStatus(t *testing.T) { func TestRootRerunLandsOnStatus(t *testing.T) {
// adminExists at start of a setup run = re-run: preflight should skip straight // adminExists at start of a setup run = re-run: preflight should skip straight
// to the "manage in panel" status screen, never touching owner/connect. // to the "manage in panel" status screen, never touching owner/connect.
+51 -5
View File
@@ -4,6 +4,7 @@ import (
"context" "context"
"errors" "errors"
"fmt" "fmt"
"os"
"strconv" "strconv"
"strings" "strings"
@@ -338,20 +339,30 @@ func applySMTPConfig(ctx context.Context, in smtpInputs) error {
if err := applyFelisConfigSecret(ctx); err != nil { if err := applyFelisConfigSecret(ctx); err != nil {
return err return err
} }
// Refresh the workload-namespace copies too (the reaper's warning path): the
// OTP path is already live in the control namespace, so a replica miss is
// reported but not fatal.
if err := replicateSMTPToWorkloadNamespace(ctx, in.password); err != nil {
fmt.Fprintf(os.Stderr, "felis setup: warning: email is configured, but refreshing the workload copies failed (pre-reap warning emails may stay suppressed): %v\n", err)
}
if err := kubectl(ctx, "-n", "felis", "rollout", "restart", "deployment/felis-api"); err != nil { if err := kubectl(ctx, "-n", "felis", "rollout", "restart", "deployment/felis-api"); err != nil {
return err return err
} }
return kubectl(ctx, "-n", "felis", "rollout", "status", "deployment/felis-api", "--timeout=180s") return kubectl(ctx, "-n", "felis", "rollout", "status", "deployment/felis-api", "--timeout=180s")
} }
// applySMTPSecret creates (or replaces) the felis-smtp Secret the felis-api // smtpSecretManifest renders the felis-smtp Secret for the given namespace, the
// Deployment injects the relay password from. Rendered in-process and piped to // one the receiving Deployment/CronJob resolves its secretKeyRef against (felis
// for felis-api, the workload namespace for the reaper's mirror). The namespace
// must be IN the manifest: kubectl rejects a manifest whose namespace conflicts
// with -n, so leaving the control namespace hardcoded made every workload-ns
// replica fail before it started. Rendered in-process and piped to
// `kubectl apply` — the password is never a command-line arg, so it never // `kubectl apply` — the password is never a command-line arg, so it never
// appears in the host process table. // appears in the host process table.
func applySMTPSecret(ctx context.Context, password string) error { func smtpSecretManifest(password, namespace string) ([]byte, error) {
secret := &corev1.Secret{ secret := &corev1.Secret{
TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Secret"}, TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Secret"},
ObjectMeta: metav1.ObjectMeta{Name: platform.SMTPSecretName, Namespace: "felis"}, ObjectMeta: metav1.ObjectMeta{Name: platform.SMTPSecretName, Namespace: namespace},
Type: corev1.SecretTypeOpaque, Type: corev1.SecretTypeOpaque,
StringData: map[string]string{ StringData: map[string]string{
platform.SMTPSecretPasswordKey: password, platform.SMTPSecretPasswordKey: password,
@@ -359,7 +370,42 @@ func applySMTPSecret(ctx context.Context, password string) error {
} }
manifest, err := yaml.Marshal(secret) manifest, err := yaml.Marshal(secret)
if err != nil { if err != nil {
return fmt.Errorf("render smtp secret: %w", err) return nil, fmt.Errorf("render smtp secret: %w", err)
}
return manifest, nil
}
func applySMTPSecret(ctx context.Context, password string) error {
manifest, err := smtpSecretManifest(password, "felis")
if err != nil {
return err
} }
return kubectlWithInput(ctx, manifest, "apply", "-f", "-") return kubectlWithInput(ctx, manifest, "apply", "-f", "-")
} }
// replicateSMTPToWorkloadNamespace refreshes the workload-namespace (minecraft)
// copy of felis-smtp after email is reconfigured. The reaper's CronJob runs
// there and resolves the password by local reference — a secretKeyRef is
// namespace-local — so without this refresh a later SMTP change would never
// reach the pre-reap warning emails. Deliberately OVERWRITES: this is a mirror
// of the control-namespace source, and a stale mirror is exactly the failure
// this closes. The felis-config mirror rides along in applyFelisConfigSecret,
// which every apply path refreshes.
func replicateSMTPToWorkloadNamespace(ctx context.Context, password string) error {
cfg, err := config.Load(hostSetupConfigPath)
if err != nil {
return err
}
ns := cfg.K8s.Namespace
if ns == "" || ns == "felis" {
return nil
}
smtpManifest, err := smtpSecretManifest(password, ns)
if err != nil {
return err
}
if err := kubectlWithInput(ctx, smtpManifest, "-n", ns, "apply", "-f", "-"); err != nil {
return fmt.Errorf("replicate %s to %s: %w", platform.SMTPSecretName, ns, err)
}
return nil
}
+37
View File
@@ -0,0 +1,37 @@
package main
import (
"strings"
"testing"
"sigs.k8s.io/yaml"
)
// TestSMTPSecretManifestCarriesTargetNamespace pins the fix for the
// workload-namespace replica: kubectl refuses a manifest whose namespace
// conflicts with -n ("the namespace from the provided object ... does not
// match"), so the mirror must render felis-smtp with the TARGET namespace —
// otherwise the "configure email" refresh fails on the first apply and the
// felis-config mirror never runs at all.
func TestSMTPSecretManifestCarriesTargetNamespace(t *testing.T) {
for _, ns := range []string{"felis", "minecraft"} {
b, err := smtpSecretManifest("pw", ns)
if err != nil {
t.Fatalf("render for %s: %v", ns, err)
}
var got struct {
Metadata struct {
Namespace string `json:"namespace"`
} `json:"metadata"`
}
if err := yaml.Unmarshal(b, &got); err != nil {
t.Fatalf("unmarshal for %s: %v", ns, err)
}
if got.Metadata.Namespace != ns {
t.Fatalf("manifest namespace = %q, want %q", got.Metadata.Namespace, ns)
}
if !strings.Contains(string(b), "name: felis-smtp") {
t.Fatalf("manifest must still name felis-smtp: %s", b)
}
}
}
+27
View File
@@ -29,6 +29,33 @@ func tuiSeparator() string {
return tuiHint.Render(strings.Repeat("─", 70)) return tuiHint.Render(strings.Repeat("─", 70))
} }
// wrapDisplayURL breaks a long URL into lines no wider than width so the TUI
// renderer never truncates it on a narrow terminal — the one-time setup URL
// carries a 43-char token and overruns 80 columns. It prefers breaking right
// after a '=' or '/' inside the window (the token then lands on its own line)
// and hard-wraps only when no boundary is available. Lines concatenate back to
// the original string.
func wrapDisplayURL(u string, width int) []string {
if width <= 0 {
width = 70
}
var lines []string
for len(u) > width {
cut := width
if i := strings.LastIndexByte(u[:width], '='); i >= 0 && i >= width/2 {
cut = i + 1
} else if i := strings.LastIndexByte(u[:width], '/'); i >= 0 && i >= width/2 {
cut = i + 1
}
lines = append(lines, u[:cut])
u = u[cut:]
}
if u != "" {
lines = append(lines, u)
}
return lines
}
// tuiStepRail renders a breadcrumb of wizard stages. Steps before `current` // tuiStepRail renders a breadcrumb of wizard stages. Steps before `current`
// render as done, `current` is highlighted, and later steps are dimmed. // render as done, `current` is highlighted, and later steps are dimmed.
func tuiStepRail(steps []string, current int) string { func tuiStepRail(steps []string, current int) string {
+26
View File
@@ -0,0 +1,26 @@
package main
import (
"strings"
"testing"
)
func TestWrapDisplayURL(t *testing.T) {
u := "https://op.console.example.net/setup?token=" + strings.Repeat("A", 43)
lines := wrapDisplayURL(u, 70)
if got := strings.Join(lines, ""); got != u {
t.Fatalf("concatenated lines = %q, want the original URL back", got)
}
for i, l := range lines {
if len(l) > 70 {
t.Errorf("line %d is %d cols wide: %q", i, len(l), l)
}
}
if len(lines) < 2 || !strings.HasSuffix(lines[0], "token=") {
t.Fatalf("want the first line to end at the 'token=' boundary, got %q", lines)
}
short := "https://a/b"
if got := wrapDisplayURL(short, 70); len(got) != 1 || got[0] != short {
t.Errorf("short URL should pass through unsplit, got %q", got)
}
}
+29 -25
View File
@@ -24,8 +24,9 @@ const updateTimeout = 60 * time.Second
// The apply side is deliberately NOT implemented in this command. Every component // The apply side is deliberately NOT implemented in this command. Every component
// here is installed by deploy/bootstrap.sh, which is idempotent, already handles the // here is installed by deploy/bootstrap.sh, which is idempotent, already handles the
// parts that are easy to get wrong (Velocity's pinned MINOR, the atomic jar install, // parts that are easy to get wrong (Velocity's pinned MINOR, the atomic jar install,
// the k3s image re-import that a byte-identical StatefulSet template will not // the image re-import + registry push that a byte-identical StatefulSet template
// trigger on its own), and is the path that gets exercised on every install. A // will not trigger on its own), and is the path that gets exercised on every
// install. A
// second installer living in this file would duplicate that policy, could drift from // second installer living in this file would duplicate that policy, could drift from
// it silently, and would be reachable only on a live node where a mistake takes the // it silently, and would be reachable only on a live node where a mistake takes the
// proxy or the control plane down. So `felis update` reports, and hands the operator // proxy or the control plane down. So `felis update` reports, and hands the operator
@@ -42,6 +43,18 @@ type updateTarget struct {
command string command string
} }
// installerRerun is the tested apply path for every planner-backed selector: re-run the
// installer. It is idempotent, and it is the only path that fetches a newer version --
// `felis setup` skips its host-bootstrap phase on a completed install (all four install
// markers already exist), so there it opens the config console and moves no component,
// and even on the bootstrap path it re-images felis-api from the binary setup is already
// running (FELIS_BOOTSTRAP_BINARY), which looks like an update and changes nothing.
//
// The URL is the same one-liner both READMEs hand out. While the repo is private it
// answers 404 (raw.githubusercontent.com hides private repos), which is why the trailer
// below points at the README's token'd form for that case.
const installerRerun = "curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash"
// updateTargets is the selector table. panel and plugins both resolve to felis-api // updateTargets is the selector table. panel and plugins both resolve to felis-api
// because they are not separately versioned: the panel is compiled into the felis // because they are not separately versioned: the panel is compiled into the felis
// binary with //go:embed, and the plugin jars are built from this same repo in the // binary with //go:embed, and the plugin jars are built from this same repo in the
@@ -51,19 +64,19 @@ var updateTargets = []updateTarget{
selector: "panel", selector: "panel",
component: "felis-api", component: "felis-api",
note: "the panel is embedded in the felis binary (//go:embed), so updating it means rebuilding the felis image and rolling felis-api", note: "the panel is embedded in the felis binary (//go:embed), so updating it means rebuilding the felis image and rolling felis-api",
command: "sudo felis setup", command: installerRerun,
}, },
{ {
selector: "velocity", selector: "velocity",
component: "velocity", component: "velocity",
note: "re-runs install_velocity: newest BUILD of the pinned minor (FELIS_VELOCITY_VERSION), atomic jar install, then restarts felis-velocity", note: "re-runs install_velocity: newest BUILD of the pinned minor (FELIS_VELOCITY_VERSION), atomic jar install, then restarts felis-velocity",
command: "sudo felis setup", command: installerRerun,
}, },
{ {
selector: "plugins", selector: "plugins",
component: "felis-api", component: "felis-api",
note: "felis-velocity.jar is a host-file swap, but felis-paper.jar and felis-limbo.jar are baked into the lobby/limbo images and need a rebuild + k3s image re-import", note: "felis-velocity.jar is a host-file swap, but felis-paper.jar and felis-limbo.jar are baked into the lobby/limbo images and need a rebuild + re-mirror into the in-cluster registry (the installer re-run does both)",
command: "sudo felis setup", command: installerRerun,
}, },
{ {
selector: "mc", selector: "mc",
@@ -219,7 +232,6 @@ func renderApplyGuidance(res updater.Result, selected map[string]bool, force boo
var b strings.Builder var b strings.Builder
var offeredCommand bool var offeredCommand bool
var offeredFelisAPI bool
for _, t := range updateTargets { for _, t := range updateTargets {
if !selected[t.selector] { if !selected[t.selector] {
continue continue
@@ -248,28 +260,20 @@ func renderApplyGuidance(res updater.Result, selected map[string]bool, force boo
} }
fmt.Fprintf(&b, " run: %s\n", t.command) fmt.Fprintf(&b, " run: %s\n", t.command)
offeredCommand = true offeredCommand = true
offeredFelisAPI = offeredFelisAPI || t.component == "felis-api"
} }
// Only explain the command when one was actually offered; a --mc-only run has // Only explain the command when one was actually offered; a --mc-only run has
// nothing to run and the trailer would be a non-sequitur. // nothing to run and the trailer would be a non-sequitur.
if offeredCommand {
b.WriteString("\nfelis setup is idempotent and re-runs the installer that owns these components;\nit does not reinstall what is already current. Restart game servers afterwards.\n")
}
// Scoped to felis-api because it is the only component setup cannot move forward.
// velocity is fine: install_velocity re-resolves the newest build of the pinned minor
// on every run. But setup hands deploy/bootstrap.sh the binary it is itself running
// (FELIS_BOOTSTRAP_BINARY), and that arm skips the release lookup entirely, so it
// rebuilds the image and rolls the deployment from the SAME binary -- a run that looks
// like a successful update and leaves the version unchanged.
// //
// The installer is the only thing that moves felis-api. It is safe to point at now // One trailer serves every selector now: setup is not an apply path at all on a
// that detect_node_ip reuses the installed root domain, so what is left to warn about // completed install (shouldRunHostBootstrapBeforeConfig only enters the host
// is the channel: FELIS_VERSION_BOOTSTRAP is not persisted anywhere and defaults to // bootstrap while an install marker is missing), so the installer re-run is the one
// release, so a bare re-run on a host tracking main quietly moves it onto releases. // worked path for all three components and there is no per-component exception left
// That is a channel change, not a broken install, which is why it is one clause and // to scope. Two caveats stay because following the advice without them bites real
// not a paragraph. // hosts: the channel is not persisted anywhere (a bare re-run on a main host quietly
if offeredFelisAPI { // moves it onto releases), and the private repo's one-liner needs the read token
b.WriteString("\nfelis-api (panel, plugins) is the exception: setup re-images it from the felis binary\nalready on this host, so it cannot install a NEWER felis-api. Re-run the bootstrap\ninstaller for that -- it keeps this install's root domain. It does default to the\nrelease channel, so pass FELIS_VERSION_BOOTSTRAP=dev if this host tracks main.\n") // back in the environment before it can resolve anything.
if offeredCommand {
b.WriteString("\nRe-running the installer applies everything above: it fetches the newest version on\nthe channel in effect and re-applies the bundle (release is the default). The channel\nis not persisted, so pass FELIS_VERSION_BOOTSTRAP=dev if this host tracks main. While\nthis repo is private, the one-liner above 404s without a token; the README's install\nsection has the token'd form that works. felis setup is not this path: on a completed\ninstall it opens the config console and installs nothing newer. Restart game servers\nafterwards.\n")
} }
return b.String() return b.String()
} }
+20 -25
View File
@@ -107,7 +107,7 @@ func TestApplyGuidanceMinecraftOffersNoCommand(t *testing.T) {
if !strings.Contains(out, "pinned by policy") { if !strings.Contains(out, "pinned by policy") {
t.Fatalf("want the pin explained:\n%s", out) t.Fatalf("want the pin explained:\n%s", out)
} }
if strings.Contains(out, "run:") || strings.Contains(out, "felis setup is idempotent") { if strings.Contains(out, "run:") || strings.Contains(out, "Re-running the installer") {
t.Fatalf("--mc must offer no command and no command trailer:\n%s", out) t.Fatalf("--mc must offer no command and no command trailer:\n%s", out)
} }
} }
@@ -133,42 +133,37 @@ func TestUpdateTargetsMatchTopology(t *testing.T) {
} }
} }
// `sudo felis setup` is the right answer for velocity and the wrong one for felis-api, // Re-running the installer is the one apply path this table may hand out. setup is NOT an
// so the caveat has to be scoped rather than appended to every run. setup hands // updater on a completed install -- its host-bootstrap phase only runs while an install
// bootstrap the binary it is already running, and that arm skips the release lookup: // marker is missing, so it opens the config console and moves no component -- and even on
// the run rebuilds the image and rolls the deployment off the SAME binary, which looks // the bootstrap path it re-images felis-api from the binary setup is already running. The
// like a successful update and changes nothing. install_velocity, by contrast, really // table used to answer with "sudo felis setup" and scope a felis-api-only exception; both
// does re-resolve the newest build on every run. // taught a model that does not survive contact with an installed host.
func TestApplyGuidanceScopesTheFelisAPICaveat(t *testing.T) { func TestApplyGuidancePointsEveryComponentAtTheInstaller(t *testing.T) {
const caveat = "cannot install a NEWER felis-api"
api := renderApplyGuidance( api := renderApplyGuidance(
planResult([]updates.Action{{Component: "felis-api", Kind: updates.ActionNotify, LatestKnown: true}}), planResult([]updates.Action{{Component: "felis-api", Kind: updates.ActionNotify, LatestKnown: true}}),
map[string]bool{"panel": true}, false) map[string]bool{"panel": true}, false)
if !strings.Contains(api, caveat) { for _, want := range []string{"deploy/bootstrap.sh", "FELIS_VERSION_BOOTSTRAP=dev", "felis setup is not this path"} {
t.Fatalf("--panel resolves to felis-api and must carry the caveat:\n%s", api) if !strings.Contains(api, want) {
t.Fatalf("--panel guidance missing %q:\n%s", want, api)
}
} }
// Naming the installer obliges us to name what a bare re-run still changes. The domain if strings.Contains(api, "run: sudo felis setup") {
// is handled -- detect_node_ip reuses the installed one -- but the channel is not t.Fatalf("setup must never be offered as the apply command:\n%s", api)
// persisted at all and defaults to release, so a host tracking main gets moved onto
// releases by following this advice.
if !strings.Contains(api, "FELIS_VERSION_BOOTSTRAP=dev") {
t.Fatalf("pointing at the installer without the channel caveat misleads a dev host:\n%s", api)
} }
// The same path serves velocity; a scoped caveat would re-teach the old model that
// setup fixes velocity.
vel := renderApplyGuidance( vel := renderApplyGuidance(
planResult([]updates.Action{{Component: "velocity", Kind: updates.ActionNotify, LatestKnown: true}}), planResult([]updates.Action{{Component: "velocity", Kind: updates.ActionNotify, LatestKnown: true}}),
map[string]bool{"velocity": true}, false) map[string]bool{"velocity": true}, false)
if strings.Contains(vel, caveat) { if !strings.Contains(vel, "run: curl -fsSL") || !strings.Contains(vel, "felis setup is not this path") {
t.Fatalf("velocity IS fixed by setup; the caveat would misdirect the operator:\n%s", vel) t.Fatalf("velocity gets the same installer path:\n%s", vel)
}
if !strings.Contains(vel, "felis setup is idempotent") {
t.Fatalf("velocity still wants the ordinary trailer:\n%s", vel)
} }
// --mc offers no command at all, so neither trailer belongs. // --mc offers no command at all, so neither trailer belongs.
mc := renderApplyGuidance(planResult(nil), map[string]bool{"mc": true}, true) mc := renderApplyGuidance(planResult(nil), map[string]bool{"mc": true}, true)
if strings.Contains(mc, caveat) { if strings.Contains(mc, "deploy/bootstrap.sh") || strings.Contains(mc, "FELIS_VERSION_BOOTSTRAP") {
t.Fatalf("--mc offers no command; the caveat is a non-sequitur:\n%s", mc) t.Fatalf("--mc offers no command; the trailer is a non-sequitur:\n%s", mc)
} }
} }
+69
View File
@@ -0,0 +1,69 @@
# Felis alert rules — plain Prometheus format (also the promtool-tested source
# for felis-prometheusrule.yaml). See docs/troubleshooting.md §14 for scraping
# and loading instructions.
#
# felis_* series come from two processes:
# - felis-operator pod :8080/metrics → felis_servers_total, felis_start_duration_seconds
# - felis-api internal :8081/metrics → felis_image_build_failures_total
# node_* / kube_* series come from node-exporter / kube-state-metrics.
groups:
- name: felis.rules
rules:
- alert: FelisImageBuildFailures
expr: increase(felis_image_build_failures_total[6h]) > 0
for: 5m
labels:
severity: warning
annotations:
summary: "modpack/image build failed in the last 6h"
description: >-
felis_image_build_failures_total increased. Inspect the failed build Job
(kubectl logs -n felis-build job/<build-job>); the same error text is on
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
- alert: FelisSlowServerStarts
expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300
for: 15m
labels:
severity: warning
annotations:
summary: "p90 server start time exceeds 5 minutes"
description: >-
Starts regularly take over five minutes (felis_start_duration_seconds,
observed when readiness is first reached). A start that never completes
records nothing — cross-check desiredState=Running servers with no ready
phase (troubleshooting §1).
- name: felis.node.rules
rules:
- alert: FelisNodeDiskSpaceLow
expr: >-
node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
/ node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15
for: 15m
labels:
severity: warning
annotations:
summary: "node filesystem {{ $labels.mountpoint }} below 15% available"
description: >-
Sustained disk pressure evicts game pods and garbage-collects images
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
- alert: FelisNodeDiskPressure
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
for: 5m
labels:
severity: critical
annotations:
summary: "kubelet reports DiskPressure on {{ $labels.node }}"
description: >-
The eviction chain is in progress: control-plane pods hold
system-cluster-critical and survive, game pods do not. Free disk now
(troubleshooting §13b).
- alert: FelisNodeMemoryLow
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
for: 15m
labels:
severity: warning
annotations:
summary: "node memory available below 10% for 15m"
description: >-
PostgreSQL, the control plane, the registry and game servers share one
node; sustained memory pressure risks OOM kills.
+107
View File
@@ -0,0 +1,107 @@
# promtool unit tests: `promtool test rules felis-alerts_test.yml`
# Proves every shipped rule actually fires on its target condition (and stays
# silent before it).
rule_files:
- felis-alerts.yaml
evaluation_interval: 1m
tests:
- name: build failure and slow starts
interval: 1m
input_series:
# counter: quiet for 5m, then one failure per step.
- series: 'felis_image_build_failures_total'
values: '0x5 1x15'
# histogram: all observations land in the (300,600] bucket.
- series: 'felis_start_duration_seconds_bucket{le="120"}'
values: '0x22'
- series: 'felis_start_duration_seconds_bucket{le="300"}'
values: '0x22'
- series: 'felis_start_duration_seconds_bucket{le="600"}'
values: '0+10x21'
- series: 'felis_start_duration_seconds_bucket{le="+Inf"}'
values: '0+10x21'
alert_rule_test:
- eval_time: 2m
alertname: FelisImageBuildFailures
exp_alerts: []
- eval_time: 20m
alertname: FelisImageBuildFailures
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "modpack/image build failed in the last 6h"
description: >-
felis_image_build_failures_total increased. Inspect the failed build Job
(kubectl logs -n felis-build job/<build-job>); the same error text is on
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
- eval_time: 20m
alertname: FelisSlowServerStarts
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "p90 server start time exceeds 5 minutes"
description: >-
Starts regularly take over five minutes (felis_start_duration_seconds,
observed when readiness is first reached). A start that never completes
records nothing — cross-check desiredState=Running servers with no ready
phase (troubleshooting §1).
- name: node disk and memory thresholds
interval: 1m
input_series:
- series: 'node_filesystem_avail_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
values: '10x26'
- series: 'node_filesystem_size_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
values: '100x26'
- series: 'kube_node_status_condition{condition="DiskPressure",node="n1",status="true"}'
values: '0x4 1x22'
- series: 'node_memory_MemAvailable_bytes{instance="node1",job="node-exporter"}'
values: '5x26'
- series: 'node_memory_MemTotal_bytes{instance="node1",job="node-exporter"}'
values: '100x26'
alert_rule_test:
- eval_time: 2m
alertname: FelisNodeDiskPressure
exp_alerts: []
- eval_time: 20m
alertname: FelisNodeDiskSpaceLow
exp_alerts:
- exp_labels:
device: /dev/vda1
fstype: xfs
instance: node1
job: node-exporter
mountpoint: /
severity: warning
exp_annotations:
summary: "node filesystem / below 15% available"
description: >-
Sustained disk pressure evicts game pods and garbage-collects images
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
- eval_time: 20m
alertname: FelisNodeDiskPressure
exp_alerts:
- exp_labels:
condition: DiskPressure
node: n1
status: "true"
severity: critical
exp_annotations:
summary: "kubelet reports DiskPressure on n1"
description: >-
The eviction chain is in progress: control-plane pods hold
system-cluster-critical and survive, game pods do not. Free disk now
(troubleshooting §13b).
- eval_time: 20m
alertname: FelisNodeMemoryLow
exp_alerts:
- exp_labels:
instance: node1
job: node-exporter
severity: warning
exp_annotations:
summary: "node memory available below 10% for 15m"
description: >-
PostgreSQL, the control plane, the registry and game servers share one
node; sustained memory pressure risks OOM kills.
+74
View File
@@ -0,0 +1,74 @@
# prometheus-operator twin of felis-alerts.yaml (kube-prometheus-stack loads
# rules through the PrometheusRule CRD, not rule_files). The plain file is the
# promtool-tested source; keep the groups in sync.
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: felis-alerts
namespace: monitoring
labels:
# Change to match your stack's ruleSelector (kube-prometheus-stack's
# default selects on the Helm release name).
release: kube-prometheus-stack
spec:
groups:
- name: felis.rules
rules:
- alert: FelisImageBuildFailures
expr: increase(felis_image_build_failures_total[6h]) > 0
for: 5m
labels:
severity: warning
annotations:
summary: "modpack/image build failed in the last 6h"
description: >-
felis_image_build_failures_total increased. Inspect the failed build Job
(kubectl logs -n felis-build job/<build-job>); the same error text is on
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
- alert: FelisSlowServerStarts
expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300
for: 15m
labels:
severity: warning
annotations:
summary: "p90 server start time exceeds 5 minutes"
description: >-
Starts regularly take over five minutes (felis_start_duration_seconds,
observed when readiness is first reached). A start that never completes
records nothing — cross-check desiredState=Running servers with no ready
phase (troubleshooting §1).
- name: felis.node.rules
rules:
- alert: FelisNodeDiskSpaceLow
expr: >-
node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
/ node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15
for: 15m
labels:
severity: warning
annotations:
summary: "node filesystem {{ $labels.mountpoint }} below 15% available"
description: >-
Sustained disk pressure evicts game pods and garbage-collects images
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
- alert: FelisNodeDiskPressure
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
for: 5m
labels:
severity: critical
annotations:
summary: "kubelet reports DiskPressure on {{ $labels.node }}"
description: >-
The eviction chain is in progress: control-plane pods hold
system-cluster-critical and survive, game pods do not. Free disk now
(troubleshooting §13b).
- alert: FelisNodeMemoryLow
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
for: 15m
labels:
severity: warning
annotations:
summary: "node memory available below 10% for 15m"
description: >-
PostgreSQL, the control plane, the registry and game servers share one
node; sustained memory pressure risks OOM kills.
+555 -99
View File
File diff suppressed because it is too large. Load diff
+879
View File
@@ -61,6 +61,885 @@ expect "an uppercase digest is the same digest" "LOG: installing the Felis-Legac
out="$(run_gate "$(printf '%s' "$want" | sed 's/../& /g')")" out="$(run_gate "$(printf '%s' "$want" | sed 's/../& /g')")"
expect "a space-separated digest is the same digest" "LOG: installing the Felis-Legacy Velocity fork" "$out" expect "a space-separated digest is the same digest" "LOG: installing the Felis-Legacy Velocity fork" "$out"
# --- papermc_latest_jar answers "url sha256" from one response --------------------------
# Fill's download URLs are content-addressed (/v1/objects/<sha256>/<name>.jar), and the
# resolver's contract is to hand both halves back from the same grep — or refuse a URL
# that carries no digest, rather than wave the download through unchecked. Run under
# bash, not sh: bootstrap.sh is bash and the function uses $'\n'.
fn="$(awk '/^papermc_latest_jar\(\)/,/^}/' "$BS")"
[ -n "$fn" ] || { echo "FAIL: no papermc_latest_jar in $BS"; exit 1; }
[ "$(printf '%s\n' "$fn" | wc -l)" -lt 30 ] \
|| { echo "FAIL: the extracted papermc_latest_jar is not just the function -- did its closing brace move?"; exit 1; }
rsha=0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef
run_resolver() { # canned-fill-response
CANNED="$1" bash -c '
curl() { printf "%s" "$CANNED"; }
'"$fn"'
if out="$(papermc_latest_jar velocity 3.5.1)"; then
printf "RESOLVED %s\n" "$out"
else
printf "REFUSED\n"
fi
'
}
out="$(run_resolver "{\"url\":\"https://fill-data.papermc.io/v1/objects/${rsha}/velocity-3.5.1-615.jar\"}")"
expect "the resolver pairs the url with its own digest" \
"RESOLVED https://fill-data.papermc.io/v1/objects/${rsha}/velocity-3.5.1-615.jar ${rsha}" "$out"
out="$(run_resolver '{"url":"https://fill-data.papermc.io/mirror/velocity-3.5.1-615.jar"}')"
expect "a URL that carries no digest is refused" "REFUSED" "$out"
# --- the resolved-Velocity digest gate --------------------------------------------------
# The download must hash to what the content-addressed URL promised, BEFORE
# atomic_install_file — the same refusal the Via plugins and the fork jar already get.
# The end pattern spells ${VELOCITY_DIR} with dots: escaped braces are literal in gawk
# and mawk but undefined in POSIX awk, and CI's awk is whatever ubuntu ships.
vblock="$(awk '/log "resolving the newest Velocity/,/atomic_install_file "\$tmp" "\$.VELOCITY_DIR.\/velocity\.jar"/' "$BS")"
[ -n "$vblock" ] || { echo "FAIL: no resolved-Velocity install block found in $BS"; exit 1; }
[ "$(printf '%s\n' "$vblock" | wc -l)" -lt 30 ] \
|| { echo "FAIL: the extracted block is not the velocity install -- did its last line move?"; exit 1; }
vdir="$(mktemp -d)"
trap 'rm -f "$jar"; rm -rf "$vdir"' EXIT
vwant="$(printf 'stand-in velocity build\n' | sha256sum | cut -d' ' -f1)"
run_velocity_install() { # digest-the-resolver-reports
WANT="$1" VELOCITY_DIR="$vdir" FELIS_VELOCITY_VERSION=3.5.1 bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { printf "LOG: %s\n" "$*"; }
remember_temp() { :; }
papermc_latest_jar() {
printf "%s %s\n" "https://fill-data.papermc.io/v1/objects/${WANT}/velocity-3.5.1-615.jar" "$WANT"
}
curl() { while [ "$#" -gt 1 ] && [ "$1" != "-o" ]; do shift; done; printf "stand-in velocity build\n" > "$2"; }
atomic_install_file() { printf "INSTALL: %s\n" "$2"; }
'"$vblock"
}
out="$(run_velocity_install deadbeef)"
expect "a download that does not hash to the promised digest is refused" \
"DIE: Velocity 3.5.1 checksum mismatch: got ${vwant}, expected deadbeef" "$out"
case "$out" in
*INSTALL:*) echo "FAIL a refused download must not reach atomic_install_file"; fails=$((fails + 1)) ;;
*) echo "PASS a refused download is not installed" ;;
esac
out="$(run_velocity_install "$vwant")"
expect "the matching download installs" "INSTALL: ${vdir}/velocity.jar" "$out"
# --- [[auth_source]] carry-forward -----------------------------------------------------
# write_felis_toml regenerates felis.toml wholesale on every run; this is what keeps the
# operator's Yggdrasil roots from being reset to the shipped default.
ablock="$(awk '/^persisted_auth_source_blocks\(\) \{/,/^}/' "$BS")"
[ -n "$ablock" ] || { echo "FAIL: no persisted_auth_source_blocks found in $BS"; exit 1; }
[ "$(printf '%s\n' "$ablock" | wc -l)" -lt 20 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
sdir="$(mktemp -d)"
trap 'rm -f "$jar"; rm -rf "$vdir" "$sdir"' EXIT
run_carry() {
STATE_DIR="$sdir" bash -c "$ablock"'
persisted_auth_source_blocks'
}
out="$(run_carry)"
expect "a first install gets the LittleSkin default" 'tag = "littleskin"' "$out"
printf '%s\n' '[server]' 'listen = "0.0.0.0:8080"' '' '[[auth_source]]' 'tag = "guild"' \
'prefix = "GD"' 'url = "https://guild.example/hasJoined"' '' '[smtp]' 'host = "mail.example"' \
> "$sdir/felis.host.toml"
out="$(run_carry)"
expect "an operator's root is carried forward" 'tag = "guild"' "$out"
case "$out" in
*littleskin*|*"[smtp]"*) echo "FAIL the carried list must be exactly the operator's tables:"; echo "$out"; fails=$((fails + 1)) ;;
*) echo "PASS the carried list stops at the next section and adds no default" ;;
esac
printf '%s\n' '[server]' 'listen = "0.0.0.0:8080"' > "$sdir/felis.host.toml"
out="$(run_carry)"
if [ -z "$out" ]; then
echo "PASS a config with no sources stays Mojang-only"
else
echo "FAIL a config with no sources must not get the default back:"; echo "$out"; fails=$((fails + 1))
fi
# The felis setup TUI re-encodes the whole file, which indents keys under each table.
rm -f "$sdir/felis.host.toml"
printf '%s\n' '[[auth_source]]' ' tag = "littleskin"' ' prefix = "LS"' ' url = "https://a.example"' \
'' '[[auth_source]]' ' tag = "guild"' ' prefix = "GD"' ' url = "https://b.example"' \
> "$sdir/felis.pod.toml"
out="$(run_carry)"
expect "both encoder-written tables are carried (first)" ' tag = "littleskin"' "$out"
expect "both encoder-written tables are carried (second)" ' tag = "guild"' "$out"
# TOML allows spaces inside the brackets and a quoted key. Each is still the operator's table.
for hdr in '[[ auth_source ]]' '[["auth_source"]]' "[['auth_source']]"; do
printf '%s\n' "$hdr" 'tag = "guild"' 'prefix = "GD"' 'url = "https://b.example"' '' \
'[smtp]' 'host = "mail.example"' > "$sdir/felis.host.toml"
out="$(run_carry)"
expect "a $hdr header is carried" "$hdr" "$out"
expect "a $hdr table keeps its keys" 'tag = "guild"' "$out"
case "$out" in
*"[smtp]"*) echo "FAIL a $hdr table must stop at the next section:"; echo "$out"; fails=$((fails + 1)) ;;
*) echo "PASS a $hdr table stops at the next section" ;;
esac
done
# --- [smtp] carry-forward does not hoard the auth_source comment block -------------------
# persisted_smtp_block used to print every line between [smtp] and the next section
# header -- which includes the generated Yggdrasil comment block that sits above
# [[auth_source]]. Each re-run re-emitted that hoard plus a fresh template copy, so both
# config files grew by one comment block per run (audit #50). The carry must be the
# section's header and keys only, and must be byte-stable when written back.
sblock="$(awk '/^persisted_smtp_block\(\) \{/,/^}/' "$BS")"
[ -n "$sblock" ] || { echo "FAIL: no persisted_smtp_block found in $BS"; exit 1; }
[ "$(printf '%s\n' "$sblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
sfn="$(mktemp)"
printf '%s\n' "$sblock" > "$sfn"
smtp_dir="$(mktemp -d)"
trap 'rm -f "$jar" "$sfn"; rm -rf "$vdir" "$sdir" "$smtp_dir"' EXIT
run_smtp() { # state-dir
STATE_DIR="$1" SBLOCK_FILE="$sfn" bash -c '. "$SBLOCK_FILE"; persisted_smtp_block'
}
cat > "$smtp_dir/felis.host.toml" <<'TOML'
[server]
listen = "0.0.0.0:8080"
[smtp]
host = "mail.example"
port = 587
from = "[email protected]"
username = "relay-user"
password_ref = "smtp-password"
# Third-party Yggdrasil sources federated by the hasJoined multiplexer. Mojang is
# always the code-owned identity anchor (premium-first), prepended in Go; sources here
# append as namespace-rewritten guests. A fresh install federates LittleSkin. Edit the
# list in /etc/felis/felis.host.toml and rerun the installer; re-runs keep it as it
# is, and with no [[auth_source]] at all the server is Mojang-only.
[[auth_source]]
tag = "littleskin"
prefix = "LS"
url = "https://littleskin.cn/api/yggdrasil/sessionserver/session/minecraft/hasJoined"
TOML
out="$(run_smtp "$smtp_dir")"
expect "a configured [smtp] relay is carried" 'host = "mail.example"' "$out"
expect "its port survives the carry" 'port = 587' "$out"
expect "its credentials reference survives" 'password_ref = "smtp-password"' "$out"
case "$out" in
*"#"*)
echo "FAIL: the carry hoards comment lines:"; printf '%s\n' "$out"; fails=$((fails + 1)) ;;
*) echo "PASS the carry is header and keys only -- no comment hoard" ;;
esac
case "$out" in
*"[["*)
echo "FAIL: the carry ran into the next section:"; printf '%s\n' "$out"; fails=$((fails + 1)) ;;
*) echo "PASS the carry stops at the next section header" ;;
esac
# Write the carry back the way write_felis_toml does (carry + one fresh template block +
# the tables) and extract again: a second re-run must add nothing.
{
printf '%s\n' "$out"
printf '\n%s\n' '# Third-party Yggdrasil sources federated by the hasJoined multiplexer. Mojang is'
printf '%s\n' '[[auth_source]]' ' tag = "littleskin"' ' prefix = "LS"' \
' url = "https://littleskin.cn/api/yggdrasil/sessionserver/session/minecraft/hasJoined"'
} > "$smtp_dir/felis.host.toml"
out2="$(run_smtp "$smtp_dir")"
printf '%s\n' "$out" > "$smtp_dir/first"
printf '%s\n' "$out2" > "$smtp_dir/second"
if cmp -s "$smtp_dir/first" "$smtp_dir/second"; then
echo "PASS a carried-forward [smtp] converges (a second re-run adds nothing)"
else
echo "FAIL: carrying [smtp] is not idempotent:"; diff "$smtp_dir/first" "$smtp_dir/second" | head
fails=$((fails + 1))
fi
# --- write_nano_config leaves the unit able to read its config ---------------------------
# felis-nano runs as a DynamicUser, so the directory must be searchable by others under a
# hardened umask too, including one an older installer left at 0750 -- but the full
# install's, locked to 0700 for its secrets, must not be widened.
wblock="$(awk '/^write_nano_config\(\) \{/,/^}/' "$BS")"
[ -n "$wblock" ] || { echo "FAIL: no write_nano_config found in $BS"; exit 1; }
[ "$(printf '%s\n' "$wblock" | wc -l)" -lt 50 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_nano_config() { # state-dir
STATE_DIR="$1" SECRETS_ENV="$1/secrets.env" BOOTSTRAP_DONE="$1/bootstrap.done" bash -c 'umask 027
ok() { printf "OK: %s\n" "$*"; }
'"$wblock"'
write_nano_config'
}
mkdir "$sdir/probe" && chmod 0700 "$sdir/probe"
if [ "$(stat -c %a "$sdir/probe")" = 700 ]; then
run_nano_config "$sdir/nano" >/dev/null
expect "a fresh config dir is searchable under umask 027" 755 "$(stat -c %a "$sdir/nano")"
mkdir "$sdir/old" && chmod 0750 "$sdir/old"
run_nano_config "$sdir/old" >/dev/null
expect "a nano-only 0750 dir is opened up" 755 "$(stat -c %a "$sdir/old")"
: > "$sdir/probe/secrets.env"
run_nano_config "$sdir/probe" >/dev/null
expect "a dir holding the full install's secrets is not widened" 700 "$(stat -c %a "$sdir/probe")"
mkdir "$sdir/done" && chmod 0700 "$sdir/done" && : > "$sdir/done/bootstrap.done"
run_nano_config "$sdir/done" >/dev/null
expect "a dir marked as a full install is not widened" 700 "$(stat -c %a "$sdir/done")"
else
echo "SKIP directory modes: this filesystem ignores chmod"
fi
# --- install_nano_service reports a unit that dies at once ------------------------------
# A config the new binary rejects leaves the unit in auto-restart; the install must say so
# instead of printing "started" over a proxy whose every login now fails.
iblock="$(awk '/^install_nano_service\(\) \{/,/^}/' "$BS")"
[ -n "$iblock" ] || { echo "FAIL: no install_nano_service found in $BS"; exit 1; }
[ "$(printf '%s\n' "$iblock" | wc -l)" -lt 50 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_nano_service() { # exit status systemctl is-active reports
ACTIVE="$1" NANO_SERVICE="$sdir/felis-nano.service" HOST_BIN=/usr/local/bin/felis \
STATE_DIR=/etc/felis FELIS_NANO_LISTEN=127.0.0.1:25580 bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
ok() { printf "OK: %s\n" "$*"; }
sleep() { :; }
systemctl() { if [ "$1" = is-active ]; then return "$ACTIVE"; fi; }
journalctl() { printf "JOURNAL: config: needs prefix\n"; }
'"$iblock"'
install_nano_service'
}
out="$(run_nano_service 3)"
expect "a unit that dies at once fails the install" "DIE: felis-nano did not stay up" "$out"
expect "the failure shows the unit's own log" "JOURNAL: config: needs prefix" "$out"
case "$out" in
*"OK: felis-nano.service"*) echo "FAIL a dead unit must not be reported as started"; fails=$((fails + 1)) ;;
*) echo "PASS a dead unit is not reported as started" ;;
esac
out="$(run_nano_service 0)"
expect "a unit that stays up is reported as started" "OK: felis-nano.service enabled and started" "$out"
# --- a re-run on a nano host keeps what that host is --------------------------------------
# Re-running the installer is how a nano host updates. It must not move the endpoint an
# off-host proxy points at, nor default a nano-only host to the full control plane. The unit
# read back here is the one install_nano_service wrote above.
rblock="$(awk '/^resolve_nano_listen\(\) \{/,/^}/' "$BS")"
[ -n "$rblock" ] || { echo "FAIL: no resolve_nano_listen found in $BS"; exit 1; }
[ "$(printf '%s\n' "$rblock" | wc -l)" -lt 20 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_listen() { # env-value unit-path
FELIS_NANO_LISTEN="$1" NANO_SERVICE="$2" bash -c "$rblock"'
resolve_nano_listen
printf "LISTEN: %s\n" "$FELIS_NANO_LISTEN"'
}
expect "a re-run keeps the unit's listen address" "LISTEN: 127.0.0.1:25580" \
"$(run_listen '' "$sdir/felis-nano.service")"
expect "the operator's address beats the unit's" "LISTEN: 10.0.0.5:8081" \
"$(run_listen 10.0.0.5:8081 "$sdir/felis-nano.service")"
expect "a first install listens on loopback" "LISTEN: 127.0.0.1:8081" \
"$(run_listen '' "$sdir/absent.service")"
expect "a first install takes the operator's address" "LISTEN: 10.0.0.5:8081" \
"$(run_listen 10.0.0.5:8081 "$sdir/absent.service")"
# Whatever the address came from, it reaches the firewall, the summary and the unit as-is.
lblock="$(awk '/^validate_listen\(\) \{/,/^}/' "$BS")"
[ -n "$lblock" ] || { echo "FAIL: no validate_listen found in $BS"; exit 1; }
[ "$(printf '%s\n' "$lblock" | wc -l)" -lt 20 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
check_listen() { # value
bash -c 'die() { printf "DIE: %s\n" "$*"; exit 1; }
'"$lblock"'
validate_listen FELIS_NANO_LISTEN "$1" && echo VALID' _ "$1" 2>&1
}
for v in 8081 127.0.0.1 127.0.0.1:0 127.0.0.1:65536 127.0.0.1:x ::1:8081; do
expect "listen address $v is refused" "DIE: FELIS_NANO_LISTEN" "$(check_listen "$v")"
done
for v in '[::1]:8081' '[::]:8081' :8081 0.0.0.0:8081 127.0.0.1:8081; do
expect "listen address $v is accepted" VALID "$(check_listen "$v")"
done
# The summary hands the operator the URL to paste into the proxy's JVM flags, so it must
# name the address nano actually binds, and the node's only for a wildcard bind.
sblock="$(awk '/^summary_nano\(\) \{/,/^}/' "$BS")"
[ -n "$sblock" ] || { echo "FAIL: no summary_nano found in $BS"; exit 1; }
[ "$(printf '%s\n' "$sblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
kblock="$(awk '/^nano_listen_is_loopback\(\) \{/,/^}/' "$BS")"
[ -n "$kblock" ] || { echo "FAIL: no nano_listen_is_loopback found in $BS"; exit 1; }
[ "$(printf '%s\n' "$kblock" | wc -l)" -lt 10 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_summary() { # listen [proxy-cidr]
FELIS_NANO_LISTEN="$1" FELIS_NANO_PROXY_CIDR="${2:-}" NODE_IP=203.0.113.9 STATE_DIR=/etc/felis bash -c '
ok() { printf "OK: %s\n" "$*"; }
log() { printf "LOG: %s\n" "$*"; }
systemctl() { :; }
'"$kblock"'
'"$sblock"'
summary_nano'
}
expect "a private bind is the address printed" "http://10.0.0.5:8081/session/minecraft/hasJoined" \
"$(run_summary 10.0.0.5:8081)"
expect "an IPv6 loopback bind is printed as bound" "http://[::1]:8081/session/minecraft/hasJoined" \
"$(run_summary '[::1]:8081')"
out="$(run_summary 127.0.0.1:8081)"
expect "a loopback bind is printed as bound" "http://127.0.0.1:8081/session/minecraft/hasJoined" "$out"
expect "a loopback bind keeps its loopback note" "Bound to loopback" "$out"
for v in 0.0.0.0:8081 '[::]:8081' :8081; do
expect "a wildcard $v bind prints the node's address" "http://203.0.113.9:8081/session/minecraft/hasJoined" \
"$(run_summary "$v")"
done
# Loopback is what keeps the firewall shut, and hasJoined takes no token: a default that
# does not classify as loopback turns every fresh nano host into a public auth relay.
run_loopback() { # listen
FELIS_NANO_LISTEN="$1" bash -c "$kblock"'
if nano_listen_is_loopback; then echo LOOPBACK; else echo ROUTABLE; fi'
}
for v in 127.0.0.1:8081 127.0.0.5:8081 localhost:8081 '[::1]:8081'; do
expect "$v is loopback" LOOPBACK "$(run_loopback "$v")"
done
for v in 0.0.0.0:8081 10.0.0.5:8081 '[::]:8081'; do
expect "$v is not loopback" ROUTABLE "$(run_loopback "$v")"
done
ndefault="$(run_listen '' "$sdir/absent.service")"
ndefault="${ndefault#LISTEN: }"
expect "the default listen address (${ndefault:-empty}) is loopback" LOOPBACK "$(run_loopback "$ndefault")"
# --- firewalld admits the proxy alone ---------------------------------------------------
# hasJoined takes no token, so a routable bind is opened only to FELIS_NANO_PROXY_CIDR, never
# to every source, and a re-run closes the port an earlier installer opened to everyone.
cblock="$(awk '/^validate_cidr\(\) \{/,/^}/' "$BS")"
[ -n "$cblock" ] || { echo "FAIL: no validate_cidr found in $BS"; exit 1; }
[ "$(printf '%s\n' "$cblock" | wc -l)" -lt 15 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
check_cidr() { # value
bash -c 'die() { printf "DIE: %s\n" "$*"; exit 1; }
'"$cblock"'
validate_cidr FELIS_NANO_PROXY_CIDR "$1" && echo VALID' _ "$1" 2>&1
}
for v in 10.0.0.7 10.0.0.7/ /32 10.0.0.0/8/9 10.0.0.7/x '10.0.0.7/32 port' '10.0.0.7/32"'; do
expect "proxy CIDR <$v> is refused" "DIE: FELIS_NANO_PROXY_CIDR" "$(check_cidr "$v")"
done
for v in '' 10.0.0.7/32 192.168.0.0/24 fd00::7/128; do
expect "proxy CIDR <$v> is accepted" VALID "$(check_cidr "$v")"
done
fblock="$(awk '/^configure_nano_firewall\(\) \{/,/^}/' "$BS")"
[ -n "$fblock" ] || { echo "FAIL: no configure_nano_firewall found in $BS"; exit 1; }
[ "$(printf '%s\n' "$fblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_fw() { # listen proxy-cidr port-already-open(0|1)
FELIS_NANO_LISTEN="$1" FELIS_NANO_PROXY_CIDR="$2" OPEN="$3" bash -c '
ok() { printf "OK: %s\n" "$*"; }
log() { printf "LOG: %s\n" "$*"; }
warn() { printf "WARN: %s\n" "$*"; }
systemctl() { return 0; }
firewall-cmd() {
case "$*" in *--query-port=*) [ "$OPEN" = 1 ]; return ;; esac
printf "FW: %s\n" "$*"
}
'"$kblock"'
'"$fblock"'
configure_nano_firewall' 2>&1
}
no_blanket_port() { # label output
case "$2" in
*--add-port*) echo "FAIL $1: the port was opened to every source:"; echo "$2"; fails=$((fails + 1)) ;;
*) echo "PASS $1" ;;
esac
}
out="$(run_fw 0.0.0.0:8081 10.0.0.7/32 0)"
expect "a proxy CIDR opens the port to that source alone" \
'FW: --permanent --add-rich-rule=rule family="ipv4" source address="10.0.0.7/32" port port="8081" protocol="tcp" accept' "$out"
no_blanket_port "a proxy CIDR never opens the port to every source" "$out"
expect "an IPv6 proxy CIDR gets an ipv6 rule" 'rule family="ipv6" source address="fd00::7/128"' \
"$(run_fw '[::]:8081' fd00::7/128 0)"
out="$(run_fw 0.0.0.0:8081 '' 0)"
expect "no proxy CIDR says the port stays closed" "WARN: no FELIS_NANO_PROXY_CIDR" "$out"
no_blanket_port "no proxy CIDR opens nothing" "$out"
case "$out" in
*--add-rich-rule*) echo "FAIL no proxy CIDR must add no rule:"; echo "$out"; fails=$((fails + 1)) ;;
*) echo "PASS no proxy CIDR adds no rule" ;;
esac
expect "a re-run closes the port an earlier install opened to everyone" "FW: --permanent --remove-port=8081/tcp" \
"$(run_fw 0.0.0.0:8081 10.0.0.7/32 1)"
case "$(run_fw 127.0.0.1:8081 10.0.0.7/32 1)" in
*FW:*) echo "FAIL a loopback bind must leave firewalld alone"; fails=$((fails + 1)) ;;
*) echo "PASS a loopback bind leaves firewalld alone" ;;
esac
expect "a routable bind with no proxy CIDR is warned about" "WARNING: bound to 10.0.0.5:8081 with no FELIS_NANO_PROXY_CIDR" \
"$(run_summary 10.0.0.5:8081)"
expect "a routable bind with a proxy CIDR names it" "admits 8081/tcp only from" \
"$(run_summary 10.0.0.5:8081 10.0.0.7/32)"
pblock="$(awk '/^prompt_install_mode\(\) \{/,/^}/' "$BS")"
[ -n "$pblock" ] || { echo "FAIL: no prompt_install_mode found in $BS"; exit 1; }
[ "$(printf '%s\n' "$pblock" | wc -l)" -lt 60 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
tblock="$(awk '/^bootstrap_from_tui\(\) \{/,/^}/' "$BS")"
[ -n "$tblock" ] || { echo "FAIL: no bootstrap_from_tui found in $BS"; exit 1; }
[ "$(printf '%s\n' "$tblock" | wc -l)" -lt 5 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
# Only the no-terminal path can run unattended, and with a terminal attached the prompt
# would sit waiting on it. setsid drops the controlling terminal, as cloud-init and CI have.
notty=""
if (: </dev/tty) 2>/dev/null; then
if command -v setsid >/dev/null 2>&1; then notty=setsid; else notty=skip; fi
fi
run_mode() { # unit-path done-marker-path [FELIS_INSTALL_MODE [FELIS_BOOTSTRAP_FROM_TUI]]
INSTALL_MODE="${3:-}" FELIS_BOOTSTRAP_FROM_TUI="${4:-}" NANO_SERVICE="$1" BOOTSTRAP_DONE="$2" \
$notty bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { printf "LOG: %s\n" "$*"; }
'"$tblock"'
'"$pblock"'
prompt_install_mode </dev/null
printf "MODE: %s\n" "$INSTALL_MODE"' 2>&1
}
if [ "$notty" = skip ]; then
echo "SKIP install-mode default: a terminal is attached and there is no setsid to drop it"
else
: > "$sdir/bootstrap.done"
expect "a nano-only host re-runs as nano" "MODE: nano" \
"$(run_mode "$sdir/felis-nano.service" "$sdir/absent.done")"
expect "a host with the full install re-runs as full" "MODE: full" \
"$(run_mode "$sdir/felis-nano.service" "$sdir/bootstrap.done")"
out="$(run_mode "$sdir/absent.service" "$sdir/absent.done")"
expect "a fresh host defaults to full" "MODE: full" "$out"
expect "no controlling terminal takes the no-prompt path" "LOG: no terminal for a prompt" "$out"
# felis setup goes on to need the control plane, so under it nano is refused, and the
# nano-only default above must not apply either.
expect "felis setup refuses FELIS_INSTALL_MODE=nano" "DIE: felis setup installs the full control plane" \
"$(run_mode "$sdir/absent.service" "$sdir/absent.done" nano 1)"
expect "felis setup installs full on a nano-only host" "MODE: full" \
"$(run_mode "$sdir/felis-nano.service" "$sdir/absent.done" "" 1)"
fi
# --- install_go_toolchain checks the tarball before it replaces anything ----------------
# The tarball is unpacked and run as root, so a download that does not hash to the pin is
# refused -- and refused before the working toolchain is removed.
# The function replaces whatever version sits at GOROOT_DIR, so that has to be a directory
# Felis owns, never an operator's /usr/local/go.
case "$(grep '^GOROOT_DIR=' "$BS")" in
'GOROOT_DIR="/opt/felis/'*) echo "PASS the Go toolchain lives under /opt/felis" ;;
*) echo "FAIL the Go toolchain must live under /opt/felis, got: $(grep '^GOROOT_DIR=' "$BS")"; fails=$((fails + 1)) ;;
esac
gblock="$(awk '/^install_go_toolchain\(\) \{/,/^}/' "$BS")"
[ -n "$gblock" ] || { echo "FAIL: no install_go_toolchain found in $BS"; exit 1; }
[ "$(printf '%s\n' "$gblock" | wc -l)" -lt 50 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
gsum="$(printf 'stand-in go toolchain\n' | sha256sum | cut -d' ' -f1)"
groot="$sdir/go"
run_go() { # FELIS_GO_VERSION pinned-amd64-digest [FELIS_GO_SHA256]
FELIS_GO_VERSION="$1" GO_PINNED_VERSION=1.26.4 GO_PINNED_SHA256_AMD64="$2" \
GO_PINNED_SHA256_ARM64=unused FELIS_GO_SHA256="${3:-}" GOROOT_DIR="$groot" TMPDIR="$sdir" bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { printf "LOG: %s\n" "$*"; }
ok() { printf "OK: %s\n" "$*"; }
remember_temp() { printf "TEMP: %s\n" "$1"; }
uname() { echo x86_64; }
curl() { while [ "$#" -gt 1 ] && [ "$1" != "-o" ]; do shift; done
printf "stand-in go toolchain\n" > "$2"; printf "CURL: %s\n" "$2"; }
tar() { printf "TAR: %s\n" "$*"; }
'"$gblock"'
install_go_toolchain'
}
mkdir -p "$groot" && : > "$groot/KEEP"
out="$(run_go 1.26.4 deadbeef)"
expect "a Go download that does not match the pin is refused" \
"DIE: Go 1.26.4 (amd64) checksum mismatch: got ${gsum}, expected deadbeef" "$out"
case "$out" in
*TAR:*) echo "FAIL a refused Go download must not be unpacked"; fails=$((fails + 1)) ;;
*) echo "PASS a refused Go download is not unpacked" ;;
esac
if [ -e "$groot/KEEP" ]; then
echo "PASS a refused Go download leaves the old toolchain in place"
else
echo "FAIL a refused Go download must not remove the old toolchain"; fails=$((fails + 1))
fi
expect "an unpinned FELIS_GO_VERSION without a digest is refused" "DIE: no pinned sha256 for Go 1.99.0" \
"$(run_go 1.99.0 "$gsum")"
expect "an unpinned FELIS_GO_VERSION installs with its own FELIS_GO_SHA256" "TAR: " \
"$(run_go 1.99.0 deadbeef "$gsum")"
out="$(run_go 1.26.4 "$gsum")"
expect "a Go download matching the pin is unpacked" "TAR: " "$out"
gtmp="$(printf '%s\n' "$out" | sed -n 's/^TEMP: //p')"
expect "the Go download is staged in a directory the cleanup removes" \
"CURL: ${gtmp:-<none>}/go1.26.4.linux-amd64.tar.gz" "$out"
# --- a private repo without a token fails with the hint instead of prompting -------------
# git asks for credentials on /dev/tty, where a piped install would sit waiting. Every
# network git call goes through git_auth, so the switch belongs there.
gablock="$(awk '/^git_auth\(\) \{/,/^}/' "$BS")"
[ -n "$gablock" ] || { echo "FAIL: no git_auth found in $BS"; exit 1; }
[ "$(printf '%s\n' "$gablock" | wc -l)" -lt 15 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_git_auth() { # token
FELIS_GITHUB_TOKEN="$1" bash -c '
unset GIT_TERMINAL_PROMPT # whatever runs this harness may have set it already
git() { printf "GIT: prompt=%s\n" "${GIT_TERMINAL_PROMPT:-<unset>}"; }
'"$gablock"'
git_auth clone https://example.invalid/felis.git'
}
expect "git never prompts without a token" "GIT: prompt=0" "$(run_git_auth '')"
expect "git never prompts with a token" "GIT: prompt=0" "$(run_git_auth ghp_example)"
fblock="$(awk '/^fetch_source\(\) \{/,/^}/' "$BS")"
[ -n "$fblock" ] || { echo "FAIL: no fetch_source found in $BS"; exit 1; }
[ "$(printf '%s\n' "$fblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_fetch() { # src-dir
SRC_DIR="$1" FELIS_REF=main FELIS_REPO_URL=https://example.invalid/felis.git bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { :; }
ok() { :; }
resolve_install_ref() { :; }
stamp_version() { :; }
git_auth() { return 128; }
git() { :; }
'"$fblock"'
fetch_source'
}
expect "a failed clone names the token" "set FELIS_GITHUB_TOKEN" "$(run_fetch "$sdir/src")"
mkdir -p "$sdir/src/.git"
expect "a failed fetch into an existing checkout names the token" "set FELIS_GITHUB_TOKEN" \
"$(run_fetch "$sdir/src")"
# --- default install keeps backups, and retention envs reach the renderer ----------------
# A default install must render the world-archive PVC (without one, backup/restore answer an
# honest 503), and FELIS_WORLDS_HOST_PATH must turn into the reaper's two flags or an
# operator's retention enablement silently renders no CronJob. Extracted, not retyped.
mblock="$(awk '/^ log "rendering \+ applying the control-plane bundle"/,/kube apply -f -/' "$BS")"
[ -n "$mblock" ] || { echo "FAIL: no manifest_args block found in $BS"; exit 1; }
[ "$(printf '%s\n' "$mblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the manifest_args block -- did it move?"; exit 1; }
run_bundle_flags() { # backup-pvc worlds-host-path
FELIS_IMAGE=reg/felis:test FELIS_PANEL_NODEPORT=30443 NODE_IP=10.0.0.5 \
FELIS_BACKUP_PVC="$1" FELIS_WORLDS_HOST_PATH="$2" FELIS_ARCHIVE_LOCAL_PATH=/var/lib/felis/archives \
HOST_BIN=myManifests bash -c '
log() { :; }
warn() { printf "WARN: %s\n" "$*"; }
kube() { cat; }
myManifests() { printf "%s\n" "$@"; }
setfacl() { printf "SETFACL %s\n" "$*"; }
run_bundle() {
'"$mblock"'
}
run_bundle'
}
out="$(run_bundle_flags felis-backups '')"
expect "a default install asks the renderer for the archive PVC" "--backup-pvc
felis-backups" "$out"
case "$out" in
*--worlds-host-path*) echo "FAIL: no reaper flags may render without FELIS_WORLDS_HOST_PATH"; fails=$((fails + 1)) ;;
esac
out="$(run_bundle_flags '' '')"
expect "an emptied FELIS_BACKUP_PVC is the explicit no-backup shape" "--backup-pvc=" "$out"
out="$(run_bundle_flags felis-backups /var/lib/rancher/k3s/storage)"
expect "enabling retention passes the worlds root" "--worlds-host-path
/var/lib/rancher/k3s/storage" "$out"
expect "enabling retention passes the archive mount that must match felis.toml" "--archive-local-path
/var/lib/felis/archives" "$out"
# The warn fires only when the root is ABSENT (hostPath type Directory would fail);
# the case above passes a path that exists on any host already running k3s, so it
# must not also demand the warning — probing the real /var/lib/rancher path made
# this suite red on exactly the hosts the installer is for. Point the warn case at
# a path guaranteed missing.
missing="/tmp/felis-worlds-root-must-not-exist-$$"
out="$(run_bundle_flags felis-backups "$missing")"
expect "a missing worlds root is warned about, not silently skipped" "WARN: worlds root $missing does not exist yet" "$out"
# The reaper pod is non-root (uid 1000) and k3s ships the storage root 0700 root:root, so
# the installer must grant traverse or every archive dies with permission denied.
wdir="$(mktemp -d)"
out="$(run_bundle_flags felis-backups "$wdir")"
expect "enabling retention grants the reaper uid traverse on the worlds root" "SETFACL -m u:1000:x $wdir" "$out"
# --- the registry mirror writer -----------------------------------------------------------
# k3s only consults registries.yaml at agent start, so a CONTENT change must restart k3s and
# an identical file (every re-run) must restart nothing. The k3s restart is the expensive,
# disruptive half of the pair -- getting the idempotence wrong bounces the whole cluster on
# every installer re-run, so both halves are pinned here against the extracted function.
cmblock="$(awk '/^configure_registry_mirror\(\) \{/,/^}/' "$BS")"
[ -n "$cmblock" ] || { echo "FAIL: no configure_registry_mirror found in $BS"; exit 1; }
run_mirror() { # scratch-file
K3S_REGISTRIES_FILE="$1" REGISTRY_URL=registry.felis.svc:5000 REGISTRY_PUSH_HOST=127.0.0.1:5000 \
bash -c '
log() { printf "LOG: %s\n" "$*"; }
ok() { printf "OK: %s\n" "$*"; }
die() { printf "DIE: %s\n" "$*"; exit 1; }
warn() { printf "WARN: %s\n" "$*"; }
remember_temp() { :; }
systemctl() { printf "SYSTEMCTL %s\n" "$*"; }
kube() { printf "n Ready \n"; }
wait_for_node_ready() { kube get nodes | grep -q " Ready " && ok "k3s node Ready"; }
'"$cmblock"'
configure_registry_mirror'
}
mfile="$(mktemp -u)"
out="$(run_mirror "$mfile")"
expect "a missing registries.yaml is written" "\"registry.felis.svc:5000\":" "$(cat "$mfile" 2>/dev/null)"
expect "the mirror endpoint is the node loopback push/pull host" "\"http://127.0.0.1:5000\"" "$(cat "$mfile" 2>/dev/null)"
expect "a content change restarts k3s" "SYSTEMCTL restart k3s" "$out"
out="$(run_mirror "$mfile")"
expect "an identical registries.yaml is recognised" "already configured" "$out"
case "$out" in
*"SYSTEMCTL restart"*) echo "FAIL: a re-run with identical content must not restart k3s"; fails=$((fails + 1)) ;;
esac
printf 'mirrors: {}\n' >"$mfile"
out="$(run_mirror "$mfile")"
expect "changed content restarts k3s again" "SYSTEMCTL restart k3s" "$out"
rm -f "$mfile"
# --- image mirroring ----------------------------------------------------------------------
# The registry keys a repository by the path AFTER the host, so the push must swap the
# registry host for the node's loopback endpoint and nothing else. A ref outside the
# registry must be warned about, not silently pushed somewhere unintended.
pblock="$(awk '/^push_image_to_registry\(\) \{/,/^}/' "$BS")"
[ -n "$pblock" ] || { echo "FAIL: no push_image_to_registry found in $BS"; exit 1; }
run_push() { # ref [docker-push-exit]
REF="$1" PUSH_EXIT="${2:-0}" \
REGISTRY_URL=registry.felis.svc:5000 REGISTRY_PUSH_HOST=127.0.0.1:5000 \
bash -c '
log() { printf "LOG: %s\n" "$*"; }
warn() { printf "WARN: %s\n" "$*"; }
die() { printf "DIE: %s\n" "$*"; exit 1; }
ok() { :; }
systemctl() { :; }
docker() {
case "$1" in
push) printf "DOCKER %s\n" "$*"; return "$PUSH_EXIT" ;;
*) printf "DOCKER %s\n" "$*" ;;
esac
}
'"$pblock"'
push_image_to_registry "$REF"'
}
out="$(run_push registry.felis.svc:5000/felis/felis:demo)"
expect "a registry ref is re-tagged onto the node loopback endpoint" \
"DOCKER tag registry.felis.svc:5000/felis/felis:demo 127.0.0.1:5000/felis/felis:demo" "$out"
expect "and pushed to exactly that endpoint" "DOCKER push 127.0.0.1:5000/felis/felis:demo" "$out"
out="$(run_push registry.felis.svc:50000/felis/felis:demo)"
expect "a ref outside the registry is refused with a warning" "WARN: not mirroring" "$out"
case "$out" in
*"DOCKER push"*) echo "FAIL: a non-registry ref must not be pushed"; fails=$((fails + 1)) ;;
esac
out="$(run_push registry.felis.svc:5000/felis/felis:demo 1)"
expect "a failed push fails the install loudly" "DIE: could not mirror" "$out"
# docker must be started ONCE for the whole batch: a start/stop pair per image trips
# systemd's start rate limit ("start-limit-hit" — observed live; the 4th image was never
# mirrored because docker.service is socket-triggered and each cycle counts twice).
wiblock="$(awk '/^push_images_to_registry\(\) \{/,/^}/' "$BS")"
[ -n "$wiblock" ] || { echo "FAIL: no push_images_to_registry found in $BS"; exit 1; }
out="$(
FELIS_IMAGE=a FELIS_LIMBO_IMAGE=b FELIS_LOBBY_IMAGE=c FELIS_PAPER_IMAGE=d bash -c '
systemctl() { printf "SYSTEMCTL %s\n" "$*"; }
push_image_to_registry() { printf "PUSH %s\n" "$1"; }
'"$wiblock"'
push_images_to_registry'
)"
starts="$(printf '%s\n' "$out" | grep -c 'SYSTEMCTL start docker')"
stops="$(printf '%s\n' "$out" | grep -c 'SYSTEMCTL stop docker')"
[ "$starts" = 1 ] && [ "$stops" = 1 ] && [ "$(printf '%s\n' "$out" | grep -c '^PUSH')" = 4 ] \
&& echo "PASS the batch wraps all four pushes in ONE docker start/stop" \
|| { echo "FAIL: expected 1 start / 1 stop / 4 pushes, got:"; printf '%s\n' "$out"; fails=$((fails + 1)); }
# --- the registry's own image must not be re-pulled on every run --------------------------
iblock="$(awk '/^import_registry_image\(\) \{/,/^}/' "$BS")"
[ -n "$iblock" ] || { echo "FAIL: no import_registry_image found in $BS"; exit 1; }
out="$(
bash -c '
log() { printf "LOG: %s\n" "$*"; }
ok() { printf "OK: %s\n" "$*"; }
warn() { printf "WARN: %s\n" "$*"; }
die() { printf "DIE: %s\n" "$*"; exit 1; }
systemctl() { :; }
k3s_cmd() { case "$*" in "ctr images ls -q") printf "docker.io/library/registry:2\n" ;; esac; }
docker() { printf "DOCKER %s\n" "$*"; return 1; }
'"$iblock"'
import_registry_image'
)"
expect "an already-imported registry:2 is left alone" "already in k3s containerd" "$out"
case "$out" in
*DOCKER*) echo "FAIL: a present registry:2 must not trigger a docker pull"; fails=$((fails + 1)) ;;
esac
# --- installer re-runs refresh the workload namespace's felis-config copy ---------------
# The backup/restore/fileedit Jobs and the reaper mount the workload namespace's own
# felis-config (a secretKeyRef is namespace-local). `felis setup` makes that replica
# create-if-absent -- right for credentials, wrong for a rendered config -- so the
# installer must refresh it every run; a stale copy keeps old DB/archive settings.
fcblock="$(awk '/^apply_felis_config_secrets\(\) \{/,/^}/' "$BS")"
[ -n "$fcblock" ] || { echo "FAIL: no apply_felis_config_secrets found in $BS"; exit 1; }
[ "$(printf '%s\n' "$fcblock" | wc -l)" -lt 20 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
kubcalls="$(mktemp)"
run_fc() {
: > "$kubcalls"
CONTROL_NS=felis MINECRAFT_NS=minecraft STATE_DIR=/tmp/fc KUBCALLS="$kubcalls" bash -c '
kube() { printf "%s\n" "$*" >> "$KUBCALLS"; }
'"$fcblock"'
apply_felis_config_secrets'
cat "$kubcalls"
}
out="$(run_fc)"
expect "the control plane's felis-config is applied" \
"-n felis create secret generic felis-config" "$out"
expect "the workload namespace's copy is applied too" \
"-n minecraft create secret generic felis-config" "$out"
expect "both copies render from the pod config" \
"felis.toml=/tmp/fc/felis.pod.toml" "$out"
applies="$(printf '%s\n' "$out" | grep -c '^apply -f -$')"
if [ "$applies" -eq 2 ]; then
echo "PASS both rendered copies are piped to kubectl apply"
else
echo "FAIL: expected 2 applies, got $applies:"; printf '%s\n' "$out"; fails=$((fails + 1))
fi
rm -f "$kubcalls"
# --- installer re-runs keep the operator's [registry] overrides --------------------------
# §15's upgrade path is re-running the installer, but the build-lane mirrors and the
# uploads backend live in [registry] as hand-written keys (docs/troubleshooting.md §8e or
# the storage wizard) that nothing in this script's inputs derives. A re-run must carry
# them forward — without letting a stale url/build_namespace survive (installer-owned).
wrblock="$(awk '/^write_felis_toml\(\) \{/,/^}/' "$BS")"
prblock="$(awk '/^persisted_registry_block\(\) \{/,/^}/' "$BS")"
pablock="$(awk '/^persisted_archive_block\(\) \{/,/^}/' "$BS")"
{ [ -n "$wrblock" ] && [ -n "$prblock" ] && [ -n "$pablock" ]; } \
|| { echo "FAIL: write_felis_toml / persisted_{registry,archive}_block not found in $BS"; exit 1; }
# The blocks quote themselves (the awk program uses single quotes), so they are
# sourced from a file instead of being spliced into a single-quoted bash -c.
fnfile="$(mktemp)"
printf '%s\n%s\n%s\n' "$prblock" "$pablock" "$wrblock" > "$fnfile"
rdir="$(mktemp -d)"
cat > "$rdir/felis.host.toml" <<'TOML'
[registry]
url = "stale.invalid:5000"
build_namespace = "stale-ns"
kaniko_image = "registry.felis.svc:5000/mirror/kaniko-executor:v1.24.0"
trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2"
trivy_java_db_repository = "registry.felis.svc:5000/mirror/trivy-java-db:1"
[registry.s3]
endpoint = "https://s3.example"
region = "us-east-1"
[archive]
store = "tarLocal"
local_path = "/stale/path"
retention = "30d"
TOML
run_write() { # out-file
STATE_DIR="$rdir" OUT_TOML="$1" FNFILE="$fnfile" bash -c '
log() { :; }
persisted_smtp_block() { :; }
persisted_auth_source_blocks() { :; }
. "$FNFILE"
FELIS_ROOT_DOMAIN=r.example.com DB_USER=u DB_PASSWORD=p DB_NAME=d MINECRAFT_NS=minecraft \
FELIS_EGRESS_MODE=nodeport FELIS_LIMBO_IMAGE=li FELIS_LOBBY_IMAGE=lo \
REGISTRY_URL=registry.felis.svc:5000 BUILD_NS=felis-build FELIS_ARCHIVE_LOCAL_PATH=/a \
write_felis_toml "$OUT_TOML" 127.0.0.1'
}
run_write "$rdir/out.toml"
out="$(cat "$rdir/out.toml")"
expect "a re-run carries the build-lane executor mirrors" \
'kaniko_image = "registry.felis.svc:5000/mirror/kaniko-executor:v1.24.0"' "$out"
expect "a re-run carries the trivy vulnerability-DB mirror" \
'trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2"' "$out"
expect "a re-run carries the trivy java-DB mirror" \
'trivy_java_db_repository = "registry.felis.svc:5000/mirror/trivy-java-db:1"' "$out"
expect "a re-run carries the [registry.s3] uploads subtable" "[registry.s3]" "$out"
expect "the carried subtable keeps its keys" 'endpoint = "https://s3.example"' "$out"
expect "url stays installer-owned" 'url = "registry.felis.svc:5000"' "$out"
expect "a re-run carries the archive retention window" 'retention = "30d"' "$out"
expect "the archive mount stays installer-owned" 'local_path = "/a"' "$out"
case "$out" in
*stale.invalid* | *stale-ns* | *stale/path*)
echo "FAIL: stale installer-owned values survived the re-run"; fails=$((fails + 1)) ;;
esac
cp "$rdir/out.toml" "$rdir/felis.host.toml"
run_write "$rdir/out2.toml"
if cmp -s "$rdir/out.toml" "$rdir/out2.toml"; then
echo "PASS a carried-forward config converges (the second re-run is a no-op)"
else
echo "FAIL: carrying [registry] overrides is not idempotent"
diff "$rdir/out.toml" "$rdir/out2.toml" | head
fails=$((fails + 1))
fi
rm -f "$fnfile"
# --------------------------------------------------------------------------------------- # ---------------------------------------------------------------------------------------
if [ "$fails" -eq 0 ]; then if [ "$fails" -eq 0 ]; then
echo "ALL PASS" echo "ALL PASS"
@@ -347,6 +347,14 @@ spec:
- type - type
type: object type: object
type: array type: array
emptySince:
description: |-
EmptySince is when the operator first observed 0 online players during
a Running phase (spec §8 idle auto-stop). It is reset when a player joins
or the server stops, so the empty-duration counter starts fresh each time
the server becomes unoccupied.
format: date-time
type: string
endpoint: endpoint:
description: Endpoint is where the proxy should route traffic. description: Endpoint is where the proxy should route traffic.
properties: properties:
+25 -74
View File
@@ -1,29 +1,28 @@
#!/bin/bash #!/bin/bash
# demo-up.sh — one-shot Felis demo bring-up. # demo-up.sh — one-shot Felis demo bring-up.
# #
# Collapses the four manual steps (bootstrap -> build/import limbo+lobby images ->
# edit felis.toml -> felis setup) into a single command:
#
# sudo bash deploy/demo-up.sh # sudo bash deploy/demo-up.sh
# #
# It ends by exec'ing the interactive `felis setup` TUI (create the Owner account) — # Every piece a demo box needs — base platform (k3s + felis + docker + cloudflared +
# that human step is the only thing this script cannot do for you. # control plane), the limbo/lobby/paper images, the felis-velocity proxy plugin, and
# the [velocity] wiring in felis.host.toml — is built by deploy/bootstrap.sh. This
# wrapper adds only the one step the installer cannot do: the interactive
# `felis setup` TUI that creates the Owner account.
# #
# Image source, in order of preference: # It used to rebuild the game images here with its own copy of that logic, written
# 1. Prebuilt tars at deploy/images/felis-limbo.tar + felis-lobby.tar (imported as-is). # before bootstrap grew the job. The copy drifted: it pinned Paper 1.21.8 while the
# 2. Otherwise built on this host with docker, resolving the LOOHP/Limbo CI jar and # installer derives one version from the Limbo login gate (both hops of a login must
# the latest stable Paper jar automatically. Override any of: # speak one protocol), never built felis-velocity.jar (so the proxy it wired had
# LIMBO_JAR_URL LIMBO_SCHEM_URL LIMBO_VERSION PAPER_JAR_URL PAPER_MC_VERSION # nowhere to route), and left docker running. The installer is the single origin.
# #
# Toggles: SKIP_BOOTSTRAP=1 (base already up), SKIP_SETUP=1 (stop before the TUI). # Toggles: SKIP_BOOTSTRAP=1 (base + game stack already installed by a full bootstrap),
# SKIP_SETUP=1 (stop before the TUI).
set -Eeuo pipefail set -Eeuo pipefail
SRC_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" SRC_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
STATE_DIR=/etc/felis STATE_DIR=/etc/felis
HOST_TOML="$STATE_DIR/felis.host.toml" HOST_TOML="$STATE_DIR/felis.host.toml"
IMG_DIR="$SRC_DIR/deploy/images" PLUGIN_JAR=/opt/felis/velocity/plugins/felis-velocity.jar
LIMBO_IMAGE="felis-limbo:demo"
LOBBY_IMAGE="felis-lobby:demo"
K3S=/usr/local/bin/k3s K3S=/usr/local/bin/k3s
FELIS=/usr/local/bin/felis FELIS=/usr/local/bin/felis
@@ -32,11 +31,11 @@ die() { printf '\033[1;31mERROR: %s\033[0m\n' "$*" >&2; exit 1; }
[ "$(id -u)" -eq 0 ] || die "run as root (sudo bash deploy/demo-up.sh)" [ "$(id -u)" -eq 0 ] || die "run as root (sudo bash deploy/demo-up.sh)"
# 1. base platform (k3s + felis + docker + cloudflared + control plane) ---------- # 1. base platform + full game stack (deploy/bootstrap.sh) -----------------------
if [ "${SKIP_BOOTSTRAP:-0}" = 1 ]; then if [ "${SKIP_BOOTSTRAP:-0}" = 1 ]; then
log "SKIP_BOOTSTRAP=1 — assuming the base platform is already up" log "SKIP_BOOTSTRAP=1 — assuming the base platform is already up"
else else
log "bringing up the base platform (deploy/bootstrap.sh)" log "bringing up the base platform and the game stack (deploy/bootstrap.sh)"
bash "$SRC_DIR/deploy/bootstrap.sh" bash "$SRC_DIR/deploy/bootstrap.sh"
fi fi
@@ -45,67 +44,19 @@ command -v "$K3S" >/dev/null 2>&1 || K3S=k3s
command -v "$K3S" >/dev/null 2>&1 || die "k3s not found — did bootstrap complete?" command -v "$K3S" >/dev/null 2>&1 || die "k3s not found — did bootstrap complete?"
command -v "$FELIS" >/dev/null 2>&1 || die "felis not found — did bootstrap complete?" command -v "$FELIS" >/dev/null 2>&1 || die "felis not found — did bootstrap complete?"
# 2. get the two game images into k3s containerd -------------------------------- # SKIP_BOOTSTRAP=1 trusts an earlier run to be complete. Check that it actually left
if [ -f "$IMG_DIR/felis-limbo.tar" ] && [ -f "$IMG_DIR/felis-lobby.tar" ]; then # the full stack behind: a base from before the game-stack installer, or one whose
log "importing prebuilt image tars from $IMG_DIR" # pieces were pruned by hand, must fail here with a pointer — not present as a proxy
"$K3S" ctr images import "$IMG_DIR/felis-limbo.tar" # that accepts logins and routes nowhere, with nothing in any log to say why.
"$K3S" ctr images import "$IMG_DIR/felis-lobby.tar"
# Optional: the plain-Paper recommended base, if a tar was staged for it.
[ -f "$IMG_DIR/felis-paper.tar" ] && "$K3S" ctr images import "$IMG_DIR/felis-paper.tar"
else
log "no prebuilt tars in $IMG_DIR — building on this host with docker"
command -v docker >/dev/null 2>&1 || die "docker not found; cannot build images"
rel=$(curl -fsSL --max-time 30 "https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/api/json" \
| grep -oE 'target/Limbo-[0-9][^"]+\.jar' | head -1) || true
: "${LIMBO_JAR_URL:=https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/artifact/$rel}"
: "${LIMBO_SCHEM_URL:=https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/artifact/spawn.schem}"
: "${LIMBO_VERSION:=$(basename "$rel" | sed -E 's/^Limbo-//; s/\.jar$//; s/-[0-9]+\.[0-9]+$//')}"
[ -n "$rel" ] || [ -n "${LIMBO_JAR_URL##*artifact/}" ] || die "could not resolve the Limbo jar; set LIMBO_JAR_URL"
log "building $LIMBO_IMAGE (Limbo $LIMBO_VERSION)"
docker build -f "$SRC_DIR/deploy/limbo/Dockerfile" \
--build-arg LIMBO_JAR_URL="$LIMBO_JAR_URL" \
--build-arg LIMBO_SCHEM_URL="$LIMBO_SCHEM_URL" \
--build-arg LIMBO_VERSION="$LIMBO_VERSION" \
-t "$LIMBO_IMAGE" "$SRC_DIR"
docker save "$LIMBO_IMAGE" | "$K3S" ctr images import -
: "${PAPER_MC_VERSION:=1.21.8}"
: "${PAPER_JAR_URL:=$(curl -fsSL --max-time 30 "https://fill.papermc.io/v3/projects/paper/versions/${PAPER_MC_VERSION}/builds/latest" | grep -oE 'https://fill-data\.papermc\.io/[^"]+\.jar' | head -1)}"
[ -n "$PAPER_JAR_URL" ] || die "could not resolve the Paper jar; set PAPER_JAR_URL"
log "building $LOBBY_IMAGE (Paper $PAPER_MC_VERSION)"
docker build -f "$SRC_DIR/deploy/lobby/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
-t "$LOBBY_IMAGE" "$SRC_DIR"
docker save "$LOBBY_IMAGE" | "$K3S" ctr images import -
# Plain Paper recommended base — same PAPER_JAR_URL, no plugins, no secret gate.
: "${PAPER_IMAGE:=felis-paper:demo}"
log "building $PAPER_IMAGE (plain Paper $PAPER_MC_VERSION, forwarding via the operator initContainer)"
docker build -f "$SRC_DIR/deploy/paper/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
-t "$PAPER_IMAGE" "$SRC_DIR"
docker save "$PAPER_IMAGE" | "$K3S" ctr images import -
fi
# 3. wire the images into the config `felis setup` reads ------------------------
log "wiring [velocity] images into $HOST_TOML"
[ -f "$HOST_TOML" ] || die "missing $HOST_TOML — did bootstrap run?" [ -f "$HOST_TOML" ] || die "missing $HOST_TOML — did bootstrap run?"
if grep -q '^\[velocity\]' "$HOST_TOML"; then grep -q '^\[velocity\]' "$HOST_TOML" \
echo " [velocity] table already present — leaving it untouched" || die "$HOST_TOML has no [velocity] section — re-run the installer without SKIP_BOOTSTRAP so the system servers get wired"
else [ -f "$PLUGIN_JAR" ] \
cat >> "$HOST_TOML" <<EOF || die "$PLUGIN_JAR missing — this base did not finish the full installer, and a proxy without it silently routes nothing; re-run the installer without SKIP_BOOTSTRAP"
[velocity] # 2. interactive Owner creation + system-server provisioning --------------------
login_image = "$LIMBO_IMAGE"
lobby_image = "$LOBBY_IMAGE"
EOF
echo " appended login_image=$LIMBO_IMAGE / lobby_image=$LOBBY_IMAGE"
fi
# 4. interactive Owner creation + system-server provisioning -------------------
if [ "${SKIP_SETUP:-0}" = 1 ]; then if [ "${SKIP_SETUP:-0}" = 1 ]; then
log "SKIP_SETUP=1 — base + images + config ready. Finish with: sudo felis setup" log "SKIP_SETUP=1 — base + game stack ready. Finish with: sudo felis setup"
else else
log "launching 'felis setup' — create the Owner account (this is the only interactive step)" log "launching 'felis setup' — create the Owner account (this is the only interactive step)"
exec "$FELIS" setup exec "$FELIS" setup
+10 -3
View File
@@ -90,14 +90,21 @@ docker build -f deploy/limbo/Dockerfile \
version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the
maven repo). maven repo).
Import into k3s and point config at it: Publish it into the cluster's registry and point config at it. On the node
itself (docker treats `127.0.0.1` as insecure by default):
``` ```
docker save felis-limbo:demo | sudo k3s ctr images import - docker tag felis-limbo:demo 127.0.0.1:5000/felis/limbo:demo
# felis.toml → [velocity] login_image = "felis-limbo:demo" docker push 127.0.0.1:5000/felis/limbo:demo
# felis.toml → [velocity] login_image = "registry.felis.svc:5000/felis/limbo:demo"
sudo felis setup sudo felis setup
``` ```
The registry keys a repository by the path after the host, so pushing through a
`kubectl -n felis port-forward svc/registry 5000:5000` from another machine is
equivalent. Hosting the image in the registry (rather than only importing it
into containerd) is what lets kubelet re-pull it after an image GC.
## Ports (handled for you) ## Ports (handled for you)
The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to
+9
View File
@@ -9,6 +9,7 @@
# Fill v3 API — api.papermc.io v2 has returned HTTP 410 since 2026-07-01): # Fill v3 API — api.papermc.io v2 has returned HTTP 410 since 2026-07-01):
# docker build -f deploy/lobby/Dockerfile \ # docker build -f deploy/lobby/Dockerfile \
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-26.2-<build>.jar \ # --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-26.2-<build>.jar \
# --build-arg PAPER_JAR_SHA256=<that same sha — the objects/ path segment> \
# --build-arg LUCKPERMS_JAR_URL="$(curl -fsSL https://metadata.luckperms.net/data/all \ # --build-arg LUCKPERMS_JAR_URL="$(curl -fsSL https://metadata.luckperms.net/data/all \
# | grep -o 'https://download.luckperms.net/[^"]*/bukkit/loader/[^"]*\.jar')" \ # | grep -o 'https://download.luckperms.net/[^"]*/bukkit/loader/[^"]*\.jar')" \
# -t felis-lobby:demo . # -t felis-lobby:demo .
@@ -42,6 +43,10 @@ RUN cd plugins/paper \
# also runs the plugin's Java-21 bytecode, so only the runtime moves. # also runs the plugin's Java-21 bytecode, so only the runtime moves.
FROM eclipse-temurin:25-jre FROM eclipse-temurin:25-jre
ARG PAPER_JAR_URL ARG PAPER_JAR_URL
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
# that shape at build time. Checking the digest after the download turns a truncated or
# tampered fetch into a failed build instead of a lobby booted on the wrong bytes.
ARG PAPER_JAR_SHA256
# LuckPerms is required, not optional: the panel's whole permission surface # LuckPerms is required, not optional: the panel's whole permission surface
# (internal/api/handlers_access.go) issues `lp user ...` over RCON, so a lobby built # (internal/api/handlers_access.go) issues `lp user ...` over RCON, so a lobby built
# without it answers every grant with "Unknown command" — a failure the operator only # without it answers every grant with "Unknown command" — a failure the operator only
@@ -55,12 +60,16 @@ RUN set -eu; \
if [ -z "${PAPER_JAR_URL:-}" ]; then \ if [ -z "${PAPER_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \ echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \
fi; \ fi; \
if [ -z "${PAPER_JAR_SHA256:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_SHA256=<paper jar sha256> is required" >&2; exit 1; \
fi; \
if [ -z "${LUCKPERMS_JAR_URL:-}" ]; then \ if [ -z "${LUCKPERMS_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg LUCKPERMS_JAR_URL=<luckperms bukkit jar> is required" >&2; exit 1; \ echo "ERROR: --build-arg LUCKPERMS_JAR_URL=<luckperms bukkit jar> is required" >&2; exit 1; \
fi; \ fi; \
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \ apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
mkdir -p /paper/plugins; \ mkdir -p /paper/plugins; \
curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \ curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \
echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \
curl -fSL "$LUCKPERMS_JAR_URL" -o /paper/plugins/LuckPerms.jar; \ curl -fSL "$LUCKPERMS_JAR_URL" -o /paper/plugins/LuckPerms.jar; \
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \ apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \
echo "eula=true" > /paper/eula.txt echo "eula=true" > /paper/eula.txt
+7 -2
View File
@@ -29,9 +29,14 @@ this at every layer:
``` ```
docker build -f deploy/lobby/Dockerfile \ docker build -f deploy/lobby/Dockerfile \
--build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \ --build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \
--build-arg PAPER_JAR_SHA256=<sha256 of that jar> \
-t felis-lobby:demo . -t felis-lobby:demo .
docker save felis-lobby:demo | sudo k3s ctr images import - # Publish into the cluster's registry (on the node; docker treats 127.0.0.1 as
# felis.toml → [velocity] lobby_image = "felis-lobby:demo" # insecure by default — or through a `kubectl -n felis port-forward svc/registry
# 5000:5000`, which is equivalent: only the path after the host matters).
docker tag felis-lobby:demo 127.0.0.1:5000/felis/lobby:demo
docker push 127.0.0.1:5000/felis/lobby:demo
# felis.toml → [velocity] lobby_image = "registry.felis.svc:5000/felis/lobby:demo"
sudo felis setup sudo felis setup
``` ```
+1 -1
View File
@@ -93,7 +93,7 @@ else
echo " injects it from the <server>-rcon Secret when spec.rcon.enabled is true." >&2 echo " injects it from the <server>-rcon Secret when spec.rcon.enabled is true." >&2
fi fi
# ponytail: rewritten whole, not merged. Paper loads this file and fills every key it does # Rewritten whole, not merged. Paper loads this file and fills every key it does
# not find with the default, then writes the full tree back — so a proxies-only file is a # not find with the default, then writes the full tree back — so a proxies-only file is a
# complete, stable input, and the lobby's other globals are simply always the defaults. # complete, stable input, and the lobby's other globals are simply always the defaults.
# That is true of a system server Felis owns end to end; if admins are ever allowed to tune # That is true of a system server Felis owns end to end; if admins are ever allowed to tune
+9
View File
@@ -18,6 +18,7 @@
# API — the SAME url the lobby build resolves, so this reuses it and adds no new dependency): # API — the SAME url the lobby build resolves, so this reuses it and adds no new dependency):
# docker build -f deploy/paper/Dockerfile \ # docker build -f deploy/paper/Dockerfile \
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-<ver>-<build>.jar \ # --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-<ver>-<build>.jar \
# --build-arg PAPER_JAR_SHA256=<that same sha — the objects/ path segment> \
# -t felis-paper:demo . # -t felis-paper:demo .
# docker save felis-paper:demo | sudo k3s ctr images import - # docker save felis-paper:demo | sudo k3s ctr images import -
# # felis.toml → recommended via 0019_recommended_paper.sql (no [velocity] key points here) # # felis.toml → recommended via 0019_recommended_paper.sql (no [velocity] key points here)
@@ -30,13 +31,21 @@
# to boot on anything older. # to boot on anything older.
FROM eclipse-temurin:25-jre FROM eclipse-temurin:25-jre
ARG PAPER_JAR_URL ARG PAPER_JAR_URL
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
# that shape at build time. Checking the digest after the download turns a truncated or
# tampered fetch into a failed build instead of a server booted on the wrong bytes.
ARG PAPER_JAR_SHA256
RUN set -eu; \ RUN set -eu; \
if [ -z "${PAPER_JAR_URL:-}" ]; then \ if [ -z "${PAPER_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \ echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \
fi; \ fi; \
if [ -z "${PAPER_JAR_SHA256:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_SHA256=<paper jar sha256> is required" >&2; exit 1; \
fi; \
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \ apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
mkdir -p /paper; \ mkdir -p /paper; \
curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \ curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \
echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/* apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*
COPY deploy/paper/entrypoint.sh /usr/local/bin/felis-entrypoint.sh COPY deploy/paper/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
@@ -1,41 +0,0 @@
# Foundational subsystems: the initial Felis import (ledger backfill)
- **Type:** feature (initial import) — retroactive ledger entry
- **Date:** 2026-06-26
- **Area:** `apis/`, `internal/` (naming, rcon, store, config, build, backup, operator,
submit, api, platform), `cmd/felis`, `plugins/`
- **Commits:**
- `7fbebfe` feat(apis): MinecraftServer CRD types (v1alpha1) — the lifecycle source of truth (§1)
- `708cdfc` feat(core): naming, RCON, store (Postgres + embedded migrations), config, image-build libraries
- `43ab921` feat(backup): archive-based world backup/restore + the retention/idle reaper
- `78b8cf6` feat(operator): MinecraftServer controller and reconcilers
- `d39605e` feat(submit): user modpack build + admin-approval pipeline (see [modpack-submission-lane](2026-06-26-modpack-submission-lane.md))
- `b508fcc` feat(api): dual-faced felis-api — permissions/LuckPerms, modpack lane, admin fleet read
- `47fcd90` feat(platform): node orchestration + the `cmd/felis` single-binary entrypoint
- `93f143f` feat(plugins): Velocity proxy + Fabric/Forge/NeoForge/Paper integration mods
- **Tasks:** #23 (permissions), #24 (modpack lane), #25 (fleet read)
## What it did
Stood up the whole backend spine in one build-order sweep: the Kubernetes CRD that is
the lifecycle source of truth, the core libraries (deterministic resource naming, the
RCON client, the Postgres store with embedded SQL migrations, config loading, container
image-build helpers), the backup/restore/reaper subsystems, the operator controller
that drives `MinecraftServer` resources, the user-modpack submit+approval pipeline, the
dual-faced (internal/external) felis-api behind a Zero-Trust guard, the platform
orchestrator that wires it all together under `cmd/felis`, and the server-side
integration plugins.
## Why
This is the project's first functional import — the substrate every later change edits.
It predates the change-ledger convention (established `fad48ff`, 2026-07-06), so it never
got a contemporaneous detail doc; this entry backfills one.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the
> change-ledger's detail-doc axis (§ Convention). This entry deliberately describes only
> what these eight commits **introduced** on 2026-06-26 — the named subsystems have been
> extended and reworked many times since (auth, passkey, metrics, quotas, updates), and
> that later work lives in its own dated detail docs, not here. Not independently
> re-verified for this doc; each subsystem was verified at its original commit and the
> current tree builds green at `9911b8c` (WSL oracle, go1.26.4).
@@ -1,30 +0,0 @@
# Modpack submission lane: build/approval pipeline + storage backends (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-06-26 – 2026-07-02
- **Area:** `internal/submit` (build/approval pipeline, storage backends), `internal/api` (submission endpoints)
- **Commits:**
- `d39605e` feat(submit): user modpack build + approval pipeline — an uploaded modpack stays `pending_review` and is never built until an admin approves; approval is a single-winner compare-and-swap handing off to the image-build Job, keeping the mandatory vulnerability scan in front of any push
- `598f3d3` feat(submit): local + S3 backends for modpack upload contexts, installer-selectable
- **Tasks:** #24 (§8 user-submitted modpack approval lane)
## What it did
Built the user-directed extension over the image-build subsystem: a player uploads a
modpack context, it sits in `pending_review`, and an admin's approval is the single-winner
gate that hands off to the build Job — with the vulnerability scan always ahead of any
registry push. `598f3d3` makes the upload-context store pluggable (local filesystem or S3),
selectable at install time.
## Why
Untrusted user content must never build or push unreviewed, and the compare-and-swap
approval guarantees exactly one build per submission even under a double-click or retry.
The storage-backend choice lets a single-node demo use local disk while a real deployment
uses S3, without a code change.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The approval
> compare-and-swap and endpoints were unit-tested at their commits; the S3 path is
> integration-configurable. The panel-side submission/approval UI is the collaborator's
> frontend work and is tracked only by its INDEX rows. Not independently re-verified for
> this doc; current tree green at `9911b8c`.
@@ -1,38 +0,0 @@
# Cloudflare Tunnel + Access edge (cfsetup) + NodePort fencing (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-01
- **Area:** `internal/cfsetup` (pure core + integration runner), `cmd/felis` (TUI edge flow), edge nftables fence
- **Commits:**
- `53a7664` feat(cfsetup): recommended Cloudflare Tunnel + Access edge (§14) — domain- and IdP-agnostic; the load-bearing `validateFailClosed` allowlist refuses any policy that could be public; fail-shut 404 catch-all; the raw game host is never proxied
- `ba13839` feat(breakglass): optional Tunnel + Access setup in the sudo TUI, an independent peer of Owner provisioning
- `a531f5e` fix(cfsetup): keep the connector install in the host apply layer only (drop the duplicate `StartConnector`)
- `2810fe8` fix(cfsetup): repoint a stale DNS record when routing a tunnel hostname
- `7d3be64` feat(cfsetup): start the tunnel connector as a setup step
- `346ec68` refactor(deploy): rework the cloudflare-edge walkthrough — restructured the edge TUI flow and added a tested `cfsetup` integration-runner path (with TUI height-measure/root tests)
- `e058a64` feat(edge): close the panel NodePort to the public after the tunnel is up — nftables at prerouting `raw` (-300), before kube-proxy's NodePort DNAT, gated on the connector actually serving; loopback accepted first so the connector origin hop is untouched
- **Tasks:** #37 (fence panel NodePort to public after tunnel)
## What it did
Stood up the optional one-click Zero-Trust edge: a Cloudflare Tunnel routing only the web
hostnames plus a fail-closed Access application, provisioned from the sudo TUI against the
operator's own Cloudflare account. `e058a64` then closes the Access-bypass hole where a
direct `https://<node-ip>:<nodeport>/` with the right Host header reached the origin
behind Access, by fencing the NodePort at the nftables raw hook so the packet is caught on
its original destination port — but only once the connector is confirmed serving, so
fencing never severs the only web path to a live origin.
## Why
Access is only a security boundary if the origin cannot be reached around it. The
fail-closed policy guard (`validateFailClosed`) and the NodePort fence are the two
load-bearing safety properties: a policy that could be public aborts the run with nothing
created, and a routable-but-unfenced NodePort would defeat the whole edge.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The policy guard,
> ingress generation, request bodies, gating, and the nftables ruleset shape / conn-count
> gate are unit-tested; the live cloudflared/Cloudflare-API and `nft` calls are
> INTEGRATION-ONLY (need a real account). KNOWN-LIMITATION: the fence targets nftables;
> firewalld-native coordination is deferred. Not independently re-verified for this doc;
> current tree green at `9911b8c`.
@@ -1,32 +0,0 @@
# Console auth: local-password login → passwordless migration (ledger backfill)
- **Type:** feature + refactor — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-04
- **Area:** `internal/api` (auth handlers, sessions), `internal/store` (users schema)
- **Commits:**
- `af14f02` feat(api): local-password authentication backend — login/logout/change-password on `op.console`; HttpOnly+Secure+SameSite=Lax host-only server-side sessions (SHA-256, 12h TTL); anti-enumeration uniform bcrypt; JSON-only credential writes (415 otherwise); fails closed unless `local_auth_enabled`
- `0c1cc59` feat(auth): migrate console login to passwordless
- `3b43f05` refactor(api): drop the dead login concurrency limiter and reconcile passwordless comments
- `c20b12c` refactor(api): drop the dead password-era `ResetMailer`, reconcile passkey-unbind docs
- **Tasks:** #27 (B1 thin thread), #79/#80/#81 (residue sweep + primitive adjudication)
## What it did
Shipped the staff local-password door (`af14f02`) as the primary web login when
Zero Trust is not in front of the API, then migrated the console to passwordless
(`0c1cc59`) once email-OTP + passkey were the intended factors. The two refactors
(`3b43f05`, `c20b12c`) then swept the password-era residue — the now-dead login
concurrency limiter and the `ResetMailer` — so no unused password machinery lingered in
the compile path, and reconciled the stale comments that referenced it.
## Why
`op.console` needs a real login even in deployments without a Cloudflare-Access edge; the
password backend was that. Once the passwordless factors landed, keeping the old password
scaffolding around was a bug farm — the sweep is the closeout evidence that the migration
was complete, not half-done.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. `af14f02` was
> covered by Go unit tests (content-type guard, anti-enumeration, forced-change lockdown)
> at its commit. Not independently re-verified for this doc; current tree green at
> `9911b8c` (WSL oracle, go1.26.4).
@@ -1,38 +0,0 @@
# Deploy: one-line bootstrap installer + demo bring-up (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-03
- **Area:** `deploy/` (bootstrap.sh, Dockerfiles, demo-up.sh), image build context
- **Commits:**
- `58fa4b0` feat(deploy): one-line bootstrap installer + distroless felis image (auto-detects apt/dnf, installs Docker/k3s/PostgreSQL, opens pg_hba to the pod CIDR, runs migrations, applies the control-plane bundle, leaves Web disabled pending `felis setup`)
- `94a3b7b` fix(deploy): harden bootstrap for RHEL-family Linux
- `deaa2f8` feat(deploy): zypper support (openSUSE/SLES)
- `318a724` feat(deploy): pacman support (Arch)
- `e5f1682` refactor(deploy)!: TUI (breaking walkthrough restructure)
- `28c3eee` refactor(deploy): improved TUI walkthrough
- `c14ed17` fix(docker): keep embedded `panel/` and `deploy/` in the image build context
- `d9e866f` fix(deploy): make the lobby image actually build (re-include `plugins/paper`, build on `gradle:8.14-jdk21`)
- `b84debf` feat(deploy): one-shot `demo-up.sh` — bootstrap → build/import limbo+lobby images → wire `[velocity]` image refs → `felis setup`, ending in the interactive Owner TUI
- **Tasks:** #26 (Phase A bootstrap verified end-to-end on the Demo VM)
## What it did
Made a bare Linux box a running Felis with one command. `bootstrap.sh` auto-detects the
host package manager across the four major families (apt/dnf/zypper/pacman), installs
whatever is missing (Docker, k3s, PostgreSQL, cloudflared), builds+imports the distroless
felis image, opens `pg_hba` to the pod CIDR, runs migrations, and applies the rendered
control-plane bundle. `demo-up.sh` wraps that plus the login-limbo/lobby image build and
`felis setup` into a single command, stopping only at the Owner-creation TUI it cannot
automate.
## Why
The spec calls for a self-hostable single-node deployment a SysAdmin can stand up without
a Kubernetes background. The package-manager fan-out and the demo wrapper are what make
"one line" true across real distros rather than only on the author's box.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the
> change-ledger's detail-doc axis. `deploy/` is shell + Dockerfiles (not Go-oracle
> verifiable); `d9e866f` records a real build+boot check (limbo `/healthz` 200, lobby
> reaches "Done"). Not independently re-verified for this doc; current tree green at
> `9911b8c`.
@@ -1,35 +0,0 @@
# felis CLI: break-glass recovery console + first-run setup (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-06-30
- **Area:** `cmd/felis` (break-glass/setup TUI, apply, migrate), `internal/api` (audit, owner store), `deploy/`
- **Commits:**
- `e108a37` feat(cli): break-glass emergency console TUI — root-only (`euid==0`), provisions/resets the Owner directly against Postgres, enables local login, prints a durable one-time-password summary
- `2d0bbb0` feat(cli): attribute break-glass recovery to the SysAdmin who runs it — bootstrap / recovery (bcrypt) / root-override, each audited with an honest `verified` flag and payload
- `a94b001` feat(deploy): break-glass Operator account provisioning
- `eb5875a` feat(felis): Operator break-glass op behind an operation menu
- `f5d00f3` feat(cli): `felis apply` for direct CRD creation
- `9c46632` feat(cli): `felis setup` first-run console (shared `runConsoleTUI` model, reclaim protection, cfsetup idempotency, `[auth].admin_hostname` respect)
- `7d91373` fix(migrate): honor `-config` placed after the `up` verb (flag.Parse stops at the first non-flag token)
- **Tasks:** #27 (B1 login→change-pw→TUI reset)
## What it did
Built the local-root recovery and first-run surface that bypasses web Zero Trust by
design. `felis breakGlass` mints or resets the Owner when the web login is unreachable;
`2d0bbb0` makes it accountable by recording *which* SysAdmin broke the glass across three
audited modes. `felis setup` is the non-emergency first-run twin sharing the same console
model. `felis apply` writes a `MinecraftServer` CRD directly, and `7d91373` fixes the
`migrate` flag parse so a configured DB path after `up` is honored.
## Why
An operator with root on the node and a kubeconfig must always be able to recover the
platform — that is break-glass's whole job, so it never refuses. Attribution
(`2d0bbb0`) closes the gap that root is machine authority, not a human identity: the root
gate is necessary but not sufficient for the audit trail.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The core logic was
> covered by Go unit tests over a fake owner store at each commit (auth match/non-match,
> the three audit modes, headless TUI drive). The bubbletea TUI glue is untested by house
> convention. Not independently re-verified for this doc; current tree green at `9911b8c`.
@@ -1,36 +0,0 @@
# Player onboarding data layer §B2: email-OTP, account-link, QR, Bind-Code (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-03
- **Area:** `internal/api` (onboarding/auth-bind handlers), `internal/store` (migrations 0004–0006)
- **Commits:**
- `dbe34a1` feat(api): player email-OTP verification (§B2) — `POST /account/email/{start,verify}`; 6-digit code, SHA-256-at-rest, 10-min TTL, 5-attempt cap enforced in the repo
- `1f8b9bb` feat(api): record account-link auth source (`mojang|thirdparty`) for the dual-Yggdrasil split (§10)
- `116595f` feat(api): QR scan-login completion poll on the internal face (`GET /internal/account/link/status/{mc_uuid}`) — read-only, reuses `UserByMCUUID`, no migration
- `fe2ece0` feat(api): public Bind-Code onboarding (`POST /auth/bind`) — the one pre-account entrypoint of `console.<root_domain>`; refuses a staff-UUID code with 403 without consuming it, so the public door provably never yields an admin principal
- `55592ed` feat(auth): public auth-bind endpoint wiring
- `6c3999a` fix(api): rate-limit email-OTP sends to close the email-bomb vector
- `879b177` fix(api): make OTP-start throttle atomic to close the concurrent-burst bypass
- **Tasks:** #29 (B2 data layer), #32 (OTP rate-limit), #35 (atomic throttle), #39 (console access model)
## What it did
Built the Go-verifiable data layer of forced web onboarding: prove control of an email
(OTP), record which Yggdrasil authenticated an in-game UUID, let a phone already signed in
to the panel complete a QR device-code link, and let an account-less player redeem a
one-time Bind Code minted in the Login Lobby to create+link+session in one public step.
The two fixes bound the OTP abuse surface — a per-target send rate limit and an atomic
reserve that closes the check-then-act race on the attempt counter.
## Why
The spec forces onboarding through the web so every account is provably email-controlled
and UUID-linked before it can operate anything. The `op.console` redline in `fe2ece0` — a
staff-UUID code is refused without being consumed — is what keeps the public console door
from ever minting an admin principal.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The account/session
> logic, single-use codes, and the op.console redline were covered by handler tests + the
> OpenAPI parity gate at each commit; the identity guarantee behind a Bind Code lives in
> velocity/Java (CODE-ONLY) and is not verifiable from this repo. Not independently
> re-verified for this doc; current tree green at `9911b8c`.
@@ -1,32 +0,0 @@
# §B3 username-collision reclaim + account migration (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-05
- **Area:** `internal/api` (internal-face reclaim/blacklist, account migrate), `internal/store` (migration 0006)
- **Commits:**
- `a29571d` feat(api): reclaim squatted usernames for Mojang-priority players (§B3, 正版优先) — `POST /internal/player/reclaim` bars the squatter UUID + stashes its data (30-day hold) in one transaction, idempotent, returns the *first* reclaim's expiry; `GET /internal/player/blacklist/{mc_uuid}` is the login-gate check
- `fdb6efb` feat(account): migrate a live account's owned servers to a new account (§B3 inherit)
- **Tasks:** #30 (B3 game-login + username-collision reclaim)
## What it did
Built the data layer of the Mojang-priority collision flow: when the configured
third-party Yggdrasil and official Mojang issue the same username under different UUIDs,
the non-genuine squatter is displaced in favour of the real Mojang owner. Both tables are
keyed by `mc_uuid`, so the genuine player — identical username, *different* UUID — is
never caught by the bar. `fdb6efb` adds the inherit half: migrating an existing account's
owned servers onto a new account.
## Why
Two players cannot hold one username across two Yggdrasils; the spec resolves it in the
genuine Mojang owner's favour with a 30-day data hold for the displaced squatter, told the
truth about how long their data is kept (the first hold's window, never a fresh `now()+30d`
on retry).
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Handlers + the
> in-memory repo contract were unit-tested at each commit; the Postgres SQL path is
> integration-only, and the velocity collision-routing / limbo prompt / authlib
> dual-backend are code-only (Java) and out of this data-layer slice. Not independently
> re-verified for this doc; current tree green at `9911b8c`. Related: the operator-facing
> `/felis migrate` command has its own doc ([felis-migrate-command](2026-07-05-felis-migrate-command.md)).
@@ -1,39 +0,0 @@
# felis-api security + robustness hardening (audit sweep) (ledger backfill)
- **Type:** fix — retroactive ledger entry
- **Date:** 2026-06-30 – 2026-07-01
- **Area:** `internal/api` (login, request-id, listeners, SSE relays, quota/claim, MyServers), `internal/operator`
- **Commits:**
- `7a51c1d` fix(api): bound concurrent login bcrypt to shed CPU-pin floods (429 `auth_busy` before the compare; a cap, not a per-account lockout) — *audit #2*
- `164ac44` fix(api): validate inbound `X-Request-Id` before echo + audit persist (≤64 bytes, log-safe charset) — *audit-integrity*
- `c6c0772` fix(api): read/idle timeouts on all three listeners via a `newAPIServer` factory (closes Slowloris via `ReadHeaderTimeout`; `WriteTimeout` left unset so SSE isn't severed) — *audit #3*
- `3c1d647` fix(api): per-principal SSE stream cap (429 `too_many_streams`) — *audit #1, blast-radius bound*
- `d6e3189` fix(api): per-write deadline on SSE relay to sever a stalled reader (the real leak close behind the cap) — *audit #1*
- `8f41a00` fix(api): clear the SSE write deadline on return so it can't leak onto a reused keep-alive connection — *audit #1*
- `6368ab1` fix(api): `COALESCE` the MyServers `owned` flag so an ownerless row doesn't 500 the listing
- `2a4a81b` fix(api): don't burn the wake cooldown when refused at capacity
- `9873904` fix(operator): populate `Status.Players` from an RCON `list` probe (so the panel doesn't report 0/0)
- **Tasks:** #33 (wake cooldown), #34 (Status.Players), #41–#46 (audit #1–#4)
## What it did
A hardening sweep across the API's abuse and robustness surface: bound the two unbounded
CPU/goroutine amplifiers (concurrent bcrypt, per-principal SSE streams), close the SSE
relay's real stalled-reader leak with a per-write deadline (and clear it so it can't leak
onto a pooled connection), validate the caller-supplied request id before it reaches the
audit trail, set listener timeouts to close Slowloris, and fix two functional bugs — the
ownerless-row 500 and the wake cooldown burned on a capacity refusal.
## Why
Each is a specific, demonstrated failure mode: a login flood pins every core in bcrypt; a
stalled SSE reader leaks a relay goroutine + its upstream kube-apiserver follow *for the
life of the process*; an unvalidated `X-Request-Id` is a CR/LF log-forgery vector. The
`WriteTimeout`-left-unset detail is load-bearing — a blanket write timeout would sever the
healthy long-lived console/build-log streams the platform depends on.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Each fix shipped a
> targeted test at its commit — notably `d6e3189`/`8f41a00` use a deadline-aware
> `ResponseWriter` that fails closed if the guard is removed. The quota-claim TOCTOU
> (audit #4) is a documented KNOWN-LIMITATION (`2c56d17`), closeable only against a real
> Postgres. Not independently re-verified for this doc; current tree green at `9911b8c`.
-25
View File
@@ -1,25 +0,0 @@
# felis_* Prometheus metrics (§23) (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-06-30
- **Area:** `internal/metrics` + the emit sites in build, platform/fleet, and the start lifecycle
- **Commits:**
- `75642d9` feat(metrics): named `felis_*` Prometheus collectors
- `2a93a9e` feat(metrics): record `felis_image_build_failures_total` on failed builds
- `79eae7f` feat(metrics): publish `felis_servers_total` from a fleet snapshot
- `8ac5e64` feat(metrics): observe `felis_start_duration_seconds` across the start lifecycle
- **Tasks:** #17 (§23 felis_* metrics decision)
## What it did
Added the named `felis_*` collector set and wired the three emit points that make it
non-empty: a counter incremented on image-build failure, a gauge published from a fleet
snapshot, and a histogram observed across the server start lifecycle.
## Why
§23 calls for first-class operational metrics under a stable `felis_` namespace rather than
ad-hoc logging, so an operator can alert on build failures, fleet size, and start latency.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Not independently
> re-verified for this doc; current tree green at `9911b8c` (WSL oracle, go1.26.4).
@@ -1,37 +0,0 @@
# Auto-update subsystem: decision core + sources + gatherer + window API (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-07-01 – 2026-07-05
- **Area:** `internal/updates` (pure decision core), `internal/updater` (release sources, gatherer), `internal/api` (window admin API)
- **Commits:**
- `c01f133` feat(updates): pure I/O-free decision core — each tracked component is Pinned (Minecraft, left alone), Notify, or Scheduled (apply only inside a SysAdmin window); never force-applied, never a downgrade, never an auto-applied prerelease
- `3673af6` feat(api): admin API for the maintenance window (`GET`/`PUT /updates/window`), stored as JSON under `platform_settings` — API + persistence only, nothing consumes it yet
- `7464fa7` fix(updates): tag `Window` JSON so the persisted window round-trips (the obvious decode is correct by construction; a zero window fails closed to notify-only)
- `96b3cc9` feat(updater): wire `updates.Run` to a caller with PaperMC v3 release discovery
- `7d27640` feat(updater): GitHub Releases source, routing felis-api/k3s/cloudflared
- `7db57b9` feat(updater): `VersionGatherer` extraction core + CLI gather seam
- **Tasks:** #38 (auto-update: Felis/k3s/components/Velocity, pin Minecraft)
## What it did
Built the auto-update spine as a pure decision core plus the release-discovery sources
(PaperMC, GitHub Releases) and the version gatherer, with a SysAdmin-set maintenance
window read/written through an admin API. Version parsing tolerates the real feeds (leading
`v`, k3s `+k3s1` suffix, calendar versions, prerelease tails) and orders by SemVer
precedence.
## Why
The red lines are `不要强制自动更新` (never force auto-update) and `能不动的就别动`
(Minecraft stays pinned). The design encodes them structurally: a component may be applied
*only* inside a window the operator explicitly set, and Minecraft is Pinned so it is never
touched. `7464fa7`'s fail-closed zero-window (decodes to notify-only, never a rogue apply)
is the safety property for the not-yet-built runner.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The load-bearing
> invariants (pinned never changes, no downgrade, no auto-prerelease, apply-only-in-window)
> and the JSON round-trip contract were unit-tested at their commits. This subsystem is
> deliberately **report-only / integration-deferred**: the concrete Notifier/Applier,
> the `felis update` CLI + CronJob, and the current-version producing seams are declared
> but not wired (see `internal/updater/doc.go`, `openapi.yaml`). Not independently
> re-verified for this doc; current tree green at `9911b8c`.
@@ -1,37 +0,0 @@
# Passkey (WebAuthn) enrollment subsystem + hardening (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-07-01 – 2026-07-02
- **Area:** `internal/passkey` (go-webauthn adapter), `internal/api` (enrollment handlers/audit), `internal/store` (migrations 0007–0009)
- **Commits:**
- `f2c916d` feat(api): passkey enrollment persistence layer
- `742f15f` feat(api): passkey enrollment endpoints
- `0261204` feat(passkey): go-webauthn enrollment verifier adapter (Oracle-verified against a virtual authenticator)
- `fce0fce` feat(passkey): wire the enrollment verifier into felis-api
- `7278cd7` feat(passkey): require + record user verification at enrollment (`UserVerification=required`; capture `user_verified`/`backup_eligible`/`backup_state` — migration 0009) — *fix (d)*
- `cdbb5ab` fix(api): record credential id in the passkey-register audit event so bind/unbind are symmetric — *fix (a)*
- `9953275` fix(api): bound `webauthn_challenges` growth by superseding *all* prior rows per (user, purpose) — *fix (b)*
- `20e31fb` fix(store): cascade-delete passkeys + challenges on user removal (recreate both FKs `ON DELETE CASCADE`, scoped to the passkey tables only) — *fix (c)*
- `54bc6ef` fix(api): clear bound passkeys on password change to close a takeover foothold — *fix (e)*
- **Tasks:** #36 (passkey bind with email-OTP fallback), #48–#52 (fixes a–e)
## What it did
Built the WebAuthn *enrollment* half — persistence, the go-webauthn crypto adapter, and
the register-begin/finish endpoints — then hardened it through the five-fix batch (a–e):
symmetric audit, a bounded challenge table, cascade cleanup, enforced+recorded user
verification, and unbinding every passkey on a password reset so a passkey planted through
a transiently-hijacked session cannot survive as a standing login foothold.
## Why
Passkeys are the phishing-resistant factor with email-OTP as the fallback. The hardening
batch closes the seams that make enrollment safe to *rely on*: without UV enforcement a
passkey proves possession but not user; without the password-reset clear, a planted
passkey outlives the very remediation meant to evict an attacker.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The adapter crypto
> was verified against a virtual authenticator (virtualwebauthn), and each fix shipped
> with a targeted test (UV-negative rejection, challenge-growth bound, cascade, symmetric
> audit) at its commit. Not independently re-verified for this doc; current tree green at
> `9911b8c`. The assertion/login half is a separate doc ([passkey-login](2026-07-01-passkey-login.md)).
-36
View File
@@ -1,36 +0,0 @@
# Passkey (WebAuthn) login: assertion, discoverable, clone-detection (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-07-01 – 2026-07-05
- **Area:** `internal/passkey` (assertion crypto), `internal/api` (login/assertion, unbind, UA-guard), `internal/store` (migrations 0013/0014)
- **Commits:**
- `e035142` feat(passkey): WebAuthn login/assertion crypto adapter (BeginLogin/FinishLogin over go-webauthn, Oracle-verified against a virtual authenticator; surfaces the signature counter as a ceremony fact)
- `ec468ba` feat(auth): discoverable (usernameless) passkey login — the from-zero door the username-first assertion couldn't key on
- `0dbd557` fix(store): renumber the discoverable-login migration 0013 → 0014
- `9e1df12` feat(passkey): advance `sign_count`, reject clone-warned assertions
- `4f59d51` feat(auth): owner-tier passkey-unbind remediation endpoint
- `a63f49d` feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey — a backend-only UA interstitial (the SPA is untouched); asset/API/health requests pass through, an `ua_ack` cookie lets a determined user continue
- **Tasks:** #40 (from-zero discoverable login), #67 (WeChat/QQ UA-guard in `internal/panel`)
## What it did
Built the assertion (login) half of the ceremony: the crypto adapter, then discoverable
credentials so a user with no typed identifier can still log in (the enrollment
identifier problem the earlier deferral doc named), clone detection via the advancing
signature counter, and the owner-tier unbind remediation. `a63f49d` guards the flow at the
transport edge — WebAuthn is unusable inside the WeChat/QQ WebViews, so those UAs get a
bilingual "open in your system browser" page instead of the passkey SPA.
## Why
Enrollment without a login path is half a feature. Discoverable credentials resolve the
blocker recorded in the earlier deferral (`users.email` is nullable/non-unique and a
player's username is their Minecraft UUID, so username-first assertion had nothing to key
on). The UA-guard stops the most common real-world dead end: a passkey prompt that can
never succeed inside an in-app browser.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The assertion crypto
> was verified against a virtual authenticator (enrollment→assertion chain, origin-mismatch
> and unbound-credential rejection); `9e1df12`'s clone policy and the UA-guard pass-through
> were unit-tested at their commits. Not independently re-verified for this doc; current
> tree green at `9911b8c`.
@@ -1,37 +0,0 @@
# System servers: login-limbo + lobby (always-on gate) (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-07-02
- **Area:** `internal/config`, `internal/naming`, `internal/api` (CRD readiness), `internal/operator`, `internal/platform`, `cmd/felis`, `plugins/limbo`, `deploy/limbo` + `deploy/lobby`
- **Commits:**
- `9bed51b` feat(config): `[velocity] login_image/lobby_image` — setup provisions the always-on system services only when set (empty = fail-loud skip; no official LOOHP/Limbo image exists)
- `9ef817f` feat(naming): reserved system-server names + service-token identifiers (single source of truth for the internal-API credential Secret)
- `159107b` feat(api): HTTP readiness knob on `MinecraftServer` + user-server fallback defaults to the login gate
- `dc23cb5` feat(operator): system-server pod HTTP readiness probe + login-only `FELIS_SERVICE_TOKEN` env (keyed off the reserved name so it can never leak into a user pod; sourced via `secretKeyRef`, never inlined)
- `3fdb3d0` feat(platform): internal-API base-URL helper + single-sourced token Secret
- `f554d52` feat(cli): provision the reaper-exempt login/lobby servers + replicate the service-token Secret into the minecraft namespace
- `241fe21` feat(limbo): felis-limbo in-game login flow (join → blacklist check → mint bind code → open book to `console.<root_domain>` → poll link-status → BungeeCord transfer to lobby; fail-closed)
- `c7315e4` feat(deploy): login-limbo + lobby images with game-port pinning (server-port pinned to GamePort 25565 on every start)
- **Tasks:** #53–#68 (system-server plumbing L1–L4, limbo plugin, operator env injection)
## What it did
Stood up the always-on authentication gate: reserved, reaper-exempt login/lobby
`MinecraftServer`s provisioned by setup, an HTTP readiness path for the RCON-less LOOHP/Limbo
loader (which reports "started" only after the first tick), and the felis-limbo plugin that
runs the whole onboarding *inside* Limbo before transferring an admitted player to the
lobby. A fresh connection always lands on the login gate, never a user backend, so
authentication is always in front.
## Why
The spec requires that a player authenticate before reaching any real server. That needs a
purpose-built always-on front server (Limbo) that speaks to the internal API — hence the
login-only service-token injection (keyed to the reserved name so it can never reach a user
pod) and the HTTP readiness knob for a loader that has no RCON.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The Go layer
> (config/naming/readiness/operator env/platform) was unit-tested at each commit; the
> felis-limbo plugin is Java verified against a real Limbo jar via podman (#65), and the
> images carry a real build+boot check (#63, limbo `/healthz` 200 on 25565). Not
> independently re-verified for this doc; current tree green at `9911b8c`.
@@ -1,83 +0,0 @@
# Break-glass "halt a running server" op (#31 B4)
- **Type:** feature (addition)
- **Date:** 2026-07-05
- **Area:** `cmd/felis` — break-glass recovery console (Go, oracle-verifiable)
- **Commit:** `c2ee21a` — feat(breakglass): add halt-a-server op to the recovery console (§B4)
- **Task:** #31 Phase B4 (felis TUI break-glass ops)
## What it does
Adds a **"Halt a running server"** operation to the root-gated break-glass console.
The operator picks a server from the live fleet and the console flips that
`MinecraftServer` CRD's `spec.desiredState` to `Stopped`, letting the operator
reconcile it into a graceful shutdown. It is the emergency "stop this now" lever for
when the panel is unreachable but the box still has `root` + a kubeconfig.
## Why
The break-glass console already provisions the Owner and adds Operators, but there
was no local, panel-independent way to **stop** a misbehaving server (runaway,
compromised, resource-pinning). Halting is a reversible state nudge — the safest
possible break-glass power — so it belongs in the same root-gated recovery surface.
## Design decisions
- **CRD write, not pod kill.** The console flips `spec.desiredState=Stopped` with a
**spec-only merge patch** (`client.MergeFrom`), never a full-object `Update`. The
operator writes `status` on the same object continuously; a merge patch of
`spec.desiredState` touches a disjoint field and cannot race/clobber the operator's
status writes. A halt is therefore exactly the CRD write the operator already knows
how to honour.
- **Authority = root + kubeconfig.** The accountable actor is the OS user who
escalated to root (`osUser`), recorded for attribution — not proof. The root gate
plus kubeconfig possession *is* the authority, so (unlike the owner/operator paths)
no credential-minting auth sub-flow is needed for a reversible state change.
- **System servers allowed but named.** Halting the `login`/`lobby` system servers
takes the shared front door down (login has no fallback). Break-glass is deliberately
full power, so the console **warns** rather than forbids: a `⚠ system` tag in the
picker and an explicit `WARNING` line in the post-exit summary.
- **Audit is best-effort.** `performHalt` mirrors `performBreakGlass`: the halt
succeeds even if the audit sink is down (break-glass must work with logging broken);
any audit error rides back in the outcome and is surfaced as a summary `WARNING`.
- **Already-stopped is a no-op** reported distinctly ("was already stopped" vs "is now
stopping"), so the console never claims a stop it didn't perform.
- **Namespace from config.** The target namespace is `cfg.K8s.Namespace`, threaded
through the console constructors — never hardcoded.
## Files
| File | Change |
|---|---|
| `cmd/felis/halt.go` | **new** — pure core (no bubbletea): `listServersForHalt`, `haltServer` (merge patch), `isSystemServer`, `performHalt`, `auditHalt` |
| `cmd/felis/halt_test.go` | **new** — table tests against a controller-runtime **fake client** (applies patches for real): running→stopped persists, already-stopped no-op, missing→error, system flag, list projection + desired-state fallback, audit success, audit-failure-still-halts |
| `cmd/felis/tui_halt.go` | **new** — bubbletea/huh shell mirroring `ownerModel` (load → pick → work → done), empty-fleet guard, `⚠ system` picker labels, outcome card |
| `cmd/felis/tui_menu.go` | `bgHaltServer` enum + "Halt a running server" menu option |
| `cmd/felis/tui_root.go` | `namespace` field; `bgHaltServer` dispatch to `newHaltModel`; `haltResultMsg` terminal handling |
| `cmd/felis/breakglass.go` | halt fields on `breakGlassResult`; `namespace` threaded through `runBreakGlassTUI`/`runSetupTUI`/`runConsoleTUI`; post-exit halt summary (stopping / already-stopped, system + audit warnings, restart hint) |
| `cmd/felis/setup.go` | pass `cfg.K8s.Namespace` into `runSetupTUI` |
| `cmd/felis/tui_root_test.go` | pass `"minecraft"` namespace into `newRootModel` test call |
## Verification
WSL oracle (go1.26.4, FedoraLinux-44), authoritative for Go:
```
go build ./... → BUILD_OK
go vet ./cmd/felis/... → VET_OK
go test ./... → all 20 packages ok, ALL_GREEN
```
The core (`halt.go`) is fully unit-tested against a real `fake.Client`, which applies
the merge patch, so the test asserts the **persisted** `spec.desiredState`, not merely
that `Patch` was called. `tui_halt.go` is thin bubbletea glue (untested by house
convention, mirrors the existing `tui_owner.go`).
## Self-review outcome
- **ponytail (over-engineering):** lean — no one-impl interface, every field consumed,
audit seam justified. Nothing cut.
- **correctness:** caught and fixed a misleading restart hint — the summary originally
pointed at `felis apply`, but that command is **create-only** (errors "already
exists" on an existing server); corrected to "restart from the panel, or set
`spec.desiredState` back to Running."
@@ -1,80 +0,0 @@
# `/felis migrate` in-game command (§B3 inherit, Velocity side)
- **Type:** feature (addition)
- **Date:** 2026-07-05
- **Area:** `plugins/velocity` + `plugins/shared` — Velocity proxy plugin (Java, compile-verified)
- **Commit:** `c1aa38b` — feat(velocity): add /felis migrate to open an account migration (§B3 inherit)
- **Task:** completes the code-only gap named in `internal/api/handlers_account_migrate.go`
## What it does
Adds the in-game `/felis migrate` command that a player runs to **open an account
migration** — the first step of handing their owned servers to another account (spec
§B3 "inherit", scenario A). The command posts the player's Mojang-verified UUID to the
backend, which puts that account into migrate mode (`state=initiated`). The player then
finishes the migration on the web console (prove it's them, name the receiving account,
redeem a one-time code).
The Go backend (`handleMigrateStart` and the web-driven steps 2–4) already existed and
was tested; its header comment explicitly named **"the `/felis migrate` command that
calls handleMigrateStart"** as the code-only gap. This change closes that gap.
## Why
Without the in-game command, the migration flow had no entry point — the backend
handler was reachable only in theory. `/felis migrate` is the trustworthy initiator:
Velocity has already established the caller's online-mode UUID, so the sensitive proof
can be deferred to the web step-up while the in-game command just opens the migration.
## Design decisions
- **Mirrors the existing command suite verbatim.** `doMigrate` follows `doClaim`;
`migrateError` follows `claimError`; `migrateStart` follows `claim`/`opLoginApprove`.
No new imports, types, or idioms — every construct already appears in the same files.
- **Identity-bound + out-of-limbo, but server-independent.** Like `claim`, it requires
a real player past the login limbo (`requirePlayer` + `ensureOutOfLimbo`). Unlike
`claim`, it acts on the caller's *account*, not the server they stand on, so there is
**no** `registry`/current-server check.
- **Expects HTTP 201.** `migrateStart` posts to
`/api/v1/internal/account/migrate/start` and expects **201 Created** (`handleMigrateStart`
returns `StatusCreated`) — not 200 like the other calls. A 201 that does not affirm
`started:true` is treated as a contract breach, not a refusal.
- **Error mapping matches the handler's refusals:** 404 `not_linked` → "Link your
account on the web console before migrating"; 409 `account_retired` → "This account
can't start a migration (already migrated or retired)"; transport (0) and default →
generic retry text.
- **Points the player to the console on success.** The command only *opens* the
migration, so on success it prints the player web console URL
(`https://console.<root_domain>`, derived from config — never a hardcoded domain) and
a one-line description of the remaining steps. A proxy-side `logger.info` records the
initiating username against the UUID (the backend audit only has the UUID).
## Files
| File | Change |
|---|---|
| `plugins/shared/.../link/FelisApiClient.java` | **+`migrateStart(UUID)`** — POST mc_uuid, expect 201, affirm `started:true` |
| `plugins/velocity/.../FelisVelocityPlugin.java` | `migrate` literal in the Brigadier tree; **`doMigrate`** handler; **`migrateError`** mapper; `/felis migrate` help line |
## Verification
Java is not oracle-verifiable via the Go suite, but it **is** compile-verifiable via
the podman gradle toolchain established in #63/#65:
```
podman run --rm -v plugins:/work -w /work/velocity \
docker.io/library/gradle:jdk17 gradle --no-daemon compileJava
→ BUILD SUCCESSFUL in 19s (compiled against real velocity-api:3.3.0-SNAPSHOT)
```
The change compiles clean against the real Velocity API jar (including the shared
`FelisApiClient` compiled straight into the velocity module). The backend contract it
speaks to (`handleMigrateStart`) is covered by `handlers_account_migrate_test.go` on
the Go side.
## Self-review outcome
- **ponytail (over-engineering):** lean — pure mirror of three existing, compiling
methods; no speculative abstraction. Nothing cut.
- **correctness:** the one contract divergence (201 vs 200) was verified against the Go
handler source before writing.
@@ -1,29 +0,0 @@
# Operator: idle auto-stop, quotas, startup/readiness timeouts, /readyz (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-07-05
- **Area:** `internal/operator` (idle stop, timeouts), `internal/api` (quotas, /readyz)
- **Commits:**
- `91bfa27` feat(operator): idle auto-stop (§8)
- `e574749` feat(api): enforce CPU/memory/storage quotas (§9.3, §22)
- `7f7e459` fix(operator): enforce startup and readiness timeouts (§5, §8)
- `7becb38` fix(api): implement `/readyz` with real DB + K8s API + CRD checks (§7)
- **Tasks:** §5/§7/§8/§9.3/§22 operator + resource-governance spec items
## What it did
Rounded out the operator's lifecycle governance: stop idle servers automatically, enforce
per-resource CPU/memory/storage quotas at claim/create, bound how long a server may sit in
startup/readiness before the operator gives up, and make `/readyz` a real dependency check
(DB, Kubernetes API, and the CRD) rather than a static 200.
## Why
An orchestrator that never reclaims idle capacity or bounds startup will accumulate stuck
and wasteful workloads; a `/readyz` that always returns 200 tells the load balancer a
broken control plane is healthy. These are the spec's resource-governance and
readiness-correctness requirements (§5/§7/§8/§9.3/§22).
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Covered by Go unit
> tests at each commit. Not independently re-verified for this doc; current tree green at
> `9911b8c` (WSL oracle, go1.26.4).
@@ -1,117 +0,0 @@
# Break-glass "back up a world now" console peer (§B4 "Sync", phase 2b)
- **Type:** feature (addition)
- **Date:** 2026-07-07
- **Area:** `cmd/felis` (Go, oracle-verified); `internal/api` + `docs/openapi.yaml`
(the internal endpoint's `os_user` accountability extension)
- **Commit:** `fc748d3`
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation. Per the user's
**"两者都要"** decision the feature was built in two halves: the felis-api endpoint
that does the real backup-Job orchestration (phase 1 `7a7c0d5` external face, phase
2a `f2fc57c` internal face) and a break-glass menu peer that calls it while the API
is alive. **This change is that peer — phase 2b — the last build step of "Sync".**
(S3 remains open in B4; this does not close the phase.)
## What it does
Adds a **"Back up a world now (Sync)"** operation to the root-gated break-glass
console. The operator picks a server from the live fleet; the peer resolves the
`felis-api-internal` ClusterIP Service and the service token from the control
namespace, then POSTs the internal backup endpoint to snapshot that server's world
while felis-api is alive. It is the on-node counterpart to the halt op: an emergency
"snapshot this now" lever for when the panel is unreachable but the box still has
`root` + a kubeconfig and the API is running.
It also extends the internal backup endpoint (phase 2a) to accept an optional
`{"os_user":"..."}` body so the audit row names the operator at the keyboard rather
than the generic `break-glass`.
## Why
The console cannot render the backup Job itself — it lacks the deployment coordinates
(`FELIS_IMAGE`, `FELIS_BACKUP_PVC`) that only felis-api holds. So, unlike the halt peer
(which writes the `MinecraftServer` CRD directly), the backup peer must go **through**
the API. The internal face exists precisely so an on-node machine caller with the
service token — no browser session, no Cloudflare-Access Principal — can reach that
orchestration. This change is the client that knocks on that door.
## Design decisions
- **Goes through the API, does not orchestrate locally.** Mirrors the phase-2a
rationale: deployment coordinates live only in felis-api. The peer's job is to
resolve the endpoint, authenticate with the token, and translate the HTTP result
into a friendly outcome card — not to build a Job.
- **ClusterIP resolution, not DNS.** `resolveInternalAPI` `Get`s the
`platform.APIInternalServiceName` Service and dials its `Spec.ClusterIP:APIInternalPort`
directly, erroring on an empty or `None` (headless) ClusterIP. The on-node console's
host resolver is not CoreDNS, so the in-cluster Service DNS name would not resolve
from the host; the ClusterIP is routable from the node and is what the `2ba9948`
dedicated ClusterIP Service exists to provide.
- **`os_user` attribution, parity with halt.** The console sends the escalated OS user
in the request body; the endpoint makes it the audit actor. The body is decoded
whenever `ContentLength != 0` — **not** gated on `Content-Type` (`decodeJSON` checks
only for unknown/trailing fields, not the header) — so a console that forgets the
header still records the operator. Absent/blank falls back to `break-glass`. The
console does **not** double-audit: the API audits at the boundary, single-sourced.
- **Stopped-gate stays server-side.** The world PVC is RWO, so the server must be
stopped. The peer does not pre-check this; it lets the API's stopped-gate return
`409 not_stopped` and renders that as a "must be stopped — halt it first" card. The
safety check is single-sourced in the API, never duplicated (and possibly drifting)
in the console.
- **A running pick ends the session with exit 1 — deliberately accepted.** The picker
lists all servers (a running pick is easy to hit), and a `409` surfaces as an error
through `backupResultMsg{err}` → root sets `m.err` → `breakglass.go` prints the
friendly card to stderr and exits non-zero. A dedicated `not_stopped` non-error path
was considered and **rejected**: every break-glass op ends the session anyway (all
`tea.Quit`), so `409`-vs-success differs only in exit code — marginal for an
interactive TUI. No error is swallowed; the friendly card is shown either way. Adding
a soft-landing state machine for one status code is complexity the interactive
surface does not earn.
- **Core/shell split, mirrors halt.** All decision logic (`resolveInternalAPI`,
`requestBackup`, `backupErrorFromResponse`, `performBackupNow`) lives in `backupnow.go`
and is unit-tested against a controller-runtime fake client + `httptest`.
`tui_backupnow.go` is thin bubbletea/huh glue (untested by house convention, mirrors
`tui_halt.go`). The picker **reuses** `listServersForHalt`/`haltableServer` rather
than cloning a second server-listing path.
## Files
| File | Change |
|---|---|
| `cmd/felis/backupnow.go` | **new** — pure core: `resolveInternalAPI` (Service ClusterIP + token secret), `requestBackup` (POST + Bearer + `os_user` body, status→outcome), `backupErrorFromResponse` (409/503/404/error-body mapping), `performBackupNow` |
| `cmd/felis/backupnow_test.go` | **new** — table tests against a fake client + `httptest`: happy resolve, headless/empty-token errors; 202 asserts Bearer + `os_user` + path; 409/503/404 + transport-failure mapping |
| `cmd/felis/tui_backupnow.go` | **new** — bubbletea/huh shell mirroring `tui_halt.go`: load → pick → work → outcome card; empty-fleet guard; friendly error card |
| `cmd/felis/tui_menu.go` | `bgSyncBackup` enum + "Back up a world now (Sync)" menu option after the halt option |
| `cmd/felis/tui_root.go` | `bgSyncBackup` dispatch to `newBackupModel`; `backupResultMsg` terminal handling into `breakGlassResult` |
| `cmd/felis/breakglass.go` | `backedUp`/`backupServer`/`backupStatus` result fields; cancel guard; post-exit backup summary |
| `internal/api/handlers_backups.go` | `handleInternalBackup` decodes the optional `os_user` body (gated on `ContentLength`, not `Content-Type`) and passes it as the audit actor; defaults `break-glass` |
| `internal/api/handlers_backup_now_test.go` | **+subtest** in `TestInternalBackup`: `os_user` body attributes the audit to the operator |
| `docs/openapi.yaml` | document the `internalBackupNow` optional `os_user` request body |
## Verification
WSL oracle (go1.26.4, FedoraLinux-44, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL_GREEN
```
`cmd/felis` and `internal/api` both re-ran (not cached), so the new `backupnow_test.go`
and the added `TestInternalBackup` subtest executed. `TestOpenAPIMatchesServedRoutes`
still passes: the `os_user` body is an addition to an already-documented operation, so
the served⇔documented route match is unchanged. The core is covered against a real
`fake.Client` (resolves the Service/Secret) + `httptest.Server` (asserts the wire
request and maps every status), so the tests exercise persisted/observable behaviour,
not merely that a call was made. `tui_backupnow.go` is thin glue, untested per the
`tui_halt.go` convention.
## Self-review outcome
- **ponytail (over-engineering):** lean — no new abstraction beyond the four core
functions two call sites (test + TUI) already justify; the picker reuses halt's
server-list core rather than cloning it; the deliberate rejection of a `not_stopped`
soft-landing path kept the state machine at four steps. Nothing cut.
- **correctness:** the `os_user` decode is gated on `ContentLength`, matching how
`decodeJSON` actually works (no `Content-Type` check), so a header-less console still
attributes correctly; the stopped-gate and audit stay single-sourced server-side, so
the peer cannot drift from the endpoint on the RWO safety check or the audit record.
@@ -1,82 +0,0 @@
# Separate ClusterIP Service for the felis-api internal face (8081)
- **Type:** bug fix (latent networking gap) + enabling change
- **Date:** 2026-07-07
- **Area:** `internal/platform` — Go (struct render oracle-verified; packet path **not**
verifiable in this environment — see Verification)
- **Commit:** `2ba9948`
- **Task:** #31 Phase B4 — surfaced while wiring the break-glass backup console peer
(phase 2b): the peer needs a routable path to the internal face, and that path was
broken for the login pod too.
## What it does
Renders a **new ClusterIP-only Service `felis-api-internal`** (control namespace)
that fronts the felis-api pod's internal port 8081, and repoints
`InternalAPIBaseURL` (the URL baked into the login pod's `FELIS_API_BASE_URL`) at
that Service name. Adds exported `APIInternalServiceName` / `APIInternalPort` so the
on-node break-glass console can resolve the Service's ClusterIP and dial it.
## Why (the latent bug)
The login limbo pod is configured with
`FELIS_API_BASE_URL = http://felis-api.<ns>.svc.cluster.local:8081` (setup.go) and
dials the internal face with the service token to mint bind codes and poll link
status. But the only Service named `felis-api` is the **external** face: a NodePort
Service that declares **only** port 443. A Service answers only on its declared
ports, so `felis-api:8081` had no backend — **every login-pod call to the internal
API silently failed to connect.** `deploy/limbo/README.md` even documented the
"login-pod → felis-api internal-port (8081) path" as reachable; it was not.
## Design decisions
- **A separate Service, not a second port on `felis-api`.** A `Type: NodePort`
Service allocates a node port for **every** declared port, with no per-port
opt-out. Folding 8081 into the NodePort `felis-api` Service would therefore publish
the internal face — which is service-token-only, explicitly **no Zero Trust** — on
every node's external IP. That violates the two-face security posture. A distinct
`ClusterIP` Service exposes 8081 **in-cluster only**: reachable by the login pod via
cross-namespace DNS, and by the on-node console via the ClusterIP (kube-proxy
programs ClusterIPs into the node's routing).
- **Repoint `InternalAPIBaseURL` to the new Service name.** The helper single-sources
the name the login pod is told to call; pointing it at `felis-api-internal` keeps
the login pod and the Service in agreement by construction.
- **Export the name + port for the console.** The break-glass backup peer (phase 2b)
resolves `APIInternalServiceName`'s ClusterIP at runtime and dials
`http://<clusterIP>:APIInternalPort` — it cannot use the cluster-DNS form because
the host's resolver is not CoreDNS.
## Files
| File | Change |
|---|---|
| `internal/platform/workloads.go` | **+`apiInternalService`** (ClusterIP, 8081→`internal`), wired into `Workloads()`; **+exported `APIInternalServiceName`/`APIInternalPort`**; `InternalAPIBaseURL` repointed at the internal Service, comment corrected |
| `internal/platform/workloads_test.go` | **+`TestAPIInternalService_ClusterIP`** — ClusterIP (never NodePort), 8081→`internal`, no nodePort, selects the api pods, name distinct from `felis-api` |
| `deploy/limbo/README.md` | document the `felis-api-internal` Service; correct the reachability note |
| `docs/troubleshooting.md` | §6 note: internal calls reached via `felis-api-internal`; a *connect* failure (not 401) points at that Service |
## Verification
WSL oracle (go1.26.4, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL GREEN
```
`TestAPIInternalService_ClusterIP` freezes the Service's shape. **This is a
code-level fix only.** `go build/vet/test` verifies the Service *struct* renders
correctly; it verifies **nothing** about packets flowing — not the login pod's
in-cluster call, not the console's host→ClusterIP dial (which relies on kube-proxy's
OUTPUT-chain DNAT, present on k3s but unverified here), not that 8081 is programmed
on a live cluster. Per the project's "Java/K8s code-only" reality, the runtime path
is **pending real-cluster verification**; the manifest-level defect (a DNS name with
no backing port) is fixed and asserted.
## Self-review outcome
- **ponytail (over-engineering):** one Service + two exported identifiers, all
load-bearing (the login pod and the console both need the routable 8081). No new
abstraction; `apiInternalService` mirrors `apiService`/`registryService`.
- **correctness / security:** the ClusterIP-not-NodePort choice is the crux — it keeps
the no-Zero-Trust internal face off every node's external interface, which a second
port on the NodePort Service could not.
@@ -1,83 +0,0 @@
# Internal-face break-glass world backup endpoint (§B4 "Sync", phase 2a)
- **Type:** feature (addition)
- **Date:** 2026-07-07
- **Area:** `internal/api`, `docs/openapi.yaml` — Go, oracle-verified
- **Commit:** `f2fc57c`
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation. Per the user's
**"两者都要"** decision the feature is built in two halves: the felis-api endpoint
that does the real backup-Job orchestration (phase 1, `7a7c0d5`) and a break-glass
menu peer that calls it while the API is alive (phase 2b, follow-up). **This change
is phase 2a: the second, internal face of that endpoint** — the door the console
peer will knock on.
## What it does
Adds `POST /api/v1/internal/servers/{name}/backup`, an **internal-face** twin of the
external `POST /api/v1/servers/{name}/backup`. The on-node break-glass console (root
on the host, holding the service token) POSTs here to snapshot a stopped world while
felis-api is alive. Same 202 `backing_up` / 409 `not_stopped` / 503
`backup_unavailable` / 404 / 400 `bad_name` surface as the external face.
## Why
The console cannot render the backup Job itself: it lacks the deployment coordinates
(`FELIS_IMAGE`, `FELIS_BACKUP_PVC`) that only felis-api holds — the same reason the
endpoint exists at all (phase 1). But the external face requires a Cloudflare-Access
Principal the console does not have. The internal face authenticates with the service
token (a trusted machine caller, no Principal), so the console can reach the same
orchestration without a browser session.
## Design decisions
- **No owner gate on the internal face.** The external handler enforces owner-or-admin
from the Principal; the internal handler has none — the service token IS the
authorization (the operator already has root on the node), so a server owned by
someone else still backs up. This mirrors how the other internal-face handlers
(op-login approve, QR poll) trust the token rather than a Principal.
- **Shared `enqueueBackup` tail.** The RWO stopped-gate, the optional-Backuper 503,
the async hand-off, and the audit+202 were refactored out of `handleBackupNow` into
a single `enqueueBackup(w, r, name, rec, actor, source)` that both faces call. The
two faces differ **only** in how the caller is authorized and in the audit
actor/source — the security-critical stopped-gate is single-sourced so the faces
cannot drift apart.
- **Audit attributed to break-glass/internal.** The internal handler audits directly
via `Repo.Audit` with `Actor:"break-glass", Source:"internal"` (the `a.audit`
helper hardcodes `Source:"external"`), so a console-initiated backup is
distinguishable in the audit log from an owner's self-service one.
- **Console does not double-audit.** Unlike the halt peer — which writes the CRD
directly and audits locally — the backup peer goes through the API, and the API
audits at the boundary. Auditing is single-sourced there; the console will not
emit its own row.
## Files
| File | Change |
|---|---|
| `internal/api/handlers_backups.go` | **+`handleInternalBackup`**, **+`enqueueBackup`**; `handleBackupNow` tail now calls `enqueueBackup(..., p.Email, "external")` |
| `internal/api/api.go` | register `POST /api/v1/internal/servers/{name}/backup` on the internal-face route table |
| `docs/openapi.yaml` | document the `internalBackupNow` operation (`x-felis-face: [internal]`, `serviceToken` security) |
| `internal/api/handlers_backup_now_test.go` | **+`TestInternalBackup`** — no-owner-gate, break-glass/internal audit, stopped-gate/503/404/400 |
## Verification
WSL oracle (go1.26.4, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL GREEN
```
`TestOpenAPIMatchesServedRoutes` gates the new route against `docs/openapi.yaml` in
both directions (served⇔documented) and passes. `TestInternalBackup` (5 subtests) and
the existing `TestBackupNow` (10) both pass — the external refactor is
behaviour-preserving (same audit actor `p.Email`/source `external`).
## Self-review outcome
- **ponytail (over-engineering):** the internal face is not a copy of the external
handler — the shared tail (`enqueueBackup`) collapses the duplication, and the two
handlers hold only their distinct auth + audit-attribution. No new abstraction
beyond the one shared function two callers already justify.
- **correctness:** the no-owner-gate difference is deliberate and matches the other
service-token handlers; the stopped-gate is unchanged and now single-sourced, so the
external and internal faces cannot diverge on the RWO safety check.
@@ -1,115 +0,0 @@
# On-demand world backup (§B4 break-glass "Sync"; felis-api endpoint + Job executor)
- **Type:** feature (addition)
- **Date:** 2026-07-07
- **Area:** `internal/backupjob` (new pkg), `internal/api`, `cmd/felis`, `docs/openapi.yaml` — Go, oracle-verified
- **Commit:** `7a7c0d5`
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation, resolved with the user as **immediate/on-demand world backup**. Per the user's "两者都要" decision this is built in two halves: **(this change) the felis-api endpoint that does the real backup-Job orchestration**, and (a follow-up) a break-glass menu peer that calls it while the API is alive.
## What it does
Adds `POST /api/v1/servers/{name}/backup`: an owner or admin snapshots a **stopped**
server's world into the archive store on demand, recorded as a first-class
`world_backups` row (reason `manual`) — restorable later by the existing restore path
and expired by the reaper's retention pass, so it never leaks as an orphan archive.
The backup runs asynchronously as a one-shot Kubernetes Job (the new
`internal/backupjob` package), mirroring how restore and image builds hand off to
Jobs. The handler answers **202 `backing_up`**.
## Why
felis-api cannot archive a world in-process: the world PVC is **RWO** and owned by the
operator's StatefulSet, so the API has nothing to mount at request time — the same
constraint that already makes `internal/restore` a Job. The break-glass console (which
runs direct-to-Postgres) likewise lacks the deployment coordinates (`FELIS_IMAGE`,
`FELIS_BACKUP_PVC`) needed to render the Job. Both point to the same home: the
orchestration belongs in felis-api, which holds those coordinates; other callers
invoke the endpoint.
## Design decisions
- **Backup Job self-records its `world_backups` row.** Unlike the restore Job — which
is deliberately DB-blind because it processes a potentially poisoned archive — the
backup Job **does** mount the felis config Secret and inserts its own backup row,
exactly like the reaper (the only other component holding both a world mount and the
database). This avoids the archive-then-async-record split that would otherwise leak
orphan archives on a crash. The security review for that one departure lives in
`internal/backupjob/jobspec.go` and is frozen by `jobspec_test.go`. Rationale: a
backup only **reads** a world the operator already owns and tars it (bytes, never
executed), so restore's poisoned-input threat does not apply; its blast radius (DB +
two PVCs) is a strict subset of the reaper's, and it never deletes a PVC nor calls
the K8s API (SA token stays un-mounted).
- **World mounted read-only, backup PVC read-write** — the mirror image of restore.
- **Stopped-gate (409 `not_stopped`).** The world PVC is RWO and held by a running
server, so a backup Job cannot double-mount it; the handler refuses unless the server
is fully stopped (`info.Ready || DesiredState != Stopped`). This also guarantees a
quiescent, non-torn archive. Mirrors `handleRestoreBackup`'s gate.
- **Authorization is restore's front half, minus the former-owner match.** Backup is
initiated by the **current** owner and records **their** ownership, so there is no
prior owner's data to leak — the leak guard that restore needs does not apply here.
An admin may back up an unowned (released) world; the recorded former owner is then
empty, exactly as the reaper records for an unowned reap.
- **Unique Job name per request.** Each backup Job is named `backup-<server>-<rand>`,
not a deterministic `backup-<server>`. A deterministic name would collide with a
just-finished Job still inside its `TTLSecondsAfterFinished` window (10m), and the
`AlreadyExists → 202` path would then silently produce **no** archive — the exact
window a user (or the console "立即备份" button) retries in. Unique names make every
request produce its own archive; `ErrAlreadyExists` remains only as a defensive
no-op on the ~impossible suffix collision. Ceiling (documented in `backup.go`): two
truly simultaneous taps may schedule two backup Pods — both mount the world PVC
read-only, so neither corrupts anything; single-flight-on-running is the upgrade
path if a double-tap storm ever appears.
- **One retention clock.** The entrypoint reuses the reaper's `reaperConfig` derivation
so a manual backup expires on the same schedule as an inactivity backup — one policy,
not two. The `"bk-"+hex` id scheme also matches, so manual and inactivity backups are
indistinguishable downstream.
- **Fail-safe on record failure.** If the row insert fails, the entrypoint deletes the
just-written archive so a failed backup leaves no unrecorded bytes.
- **Optional executor, honest 503.** Wired only when `FELIS_IMAGE` + `FELIS_BACKUP_PVC`
are supplied (same gate as restore); otherwise `API.Backuper` is nil and the endpoint
returns 503 `backup_unavailable`, so the authorization boundary is exercised before
the Job executor is deployable.
## Files
| File | Change |
|---|---|
| `internal/backupjob/jobspec.go` | **new** — `BackupJob` renderer + `BackupJobName`; weak SA, token off, hardened container, world RO / backup RW, config-Secret mount |
| `internal/backupjob/backup.go` | **new** — `Backuper` (idempotent enqueue) + `Config`/`withDefaults` |
| `internal/backupjob/k8sjobs.go` | **new** — controller-runtime `CreateBackupJob` (AlreadyExists → idempotent) |
| `internal/backupjob/jobspec_test.go` | **new** — freezes the Job's security shape incl. the deliberate config-Secret mount |
| `internal/backupjob/backup_test.go` | **new** — asserts each `Backup` call mints a unique Job name (repeat-tap must not silently no-op) |
| `cmd/felis/backup.go` | **new** — `felis backup` in-Pod entrypoint: archive + self-record + orphan-cleanup |
| `cmd/felis/run.go` | dispatch `case "backup"` + usage line |
| `internal/api/backuper.go` | **new** — the narrow `Backuper` port |
| `internal/api/handlers_backups.go` | **+`handleBackupNow`** |
| `internal/api/api.go` | `Backuper` field + `POST /servers/{name}/backup` route |
| `internal/api/handlers_backup_now_test.go` | **new** — `fakeBackuper` + handler subtests |
| `internal/api/backuper_wire_test.go` | **new** — compile-time `Backuper = (*backupjob.Backuper)(nil)` |
| `cmd/felis/api.go` | wire `backuper` under the `FELIS_IMAGE`+`FELIS_BACKUP_PVC` gate; `backupConfig` helper |
| `docs/openapi.yaml` | document the `backupNow` operation |
## Verification
WSL oracle (go1.26.4, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL GREEN
```
The `internal/api` OpenAPI served-route contract test (`TestOpenAPIMatchesServedRoutes`)
initially failed — the new route was served but undocumented — and passes after adding
the `backupNow` operation to `docs/openapi.yaml`. `internal/backupjob` and the new
handler subtests pass. The controller-runtime `K8sJobs` binding is integration-only
(needs a live cluster) and is exercised only by the interface conformance test.
## Self-review outcome
- **ponytail (over-engineering):** the backup Job is a near-mirror of the restore Job,
not a shared parameterization — deliberate, because its security shape differs (it
holds DB creds) and must be asserted independently, not hidden behind a shared knob.
No speculative config; `Config.withDefaults` fills only real deployment values.
- **correctness:** the RWO stopped-gate and the self-recording atomicity were traced to
the reaper and restore before writing; the former-owner asymmetry vs restore is
justified above.
@@ -1,155 +0,0 @@
# Test-quality integrity audit — do the verifications verify FUNCTION, or just go green?
- **Type:** audit / verification evidence (no code changed)
- **Date:** 2026-07-07
- **Method:** mutation testing on the WSL oracle (go1.26.4) + per-function coverage backbone
- **Scope:** the load-bearing safety invariants the backfilled change ledger *claims* were tested
- **Tree state:** every mutation reverted; authoritative Windows-git working tree clean at `5a7cd5a`
- **Point verified:** each gate is broken at `HEAD` (`5a7cd5a`), not per-commit — this is the
right reading of "does the verification verify the FUNCTION": the current test pins the
current implementation. A per-commit sweep would audit history hygiene, a different question.
## Why this audit exists
The ledger backfill asserts, per subsystem, that a set of load-bearing safety
properties are "unit-tested". A passing suite proves the tests are GREEN; it does
not prove they would go RED if the behaviour broke. Those are different claims —
"passing ≠ verifying". This audit closes that gap the only way that earns the word
*verified*: **break the implementation, confirm the specific test turns red.** A
subagent (or a human) *reading* a test and judging it "looks thorough" reproduces
the exact error being audited (looks-right ≠ verifies), so reading was used only to
locate the gate line; the verdict is always the mutation result.
## Result: 18 / 18 crown-jewel invariants mutation-verified
Each row is a one-line break of the implementation, run against its own package on
the oracle. **CAUGHT = the suite went red** = the test genuinely pins the behaviour.
| # | Invariant (claimed tested) | Impl gate mutated | Verdict |
|---|---|---|---|
| 1 | Pinned component is NEVER changed | `plan.go` pin branch → fall through | CAUGHT |
| 2 | A downgrade is NEVER proposed | `plan.go` `lv.After(current)` → `true` | CAUGHT |
| 3 | A prerelease is NEVER auto-applied | `plan.go` `!lv.IsPrerelease()` → `true` | CAUGHT |
| 4 | Apply ONLY inside the SysAdmin window | `plan.go` `Window.Contains(now)` → `true` | CAUGHT |
| 5 | `After` is strict (no equal-version churn) | `version.go` `> 0` → `>= 0` | CAUGHT |
| 6 | Clone-warned assertion refused fail-closed | `handlers_passkey.go` `if va.CloneWarning` → `if false` | CAUGHT |
| 7 | Approval CAS builds exactly once | `submit.go` `if !won` → `if false` | CAUGHT |
| 8 | An `everyone` base is not fail-open | `cfsetup.go` `if !includeHasEveryone` → `if true` | CAUGHT |
| 9 | Scoped-identity recognition actually admits | `cfsetup.go` `scoped = true` → `scoped = false` | CAUGHT |
| 10 | SSE per-principal stream cap holds | `api.go` cap-disable threshold | CAUGHT |
| 11 | OTP atomic reserve → one winner per burst | `api.go` `Sub(last) < window` → `< 0` | CAUGHT |
| 12 | Idle server is auto-stopped | `reconciler.go` `AutoStopEnabled &&` → `false &&` | CAUGHT |
| 13 | Startup/readiness timeout fires | `reconciler.go` `>= timeout` → `>= timeout + 1h` | CAUGHT |
| 14 | `/readyz` 503s when DB/K8s is down | `handlers_internal.go` dep-check `err != nil` → `false` | CAUGHT |
| 15 | NodePort fence only fires once connector serves | `tui_edge_apply.go` `connectorConnCount` parse-fail `return 0` → `1` | CAUGHT |
| 16 | CRITICAL-CVE build is NEVER admitted | `build.go` scan-gate `JobFailed`→`StatusFailed` → `StatusSucceeded` | CAUGHT |
| 17 | A user can NEVER claim a reserved system name | `naming.go` `reserved[name]` → `reserved["__nomatch__"]` | CAUGHT |
| 18 | Service token reaches ONLY the login pod | `builders.go` `Name == SystemLoginServer` → `true` | CAUGHT |
Rows 15–18 close the gap a review of this audit surfaced: the first pass verified a
*subset* and worded the verdict as the whole set. They are the four remaining
load-bearing safety properties the ledger docs name as "unit-tested" (§ *Documented-tested
claim reconciliation* below). Each mutation produced a real `--- FAIL` on the specifically
named test — e.g. #15 reddened `TestConnectorConnCount/garbage_is_not_a_healthy_tunnel`,
#16 `TestSyncFailedDoesNotAdmitImage`, #18 `TestBuildEnvWithholdsServiceTokenFromUserServers`
— i.e. an assertion failure, not a compile break.
Not one crown-jewel test was vacuous. The `cfsetup` fail-closed test additionally
feeds five distinct *violating* policies (bare-everyone, everyone-OR-identity,
unrecognized `ip` type, empty rule, wrong decision) and asserts each is rejected —
strong negative-path coverage, confirmed by mutations #8–#9.
## Coverage backbone — what no oracle test executes (failure-mode B)
Coverage triages code that no test even runs (so it cannot be verified). It does NOT
itself earn "verified" — high coverage with weak asserts is the same green-number
trap. Per-function scan of the security packages:
**Integration-only by design (0% on the oracle — honest, NOT a gap).** The real
adapters run only against live infra; unit tests exercise the ports through fakes:
- `pgrepo.go` — all SQL, **including `QuotaAvailable`/`QuotaCheck` (the quota TOCTOU
atomic claim)**. This matches task #45's own "ENV-blocked" note: the atomic claim
is a Postgres `INSERT … WHERE`, verifiable only against a real DB.
- `k8scluster.go`, the K8s console/log-stream adapters — real Kubernetes/RCON I/O.
- `tui_edge_apply.go` `verifyConnectorServing` + the nftables fence apply — shell out to
live `cloudflared`/`nft`. **Correction from the first pass:** the doc splits this from the
*pure* `connectorConnCount` decision gate, which IS unit-tested and is now mutation-proven
(#15). The first pass wrongly folded the whole fence into "integration-only"; only the live
calls are. The gate that decides *whether* to fence is verified.
**Genuine coverage gap (untested at the HTTP layer — "not verified").** These are
*missing* tests, not fake-passing ones:
- `handlers_users.go` — the P5 SysAdmin account-management suite: `handleCreateUser`,
`handleGetQuotas`, `handleSetQuotas`, `handleListUsers`, `handleGetUser`,
`handlePatchUser`, `handleDeleteUser`, `handleDisableUser`, `handleLinkAccount`,
`handleUnlinkAccount`, `handleListUserSessions`, `handleRevokeUserSessions`, and
`validateUsername`. All 0%; no `handlers_users*_test.go` exists. (The adjacent
`DELETE …/passkeys` remediation handler *is* tested by `TestUnbindUserPasskeys`.)
- `handleReady` — the internal-face "server is up" push (distinct from the tested
`handleReadyz`); 0%.
These handlers are owner/operator-role-gated, so the blast radius is bounded, but
`validateUsername` is load-bearing input validation and is the highest-value target
for a follow-up test. **Recommendation:** add an `handlers_users_test.go` covering
create/quota/link + `validateUsername` negative paths. Filed as a proposed change,
not made here (this is a read-and-verify audit — no test/impl was modified).
## Cheap tells (static pre-pass)
- 3 `t.Skip` sites, all benign: RNG-collision reruns (a 1-in-10^6 OTP code clash),
not coverage-gating skips.
- No test file falls below 2 assertions per test function.
## Documented-tested claim reconciliation
To avoid the subset-verified/whole-worded trap a second time, every "unit-tested"
string in the ledger docs was enumerated (`grep -niE "unit-tested" docs/changes/*.md`)
and mapped to a verdict — verified fail-open gates get a mutation; behavioural/contract
claims are scoped, not silently dropped:
| Doc claim | Verdict |
|---|---|
| modpack: approval CAS builds once | mutation #7 |
| modpack: **scan in front of any push** | mutation #16 |
| cloudflare: `validateFailClosed` refuses public policy | mutations #8–#9 |
| cloudflare: **conn-count fence gate** | mutation #15 |
| system-servers: **naming reservation** | mutation #17 |
| system-servers: **service-token → login pod only** | mutation #18 |
| operator: idle stop / startup+readiness timeout / `/readyz` | mutations #12 / #13 / #14 |
| auto-update: pin / no-downgrade / no-prerelease / window / strict-`After` | mutations #1–#5 |
| passkey: clone-warned assertion refused | mutation #6 |
**Scoped, NOT individually mutation-proven** (behavioural/contract-level, not fail-open
safety gates — they rest on the green suite + the coverage backbone, and are called out here
rather than folded into the verdict):
- console-auth: content-type guard, anti-enumeration, forced-change lockdown. Anti-enumeration
is the one with security weight; the current public login door is email-OTP/passkey, and its
anti-enumeration behaviour is a candidate for a future mutation pass.
- break-glass setup: owner-auth match/non-match over the fake store.
- username-reclaim + auto-update JSON round-trip: in-memory repo contract / serialization.
## Toolchain honesty
Only Go runs on the oracle. The felis-limbo plugin (Java) is podman-verified against
a real Limbo jar (#65); the limbo/lobby images carry a build+boot check (#63). Those
completions were never a green-Go-tests claim and are not audited as if they were.
## Verdict
The commit history's verification claims are **accurate**: all 18 load-bearing *fail-open
safety gates* the ledger names as tested — spanning every subsystem, reconciled one-for-one
against the docs' "unit-tested" claims above — are mutation-proven to pin behaviour, not
merely to pass. No crown-jewel test was vacuous. The shortfalls are (a) integration seams
unrunnable on the oracle by design (honestly classified — including the live
`cloudflared`/`nft` fence-apply, whose *decision* gate is nonetheless verified); (b) one
untested cluster of admin user-management handlers — a missing test, not a false green; and
(c) a residue of behavioural/contract-level "unit-tested" claims (console anti-enumeration,
break-glass owner-auth, repo/JSON contracts) that rest on the green suite plus coverage and
are scoped above rather than individually mutation-proven — the honest boundary of this pass.
**Method note.** The first pass mutation-verified 14 gates but worded its verdict as "every"
invariant; a review caught that 4 documented safety gates (fence, scan, naming, service-token)
were named-as-tested yet unverified, and one (the fence) was mis-classified as integration-only.
Those four are now mutation-proven (#15–#18) and the classification corrected. The lesson is
the audit's own thesis turned on itself: *reading a scope and judging it complete* reproduces
the *looks-right ≠ verifies* error — only the enumerate-and-mutate reconciliation earns the word.
@@ -1,129 +0,0 @@
# Adversarial input-validation audit — every dangerous sink fails closed
- **Type:** negative-path / input-validation audit (no production code change)
- **Date:** 2026-07-08
- **Area:** `internal/api`, `internal/naming`, `internal/submit`, `internal/rcon`
- **Task:** a different question from the #82 / round-2 mutation audits. Those asked
*do the TESTS catch a gate regression?* This asks *does the CODE reject hostile
INPUT, or does bad data PASS?* — feed the real endpoints malformed, boundary, and
hostile bodies (故意加错误数据) and confirm they fail closed (4xx) rather than letting
the garbage reach a sink.
## Method — sink-first, not fuzz-everything
The low-hanging garbage (oversized body, unknown field, wrong content-type) is already
caught by the universal body guards, so a blanket "fuzz all ~80 handlers" would burn
effort where the answer is known. The real "can bad data PASS?" risk lives at the
**sinks** — the few places a request string is concatenated into an RCON command, used
as a K8s object name, joined into a filesystem/archive path, put in a SQL query, or
accepted as an enum/quantity **without a validator in front**. So the audit traces each
dangerous sink class from its handler entry to the sink, and for the crown-jewel class
(text → RCON) **mutation-verifies** the guard is non-vacuously pinned: loosen the guard
in source, run the package tests, confirm the specifically-named negative test reddens
(`--- FAIL: <subtest>`), then revert. Oracle: WSL Fedora-44, go1.26.4.
## The universal belt (caught before any sink)
`decodeJSON` (`internal/api/util.go`) wraps every body in
`http.MaxBytesReader(w, r.Body, 1<<20)` (1 MiB cap), sets `DisallowUnknownFields()`, and
rejects trailing data after the first JSON value — all → `400 bad_request`.
`requireJSONContentType` returns `415 unsupported_media_type` on credential writes
(a CSRF belt). So oversized, unknown-field, multi-document, and wrong-type bodies never
reach a handler body at all.
## The dangerous sinks — each traced fail-closed
### 1. Text → RCON (the lead). Two vectors, both fenced.
**(a) Structured commands** — `handlers_access.go` (LuckPerms permission/group,
whitelist/ban/kick). Every operand is validated against an anchored allow-list charset
*before* it is concatenated: `mcNameRe = ^[A-Za-z0-9_]{1,16}$` (player),
`lpNodeRe = ^[A-Za-z0-9_.*-]{1,64}$` (node), `lpCtxRe = ^[A-Za-z0-9_-]{1,48}$`
(world/group). No space, separator, or control character can appear in a validated
operand, and **there is no free-text field anywhere** — a ban/kick deliberately carries
no reason string (that would be the one splice vector). Go's `$` is `\z` (absolute end,
not `\Z`), so even a single trailing `\n` is rejected.
*Mutation-verified:* loosening `mcNameRe` to admit a space
(`^[A-Za-z0-9_ ]{1,16}$`) reddens
`TestAccessInjectionRejected/{whitelist,ban,kick,permission}_player_space` — the guard
is real, not vacuous.
**(b) Free-text passthrough** — `handlers_console.go` `handleCommand`, POST
`/servers/{name}/command`. This is the *one deliberate* free-text → RCON vector, and it
is **owner/admin-gated** (403 for a stranger, 404 for an unknown server). Its input
fence: trim + strip a single leading `/`, reject empty, cap at 1024 bytes, and
`strings.IndexFunc(command, func(c rune) bool { return c < 0x20 }) >= 0 → 400` — every
C0 control (incl. `\n`) is rejected so one request cannot splice a second command.
*Mutation-verified:* disabling the scan (`c < 0x20` → `c < 0x00`) reddens
`TestConsoleCommand/control_character_(newline)_->_400,_no_RCON_call`, whose input is
literally `{"command":"say hi\nop attacker"}` and whose assertion is 400 **and**
`console.calls == 0`. (Severity note: because this vector is owner-gated by design, the
scan is an audit-integrity measure — one request = one command — not a privilege
boundary; a splice on your *own* server escalates nothing, since the owner may already
run any RCON command. The fence exists regardless.)
### 2. Break-glass / internal-face (the newest code — scrutinised specifically)
This surface runs under a "service-token-authed / local-root, inputs trusted" posture,
the classic place a field-level guard gets skipped. Traced end to end:
- **Server name** — every internal handler (`handleReady`, `handleJoinEvent`,
`handleInternalWake`, `handleInternalClaim`, `handleInternalMenuStatus`) validates the
path name with `naming.ValidateServerName` before use.
- **`mc_uuid`** (join/wake/claim bodies) — checked non-empty, then flows *only* to
DB-parameterized calls (`RecordJoin`, `UserByMCUUID`, `UUIDInAllowlist`). The
"allowlist" is a **DB table**, not a live RCON `whitelist add` — there is no
`mc_uuid` → RCON path.
- **Break-glass "OP-create"** (`performAddOperator` → `provisionOperator` →
`InsertOperator(ctx, id, username, email)`) is a **parameterized DB INSERT** creating a
*panel staff account*, **not** a Minecraft `op` RCON command. The hypothesised
name → RCON `op` sink was checked and **does not exist** in this shape; the username is
`TrimSpace`d and reaches only `$N`-parameterized SQL, from a local-root caller.
- **`os_user`** (internal backup attribution) — `TrimSpace`d, sets only the audit actor
(a DB row); shown non-vacuous in the round-2 backup/restore audit (`b7b4a3b`).
### 3. SQL injection — dismissed.
`pgrepo.go` uses uniform `$1/$2/$3` parameterization throughout
(`QueryRowContext`/`ExecContext(ctx, q, args…)`); no request string is `Sprintf`'d into a
query.
### 4. Path traversal (submit) — dismissed.
`internal/submit` validates the submission id (rejects `..`, path separators, uppercase,
space, empty), and the on-disk blob name is a **fixed** constant (`contextBlobName`) — no
attacker-supplied filename is ever joined. The hostile-id matrix
`{"../evil","sub/../../etc","SUB-UPPER","has space","","a/b"}` is test-pinned in both the
local and S3 backends. The one free-form field a submission carries (`DisplayName`) is
charset-constrained by `displayNameRE` and rejects control chars
(`submit_test.go` "control chars" case).
### 5. K8s object names — validated at every cluster write.
`ValidateServerName` / `ValidateSystemServerName` (`^[a-z0-9-]{3,32}$`, no
leading/trailing dash, reserved-name set) and `ValidateHostname`
(`dnsLabelRE`, single label under the configured root domain) gate every create/patch.
`cluster.go` documents the invariant: "there is no free-form YAML path — every field is a
typed, validated value," so no raw CRD field can be smuggled through a create/patch body.
### 6. Numeric / enum — fail-closed.
`resolveResources` routes **every** quantity (memory, `resources.cpu/memory`, requests)
through `parsePositiveQuantity`, which rejects `q.Sign() <= 0` (negative *and* zero) with
a field-named 400, plus a request>limit guard; `parseStorageSize` carries the same guard.
Enums are closed sets: `parseAutostartPolicy` (ownerOnly/public/allowlist), the
access-action switch, and the image `ImageAdmitted` allow-list (no free image string).
## Verdict
**Bad data does not pass.** Every dangerous sink is fail-closed — including the
break-glass / internal-face surface, where the hypothesised `mc_uuid`/OP-create → RCON
paths were traced and found not to exist (parameterized DB, not RCON). Both text → RCON
vectors — structured (`handlers_access`) and free-text (`handleCommand`) — are
mutation-pinned by their named negative tests. No gap was found and no production code
changed; the honest result of "故意加错误数据" is that the input surface rejects it.
Scope is deliberately bounded to the dangerous sinks and the newest (break-glass) code,
not an exhaustive fuzz of all ~80 handlers — the claim proven is "every place user input
reaches a dangerous sink validates before the sink," by trace plus two mutations, not
"every handler was fuzzed."
@@ -1,57 +0,0 @@
# §B4 phase close — S3 archive backend deferred by design (decision record)
- **Type:** decision / scope record (no code changed)
- **Date:** 2026-07-08
- **Area:** `internal/config`, `internal/backup` — the archive-store backend selection
- **Task:** #31 Phase B4 break-glass ops — **closes the phase**, superseding the phase-2b
note in [break-glass-backup-peer](2026-07-07-break-glass-backup-peer.md) ("S3 remains
open in B4; this does not close the phase").
## Decision
The three operational B4 break-glass ops are built, wired into the recovery console menu,
and oracle-verified:
| Op | Menu enum | Commit |
|---|---|---|
| Provision/reset Owner | `bgProvisionOwner` | (Phase B1 lineage) |
| Add Operator ("OP create") | `bgAddOperator` | recovery-console menu |
| Halt a running server | `bgHaltServer` | `c2ee21a` |
| Back up a world now ("Sync") | `bgSyncBackup` | `7a7c0d5` / `f2fc57c` / `fc748d3` |
The fourth B4 line item — **S3 archive backend (`tarS3`)** — is **deferred by design**, not
left as a silent gap. It is closed as a documented deferral and **#31 is done**.
## Why deferring is safe (not a loose end)
- **Fail-closed at config load, frozen by a test.** `config.Load` rejects
`store = "tarS3"` (and `volumeSnapshot`, `longhorn`) with an error that points the
operator at the `tarLocal` remediation. `TestLoadRejectsUnimplementedArchiveStore`
freezes exactly this: a config naming an unimplemented backend fails at load, so
felis-api can never boot green while the reaper CronJob fails every run and restore
silently 503s. tarS3 cannot be selected into a broken state.
- **Peer to two other deferred backends.** `tarS3` sits beside `volumeSnapshot` and
`longhorn` as recognized-but-unimplemented store names. The spec's own phasing is
tarLocal-first ("起步 `tarLocal` … 要异地/跨集群 → `tarS3`"): the baseline single-node
path is `tarLocal` (tar → backup PVC), which is implemented, tested, and the backend
every built backup/restore path uses today.
- **Offsite/cross-cluster DR is the only capability gap**, and it is opt-in future work,
not a correctness hole in the shipped baseline.
## The build path, when offsite DR is wanted
`minio-go/v7` is already vendored (the modpack upload lane's `internal/submit/s3store.go`),
so tarS3 adds no dependency. A future build is bounded:
1. `internal/backup/tars3.go` — a `WorldArchiver` reusing the existing package-level
`writeTarGz`/`readTarGz`/`pruneToManifest`, streaming the tar to an object via
`PutObject` (size −1, multipart) and reading it back via `GetObject`, mirroring
`s3store.go`'s fakeable-client testability.
2. Store-selection factory in the backup/reaper entrypoint (`store = "tarS3"` → construct
the minio-backed archiver) + remove tarS3 from the config fail-closed list (update
`TestLoadRejectsUnimplementedArchiveStore` to keep only volumeSnapshot/longhorn).
3. Inject the S3 Secret into the backup/reaper Job Pods (integration wiring, like the
Kaniko S3-context credential path).
The live-S3 upload + Secret-into-Pod would be integration-only verified, exactly as the
modpack S3 backend and the pgrepo SQL are.
@@ -1,93 +0,0 @@
# Backup/restore data-safety mutation audit (round 2) — 7 gates pinned, 1 gap closed
- **Type:** test-quality audit + one test added (code change: `internal/api/handlers_backups_test.go`)
- **Date:** 2026-07-08
- **Area:** `internal/api` (`handlers_backups.go`), `internal/backupjob`
- **Task:** continues the #82 test-quality integrity audit onto the on-demand
backup/restore surface, which postdates the 18-gate round-1 audit
([mutation audit `4626ab5`](2026-07-07-test-quality-mutation-audit.md)). The
backup/restore endpoints (`7a7c0d5` / `f2fc57c` / `fc748d3`) were not in that pass.
## Method
Same as round 1: apply a one-line mutation to a fail-open gate in the source, run the
package tests, confirm the **specifically-named** test reddens with an *assertion*
failure (`--- FAIL: <subtest>`), then revert. A build break (`declared and not used`,
`undefined`) is not a valid verdict, so mutations are operator-flips that keep every
operand referenced (`!=`→`==`, drop a `!`, a literal→`true`, or `if false && <orig>` to
disable a gate without orphaning its variables). Airtightness: each mutation is re-run
with `-run` scoped to the intended subtest and `grep -- "--- FAIL: <subtest>"`, so a
reddening sibling can't be mistaken for the gate under test. Oracle: WSL Fedora-44,
go1.26.4.
## Fail-open gates mutation-verified (all CAUGHT at the named subtest)
| # | Gate (file:line) | What it guards | Mutation | Subtest that reddened |
|---|---|---|---|---|
| A | `handlers_backups.go:282` enqueueBackup stopped-gate | RWO double-mount / torn archive while the world is up | `!=`→`==` | `TestBackupNow/starting_server_->_409_not_stopped` |
| B | `:151` restore stopped-gate | restore Job can't mount a live world's RWO PVC | `!=`→`==` | `TestRestoreBackup/starting_server_->_409_not_stopped` |
| C | `:136` restore former-owner match | a fresh claimant resurrecting the previous owner's world | `!=`→`==` | `TestRestoreBackup/current_owner_who_is_not_former_owner_->_403` |
| D | `:116` restore cross-server guard | restoring server A's backup onto server B | `!=`→`==` | `TestRestoreBackup/restore_by_backup_id_cross-server_->_403` |
| E | `:214` backup owner-or-admin authz | a stranger backing up someone else's world | drop `!` | `TestBackupNow/non-owner_->_403,_no_backup` |
| F | `:26` list cross-user scope | a user seeing other tenants' backups | `p.IsAdmin()`→`true` | `TestListBackups/user_sees_only_own_former-owned_present_backups` |
## The gap this audit found — and closed
**Restore's owner-or-admin gate (`handlers_backups.go:85`) was not pinned by any test.**
It is the twin of gate E, but the two are *not* symmetric. Disabling it
(`if false && !a.isOwnerOrAdmin(p, rec)`, which keeps `rec` referenced so the package
still builds) reddened **nothing** — `go test ./internal/api/` stayed `ok`. The same
disable applied to backup's L214 (gate E) reddened `non-owner` immediately, proving the
technique valid and the asymmetry real.
Root cause: the former-owner gate at L136 backstops every non-owner case the suite
exercised (a stranger and a wrong-backup current owner both fail L136 *and* L85, so
L136's 403 masks a broken L85). The one case only L85 catches went untested: a
**superseded former owner** — a user who owned a server, took this backup
(`FormerOwner=them`), then released it to a *new* owner. They still pass L136 (they *are*
the former owner) but must be stopped by L85, or they could roll the new owner's live
server back onto their old world (cross-tenant clobber). The handler comment names this
the "must re-claim first" rule (`handlers_backups.go:77-79`).
**Fix (code):** added `TestRestoreBackup/former owner after release -> 403, no restore`,
the mirror of the existing L136 test. Verified both directions: green on the clean tree,
and it is the sole subtest that reddens when L85 is disabled — so it now pins the owner
gate specifically, not L136. No production code changed; `handlers_backups.go` is a pure
test addition away from where it was.
## Enumeration — covered vs. scoped (so "the gates" means all of them)
- **Fail-open data-safety gates — all pinned:** A–F above, plus L85 (now closed). 7/7.
- **Accountability, not fail-open (verified non-vacuous):** the `os_user` attribution at
`handlers_backups.go:254` — a supplied operator name overrides the default `break-glass`
audit actor. Mutating `u != ""`→`u == ""` reddens
`TestInternalBackup/os_user_body_attributes_the_audit_to_the_operator`, so the
attribution test isn't vacuous. A failure here degrades the audit actor; it grants no
bypass, so it is out of the fail-open bucket.
- **Contract/behavioral (tested, out of mutation scope):** nil `Backuper`/`Restorer` → 503;
a failed backup/restore → 500 **not** audited; `backup_ref` never serialized to the wire;
the internal-face actor defaults to `break-glass`. Each has a direct test; none is a
fail-open safety gate.
- **`internal/backupjob` (glanced, not mutated):** orchestration only — each backup gets a
unique Job name (`BackupJobName` + random suffix) so a repeat "立即备份" tap can't collide
with a just-finished Job still inside its TTL; `ErrAlreadyExists` is a defensive no-op;
`Backup` returns once the Job is created (the async 202 is honest). Unit-tested against a
fake `Jobs`; the controller-runtime `k8sjobs.go` and the `jobspec.go` Pod shape (weak SA
with its token un-mounted, config Secret mounted for the self-recorded row, read-only
world mount) are integration-verified per the package doc — not fail-open API gates.
## Coverage nuance (documented, not a gate failure)
The stopped-gate is `if info.Ready || info.DesiredState != DesiredStopped`. The `!=`→`==`
mutation pins the `DesiredState` operand (both A and B reddened), but `info.Ready` is not
*independently* pinned: no test sets `Ready=true` together with `DesiredState=Stopped` —
the stopping-but-still-up race. Low risk because in practice `Ready` drops as
`DesiredState` leaves `Stopped`, but the belt-and-suspenders `Ready` operand rides on
coverage of the operand beside it rather than its own case.
## Verdict
The backup/restore data-safety surface is a coherent unit, and this closes it: **7/7
fail-open gates pinned** (6 pre-existing, 1 added this round), one accountability gate
shown non-vacuous, one coverage edge documented. Not extended to every handler — that
would be an unbounded "continue the audit."
@@ -1,110 +0,0 @@
# Felis-nano — the federating `hasJoined` multiplexer (step 1: the Go resolver)
- **Type:** feature (new endpoint) — the verifiable "brain" of Felis-nano
- **Date:** 2026-07-12
- **Area:** `internal/api` (`handlers_hasjoined.go` + test), `internal/api/api.go`
(route + `AuthSources` field), `docs/openapi.yaml`, `go.mod`
- **Task:** Felis-nano provides a MultiLogin-like capability — one Velocity proxy that
accepts logins verified by **several** Yggdrasil auth servers at once (Mojang + N
third-party roots), 正版优先 (Mojang-first). This change builds **step 1**: the Go
`hasJoined` multiplexer that does the federating verification. It is the only part of
the plan that produces immediate verifiable hard evidence (a unit-tested HTTP endpoint);
the two delivery shells that point Velocity at it (a JVM `-Dmojang.sessionserver` flag,
and a thin reflection-hook plugin for third-party servers) are later steps.
## What Velocity asks for, and what this answers
On a Minecraft login Velocity's authlib computes the `serverId` hash and issues
`GET /session/minecraft/hasJoined?username=<name>&serverId=<hash>[&ip=<ip>]` against
whatever URL its `mojang.sessionserver` system property names. A 200 with a game profile
means "verified"; a 204 means "not verified" and authlib rejects the login. Vanilla points
this at Mojang alone. Felis-nano points it **here**, and this endpoint fans the same query
out to the configured Yggdrasil roots **in priority order**, returning the first source
that validates. Each upstream Yggdrasil runs its own `serverId`-hash check — the
multiplexer only relays, it computes no hashes.
## The one non-negotiable transform — per-source UUID namespacing
A third-party Yggdrasil's UUIDs are **self-asserted**: nothing stops a malicious source
from answering with a *genuine Mojang player's* UUID. If that UUID were emitted as-is, the
third-party could impersonate any Mojang player with full UUID fidelity — and the reclaim/
blacklist layer could never catch it, because its whole invariant is "the genuine Mojang
player has a **different** UUID from any squatter." That invariant would simply be false.
So the resolver rewrites every non-identity source's profile into a per-source namespace
**before it leaves the resolver** — the single entry point every login crosses:
```
canonical = UUIDv3(felisAuthNS, tag + ":" + nativeID) // third-party
canonical = the source's UUID verbatim // Mojang (Identity: true)
```
MD5 (UUIDv3) preimage resistance means no third-party can mint a value inside Mojang's
UUID space; the per-`tag` prefix means two sources can't collide onto one identity. Every
downstream key — `account_links`, `username_blacklist`, owner checks — then sees exactly
one canonical UUID per real identity, so the reclaim invariant is true **by construction**,
not by assumption.
## Fail-closed details that bite if wrong
- **Bar gate at the chokepoint.** The canonical UUID is checked against
`Repo.IsUsernameBlacklisted` *before* the profile is returned, so a reclaimed squatter
stays out even on a consumer that has no limbo plugin. Keyed on the **dashed** canonical
(`.String()`) — the exact form `Repo.ReclaimUsername` stores. A DB error there fails
closed (non-200 → authlib rejects), matching the existing `handleCheckBlacklist` pattern.
- **Emit undashed.** authlib's `GameProfile` expects the 32-hex undashed `id`
(`hex.EncodeToString(u[:])`); the DB/reclaim/blacklist keys are dashed. The resolver
**checks** on the dashed string and **emits** the undashed one. Mixing the two forms is a
silent gate miss — pinned by the tests below.
- **`properties` relayed verbatim** (`[]json.RawMessage`) so a source's signed textures
survive the multiplexer untouched.
- **Inert by default.** `AuthSources` is nil until `cmd/felis` wires configured sources,
so the endpoint 204s every login until deliberately configured — it ships off.
- **Public internal-face route.** authlib sends no service token, so the route is mounted
`Public: true` on the internal face (like `/healthz`); no third face is introduced. The
OpenAPI parity test enforces `x-felis-face: [internal]` + `x-felis-tier: public`.
## Files
| File | Change |
|---|---|
| `internal/api/handlers_hasjoined.go` | **new** — `handleHasJoined` + `resolveHasJoined` + `AuthSource`/`sessionProfile` types + `felisAuthNS` |
| `internal/api/handlers_hasjoined_test.go` | **new** — `TestHasJoined`, 7 subtests over `httptest` fake Yggdrasil roots |
| `internal/api/api.go` | `AuthSources []AuthSource` field (nil = inert) + `GET /session/minecraft/hasJoined` `Public` internal route |
| `docs/openapi.yaml` | `/session/minecraft/hasJoined` path — `x-felis-face: [internal]`, `x-felis-tier: public`, `security: []` |
| `go.mod` | promote `github.com/google/uuid` indirect→direct (first direct importer) |
## Verification evidence
Oracle: WSL Fedora-44, go1.26.4. `go build ./...` → `BUILD-OK`. Full `internal/api`
package green (`ok felis.lolicon.best/internal/api`), `go vet ./internal/api/` clean. The
full package (not a `-run` filter) was run because this change edits two shared surfaces —
the `API` struct and the `internalAPIRoutes()` table — where a route that isn't under
`/api/v1/` is exactly what a route-table-driven invariant test would trip; nothing
reddened.
`TestHasJoined` — 7 subtests, all PASS:
1. `mojang identity passthrough` — Mojang UUID unchanged, `properties` relayed.
2. **`thirdparty UUID rewritten, never emitted as-is`** — the security invariant: an evil
source returns real Notch's Mojang UUID; the resolver emits neither that UUID nor any
Mojang-space value, but the deterministic `UUIDv3(felisAuthNS, "littleskin:"+id)`.
3. `mojang priority wins over thirdparty` — Mojang-first ordering.
4. `fallthrough to thirdparty when mojang 204s` — priority scan continues past a 204.
5. `no source validates -> 204`.
6. `barred canonical UUID -> 204` — the reused reclaim bar gate holds at the resolver.
7. `missing username -> 204` — no source touched on a malformed query.
`TestOpenAPIMatchesServedRoutes` PASS — the new route's served facets match its
`docs/openapi.yaml` entry in both directions.
## Deferred (not in this change)
- **Source configuration** (step 2): a `tag`/`type`/`url`/`priority` schema and
`cmd/felis` wiring that populates `AuthSources`. Until then the endpoint is inert.
- **Delivery shells** (step 3): Shell 1 = the `-Dmojang.sessionserver` JVM flag on
Felis-managed proxies; Shell 2 = the thin reflection-hook Velocity plugin for
third-party servers, plus a Velocity verification runbook. The user has a real
server to test Shell 2 against.
- **Name-match / textures-signature enforcement** — deliberately out of scope: identity
is `tag:nativeID`, not the name, and textures are the cosmetic bucket relayed verbatim.
-255
View File
@@ -1,255 +0,0 @@
# Felis change ledger
The index of every functional change to Felis — what it did and which commit records
it. This is the durable, in-repo map that `git log` alone doesn't give: it links
substantial changes to their detail docs and flags work that is built and verified but
not yet committed.
## Convention
- **Every functional change** (a feature addition, a behaviour change, a bug fix) gets:
1. a dated detail doc in this directory — `docs/changes/YYYY-MM-DD-<slug>.md`, covering
_what it did, why, the files touched, and the verification evidence_; and
2. a row in the ledger below, carrying its **commit record** (the short SHA).
- A change that is **built and verified but not yet committed** (e.g. while PGP signing
is locked) sits in **Pending** with `commit: pending`, and moves into the ledger with
its real SHA once committed.
- Pure-cosmetic or non-functional commits (docs, style) still appear in the ledger table
for completeness, but do not require a dedicated detail doc.
- The ledger table is generated losslessly from git history and can be regenerated:
```
git log --reverse --pretty=format:'| %h | %ad | %s |' --date=short
```
## Pending (built + verified, not yet committed)
_None._
## Detail docs
Depth docs for substantial changes, keyed to the commit(s) they cover. The committed
ledger table below stays a lossless mirror of `git log` (so it can be regenerated); this
section is where a row's detail doc, when it has one, is found. Most rows — panel/UI,
docs, chore, style — have no detail doc by convention and are recorded by their table row
alone. Entries marked *(backfill)* were reconstructed retroactively on 2026-07-07 from git
history to close the ledger's detail-doc axis for the pre-convention functional commits;
each carries a backfill note stating it was not independently re-verified. Frontend/`panel`
commits are the collaborator's UI work and are not given detail docs here.
| Detail doc | Commit(s) | Scope |
|---|---|---|
| [foundational-subsystems](2026-06-26-foundational-subsystems.md) *(backfill)* | `7fbebfe` `708cdfc` `43ab921` `78b8cf6` `d39605e` `b508fcc` `47fcd90` `93f143f` | initial import: CRD, core libs, backup, operator, submit, api, platform, plugins |
| [modpack-submission-lane](2026-06-26-modpack-submission-lane.md) *(backfill)* | `d39605e` `598f3d3` | §8 modpack build/approval pipeline + local/S3 backends |
| [deploy-bootstrap-installer](2026-06-27-deploy-bootstrap-installer.md) *(backfill)* | `58fa4b0` `94a3b7b` `deaa2f8` `318a724` `e5f1682` `28c3eee` `c14ed17` `d9e866f` `b84debf` | one-line bootstrap installer + demo bring-up |
| [console-auth-passwordless](2026-06-27-console-auth-passwordless.md) *(backfill)* | `af14f02` `0c1cc59` `3b43f05` `c20b12c` | local-password login → passwordless migration + residue sweep |
| [felis-cli-break-glass-setup](2026-06-27-felis-cli-break-glass-setup.md) *(backfill)* | `e108a37` `2d0bbb0` `a94b001` `eb5875a` `f5d00f3` `9c46632` `7d91373` | break-glass recovery console + first-run setup + apply/migrate |
| [cloudflare-tunnel-access-edge](2026-06-27-cloudflare-tunnel-access-edge.md) *(backfill)* | `53a7664` `ba13839` `a531f5e` `2810fe8` `7d3be64` `346ec68` `e058a64` | §14 Tunnel + fail-closed Access edge + NodePort fence |
| [player-onboarding-b2](2026-06-27-player-onboarding-b2.md) *(backfill)* | `dbe34a1` `1f8b9bb` `116595f` `fe2ece0` `55592ed` `6c3999a` `879b177` | §B2 email-OTP, account-link, QR, Bind-Code + OTP throttle |
| [username-reclaim-b3](2026-06-27-username-reclaim-b3.md) *(backfill)* | `a29571d` `fdb6efb` | §B3 Mojang-priority reclaim + account migration |
| [felis-metrics](2026-06-30-felis-metrics.md) *(backfill)* | `75642d9` `2a93a9e` `79eae7f` `8ac5e64` | §23 felis_* Prometheus collectors |
| [felis-api-hardening](2026-06-30-felis-api-hardening.md) *(backfill)* | `7a51c1d` `164ac44` `c6c0772` `3c1d647` `d6e3189` `8f41a00` `6368ab1` `2a4a81b` `9873904` | audit #1–#3 + robustness fixes |
| [passkey-enrollment](2026-07-01-passkey-enrollment.md) *(backfill)* | `f2c916d` `742f15f` `0261204` `fce0fce` `7278cd7` `cdbb5ab` `9953275` `20e31fb` `54bc6ef` | WebAuthn enrollment + hardening a–e |
| [passkey-login](2026-07-01-passkey-login.md) *(backfill)* | `e035142` `ec468ba` `0dbd557` `9e1df12` `4f59d51` `a63f49d` | WebAuthn assertion/discoverable login + UA-guard |
| [auto-update-subsystem](2026-07-01-auto-update-subsystem.md) *(backfill)* | `c01f133` `3673af6` `7464fa7` `96b3cc9` `7d27640` `7db57b9` | update decision core + sources + gatherer + window API (report-only) |
| [system-servers-login-limbo-lobby](2026-07-02-system-servers-login-limbo-lobby.md) *(backfill)* | `9bed51b` `9ef817f` `159107b` `dc23cb5` `3fdb3d0` `f554d52` `241fe21` `c7315e4` | always-on login-limbo + lobby auth gate |
| [operator-idle-quota-readiness](2026-07-05-operator-idle-quota-readiness.md) *(backfill)* | `91bfa27` `e574749` `7f7e459` `7becb38` | idle auto-stop, quotas, timeouts, /readyz |
| [break-glass-halt](2026-07-05-break-glass-halt.md) | `c2ee21a` | §B4 break-glass halt-a-server op |
| [felis-migrate-command](2026-07-05-felis-migrate-command.md) | `c1aa38b` | §B3 `/felis migrate` account migration |
| [on-demand-world-backup](2026-07-07-on-demand-world-backup.md) | `7a7c0d5` | §B4 Sync phase 1 — external backup endpoint + Job |
| [internal-backup-endpoint](2026-07-07-internal-backup-endpoint.md) | `f2fc57c` | §B4 Sync phase 2a — internal-face backup endpoint |
| [internal-api-clusterip-service](2026-07-07-internal-api-clusterip-service.md) | `2ba9948` | felis-api internal-face ClusterIP Service |
| [break-glass-backup-peer](2026-07-07-break-glass-backup-peer.md) | `fc748d3` | §B4 Sync phase 2b — console backup peer |
| [adversarial-input-audit](2026-07-08-adversarial-input-audit.md) | `667c6d3` | adversarial input-validation audit — sink-first negative-path, two text→RCON guards mutation-pinned, no code change |
| [felis-nano-hasjoined-resolver](2026-07-12-felis-nano-hasjoined-resolver.md) | `ff550c4` | Felis-nano §B3 — federating hasJoined multiplexer + per-source UUID namespacing |
## Committed change ledger
Oldest first (project build order). Commit = short SHA on `main`. Frontend/`panel`
commits are the collaborator's UI work; backend (Go/Java/K8s) is tracked here as the
primary record.
| Commit | Date | Change |
|---|---|---|
| 5b7b38d | 2026-06-26 | chore: add Go module manifest and ignore rules |
| 5a30aa5 | 2026-06-26 | docs: add OpenAPI 3.1 served-route contract |
| 7fbebfe | 2026-06-26 | feat(apis): add MinecraftServer CRD types (v1alpha1) |
| 708cdfc | 2026-06-26 | feat(core): add naming, RCON, store, config, and image-build libraries |
| 43ab921 | 2026-06-26 | feat(backup): add backup, restore, and reaper subsystems |
| 78b8cf6 | 2026-06-26 | feat(operator): add MinecraftServer controller and reconcilers |
| d39605e | 2026-06-26 | feat(submit): add user modpack build and approval pipeline |
| b508fcc | 2026-06-26 | feat(api): add felis-api service with permissions, modpack lane, and fleet read |
| 47fcd90 | 2026-06-26 | feat(platform): add node orchestration and the felis entrypoint |
| eee00c2 | 2026-06-26 | feat(panel): add three-sided web console (User, Admin, SysAdmin) |
| 93f143f | 2026-06-26 | feat(plugins): add Velocity proxy and Fabric/Forge/NeoForge/Paper integration mods |
| ce0ba76 | 2026-06-27 | chore(api): add kubebuilder object-generation markers to v1alpha1 |
| 7d91373 | 2026-06-27 | fix(migrate): honor -config flag placed after the up verb |
| 99de43f | 2026-06-27 | chore: ignore plugin build artifacts and editor config |
| 58fa4b0 | 2026-06-27 | feat(deploy): add one-line bootstrap installer and container image |
| af14f02 | 2026-06-27 | feat(api): local-password authentication backend |
| e108a37 | 2026-06-27 | feat(cli): break-glass emergency console TUI |
| 885c4a9 | 2026-06-27 | feat(panel): local-password login and forced password change |
| 2d0bbb0 | 2026-06-27 | feat(cli): attribute break-glass recovery to the SysAdmin who runs it |
| dbe34a1 | 2026-06-27 | feat(api): add player email OTP verification (spec §B2 onboarding) |
| 1f8b9bb | 2026-06-27 | feat(api): record account-link auth source (mojang/thirdparty) |
| a29571d | 2026-06-27 | feat(api): reclaim squatted usernames for Mojang-priority players (spec §B3) |
| 53a7664 | 2026-06-27 | feat(cfsetup): recommended Cloudflare Tunnel + Access edge setup |
| ba13839 | 2026-06-27 | feat(breakglass): optional Cloudflare Tunnel + Access setup in the TUI |
| a5a6482 | 2026-06-27 | feat: dev mock |
| dd2fc6f | 2026-06-27 | feat(panel): i18n |
| 7b458da | 2026-06-27 | feat(panel): light/dark theme |
| 51c9eab | 2026-06-27 | docs: CONTRIBUTOR.md |
| 74e7e6e | 2026-06-27 | docs: CONTRIBUTING.md |
| b81b334 | 2026-06-27 | Merge branch 'main' of https://github.com/MliroLirrorsIngenuity/Felis |
| 9d13787 | 2026-06-27 | refactor(panel): dashboard |
| f3521e3 | 2026-06-27 | refactor(panel): uniform margins |
| 18b4be0 | 2026-06-27 | refactor(panel): uniform title icon styles |
| 665841b | 2026-06-27 | fix(panel): remove internal spec references from user-facing text |
| cf88bcc | 2026-06-27 | fix(panel): extract hardcoded security note into i18n keys |
| 061482d | 2026-06-27 | fix(panel): extract hardcoded Chinese text to i18n keys |
| cfe126f | 2026-06-27 | feat(panel): add RCON command input to server console |
| ae91133 | 2026-06-27 | style(panel): refine button styles with shadow, active scale, toned-down colors |
| e189265 | 2026-06-27 | refactor(panel): compact server card layout, denser grid |
| f4df3e2 | 2026-06-27 | feat(panel): pagination for server lists |
| b512a18 | 2026-06-27 | refactor(panel): adjust margins |
| a49443c | 2026-06-27 | refactor(panel): simplify sidebar |
| 86f2ae4 | 2026-06-28 | feat(panel): sidebar foot shows current account + sign-out; reorder Account cards |
| f5d00f3 | 2026-06-28 | feat(cli): implement felis apply command for direct CRD creation |
| 832b200 | 2026-06-28 | fix(panel): reactive system theme detection |
| c9cd9dc | 2026-06-28 | style(panel): unify dialog animation to fade and scale from center |
| 94a3b7b | 2026-06-28 | fix(deploy): harden bootstrap for RHEL-family Linux |
| 9c46632 | 2026-06-28 | feat(cli): add felis setup first-run console with reclaim protection and cfsetup idempotency |
| 5450c26 | 2026-06-28 | chore: normalize line endings and apply formatting |
| deaa2f8 | 2026-06-28 | feat(deploy): add zypper support for openSUSE/SLES |
| 318a724 | 2026-06-28 | feat(deploy): add pacman support for Arch Linux |
| e5f1682 | 2026-06-29 | refactor(deploy)!: TUI |
| 28c3eee | 2026-06-30 | refactor(deploy): improved TUI walkthrough |
| 116595f | 2026-06-30 | feat(api): add QR scan-login completion poll on the internal face |
| a94b001 | 2026-06-30 | feat(deploy): add break-glass Operator account provisioning |
| 346ec68 | 2026-06-30 | refactor(deploy): improved cloudflare walkthrough |
| 563041a | 2026-06-30 | feat(panel): add fail-closed role-switcher view-mode logic |
| 50b8487 | 2026-06-30 | feat(panel): wire role-switcher into the app shell |
| 75642d9 | 2026-06-30 | feat(metrics): add named felis_* Prometheus collectors |
| 2a93a9e | 2026-06-30 | feat(metrics): record felis_image_build_failures_total on failed builds |
| 79eae7f | 2026-06-30 | feat(metrics): publish felis_servers_total from a fleet snapshot |
| 8ac5e64 | 2026-06-30 | feat(metrics): observe felis_start_duration_seconds across the start lifecycle |
| 676407d | 2026-06-30 | docs(diagrams): align §28 sequence diagrams with implemented routes |
| ac02c69 | 2026-06-30 | docs(troubleshooting): add operator failure-mode checklist |
| eb5875a | 2026-06-30 | feat(felis): add Operator break-glass op behind an operation menu |
| c14ed17 | 2026-06-30 | fix(docker): keep embedded panel/ and deploy/ in the image build context |
| 2a4a81b | 2026-06-30 | fix(api): don't burn wake cooldown when refused at capacity |
| 6c3999a | 2026-06-30 | fix(api): rate-limit email-OTP sends to close the email-bomb vector |
| 9873904 | 2026-06-30 | fix(operator): populate Status.Players from an RCON list probe |
| 879b177 | 2026-06-30 | fix(api): make OTP-start throttle atomic to close concurrent-burst bypass |
| 29f5341 | 2026-07-01 | docs(api): correct cooldownLimiter doc for its OTP reuse |
| 7507cfa | 2026-07-01 | Revert "feat(panel): wire role-switcher into the app shell" |
| f2c916d | 2026-07-01 | feat(api): add passkey enrollment persistence layer |
| 742f15f | 2026-07-01 | feat(api): add passkey enrollment endpoints |
| d2de11a | 2026-07-01 | feat(panel): fleet |
| 0261204 | 2026-07-01 | feat(passkey): add go-webauthn enrollment verifier adapter |
| fce0fce | 2026-07-01 | feat(passkey): wire enrollment verifier into felis-api |
| 2810fe8 | 2026-07-01 | fix(cfsetup): repoint stale DNS record when routing a tunnel hostname |
| 7d3be64 | 2026-07-01 | feat(cfsetup): start the tunnel connector as a setup step |
| a531f5e | 2026-07-01 | fix(cfsetup): keep connector install in the host apply layer only |
| e058a64 | 2026-07-01 | feat(edge): close the panel NodePort to the public after the tunnel is up |
| c01f133 | 2026-07-01 | feat(updates): add pure decision core for component self-update |
| fe2ece0 | 2026-07-01 | feat(api): add public Bind-Code onboarding for the player console |
| 3673af6 | 2026-07-01 | feat(api): add admin API for the SysAdmin-set auto-update maintenance window |
| 7464fa7 | 2026-07-01 | fix(updates): tag Window JSON so the persisted maintenance window round-trips |
| e035142 | 2026-07-01 | feat(passkey): add WebAuthn login/assertion crypto adapter |
| f34711c | 2026-07-01 | docs(api): record passkey login-handler deferral rationale |
| 7a51c1d | 2026-07-01 | fix(api): bound concurrent login bcrypt to shed CPU-pin floods |
| 164ac44 | 2026-07-01 | fix(api): validate inbound X-Request-Id before echo and audit persist |
| c6c0772 | 2026-07-01 | fix(api): set read/idle timeouts on the felis-api listeners |
| 3c1d647 | 2026-07-01 | fix(api): cap concurrent SSE streams per principal |
| d6e3189 | 2026-07-01 | fix(api): bound SSE relay writes with a deadline to sever stalled readers |
| 2c56d17 | 2026-07-01 | docs(api): record the quota-claim TOCTOU as a KNOWN-LIMITATION (audit #4) |
| 8f41a00 | 2026-07-01 | fix(api): clear the SSE write deadline on return so it can't leak to a reused connection |
| 15c58d9 | 2026-07-01 | feat(panel): player management |
| 8ae65ae | 2026-07-02 | feat(panel): backup management |
| a15ff55 | 2026-07-02 | refactor(panel): optimize player list layout and horizontal operations |
| 149f01a | 2026-07-02 | fix(panel): change console button to outline variant on my servers page |
| 4ecaf3c | 2026-07-02 | refactor(panel): set defaultOpen parameter of whitelist card to false |
| a9dbc8b | 2026-07-02 | feat(panel): add search and status filtering to my servers page |
| fa7bab5 | 2026-07-02 | style(panel): refine search and filter layout to align with header |
| 70d17a0 | 2026-07-02 | feat(panel): align my servers page search layout with fleet table |
| 5a8eff1 | 2026-07-02 | feat(panel): remove developer comment footer cards from my servers and server admin pages |
| 92770ea | 2026-07-02 | style(panel): adjust pagination padding to pt-3 for balanced spacing |
| 6e43a46 | 2026-07-02 | fix(panel): pin sidebar navigation and enable independent content scroll |
| 3b4298d | 2026-07-02 | refactor(panel): unify servers cockpit layout, resolve duplicate pages and adjust spacing |
| c0d333b | 2026-07-02 | feat(panel): support full server config edit dialog with status prefilling |
| 6368ab1 | 2026-07-02 | fix(api): coalesce MyServers owned flag so ownerless rows do not 500 |
| cdbb5ab | 2026-07-02 | fix(api): record credential id in passkey-register audit event |
| 9953275 | 2026-07-02 | fix(api): bound webauthn_challenges growth by superseding all prior rows |
| 20e31fb | 2026-07-02 | fix(store): cascade-delete passkeys and challenges on user removal |
| 7278cd7 | 2026-07-02 | feat(passkey): require and record user verification at enrollment |
| 54bc6ef | 2026-07-02 | fix(api): clear bound passkeys on password change to close a takeover foothold |
| 19f500b | 2026-07-02 | style(panel): update destructive red color and rename wake to start |
| e0bc288 | 2026-07-02 | feat(panel): implement image build pipeline and admin whitelist with mock dev api |
| 8594622 | 2026-07-02 | feat(panel): implement email OTP verification and passkey registration management |
| 9bed51b | 2026-07-02 | feat(config): add [velocity] login_image/lobby_image for system servers |
| 9ef817f | 2026-07-02 | feat(naming): system-server names, validation, and service-token identifiers |
| 159107b | 2026-07-02 | feat(api): HTTP readiness knob on MinecraftServer and login-gate fallback default |
| dc23cb5 | 2026-07-02 | feat(operator): system-server pod readiness probe and login service-token env |
| 3fdb3d0 | 2026-07-02 | feat(platform): internal API base-URL helper and single-sourced token secret |
| f554d52 | 2026-07-02 | feat(cli): provision login/lobby system servers with login env and token replica |
| a63f49d | 2026-07-02 | feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey |
| 241fe21 | 2026-07-02 | feat(limbo): felis-limbo in-game login flow over the shared account-link client |
| c7315e4 | 2026-07-02 | feat(deploy): login-limbo and lobby images with game-port pinning |
| 191640c | 2026-07-02 | feat(panel): implement admin submission approval and reject queue |
| 598f3d3 | 2026-07-02 | feat(submit): local + S3 backends for modpack upload contexts, installer-selectable |
| adf0d99 | 2026-07-02 | feat(panel): implement user-side modpack submissions with drag & drop context upload |
| d9e866f | 2026-07-03 | fix(deploy): make the lobby image actually build |
| b84debf | 2026-07-03 | feat(deploy): one-shot demo bring-up wrapper |
| 73d6ec1 | 2026-07-03 | feat(mock): add mock submissions for owner account |
| 5427bc7 | 2026-07-03 | feat(servers): support claiming servers directly from ServersPage list |
| 55592ed | 2026-07-03 | feat(auth): support public auth bind endpoint |
| 804459c | 2026-07-03 | feat(panel): implement admin maintenance window settings page |
| dd7dff6 | 2026-07-03 | feat(panel): support importing parameters from submission with owner-restricted unapproved entries |
| a8701c1 | 2026-07-03 | fix(panel): prevent automatic wake during server claim in mock api |
| f749c2c | 2026-07-03 | style(panel): resolve double borders and uneven padding in server console |
| 5f402b1 | 2026-07-03 | fix(panel): force dark mode and pure black bg on server console card |
| b7d8000 | 2026-07-03 | feat(panel): implement dedicated LuckPerms permissions and groups management sub-page |
| e60784b | 2026-07-04 | fix(panel): eliminate page collapse and scroll shifts during LuckPerms query reload |
| 439f19e | 2026-07-04 | fix(panel): prevent page collapse and scroll shifts in players and bans management sections during reload |
| 83e57b4 | 2026-07-04 | fix(panel): implement two-step confirmation for claiming a server to prevent accidental operations |
| 3347cc0 | 2026-07-04 | feat(panel): implement user management administration panel with sessions and minecraft link support |
| 67b4e19 | 2026-07-04 | fix(panel): override generic already_exists error message during user creation and profile editing |
| 2c95da8 | 2026-07-04 | fix(panel/i18n): add missing users_col_user key to translation files |
| 627883e | 2026-07-04 | fix(panel): refine reset password messages and fix empty email placeholder in mock api response |
| 0c1cc59 | 2026-07-04 | feat(auth): migrate console login to passwordless |
| 3b43f05 | 2026-07-04 | refactor(api): drop dead login concurrency limiter and reconcile passwordless comments |
| 4f59d51 | 2026-07-04 | feat(auth): add owner-tier passkey-unbind remediation endpoint |
| c20b12c | 2026-07-04 | refactor(api): drop dead password-era ResetMailer, reconcile passkey-unbind docs |
| 0a2accd | 2026-07-04 | chore: stop tracking Autohand-generated AGENTS.md |
| 96b3cc9 | 2026-07-04 | feat(updater): wire updates.Run to a caller with PaperMC v3 release discovery |
| 9896fe1 | 2026-07-05 | docs(updater): correct PaperMC UA/fixture overclaims, re-tier the boundary |
| 7d27640 | 2026-07-05 | feat(updater): add GitHub Releases source and route felis-api/k3s/cloudflared |
| bd49313 | 2026-07-05 | refactor(panel): 抽取 10 个公共组件,消除 ~150 处重复代码 |
| 91bfa27 | 2026-07-05 | feat(operator): implement idle auto-stop (spec §8) |
| e574749 | 2026-07-05 | feat(api): enforce CPU/memory/storage quotas (spec §9.3, §22) |
| 7db57b9 | 2026-07-05 | feat(updater): add VersionGatherer extraction core and CLI gather seam |
| ec468ba | 2026-07-05 | feat(auth): add discoverable (usernameless) passkey login |
| 154002e | 2026-07-05 | docs(auth): cite MultiLogin reference for UUID-keyed reclaim split |
| 0dbd557 | 2026-07-05 | fix(store): renumber discoverable-login migration 0013 -> 0014 |
| 9e1df12 | 2026-07-05 | feat(passkey): advance sign_count, reject clone-warned assertions |
| 7f7e459 | 2026-07-05 | fix(operator): enforce startup and readiness timeouts (§5, §8) |
| 7becb38 | 2026-07-05 | fix(api): implement /readyz with real DB + K8s API + CRD checks (§7) |
| 9079a2c | 2026-07-05 | feat(panel): implement email otp and passkey login interface |
| cfe68ae | 2026-07-05 | fix(panel): align status distribution order to put Stopped at the end |
| bbcfaeb | 2026-07-05 | refactor(panel): remove redundant voxel network topology description subtitle |
| fdb6efb | 2026-07-05 | feat(account): migrate a live account's owned servers to a new account (§B3 inherit) |
| abad137 | 2026-07-06 | style(panel): unify vertical spacing below PageHeader across pages |
| c2ee21a | 2026-07-06 | feat(breakglass): add halt-a-server op to the recovery console (§B4) |
| c1aa38b | 2026-07-06 | feat(velocity): add /felis migrate to open an account migration (§B3 inherit) |
| 7a7c0d5 | 2026-07-07 | feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync) |
| f2fc57c | 2026-07-07 | feat(api): add internal-face break-glass world backup endpoint (§B4 Sync) |
| 2ba9948 | 2026-07-07 | fix(platform): front the felis-api internal face on its own ClusterIP Service |
| fc748d3 | 2026-07-07 | feat(breakglass): add "back up a world now" console peer (§B4 Sync) |
| 9911b8c | 2026-07-07 | docs(changes): record the break-glass backup console peer (§B4 Sync phase 2b) |
| 096d597 | 2026-07-07 | docs(changes): backfill detail docs for pre-ledger functional commits |
| 5a7cd5a | 2026-07-07 | docs(changes): fold 346ec68 cloudflare-edge walkthrough into its detail doc |
| 4626ab5 | 2026-07-07 | docs(changes): mutation-audit the ledger's "unit-tested" safety claims |
| 729bd7b | 2026-07-07 | docs(changes): index the mutation audit and two lagging ledger rows |
| c67a4d3 | 2026-07-08 | docs(changes): close §B4 with the S3 archive backend deferred by design |
| 85b8a92 | 2026-07-08 | test(api): pin restore's owner gate against a superseded former owner |
| b7b4a3b | 2026-07-08 | docs(changes): record the round-2 backup/restore mutation audit and index the owner-gate test |
+33 -7
View File
@@ -42,9 +42,27 @@ A grep across `*.md` and `*.go` returns both sets; only the Go ones are seams.
does not exist. `felis update` runs with a zero window, under which every does not exist. `felis update` runs with a zero window, under which every
`Scheduled` component degrades to a notify, so no path can currently claim an `Scheduled` component degrades to a notify, so no path can currently claim an
apply is under way. apply is under way.
- `internal/submit/blobstore.go:40` — the uploads PVC is mounted into felis-api but - `internal/submit/blobstore.go` — CLOSED 2026-09-22. The uploads PVC still cannot
not into the Kaniko build Pod, so a submitted context is durable at the derived cross namespaces, so the transport went through the API instead of a mount: the
location without yet being readable by the build that consumes it. derived context ref is now the internal-face URL
(`/api/v1/internal/submissions/{id}/context`, service-token gated), the build
Job's `context-fetch` initContainer streams it with `felis fetch-context` and
extracts under a zip-slip guard into a size-limited emptyDir, and Kaniko builds
`--context=/context`. The token reaches the build namespace through the same
Secret-replica mechanism the login gate uses (bootstrap + `felis setup`), and the
build egress lock allows exactly the control namespace on the internal port.
Uniform for local and s3:// stores — neither hands the sandboxed build Pod a
filesystem view or object-store credentials. Kaniko/Trivy images are
external-only by default; `[registry] kaniko_image / trivy_image /
build_cpu_limit / build_mem_limit` override them for mirrored or air-gapped
installs. Trivy's vulnerability DB is the same story, and now has its own knob:
`[registry] trivy_db_repository` points `--db-repository` at an internal mirror
(recipe in docs/troubleshooting.md §8e); `trivy_java_db_repository` does the
same for the Java DB, which Trivy fetches so soon as the scanned image contains
a jar — i.e. for every real modpack build. Left unset on an egress-locked box
the scan step fails closed — Kaniko pushes, Trivy exits on the DB download —
which is the correct fail direction but leaves the build unfinished, so the
mirrors are part of a production build install.
## Built; only its I/O is unverifiable from this repo ## Built; only its I/O is unverifiable from this repo
@@ -58,17 +76,25 @@ or a real upstream account to run it against — not an implementation.
drives the whole flow through a fake. drives the whole flow through a fake.
- `internal/api/console.go:39`, `internal/api/logstream.go:236`, - `internal/api/console.go:39`, `internal/api/logstream.go:236`,
`internal/api/logstream.go:306`, `internal/fileedit/k8sjobs.go:45` — each needs a `internal/api/logstream.go:306`, `internal/fileedit/k8sjobs.go:45` — each needs a
live cluster (RCON, `pods/log` follow, a Job). live cluster (RCON, `pods/log` follow, a Job). **Verified live 2026-09-22/23
(auditfix7–25):** the RCON command spine (wake → probe → `command`/access
mutations/stop), the log SSE stream, and the fileedit Job have each run
end-to-end on the drill cluster.
- `internal/api/handlers_access.go:170,490` — parsing real vanilla and LuckPerms - `internal/api/handlers_access.go:170,490` — parsing real vanilla and LuckPerms
command output. command output. **Verified live 2026-09-23:** players / whitelist / banlist
parses matched a live Paper server's replies (LuckPerms not installed → the raw
reply falls through as documented; the input guards held on four negative cases).
## Deliberately accepted, not scheduled to close ## Deliberately accepted, not scheduled to close
These are decisions, not backlog. Each names the condition under which it would be These are decisions, not backlog. Each names the condition under which it would be
worth revisiting. worth revisiting.
- `internal/api/pgrepo.go:281` — the quota check and `ClaimServer` are two statements - ~~`internal/api/pgrepo.go:281` — the quota check and `ClaimServer` are two statements
(audit #4 TOCTOU). Closeable only against a real Postgres. (audit #4 TOCTOU). Closeable only against a real Postgres.~~ **Closed** — the gate
moved inside `ClaimServer` (advisory lock + re-check + UPDATE in one transaction),
red-then-green in the pgint suite, which is exactly the real-Postgres harness this
line was waiting for.
- `internal/api/api.go:671` — `cooldownLimiter` is process-local, so across N api - `internal/api/api.go:671` — `cooldownLimiter` is process-local, so across N api
replicas a caller could draw up to N OTP codes per window. The intra-replica burst replicas a caller could draw up to N OTP codes per window. The intra-replica burst
is closed; cross-replica bounding needs a shared store, out of scope for a is closed; cross-replica bounding needs a shared store, out of scope for a
+303 -24
View File
@@ -337,6 +337,19 @@ components:
build_id: build_id:
type: string type: string
description: image_builds.id, set only after the build hand-off succeeds. description: image_builds.id, set only after the build hand-off succeeds.
build_status:
type: string
enum: [pending, building, succeeded, failed, cancelled]
description: >-
The linked build's outcome, attached by the LIST routes
(/me/submissions, /submissions) — for a submitter this is the only
visible outlet for a failed build. Omitted until a build is linked
and its row is readable.
build_error:
type: string
description: >-
The build's recorded failure text (e.g. a CRITICAL CVE scan
failure), attached alongside build_status.
reviewed_by: { type: string } reviewed_by: { type: string }
reject_reason: { type: string } reject_reason: { type: string }
created_at: { type: string, format: date-time } created_at: { type: string, format: date-time }
@@ -454,26 +467,47 @@ paths:
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
/metrics:
get:
tags: [metrics]
operationId: metrics
summary: Prometheus metrics (felis_* collectors) on the internal face.
description: >-
Scrape-only infrastructure route, not a product API: the internal listener is
ClusterIP-only and a Prometheus scrape carries no token, the same stance as the
probes. Serves the felis_* exposition documented in troubleshooting §14; the
external face never serves it.
x-felis-face: [internal]
x-felis-tier: public
security: []
responses:
'200':
description: Prometheus text exposition format.
content:
text/plain:
schema: { type: string }
/session/minecraft/hasJoined: /session/minecraft/hasJoined:
get: get:
tags: [nano] tags: [nano]
operationId: hasJoined operationId: hasJoined
summary: Multi-source session verifier (Felis-nano hasJoined multiplexer). summary: Multi-source session verifier (Felis-nano hasJoined multiplexer).
description: >- description: >-
Velocity's authlib is pointed here via -Dmojang.sessionserver or a thin login Velocity is pointed here with -Dmojang.sessionserver and sends the request itself.
hook. Unauthenticated — the vanilla sessionserver protocol carries no token. The Unauthenticated — the vanilla sessionserver protocol carries no token. The query
query is fanned out to the configured Yggdrasil roots in priority order (the is fanned out to the configured Yggdrasil roots in priority order (the Mojang
Mojang identity source first); the first source to validate the serverId hash identity source first); the first source to validate the serverId hash wins. A
wins. A non-identity source's self-asserted UUID is rewritten into a per-source non-identity source's self-asserted UUID is rewritten into a per-source namespace
namespace (UUIDv3) before return, so it can never land in Mojang's UUID space. (UUIDv3) before return, so it can never land in Mojang's UUID space. A rejected
A rejected or barred login is 204, which authlib maps to a verify failure. or barred login is 204, which Velocity answers with its online-mode-only kick.
Any other non-200 status makes Velocity report the auth servers as down.
x-felis-face: [internal] x-felis-face: [internal]
x-felis-tier: public x-felis-tier: public
security: [] security: []
parameters: parameters:
- { name: username, in: query, required: true, schema: { type: string } } - { name: username, in: query, required: true, schema: { type: string, maxLength: 64 } }
- { name: serverId, in: query, required: true, schema: { type: string } } - { name: serverId, in: query, required: true, schema: { type: string, maxLength: 64 } }
- { name: ip, in: query, required: false, schema: { type: string } } - { name: ip, in: query, required: false, schema: { type: string, maxLength: 64 } }
responses: responses:
'200': '200':
description: A source validated the session; the canonical game profile. description: A source validated the session; the canonical game profile.
@@ -484,10 +518,33 @@ paths:
required: [id, name] required: [id, name]
properties: properties:
id: { type: string, description: Canonical UUID, undashed 32-hex. } id: { type: string, description: Canonical UUID, undashed 32-hex. }
name: { type: string } name:
type: string
description: >-
The name the source returned. A third-party player whose name is
registered to a Mojang account gets it back as PREFIX_name, cut to
16 characters.
properties: { type: array, items: { type: object } } properties: { type: array, items: { type: object } }
'204': '204':
description: No source validated the session, or the resolved UUID is barred. description: >-
Not admitted, with no source asked when username or serverId is missing or a
parameter is over 64 bytes. Otherwise no source validated the session, the
canonical UUID is barred, a third-party source returned a name that is not a
legal Minecraft username, or the identity source returned an unparseable id.
'400':
description: >-
The request declared a body. No body is sent back, and the connection is
closed.
'500':
description: The bar-list lookup failed, so the login is not admitted.
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'503':
description: >-
No source validated the session and at least one source failed (transport
error, redirect, unexpected status, or a 200 without a usable profile). Its
player may be the one logging in, so this is not answered as a 204. No body.
# -------------------------------------------------- internal: servers ------ # -------------------------------------------------- internal: servers ------
/api/v1/servers: /api/v1/servers:
@@ -583,6 +640,34 @@ paths:
'404': '404':
$ref: '#/components/responses/NotFound' $ref: '#/components/responses/NotFound'
/api/v1/internal/submissions/{id}/context:
get:
tags: [submissions-internal]
operationId: internalSubmissionContext
summary: Stream a submission's stored build-context tarball to the build Pod.
description: >-
The build Job's fetch initContainer cannot mount the control-plane uploads
PVC (a PVC does not cross namespaces) and holds no object-store
credentials, so the API that stored the blob streams it here. Served on
the internal face (service token, no Zero Trust).
x-felis-face: [internal]
x-felis-tier: service
security: [{ serviceToken: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The stored gzip tarball, verbatim.
content:
application/gzip:
schema: { type: string, format: binary }
'401':
$ref: '#/components/responses/Unauthorized'
'404':
$ref: '#/components/responses/NotFound'
'503':
$ref: '#/components/responses/ServiceUnavailable'
/api/v1/internal/servers/{name}/join-event: /api/v1/internal/servers/{name}/join-event:
post: post:
tags: [servers-internal] tags: [servers-internal]
@@ -2666,6 +2751,13 @@ paths:
The owner's display identity (email, or username when The owner's display identity (email, or username when
the address is absent). Absent for an unclaimed server the address is absent). Absent for an unclaimed server
or when the best-effort owner lookup failed. or when the best-effort owner lookup failed.
system:
type: boolean
description: >-
True for a platform-provisioned system service (the login
gate, the lobby). Their reserved names are rejected by
every per-server route, so the cockpit renders them
read-only instead of offering actions that would 400.
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'403': '403':
@@ -2747,11 +2839,12 @@ paths:
post: post:
tags: [backups] tags: [backups]
operationId: backupNow operationId: backupNow
summary: Back up a server's world on demand (owner-or-admin; server must be stopped). summary: Back up a server's data volume on demand (owner-or-admin; server must be stopped).
description: >- description: >-
Snapshots the server's world into the archive store as a first-class Snapshots the server's whole data volume (worlds, config, plugins/mods,
world_backups row (reason "manual"), restorable later like an inactivity jars, libraries — not just world folders) into the archive store as a
backup. The world PVC is RWO and held by a running server, so the server must first-class world_backups row (reason "manual"), restorable later like an
inactivity backup. A restore replaces the volume with the archive. The world PVC is RWO and held by a running server, so the server must
be fully stopped first (409 not_stopped otherwise). The backup runs be fully stopped first (409 not_stopped otherwise). The backup runs
asynchronously as a Job, so success is 202 (backing_up). asynchronously as a Job, so success is 202 (backing_up).
x-felis-face: [external] x-felis-face: [external]
@@ -2787,6 +2880,56 @@ paths:
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
# -------------------------------------------------- async job status (app) ---
/api/v1/servers/{name}/jobs:
get:
tags: [backups]
operationId: listServerJobs
summary: Latest async world operations (backup/restore) for a server (owner-or-admin).
description: >-
Backup and restore run as cluster Jobs, so a 202 that later failed left
its only trace in the Job object. This route projects the newest such
Jobs, newest first, so failures are observable without kubectl. State is
"running" | "succeeded" | "failed".
x-felis-face: [external]
x-felis-tier: app
security: [{ accessJWT: [] }]
parameters:
- { name: name, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The server's newest backup/restore jobs.
content:
application/json:
schema:
type: object
required: [server, jobs]
properties:
server: { type: string }
jobs:
type: array
items:
type: object
required: [name, kind, state]
properties:
name: { type: string }
kind: { type: string, enum: [backup, restore] }
state: { type: string, enum: [running, succeeded, failed] }
message: { type: string }
started_at: { type: string, format: date-time }
finished_at: { type: string, format: date-time }
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/Forbidden'
'404':
description: Unknown server.
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'503':
$ref: '#/components/responses/ServiceUnavailable'
# ------------------------------------------------- server file editor (app) --- # ------------------------------------------------- server file editor (app) ---
/api/v1/servers/{name}/files: /api/v1/servers/{name}/files:
get: get:
@@ -4123,18 +4266,30 @@ paths:
$ref: '#/components/responses/BadRequest' $ref: '#/components/responses/BadRequest'
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'403':
description: >-
The per-user submission allowance is spent — too many of the
caller's submissions are awaiting review, or their stored-upload
budget is full (submission_quota_exceeded).
'429':
description: >-
A submission was created within the per-user cooldown window
(submission_cooldown).
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
get: get:
tags: [submissions] tags: [submissions]
operationId: mySubmissions operationId: mySubmissions
summary: List the caller's own modpack submissions (user-directed lane over §16). summary: List the caller's own modpack submissions with each linked build's outcome (user-directed lane over §16).
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: app x-felis-tier: app
security: [{ accessJWT: [] }] security: [{ accessJWT: [] }]
responses: responses:
'200': '200':
description: The caller's submissions, newest first. description: >-
The caller's submissions, newest first; rows with a linked build
additionally carry build_status/build_error so the submitter can see
whether their build succeeded or failed (and why).
content: content:
application/json: application/json:
schema: schema:
@@ -4161,8 +4316,10 @@ paths:
principal; a submission the caller does not own is reported as 404, so principal; a submission the caller does not own is reported as 404, so
this endpoint cannot upload to or probe another user's submission. Only a this endpoint cannot upload to or probe another user's submission. Only a
pending_review submission accepts a context (409 otherwise); a wrong-format pending_review submission accepts a context (409 otherwise); a wrong-format
or oversize body is rejected with 400. Returns 503 when the deployment's or oversize body is rejected with 400, and an upload that would push the
context store has no implemented upload transport. caller past their per-user stored-context budget is refused with 403
before the excess is persisted. Returns 503 when the deployment's context
store has no implemented upload transport.
x-felis-face: [external] x-felis-face: [external]
x-felis-tier: app x-felis-tier: app
security: [{ accessJWT: [] }] security: [{ accessJWT: [] }]
@@ -4183,10 +4340,53 @@ paths:
$ref: '#/components/responses/BadRequest' $ref: '#/components/responses/BadRequest'
'401': '401':
$ref: '#/components/responses/Unauthorized' $ref: '#/components/responses/Unauthorized'
'403':
description: >-
The upload would exceed the caller's per-user stored-context budget
(submission_quota_exceeded).
'404': '404':
$ref: '#/components/responses/NotFound' $ref: '#/components/responses/NotFound'
'409': '409':
$ref: '#/components/responses/Conflict' $ref: '#/components/responses/Conflict'
'429':
description: >-
An upload was accepted within the per-user cooldown window
(submission_cooldown).
'503':
$ref: '#/components/responses/ServiceUnavailable'
/api/v1/me/submissions/{id}:
delete:
tags: [submissions]
operationId: withdrawSubmission
summary: Withdraw your own pending submission (user side; user-directed lane over §16).
description: >-
Retracts the caller's own submission while it is still pending review:
the row and its uploaded build context are deleted, freeing the pending
slot and the per-user storage budget for a fresh submission. A reviewed
submission is frozen (409 — its build may already be consuming the
context), and a submission the caller does not own reads back as 404, so
this endpoint cannot probe or clear another user's uploads.
x-felis-face: [external]
x-felis-tier: app
security: [{ accessJWT: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The withdrawn submission, as it was before the deletion.
content:
application/json:
schema: { $ref: '#/components/schemas/Submission' }
'401':
$ref: '#/components/responses/Unauthorized'
'404':
$ref: '#/components/responses/NotFound'
'409':
description: Submission has already been reviewed and cannot be withdrawn.
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
@@ -4261,10 +4461,24 @@ paths:
type: object type: object
required: [image_ref, dockerfile, context_ref] required: [image_ref, dockerfile, context_ref]
properties: properties:
image_ref: { type: string } image_ref:
dockerfile: { type: string } type: string
context_ref: { type: string } description: Push target under the internal registry (e.g. registry.felis.svc:5000/foo:1.0).
base_image: { type: string } dockerfile:
type: string
description: >-
Audit archive of the recipe, recorded on the build row and shown in the
panel — the executed Dockerfile is the file named `Dockerfile` at the
root of the context tarball (Kaniko runs --dockerfile=Dockerfile), so
this field is never executed.
context_ref:
type: string
description: >-
Location of the uploaded gzip build context; its root must contain the
Dockerfile that gets executed.
base_image:
type: string
description: Resolved FROM, recorded for audit only — not a build gate.
responses: responses:
'202': '202':
description: Build accepted. description: Build accepted.
@@ -4498,6 +4712,39 @@ paths:
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
/api/v1/submissions/{id}/context:
get:
tags: [submissions]
operationId: downloadSubmissionContext
summary: Download a submission's uploaded build context (admin; user-directed lane over §16).
description: >-
The reviewer's read path to the artifact they are about to approve: the
executed Dockerfile lives inside this tarball (Kaniko runs the context's
root `Dockerfile`), so without it the human gate would be blind. Streams
the stored context.tar.gz verbatim with an attachment disposition — the
same bytes the build Pod fetches over the internal face. 404 when the
submission is unknown or has no uploaded context; 503 when the
deployment's context store has no implemented transport.
x-felis-face: [external]
x-felis-tier: admin
security: [{ accessJWT: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The stored build context (gzip tarball), served as an attachment.
content:
application/gzip:
schema: { type: string, format: binary }
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/Forbidden'
'404':
$ref: '#/components/responses/NotFound'
'503':
$ref: '#/components/responses/ServiceUnavailable'
/api/v1/submissions/{id}/reject: /api/v1/submissions/{id}/reject:
post: post:
tags: [submissions] tags: [submissions]
@@ -4538,3 +4785,35 @@ paths:
schema: { $ref: '#/components/schemas/Error' } schema: { $ref: '#/components/schemas/Error' }
'503': '503':
$ref: '#/components/responses/ServiceUnavailable' $ref: '#/components/responses/ServiceUnavailable'
/api/v1/submissions/{id}:
delete:
tags: [submissions]
operationId: deleteSubmission
summary: Retire a submission outright — row and uploaded context (admin; user-directed lane over §16).
description: >-
Removes the submission and its uploaded build context, any status — the
lane's only lifecycle valve, and the path that reclaims a rejected or
consumed upload from the uploads PVC. The reviewer identity is recorded
in the audit event, not on the (now deleted) row. Deleting an approved
submission whose build is still running fails that build's context
fetch; the admin has explicitly chosen to retire the artifact.
x-felis-face: [external]
x-felis-tier: admin
security: [{ accessJWT: [] }]
parameters:
- { name: id, in: path, required: true, schema: { type: string } }
responses:
'200':
description: The deleted submission, as it was before the deletion.
content:
application/json:
schema: { $ref: '#/components/schemas/Submission' }
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/Forbidden'
'404':
$ref: '#/components/responses/NotFound'
'503':
$ref: '#/components/responses/ServiceUnavailable'
+4 -4
View File
@@ -87,7 +87,7 @@ sequenceDiagram
else quota available else quota available
Repo-->>API: true Repo-->>API: true
API->>Repo: ClaimServer(name, user_id) API->>Repo: ClaimServer(name, user_id)
Note over Repo: SELECT EXISTS(server); then atomic UPDATE servers SET owner_id=$2, claimed_at=now() WHERE name=$1 AND owner_id IS NULL AND deleted_at IS NULL Note over Repo: ONE transaction: pg_advisory_xact_lock(user_id) serializes this user's claim lane; SELECT FROM servers WHERE name=$1 AND deleted_at IS NULL FOR UPDATE; re-run the four-dimension quota gate (authoritative — the pre-check above is a fast path); then UPDATE servers SET owner_id=$2, claimed_at=now() WHERE name=$1 AND owner_id IS NULL AND deleted_at IS NULL
alt server missing alt server missing
Repo-->>API: ErrNotFound Repo-->>API: ErrNotFound
API-->>Panel: 404 not_found API-->>Panel: 404 not_found
@@ -137,14 +137,14 @@ sequenceDiagram
Panel->>APIExternal: POST /api/v1/account/link/verify {code} Panel->>APIExternal: POST /api/v1/account/link/verify {code}
APIExternal->>APIExternal: trim and uppercase code APIExternal->>APIExternal: trim and uppercase code
APIExternal->>Repo: VerifyLinkCode(user_id, code, now) APIExternal->>Repo: VerifyLinkCode(user_id, code, now)
Repo->>Repo: SELECT non-expired code Repo->>Repo: SELECT mc_uuid, auth_source FROM account_link_codes WHERE code=$1 AND expires_at>$2 FOR UPDATE
alt missing or expired code alt missing or expired code
Repo-->>APIExternal: ErrLinkCodeInvalid Repo-->>APIExternal: ErrLinkCodeInvalid
APIExternal-->>Panel: 400 invalid_code APIExternal-->>Panel: 400 invalid_code
else UUID linked to another user else UUID linked to a different, live user
Repo-->>APIExternal: ErrConflict Repo-->>APIExternal: ErrConflict
APIExternal-->>Panel: 409 already_linked APIExternal-->>Panel: 409 already_linked
else valid code else valid code (re-verify by the same user is idempotent; a retired/soft-deleted owner's link is taken over)
Repo->>Repo: INSERT account_links(user_id, mc_uuid, auth_source) ON CONFLICT (user_id, mc_uuid) DO UPDATE auth_source Repo->>Repo: INSERT account_links(user_id, mc_uuid, auth_source) ON CONFLICT (user_id, mc_uuid) DO UPDATE auth_source
Repo->>Repo: DELETE account_link_codes WHERE code=$1 Repo->>Repo: DELETE account_link_codes WHERE code=$1
Repo-->>APIExternal: mc_uuid, auth_source Repo-->>APIExternal: mc_uuid, auth_source
+336 -3
View File
@@ -20,8 +20,11 @@ graded for how far the in-repo Go test suite proves the behaviour:
containerd, Postgres, or the network, not by Felis Go code; you will see it in containerd, Postgres, or the network, not by Felis Go code; you will see it in
`kubectl describe` / pod logs, never in `MinecraftServer.status`. `kubectl describe` / pod logs, never in `MinecraftServer.status`.
- **[INERT]** — the configuration field exists in the CRD but no controller - **[INERT]** — the configuration field exists in the CRD but no controller
reads it. Tuning it does nothing. §12 lists the one field this still applies reads it. Tuning it does nothing. **No CRD field carries this status today**;
to, alongside the fields that *are* read and the condition each depends on. the last one, `spec.storage.retainOnDelete`, was removed rather than
implemented (§13 records why). A field whose change seems ignored is almost
always a condition instead — §12 lists the fields that *are* read and the
condition each depends on.
The operator never invents the parent domain; routing identity is The operator never invents the parent domain; routing identity is
`spec.subdomain` under the deployment zone. Examples below use `spec.subdomain` under the deployment zone. Examples below use
@@ -391,6 +394,95 @@ Inspect:
kubectl logs -n felis-build job/<build-job> kubectl logs -n felis-build job/<build-job>
``` ```
### 8e. Build Pods never start: executor images and air-gapped installs
The build Job runs Kaniko and Trivy from external registries by default
(`gcr.io/kaniko-project/executor:latest`, `aquasec/trivy:latest`). On a box whose
build namespace cannot reach those registries (the egress policy allows only
DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box
has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the
build stays `building` until its deadline. Point the overrides at images **in
the internal registry** — the one pull source that survives an image GC (a bare
node-containerd import does not: kubelet's image GC collects unused images under
disk pressure, and an air-gapped box then has nothing to restore them from) —
in `felis.toml`:
```toml
[registry]
url = "registry.felis.svc:5000"
build_namespace = "felis-build"
kaniko_image = "registry.felis.svc:5000/mirror/kaniko-executor:v1.24.0"
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.74.0"
trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2"
trivy_java_db_repository = "registry.felis.svc:5000/mirror/trivy-java-db:1"
build_cpu_limit = "2"
build_mem_limit = "4Gi"
```
Mirror the executor images into the registry once. On the node itself, push
through the loopback hostPort the registry Deployment binds (docker treats
`127.0.0.1` as insecure by default; the installer leaves the daemon stopped, so
`sudo systemctl start docker` first):
```sh
docker pull gcr.io/kaniko-project/executor:v1.24.0 # any versions you pin
docker pull aquasec/trivy:0.74.0
docker pull mirror.gcr.io/aquasec/trivy-java-db:1
docker tag gcr.io/kaniko-project/executor:v1.24.0 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0
docker tag aquasec/trivy:0.74.0 127.0.0.1:5000/mirror/trivy:0.74.0
docker tag mirror.gcr.io/aquasec/trivy-java-db:1 127.0.0.1:5000/mirror/trivy-java-db:1
docker push 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0
docker push 127.0.0.1:5000/mirror/trivy:0.74.0
docker push 127.0.0.1:5000/mirror/trivy-java-db:1
```
From another machine, port-forward the registry instead (`kubectl -n felis
port-forward svc/registry 5000:5000`) and push to `localhost:5000/...` — the
registry keys a repository by the path after the host, so pushes through either
door land in the same place the build Pods will pull from.
Put them in **both** `/etc/felis/felis.host.toml` (host-side CLI) and
`/etc/felis/felis.pod.toml` (the file rendered into the API's `felis-config`
Secret — the two differ only in the database URL; the setup screens re-render
the Secret from the pod file, so edits made only through `kubectl` on the live
Secret are lost at the next reconfigure). A Deployment restart alone is NOT
enough — the API Pod mounts the Secret, never the host file. Re-render the
Secret from the pod file, then roll `felis-api`:
```sh
kubectl -n felis create secret generic felis-config \
--from-file=felis.toml=/etc/felis/felis.pod.toml --dry-run=client -o yaml | kubectl apply -f -
kubectl -n felis rollout restart deployment/felis-api
```
Unset fields keep the defaults.
`trivy_db_repository` is not optional on an egress-locked box. Trivy fetches its
vulnerability DB from `mirror.gcr.io`/`ghcr.io` unless told otherwise, and the
build egress policy denies those hosts — so the scan step fails closed
(`failed to download vulnerability DB`) and NO build ever completes, even though
Kaniko pushed the image. Mirror the DB into the internal registry once:
```
# On the node (docker treats 127.0.0.1 as insecure by default), or through the
# port-forward above:
# docker pull mirror.gcr.io/aquasec/trivy-db:2
# docker tag mirror.gcr.io/aquasec/trivy-db:2 127.0.0.1:5000/mirror/trivy-db:2
# docker push 127.0.0.1:5000/mirror/trivy-db:2
```
The Job's Trivy container already runs with `--insecure`, so the internal
registry's plain HTTP works for the DB pull exactly as it does for the scanned
image. Re-mirror the tag periodically (Trivy refreshes the DB several times a
day upstream; a stale mirror only means stale CVE data, never a failed gate).
`trivy_java_db_repository` is the same story one step lazier: Trivy downloads
the Java DB on demand the first time it scans an image containing Java
artifacts — every real modpack — and that download fails closed too. Mirror
`mirror.gcr.io/aquasec/trivy-java-db:1` alongside the vulnerability DB (commands
above); the Java DB refreshes far less often than the vulnerability DB, so a
one-off mirror is usually fine.
--- ---
## 9. Registry push/pull failures (spec §15) ## 9. Registry push/pull failures (spec §15)
@@ -408,6 +500,16 @@ control namespace (or `--registry-namespace`):
- **Storage:** PVC is RWO, `10Gi`, mounted at `/var/lib/registry`, **no - **Storage:** PVC is RWO, `10Gi`, mounted at `/var/lib/registry`, **no
`storageClassName`** → binds the cluster default class. If the cluster has no `storageClassName`** → binds the cluster default class. If the cluster has no
default StorageClass the PVC stays `Pending` and the registry never starts. default StorageClass the PVC stays `Pending` and the registry never starts.
- **Node-side pulls:** containerd cannot dial the Service VIP (the live stack
answered "Empty reply"), so the registry Deployment binds a loopback hostPort
(`127.0.0.1:<port>`) and the installer writes a `/etc/rancher/k3s/registries.yaml`
mirror relaying `registry.<ns>.svc:<port>` onto it. That pair is what lets
kubelet re-pull a garbage-collected image; both halves must survive together
(remove either and every pull after an image GC fails).
- **Memory:** the registry's limit is 2Gi, deliberately larger than the other
control-plane pods' 256Mi — a live 475MB-layer push OOM-killed the 256Mi
template mid-upload (audit #46). Very large layers need headroom here, not
more CPU.
- **Selector quirk worth knowing:** the registry Service selector is only - **Selector quirk worth knowing:** the registry Service selector is only
`name + component=registry` — it deliberately lacks the `name + component=registry` — it deliberately lacks the
`part-of=felis-control-plane` label, so the registry is *invisible* to the `part-of=felis-control-plane` label, so the registry is *invisible* to the
@@ -426,6 +528,16 @@ reaps a world only when `now - last_active_at > 15d` (`inactive_15d`); the 15-da
deadline is **hard-fixed in code** (only `warn_before` / `retention` / deadline is **hard-fixed in code** (only `warn_before` / `retention` /
`max_local_bytes` are configurable from `felis.toml [archive]`). `max_local_bytes` are configurable from `felis.toml [archive]`).
### What a "backup" contains
A backup tars the server's ENTIRE data volume — the same volume the server mounts
at `/data`: world folders, `server.properties`, plugins/mods, configs, jars,
libraries, logs and cache, not just the `world/` directory. A restore replaces the
volume's contents with the archive (files added since the backup are pruned), so a
restore also rolls config/plugin changes back. Sizes are dominated by
libraries/cache on stock Paper servers (~170MB for a fresh instance before any
world growth) — do not size the archive PVC as if only world data were stored.
### The backup-before-delete invariant ### The backup-before-delete invariant
The reap sequence (all [GO-TESTED] hermetically) preserves the world unless a The reap sequence (all [GO-TESTED] hermetically) preserves the world unless a
@@ -449,6 +561,29 @@ So a missing backup never results in a deleted world. [GO-TESTED:
- CRD missing → logs `reaper: CRD missing, skipping`, skipped. - CRD missing → logs `reaper: CRD missing, skipping`, skipped.
- Idle `≤ 15d` → not yet eligible. - Idle `≤ 15d` → not yet eligible.
### Pre-reap warnings (the `warn_before` offsets)
An OWNED server inside a warning window gets an email notice (`3d`/`1d` before
the deadline, `warn_before` from `[archive]`) to the owner's **verified** email —
the same `[smtp]` relay felis-api uses. The `warned_3d_at` / `warned_1d_at`
stamps record a **delivered** notice:
- No `[smtp]` configured (or owner has no verified address): the run logs
`reaper: warning suppressed — no warner wired` / a delivery error and does
NOT stamp. Nothing is falsely recorded as sent, and the day SMTP is
configured the pending warning can still go out.
- Delivery failure (relay down): logged and retried on the next daily run —
bounded by the warning window, since the reap removes the candidate anyway.
- `warned=` in the run output counts DELIVERED notices, not attempts.
The reaper runs in the minecraft namespace and reads the **mirrors** of
`felis-smtp` and `felis-config` there (a `secretKeyRef` is namespace-local). The
installed `felis setup`'s "configure email" screen refreshes both mirrors when it
applies, so configuring SMTP after install is enough; a manual edit of the
control-namespace Secret alone is not. [GO-TESTED: the delivered/retried/
suppressed matrix in `internal/reaper`; live-drilled end to end against a local
SMTP sink.]
### Genuine false-delete risk vectors ### Genuine false-delete risk vectors
- **Stale `last_active_at`.** The keep-alive is `RecordJoin`, called from the - **Stale `last_active_at`.** The keep-alive is `RecordJoin`, called from the
@@ -468,6 +603,32 @@ Only `TarLocal` (tar+gzip) archiving is implemented; VolumeSnapshot/Longhorn
backends return `not implemented in this build`. The live PVC delete / Postgres backends return `not implemented in this build`. The live PVC delete / Postgres
store paths are [INTEGRATION-ONLY]. store paths are [INTEGRATION-ONLY].
### Where worlds are read from (hostPath resolution)
The CronJob mounts `--worlds-host-path` read-only at `/worlds`; the resolver
runs `cmd/felis/reaper.resolveWorldDir`: it looks for `<root>/<pvc>`, then for
the stock local-path directory `<root>/<pv-name>_<ns>_<pvc-name>` derived from
the live PVC's `spec.volumeName` (never a glob — a leftover directory of a
deleted PV must not stand in for the world the PVC currently binds). Pointing
the flag at k3s's storage root (`/var/lib/rancher/k3s/storage`) is therefore the
supported way to enable retention on a stock install. Two deployment facts the
resolver cannot fix:
- **Permissions.** The reaper Pod runs as **root** and carries `DAC_OVERRIDE`:
worlds are written by the game image's own UID (root for every Paper image we
ship), and Paper saves `level.dat` mode-0600, so any fixed non-root identity
(the previous uid-1000 convention, and the ACL setup that went with it) could
neither walk the tree nor read the files — every archive failed
`open …/level.dat: permission denied` and the same defect failed on-demand
backups/restores. Root is the same identity the game container itself runs as
(see the operator's forwarding-init note); `DAC_OVERRIDE` extends the archive
to game images with a different UID. If a world is still **preserved** while a
reap was expected, it is now a different cause: check the run's ERROR logs for
the resolver's `lstat` messages before suspecting permissions.
- **Node placement.** Multi-node clusters: the world's directory exists only on
the node holding its volume, and the CronJob sets no `nodeSelector`, so add
one (single-node starters are pinned implicitly).
--- ---
## 11. Idle auto-stop never fires; player count always shows 0 ## 11. Idle auto-stop never fires; player count always shows 0
@@ -502,6 +663,32 @@ Both fields set and still nothing happens? Then the probe is failing rather than
disabled: the server would be stuck in `Starting` with `RconNotReachable` disabled: the server would be stuck in `Starting` with `RconNotReachable`
(`reconciler.go:156`), which is §1's symptom, not this one. (`reconciler.go:156`), which is §1's symptom, not this one.
This path used to fail even with everything configured correctly, through three
stacked defects proven and fixed on a live cluster (auditfix21/22): the
`emptySince` stamp was pruned by a missing CRD status field, a quiescent empty
server produced no watch events to re-check the timer, and the Role lacked the
`minecraftservers:patch` grant the stop write needs. If auto-stop ever looks
dead again, check these three in order (each is now pinned by a test):
```sh
# ① The stamp must persist — should print a timestamp, not an empty string,
# a few seconds after a server goes Ready with zero players.
kubectl get minecraftserver <name> -o jsonpath='{.status.emptySince}'
# ② The operator must be able to write spec.desiredState (403 in the operator
# log = missing patch grant on Role felis-operator).
kubectl auth can-i patch minecraftservers -n <ns> --as=system:serviceaccount:<ctl-ns>:felis-operator
# ③ A wake-up must be scheduled: while empty, expect whatever you set
# as emptySecondsBeforeStop to elapse and the box to flip to Stopped without
# any external action.
```
While players are online the operator re-probes on a 30s cadence so it notices
the moment the last one leaves; while empty it schedules a wake-up exactly at
the deadline. Quiet operator logs on an idle server are normal — the action is
the scheduled wake-up, not a stream of reconciles.
Note the reaper's `last_active_at` (§10) is a *different* subsystem (Postgres Note the reaper's `last_active_at` (§10) is a *different* subsystem (Postgres
business layer, bumped by join events) — it keeps worlds alive against the business layer, bumped by join events) — it keeps worlds alive against the
reaper, but it does **not** auto-stop empty running servers. reaper, but it does **not** auto-stop empty running servers.
@@ -526,6 +713,30 @@ deadline for the whole start, applied on the RCON-probe branch.
--- ---
## 12b. A newer field never reaches an already-installed system server (`felis converge`)
Provisioning is create-if-absent: `felis setup` never rewrites an existing
`login`/`lobby` `MinecraftServer` beyond the config-derived env it owns, so a
field the desired spec gained after your install sits absent forever — this is
how a deployment ends up with a lobby that has no `spec.rcon` (a dead console
and an online-player count that is always 0) and a login gate without
`spec.startup.healthHTTPPort`. `felis converge` is the explicit pass that fills
exactly those zero-valued fields (and re-adds a derived env key that is
missing). It never overwrites a value that already holds one — an operator's
RCON secretRef or tuning survives.
```
sudo felis converge
```
Run it **after the images are in place**. Enabling RCON, or the HTTP readiness
gate, on a server whose image predates the listener would hold that server in
`Starting` until the operator marks it `Failed` — that ordering is the reason
this is a command you run rather than something setup does on every re-run.
System servers that are already current report `already converged`.
---
## 13. World PVC survives after I deleted the MinecraftServer ## 13. World PVC survives after I deleted the MinecraftServer
This is expected. The world PVC is a StatefulSet `VolumeClaimTemplate`. There is This is expected. The world PVC is a StatefulSet `VolumeClaimTemplate`. There is
@@ -557,6 +768,69 @@ changes, because retention was never conditional in the first place.
--- ---
## 13b. Node runs out of disk: what survives, and how to recover
A full disk is the most destructive failure this stack sees: kubelet evicts game
pods (the control plane is protected below), and its image GC then collects
images nothing is running. The images have a pull source now — the in-cluster
registry — so they come back without an operator re-import; freeing space is
what completes the recovery.
**Eviction.** Every control-plane pod (api, operator, reaper, registry) runs
under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
node-pressure eviction refuses to touch those pods — the log shows
*"Eviction manager: cannot evict a critical pod"* for each of them — while
game-server pods at the default priority 0 are evicted first. A drill that filled
the disk to 1.7G free saw exactly this: login/lobby evicted, the whole control
plane still Running (before the fix the same drill evicted the api, operator and
registry too, and the image-GC stage below followed). User-defined
PriorityClasses cannot substitute: the API caps them at 1e9, below kubelet's
critical threshold. The built-in class allows preemption (its policy is fixed),
so a control-plane pod that cannot fit may preempt a game pod — deliberate: the
management plane must be placeable.
**The pressure condition clears slowly.** After you free space, the node can stay
`DiskPressure:True` for up to ~5 minutes
(`--eviction-pressure-transition-period` defaults to 5m, to stop the condition
flapping); pods that need scheduling wait for it. This is the bulk of the
"recovery takes minutes" observation, not a stuck node.
**The images may be gone — they come back on their own.** If pods were evicted,
the kubelet can garbage-collect their images (unused > 2 minutes under imagefs
pressure). Every image this platform runs is ALSO hosted in the in-cluster
registry: the installer builds each one as `registry.<ns>.svc:5000/felis/…`
and mirrors it there, and the node's containerd is configured (a registries.yaml
mirror onto the registry's loopback hostPort) to relay those refs back through
it. So a GC'd image is re-pulled on the next attempt with no operator action —
delete the stuck pod to force an immediate retry (or wait out the backoff), and
the workload converges.
If a pull does NOT come back:
1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are
`k3s ctr images ls -q` and the world/backup PVCs under
`/var/lib/rancher/k3s/storage`).
2. Check the registry: `kubectl -n felis get pods -l
app.kubernetes.io/component=registry` and, on the node,
`curl -s http://127.0.0.1:5000/v2/` (expect `{}`).
3. Check the mirror file: `/etc/rancher/k3s/registries.yaml` must map
`registry.felis.svc:5000` to `http://127.0.0.1:5000`. Missing or changed:
re-run the installer (it rewrites the file and restarts k3s only when the
content changed).
4. Re-mirror a tag the registry does not have (hand-built images were never
pushed): `sudo systemctl start docker` (the installer leaves the daemon
stopped), then `docker tag <ref> 127.0.0.1:5000/<repo>:<tag> && docker push
127.0.0.1:5000/<repo>:<tag>`.
For an image that is in neither place, the old fallback still stands: re-run the
installer (it rebuilds/re-imports from the local Docker store AND mirrors into
the registry), or for a single image
`docker save felis:<tag> | k3s ctr images import -`. The Docker store remains a
deliberate second copy on the node; treat it as the recovery path, not as free
space.
---
## 14. Metrics for diagnosis (spec §23) ## 14. Metrics for diagnosis (spec §23)
All four mandated metrics have real producers; scrape them when triaging: All four mandated metrics have real producers; scrape them when triaging:
@@ -571,8 +845,65 @@ All four mandated metrics have real producers; scrape them when triaging:
actually deleted post-backup (§10); a spike here means worlds crossed the 15d actually deleted post-backup (§10); a spike here means worlds crossed the 15d
idle line — cross-check that join events are flowing (§10 risk vectors). idle line — cross-check that join events are flowing (§10 risk vectors).
### Scraping
The series come from two processes:
- `felis-operator` pod `:8080/metrics` — `felis_servers_total`,
`felis_start_duration_seconds` (no Service; scrape pod-scoped, e.g. a
PodMonitor targeting port `metrics`).
- `felis-api` internal face `:8081/metrics` (Service `felis-api-internal`) —
`felis_image_build_failures_total`. Unauthenticated like the probes;
ClusterIP-only, and the external face never serves it.
- `felis_reaper_worlds_deleted_total` is produced inside the one-shot reaper
CronJob, which exits long before any scrape interval — without a pushgateway
it has no scrape path. Read the reaper Pod log or the `world_backups` table
for deletions instead.
### Alert rules
`deploy/alerts/` ships ready-made rules: build failures, slow starts, node
disk/memory thresholds, and the kubelet `DiskPressure` condition.
- Plain Prometheus: add `felis-alerts.yaml` to `rule_files`. Check and unit-test
it standalone with `promtool check rules felis-alerts.yaml` and
`promtool test rules felis-alerts_test.yml` (the tests pin exactly when each
alert fires).
- kube-prometheus-stack / prometheus-operator: `kubectl apply -f
felis-prometheusrule.yaml` (adjust its `release:` label to your stack's
ruleSelector).
--- ---
## 15. Control-plane upgrades, and rolling back a bad one
There is no in-place updater: an upgrade is re-running the installer
(`curl -fsSL <installer URL> | sudo bash`), which rebuilds/re-imports the image
and re-applies the bundle. (`sudo felis setup` is not this path; on a completed
install it only opens the config console.) The channel is not persisted across
the re-run, so pass `FELIS_VERSION_BOOTSTRAP=dev` on a host that tracks main.
Two properties of the control plane matter when you do:
- Both Deployments use strategy **Recreate** (single replica, no leader election:
two overlapping instances would fight over the same cluster). An upgrade takes
the panel/API down for the rollout window — seconds normally, longer if the new
image still has to be imported.
- If the new pod cannot start (bad tag, missing image), the installer's rollout
wait fails after 180s and prints `kubectl describe` diagnostics: you see
`ErrImagePull`/`ImagePullBackOff` there instead of a silent hang.
Roll back with:
```
kubectl -n felis rollout undo deploy/felis-api
kubectl -n felis rollout status deploy/felis-api
```
(the same for `felis-operator` and `registry`). `rollout undo` returns to the
previous ReplicaSet, whose image is normally still on the node; if the image GC
collected it, the registry re-serves it automatically (§13b) for every tag the
installer built — only hand-built tags need a manual re-mirror.
## Quick reference: symptom → section ## Quick reference: symptom → section
| Symptom | Section | | Symptom | Section |
@@ -588,10 +919,12 @@ All four mandated metrics have real producers; scrape them when triaging:
| Local password login rejected | §5c | | Local password login rejected | §5c |
| Internal callers 401 (service token) | §6 | | Internal callers 401 (service token) | §6 |
| Link/claim 400/409/412/403/404 | §7 | | Link/claim 400/409/412/403/404 | §7 |
| Build push 400 / SA denied / egress hang / Failed | §8 | | Build push 400 / SA denied / egress hang / Failed / executor ImagePullBackOff | §8, §8e |
| Registry push/pull unreachable | §9 | | Registry push/pull unreachable | §9 |
| World deleted unexpectedly / backup skipped | §10 | | World deleted unexpectedly / backup skipped | §10 |
| Idle auto-stop not firing; player count 0 | §11 | | Idle auto-stop not firing; player count 0 | §11 |
| A config field seems ignored | §12 | | A config field seems ignored | §12 |
| PVC left behind after delete | §13 | | PVC left behind after delete | §13 |
| Node out of disk; pods evicted / ImagePullBackOff | §13b |
| Which metric to scrape | §14 | | Which metric to scrape | §14 |
| Upgrade / roll back a bad control-plane image | §15 |
+5 -5
View File
@@ -9,10 +9,11 @@ require (
github.com/charmbracelet/huh v1.0.0 github.com/charmbracelet/huh v1.0.0
github.com/charmbracelet/lipgloss v1.1.0 github.com/charmbracelet/lipgloss v1.1.0
github.com/descope/virtualwebauthn v1.0.5 github.com/descope/virtualwebauthn v1.0.5
github.com/go-logr/logr v1.4.2
github.com/go-webauthn/webauthn v0.17.4 github.com/go-webauthn/webauthn v0.17.4
github.com/golang-jwt/jwt/v5 v5.3.1 github.com/golang-jwt/jwt/v5 v5.3.1
github.com/google/uuid v1.6.0 github.com/google/uuid v1.6.0
github.com/jackc/pgx/v5 v5.7.1 github.com/jackc/pgx/v5 v5.9.2
github.com/minio/minio-go/v7 v7.2.1 github.com/minio/minio-go/v7 v7.2.1
github.com/prometheus/client_golang v1.19.1 github.com/prometheus/client_golang v1.19.1
github.com/prometheus/client_model v0.6.1 github.com/prometheus/client_model v0.6.1
@@ -42,7 +43,6 @@ require (
github.com/erikgeiser/coninput v0.0.0-20211004153227-1c3628e74d0f // indirect github.com/erikgeiser/coninput v0.0.0-20211004153227-1c3628e74d0f // indirect
github.com/evanphx/json-patch/v5 v5.9.0 // indirect github.com/evanphx/json-patch/v5 v5.9.0 // indirect
github.com/fxamacker/cbor/v2 v2.9.2 // indirect github.com/fxamacker/cbor/v2 v2.9.2 // indirect
github.com/go-logr/logr v1.4.2 // indirect
github.com/go-openapi/jsonpointer v0.19.6 // indirect github.com/go-openapi/jsonpointer v0.19.6 // indirect
github.com/go-openapi/jsonreference v0.20.2 // indirect github.com/go-openapi/jsonreference v0.20.2 // indirect
github.com/go-openapi/swag v0.22.4 // indirect github.com/go-openapi/swag v0.22.4 // indirect
@@ -92,12 +92,12 @@ require (
go.yaml.in/yaml/v3 v3.0.4 // indirect go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/crypto v0.52.0 // indirect golang.org/x/crypto v0.52.0 // indirect
golang.org/x/exp v0.0.0-20231006140011-7918f672742d // indirect golang.org/x/exp v0.0.0-20231006140011-7918f672742d // indirect
golang.org/x/net v0.54.0 // indirect golang.org/x/net v0.55.0 // indirect
golang.org/x/oauth2 v0.21.0 // indirect golang.org/x/oauth2 v0.21.0 // indirect
golang.org/x/sync v0.20.0 // indirect golang.org/x/sync v0.21.0 // indirect
golang.org/x/sys v0.45.0 // indirect golang.org/x/sys v0.45.0 // indirect
golang.org/x/term v0.43.0 // indirect golang.org/x/term v0.43.0 // indirect
golang.org/x/text v0.37.0 // indirect golang.org/x/text v0.39.0 // indirect
golang.org/x/time v0.3.0 // indirect golang.org/x/time v0.3.0 // indirect
gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect
google.golang.org/protobuf v1.36.10 // indirect google.golang.org/protobuf v1.36.10 // indirect
+10 -10
View File
@@ -116,8 +116,8 @@ github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsI
github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg= github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo= github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM= github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM=
github.com/jackc/pgx/v5 v5.7.1 h1:x7SYsPBYDkHDksogeSmZZ5xzThcTgRz++I5E+ePFUcs= github.com/jackc/pgx/v5 v5.9.2 h1:3ZhOzMWnR4yJ+RW1XImIPsD1aNSz4T4fyP7zlQb56hw=
github.com/jackc/pgx/v5 v5.7.1/go.mod h1:e7O26IywZZ+naJtWWos6i6fvWK+29etgITqrqHLfoZA= github.com/jackc/pgx/v5 v5.9.2/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo= github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4= github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/josharian/intern v1.0.0 h1:vlS4z54oSdjm0bgjRigI+G1HpF+tI+9rE5LLzOg8HmY= github.com/josharian/intern v1.0.0 h1:vlS4z54oSdjm0bgjRigI+G1HpF+tI+9rE5LLzOg8HmY=
@@ -245,15 +245,15 @@ golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn
golang.org/x/net v0.0.0-20190620200207-3b0461eec859/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s= golang.org/x/net v0.0.0-20190620200207-3b0461eec859/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s=
golang.org/x/net v0.0.0-20200226121028-0de0cce0169b/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s= golang.org/x/net v0.0.0-20200226121028-0de0cce0169b/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s=
golang.org/x/net v0.0.0-20201021035429-f5854403a974/go.mod h1:sp8m0HH+o8qH0wwXwYZr8TS3Oi6o0r6Gce1SSxlDquU= golang.org/x/net v0.0.0-20201021035429-f5854403a974/go.mod h1:sp8m0HH+o8qH0wwXwYZr8TS3Oi6o0r6Gce1SSxlDquU=
golang.org/x/net v0.54.0 h1:2zJIZAxAHV/OHCDTCOHAYehQzLfSXuf/5SoL/Dv6w/w= golang.org/x/net v0.55.0 h1:bcvxaJn3e1U6InsFWt1JUq1aSjnRxLzT2rtD2KfkDF8=
golang.org/x/net v0.54.0/go.mod h1:Sj4oj8jK6XmHpBZU/zWHw3BV3abl4Kvi+Ut7cQcY+cQ= golang.org/x/net v0.55.0/go.mod h1:L5U2KuzuOe1lY7Z+aWVIKK6qEeJXnXV9yzGA+WCHJww=
golang.org/x/oauth2 v0.21.0 h1:tsimM75w1tF/uws5rbeHzIWxEqElMehnc+iW793zsZs= golang.org/x/oauth2 v0.21.0 h1:tsimM75w1tF/uws5rbeHzIWxEqElMehnc+iW793zsZs=
golang.org/x/oauth2 v0.21.0/go.mod h1:XYTD2NtWslqkgxebSiOHnXEap4TF09sJSc7H1sXbhtI= golang.org/x/oauth2 v0.21.0/go.mod h1:XYTD2NtWslqkgxebSiOHnXEap4TF09sJSc7H1sXbhtI=
golang.org/x/sync v0.0.0-20190423024810-112230192c58/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= golang.org/x/sync v0.0.0-20190423024810-112230192c58/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20190911185100-cd5d95a43a6e/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= golang.org/x/sync v0.0.0-20190911185100-cd5d95a43a6e/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20201020160332-67f06af15bc9/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= golang.org/x/sync v0.0.0-20201020160332-67f06af15bc9/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.20.0 h1:e0PTpb7pjO8GAtTs2dQ6jYa5BWYlMuX047Dco/pItO4= golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM=
golang.org/x/sync v0.20.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0= golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY= golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20200930185726-fdedc70b468f/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= golang.org/x/sys v0.0.0-20200930185726-fdedc70b468f/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
@@ -265,16 +265,16 @@ golang.org/x/term v0.43.0 h1:S4RLU2sB31O/NCl+zFN9Aru9A/Cq2aqKpTZJ6B+DwT4=
golang.org/x/term v0.43.0/go.mod h1:lrhlHNdQJHO+1qVYiHfFKVuVioJIheAc3fBSMFYEIsk= golang.org/x/term v0.43.0/go.mod h1:lrhlHNdQJHO+1qVYiHfFKVuVioJIheAc3fBSMFYEIsk=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ= golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.3/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ= golang.org/x/text v0.3.3/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.37.0 h1:Cqjiwd9eSg8e0QAkyCaQTNHFIIzWtidPahFWR83rTrc= golang.org/x/text v0.39.0 h1:UbZz4pLOvn600D6Oh6GGEI6VAmndrEBLv8/6BEXzyus=
golang.org/x/text v0.37.0/go.mod h1:a5sjxXGs9hsn/AJVwuElvCAo9v8QYLzvavO5z2PiM38= golang.org/x/text v0.39.0/go.mod h1:3UwRclnC2g0TU9x8PZiyfOajCd1zaUNHF9cvqcQZ+ZM=
golang.org/x/time v0.3.0 h1:rg5rLMjNzMS1RkNLzCG38eapWhnYLFYXDXj2gOlr8j4= golang.org/x/time v0.3.0 h1:rg5rLMjNzMS1RkNLzCG38eapWhnYLFYXDXj2gOlr8j4=
golang.org/x/time v0.3.0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ= golang.org/x/time v0.3.0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ= golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
golang.org/x/tools v0.0.0-20191119224855-298f0cb1881e/go.mod h1:b+2E5dAYhXwXZwtnZ6UAqBI28+e2cm9otk0dWdXHAEo= golang.org/x/tools v0.0.0-20191119224855-298f0cb1881e/go.mod h1:b+2E5dAYhXwXZwtnZ6UAqBI28+e2cm9otk0dWdXHAEo=
golang.org/x/tools v0.0.0-20200619180055-7c47624df98f/go.mod h1:EkVYQZoAsY45+roYkvgYkIh4xh/qjgUK9TdY2XT94GE= golang.org/x/tools v0.0.0-20200619180055-7c47624df98f/go.mod h1:EkVYQZoAsY45+roYkvgYkIh4xh/qjgUK9TdY2XT94GE=
golang.org/x/tools v0.0.0-20210106214847-113979e3529a/go.mod h1:emZCQorbCU4vsT4fOWvOPXz4eW1wZW4PmDk9uLelYpA= golang.org/x/tools v0.0.0-20210106214847-113979e3529a/go.mod h1:emZCQorbCU4vsT4fOWvOPXz4eW1wZW4PmDk9uLelYpA=
golang.org/x/tools v0.44.0 h1:UP4ajHPIcuMjT1GqzDWRlalUEoY+uzoZKnhOjbIPD2c= golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
golang.org/x/tools v0.44.0/go.mod h1:KA0AfVErSdxRZIsOVipbv3rQhVXTnlU6UhKxHd1seDI= golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0= golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191011141410-1b5146add898/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0= golang.org/x/xerrors v0.0.0-20191011141410-1b5146add898/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0= golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
+64 -14
View File
@@ -63,6 +63,11 @@ type API struct {
// authorization boundary is exercised before the backup-Job executor is wired. // authorization boundary is exercised before the backup-Job executor is wired.
Backuper Backuper Backuper Backuper
// JobStatus reads the latest backup/restore Job outcomes for GET
// /servers/{name}/jobs — the status outlet for async failures (the enqueue
// endpoints only answer 202). Optional: nil → that route reports 503.
JobStatus JobStatusReader
// Files is the server file editor (list / read / write a file in a stopped // Files is the server file editor (list / read / write a file in a stopped
// server's world volume — the "one wrong line in server.properties" repair). // server's world volume — the "one wrong line in server.properties" repair).
// Like Restorer and Backuper it is optional: when nil the file routes report // Like Restorer and Backuper it is optional: when nil the file routes report
@@ -118,6 +123,15 @@ type API struct {
// on the wake lever). Zero disables throttling. // on the wake lever). Zero disables throttling.
WakeCooldown time.Duration WakeCooldown time.Duration
// SubmitCreateCooldown / SubmitUploadCooldown throttle the user-modpack
// submission lane per user: create bounds how quickly review-queue rows can
// appear, upload bounds how often a user may stream a (up to 1 GiB) build
// context. The keys are separate, so the lane's normal shape — create, then
// upload — is never blocked by its own throttle. Zero disables each lever
// (the same idiom as WakeCooldown); cmd/felis wires positive values.
SubmitCreateCooldown time.Duration
SubmitUploadCooldown time.Duration
// MaxRunningServers caps how many servers may be desired-Running cluster-wide // MaxRunningServers caps how many servers may be desired-Running cluster-wide
// (spec §9.1: the concurrency-上限 lever hanging on the same wake chokepoint as // (spec §9.1: the concurrency-上限 lever hanging on the same wake chokepoint as
// cooldown and autostartPolicy). Zero — the default — disables it: §9.2 wires // cooldown and autostartPolicy). Zero — the default — disables it: §9.2 wires
@@ -139,9 +153,9 @@ type API struct {
// AuthSources is the Felis-nano multi-source hasJoined multiplexer's upstream // AuthSources is the Felis-nano multi-source hasJoined multiplexer's upstream
// Yggdrasil list, in priority order (config order; the Mojang Identity source // Yggdrasil list, in priority order (config order; the Mojang Identity source
// first for 正版优先). Nil — the default — makes the session verifier reject every // first for 正版优先). Nil makes the session verifier reject every login (204);
// login (204), so the endpoint ships inert until cmd/felis wires configured // cmd/felis always wires at least the Mojang source through authSourcesFromConfig.
// sources. Consumed by handleHasJoined (handlers_hasjoined.go). // Consumed by handleHasJoined (handlers_hasjoined.go).
AuthSources []AuthSource AuthSources []AuthSource
// Now is the clock, injectable for tests. Defaults to time.Now. // Now is the clock, injectable for tests. Defaults to time.Now.
@@ -153,6 +167,9 @@ type API struct {
otpCooldownOnce sync.Once otpCooldownOnce sync.Once
otpCooldown *cooldownLimiter otpCooldown *cooldownLimiter
submitCooldownOnce sync.Once
submitCooldown *cooldownLimiter
streamCapOnce sync.Once streamCapOnce sync.Once
streamCap *streamLimiter streamCap *streamLimiter
} }
@@ -199,6 +216,20 @@ func (a *API) otpLimiter() *cooldownLimiter {
return a.otpCooldown return a.otpCooldown
} }
// submitLimiter lazily builds a SEPARATE cooldown limiter for the user-modpack
// submission lane, so its throttles never share state with the wake or OTP
// keyspaces. One limiter backs both levers with prefixed keys (see the
// submissionCreateKey/UploadKey constants), so create and upload never contend
// with each other. Like the other cooldowns it is process-local; with multiple
// api replicas the effective spacing is per-replica, the same accepted
// KNOWN-LIMITATION the OTP resend throttle carries.
func (a *API) submitLimiter() *cooldownLimiter {
a.submitCooldownOnce.Do(func() {
a.submitCooldown = &cooldownLimiter{now: a.now, last: map[string]time.Time{}}
})
return a.submitCooldown
}
// streamGate lazily builds the per-principal SSE stream cap bound to // streamGate lazily builds the per-principal SSE stream cap bound to
// MaxStreamsPerPrincipal. A zero cap yields a disabled limiter that admits every // MaxStreamsPerPrincipal. A zero cap yields a disabled limiter that admits every
// stream, so a deployment (or test) that leaves it unset pays nothing. // stream, so a deployment (or test) that leaves it unset pays nothing.
@@ -259,13 +290,20 @@ type apiRoute struct {
} }
// internalAPIRoutes is the internal face's served route table (spec §7, §14): // internalAPIRoutes is the internal face's served route table (spec §7, §14):
// service-token auth, never Zero Trust. It carries both health probes. // service-token auth, never Zero Trust. It carries both health probes and the
// metrics scrape.
func (a *API) internalAPIRoutes() []apiRoute { func (a *API) internalAPIRoutes() []apiRoute {
return []apiRoute{ return []apiRoute{
{Method: "GET", Pattern: "/healthz", Public: true, h: a.handleHealthz}, {Method: "GET", Pattern: "/healthz", Public: true, h: a.handleHealthz},
{Method: "GET", Pattern: "/readyz", Public: true, h: a.handleReadyz}, {Method: "GET", Pattern: "/readyz", Public: true, h: a.handleReadyz},
// Prometheus scrape (felis_* collectors); public because a scrape carries
// no token, internal-only so it is never exposed off-cluster.
{Method: "GET", Pattern: "/metrics", Public: true, h: a.handleMetrics},
{Method: "GET", Pattern: "/api/v1/servers", h: a.handleListServers}, {Method: "GET", Pattern: "/api/v1/servers", h: a.handleListServers},
// The build Pod's context-fetch initContainer streams a submission's stored
// modpack through this route (build namespace cannot mount the uploads PVC).
{Method: "GET", Pattern: "/api/v1/internal/submissions/{id}/context", h: a.handleInternalSubmissionContext},
{Method: "POST", Pattern: "/api/v1/internal/servers/{name}/ready", h: a.handleReady}, {Method: "POST", Pattern: "/api/v1/internal/servers/{name}/ready", h: a.handleReady},
{Method: "POST", Pattern: "/api/v1/internal/servers/{name}/join-event", h: a.handleJoinEvent}, {Method: "POST", Pattern: "/api/v1/internal/servers/{name}/join-event", h: a.handleJoinEvent},
// Domain-autostart (spec §9.1, §14): velocity drives the wake lever and polls // Domain-autostart (spec §9.1, §14): velocity drives the wake lever and polls
@@ -303,12 +341,12 @@ func (a *API) internalAPIRoutes() []apiRoute {
// Mojang player (same name, different UUID) always passes. // Mojang player (same name, different UUID) always passes.
{Method: "POST", Pattern: "/api/v1/internal/player/reclaim", h: a.handleReclaimUsername}, {Method: "POST", Pattern: "/api/v1/internal/player/reclaim", h: a.handleReclaimUsername},
{Method: "GET", Pattern: "/api/v1/internal/player/blacklist/{mc_uuid}", h: a.handleCheckBlacklist}, {Method: "GET", Pattern: "/api/v1/internal/player/blacklist/{mc_uuid}", h: a.handleCheckBlacklist},
// Felis-nano multi-source session verifier (spec §B3 player game-login). // Felis-nano multi-source session verifier, behind player game-login. Velocity is
// Velocity's authlib is pointed here (-Dmojang.sessionserver or a thin login // pointed here with -Dmojang.sessionserver and issues the request itself; it speaks
// hook); it speaks the vanilla sessionserver protocol and carries no token, so // the vanilla sessionserver protocol and carries no token, so this is Public. It
// this is Public. It fans hasJoined out to the configured Yggdrasil roots // fans hasJoined out to the configured Yggdrasil roots (Mojang-first) and rewrites
// (Mojang-first) and rewrites third-party UUIDs into a per-source namespace // third-party UUIDs into a per-source namespace before returning the canonical
// before returning the canonical profile (handlers_hasjoined.go). // profile (handlers_hasjoined.go).
{Method: "GET", Pattern: "/session/minecraft/hasJoined", Public: true, h: a.handleHasJoined}, {Method: "GET", Pattern: "/session/minecraft/hasJoined", Public: true, h: a.handleHasJoined},
// Op-login (passwordless op.console login): a staff member starts the login // Op-login (passwordless op.console login): a staff member starts the login
// on the web, and an ONLINE in-game admin vouches for it via velocity's // on the web, and an ONLINE in-game admin vouches for it via velocity's
@@ -407,6 +445,7 @@ func (a *API) externalAPIRoutes() []apiRoute {
// owned), and restore is gated by owner-or-admin PLUS a former-owner match, so // owned), and restore is gated by owner-or-admin PLUS a former-owner match, so
// neither sits behind adminOnly. // neither sits behind adminOnly.
{Method: "GET", Pattern: "/api/v1/backups", h: a.handleListBackups}, {Method: "GET", Pattern: "/api/v1/backups", h: a.handleListBackups},
{Method: "GET", Pattern: "/api/v1/servers/{name}/jobs", h: a.handleServerJobs},
{Method: "POST", Pattern: "/api/v1/servers/{name}/restore-backup", h: a.handleRestoreBackup}, {Method: "POST", Pattern: "/api/v1/servers/{name}/restore-backup", h: a.handleRestoreBackup},
{Method: "POST", Pattern: "/api/v1/servers/{name}/backup", h: a.handleBackupNow}, {Method: "POST", Pattern: "/api/v1/servers/{name}/backup", h: a.handleBackupNow},
// Server file editor: list / read / write a file in a STOPPED server's world // Server file editor: list / read / write a file in a STOPPED server's world
@@ -479,6 +518,11 @@ func (a *API) externalAPIRoutes() []apiRoute {
// App-tier and owner-scoped (the id must belong to the principal), exactly // App-tier and owner-scoped (the id must belong to the principal), exactly
// like the create/list routes above. // like the create/list routes above.
{Method: "POST", Pattern: "/api/v1/me/submissions/{id}/context", h: a.handleUploadSubmissionContext}, {Method: "POST", Pattern: "/api/v1/me/submissions/{id}/context", h: a.handleUploadSubmissionContext},
// Withdraw the caller's OWN pending submission: the row and its uploaded
// context are deleted, freeing the pending slot and storage budget. Same
// owner-scoping as the upload route — a reviewed submission is frozen (409)
// and another user's id is invisible (404).
{Method: "DELETE", Pattern: "/api/v1/me/submissions/{id}", h: a.handleWithdrawSubmission},
// Admin (Zero-Trust) tier: create / mutate spec / image admission. These gate // Admin (Zero-Trust) tier: create / mutate spec / image admission. These gate
// on Principal.IsAdmin() inside the handler via the adminOnly wrapper, so the // on Principal.IsAdmin() inside the handler via the adminOnly wrapper, so the
// boundary is exercised even where the body is a later-phase stub. // boundary is exercised even where the body is a later-phase stub.
@@ -508,6 +552,13 @@ func (a *API) externalAPIRoutes() []apiRoute {
{Method: "GET", Pattern: "/api/v1/submissions", Admin: true, h: a.handleListSubmissions}, {Method: "GET", Pattern: "/api/v1/submissions", Admin: true, h: a.handleListSubmissions},
{Method: "POST", Pattern: "/api/v1/submissions/{id}/approve", Admin: true, h: a.handleApproveSubmission}, {Method: "POST", Pattern: "/api/v1/submissions/{id}/approve", Admin: true, h: a.handleApproveSubmission},
{Method: "POST", Pattern: "/api/v1/submissions/{id}/reject", Admin: true, h: a.handleRejectSubmission}, {Method: "POST", Pattern: "/api/v1/submissions/{id}/reject", Admin: true, h: a.handleRejectSubmission},
// Retire a submission outright (row + uploaded context), any status. The
// lane's lifecycle valve: without it, rejected/consumed uploads accumulated
// on the uploads PVC forever — there is no other delete path.
{Method: "DELETE", Pattern: "/api/v1/submissions/{id}", Admin: true, h: a.handleDeleteSubmission},
// The reviewer's read path to the uploaded blob: the executed Dockerfile
// lives inside it, so approval would otherwise be blind.
{Method: "GET", Pattern: "/api/v1/submissions/{id}/context", Admin: true, h: a.handleAdminSubmissionContext},
// Auto-update maintenance window (spec §B; decision core internal/updates). // Auto-update maintenance window (spec §B; decision core internal/updates).
// Admin-tier: it governs whether Felis may apply an update to itself, so setting // Admin-tier: it governs whether Felis may apply an update to itself, so setting
// it requires the admin Zero-Trust path, not a mere session. API+persistence // it requires the admin Zero-Trust path, not a mere session. API+persistence
@@ -674,10 +725,9 @@ func principalFromContext(ctx context.Context) *Principal {
// #35); cross-replica bounding would need a shared store (out of scope for the // #35); cross-replica bounding would need a shared store (out of scope for the
// single-replica demo). // single-replica demo).
type cooldownLimiter struct { type cooldownLimiter struct {
mu sync.Mutex mu sync.Mutex
now func() time.Time now func() time.Time
last map[string]time.Time last map[string]time.Time
window time.Duration
} }
// allowed reports whether name may wake now WITHOUT recording the attempt. A // allowed reports whether name may wake now WITHOUT recording the attempt. A
+250 -16
View File
@@ -38,8 +38,16 @@ type fakeRepo struct {
owners map[string]string owners map[string]string
ownersErr error ownersErr error
claimOK map[string]bool // name -> claim succeeds; absent name -> ErrNotFound claimOK map[string]bool // name -> claim succeeds; absent name -> ErrNotFound
audits []AuditEntry // claimQuotaRefuse simulates ClaimServer's atomic quota gate (audit #4)
joins []string // refusing a name whose advisory pre-check already passed.
claimQuotaRefuse map[string]bool
// serverResources / resourceUpdates mirror the cached resource columns:
// ServerResources is what the resize path reads (to preserve storage), and
// UpdateServerResources records the write for assertions.
serverResources map[string]ResourceSpec
resourceUpdates map[string]ResourceSpec
audits []AuditEntry
joins []string
// create-server seeding (spec §15) // create-server seeding (spec §15)
seeded map[string]bool // name -> servers row exists seeded map[string]bool // name -> servers row exists
aliases map[string]string // subdomain -> bound server name aliases map[string]string // subdomain -> bound server name
@@ -59,6 +67,10 @@ type fakeRepo struct {
staff map[string]*StaffUser // username -> staff login row staff map[string]*StaffUser // username -> staff login row
sessions map[string]*fakeSession // token_hash -> session sessions map[string]*fakeSession // token_hash -> session
settings map[string][]byte // key -> jsonb value settings map[string][]byte // key -> jsonb value
// failSessionUser / failGetSetting force those reads to fail with a generic
// (non-ErrNotFound) error, simulating a store outage for the 503 auth path.
failSessionUser error
failGetSetting error
// player email OTPs (spec §B2). Keyed by row id; the verify path scans for the // player email OTPs (spec §B2). Keyed by row id; the verify path scans for the
// newest live (user, purpose) just as the PG query does. // newest live (user, purpose) just as the PG query does.
otps map[string]*fakeEmailOTP otps map[string]*fakeEmailOTP
@@ -91,6 +103,10 @@ type fakeRepo struct {
// user admin fakes // user admin fakes
seededUsers []seededUser seededUsers []seededUser
fakeQuotas map[string]*QuotaView fakeQuotas map[string]*QuotaView
// deletedIDs remembers soft-deleted user ids: DeleteUser drops the row from
// seededUsers (so listings hide it, mirroring the WHERE deleted_at IS NULL
// query), and this set keeps the account dead for the liveness guards.
deletedIDs map[string]bool
// pingErr, when non-nil, is returned by Ping to simulate DB liveness check // pingErr, when non-nil, is returned by Ping to simulate DB liveness check
// failures in /readyz tests. // failures in /readyz tests.
pingErr error pingErr error
@@ -202,8 +218,9 @@ func newFakeRepo() *fakeRepo {
allowlist: map[string]map[string]bool{}, allowUUID: map[string]map[string]bool{}, allowlist: map[string]map[string]bool{}, allowUUID: map[string]map[string]bool{},
mine: map[string][]MyServerView{}, mine: map[string][]MyServerView{},
owners: map[string]string{}, owners: map[string]string{},
claimOK: map[string]bool{}, claimOK: map[string]bool{}, claimQuotaRefuse: map[string]bool{},
seeded: map[string]bool{}, aliases: map[string]string{}, serverResources: map[string]ResourceSpec{}, resourceUpdates: map[string]ResourceSpec{},
seeded: map[string]bool{}, aliases: map[string]string{},
linkCodes: map[string]fakeLinkCode{}, links: map[string]string{}, linkCodes: map[string]fakeLinkCode{}, links: map[string]string{},
linkAuthSource: map[string]string{}, linkAuthSource: map[string]string{},
staff: map[string]*StaffUser{}, staff: map[string]*StaffUser{},
@@ -220,6 +237,7 @@ func newFakeRepo() *fakeRepo {
discoverableChallenges: map[string]*fakeDiscoverableChallenge{}, discoverableChallenges: map[string]*fakeDiscoverableChallenge{},
fakeQuotas: map[string]*QuotaView{}, fakeQuotas: map[string]*QuotaView{},
migrations: map[string]*fakeMigration{}, migrations: map[string]*fakeMigration{},
deletedIDs: map[string]bool{},
} }
} }
@@ -242,13 +260,16 @@ func (f *fakeRepo) QuotaCheck(_ context.Context, userID string, _ string, _ Reso
// For hermetic tests, QuotaCheck delegates to the same QuotaAvailable // For hermetic tests, QuotaCheck delegates to the same QuotaAvailable
// store — tests that care about per-dimension checks should use // store — tests that care about per-dimension checks should use
// fakeQuotas with direct inspection. // fakeQuotas with direct inspection.
return f.QuotaAvailable(nil, userID) return f.QuotaAvailable(context.TODO(), userID)
} }
func (f *fakeRepo) UpdateServerResources(_ context.Context, _ string, _, _, _ int) error { return nil } func (f *fakeRepo) UpdateServerResources(_ context.Context, name string, cpu, mem, stor int) error {
f.resourceUpdates[name] = ResourceSpec{CPUMilli: cpu, MemoryMB: mem, StorageMB: stor}
return nil
}
func (f *fakeRepo) ServerResources(_ context.Context, _ string) (ResourceSpec, error) { func (f *fakeRepo) ServerResources(_ context.Context, name string) (ResourceSpec, error) {
return ResourceSpec{}, nil return f.serverResources[name], nil
} }
func (f *fakeRepo) CreateLinkCode(_ context.Context, code, mcUUID, authSource string, expiresAt time.Time) error { func (f *fakeRepo) CreateLinkCode(_ context.Context, code, mcUUID, authSource string, expiresAt time.Time) error {
f.linkCodes[code] = fakeLinkCode{mcUUID: mcUUID, authSource: authSource, expiresAt: expiresAt} f.linkCodes[code] = fakeLinkCode{mcUUID: mcUUID, authSource: authSource, expiresAt: expiresAt}
@@ -266,7 +287,13 @@ func (f *fakeRepo) VerifyLinkCode(_ context.Context, userID, code string, now ti
return "", "", ErrLinkCodeInvalid return "", "", ErrLinkCodeInvalid
} }
if existing, ok := f.links[rec.mcUUID]; ok && existing != userID { if existing, ok := f.links[rec.mcUUID]; ok && existing != userID {
return "", "", ErrConflict // do not consume another user's pending code // A soft-deleted link's identity is unclaimed: the fresh in-game code lets a
// live caller take it over (mirrors PGRepo). Disabled-but-not-deleted stays a
// conflict — takeover there would bypass the lockout. Neither arm consumes
// the code.
if !f.seededDeleted(existing) {
return "", "", ErrConflict
}
} }
f.links[rec.mcUUID] = userID f.links[rec.mcUUID] = userID
f.linkAuthSource[rec.mcUUID] = rec.authSource // copy/refresh, mirrors DO UPDATE f.linkAuthSource[rec.mcUUID] = rec.authSource // copy/refresh, mirrors DO UPDATE
@@ -299,6 +326,9 @@ func (f *fakeRepo) RedeemPlayerBindCode(_ context.Context, newUserID, code strin
return "", "", "", ErrPlayerBindForbidden // staff must use op.console; do not consume return "", "", "", ErrPlayerBindForbidden // staff must use op.console; do not consume
} }
} }
if f.seededDead(existing) {
return "", "", "", ErrPlayerAccountRetired // dead account; do not consume
}
delete(f.linkCodes, code) delete(f.linkCodes, code)
return existing, rec.mcUUID, rec.authSource, nil return existing, rec.mcUUID, rec.authSource, nil
} }
@@ -350,6 +380,13 @@ func (f *fakeRepo) VerifyEmailOTP(_ context.Context, userID, purpose, codeHash s
live.attempts++ // a typo costs an attempt but does not consume the code live.attempts++ // a typo costs an attempt but does not consume the code
return "", ErrOTPInvalid return "", ErrOTPInvalid
} }
// A DIFFERENT verified holder of the same address → ErrEmailTaken, code left
// live — mirrors PGRepo's guard + the users_verified_email_unique index.
for _, u := range f.staff {
if u.ID != userID && u.EmailVerified && strings.EqualFold(u.Email, live.email) {
return "", ErrEmailTaken
}
}
live.consumed = true live.consumed = true
for _, u := range f.staff { // flip the user row verified (UPDATE users ...) for _, u := range f.staff { // flip the user row verified (UPDATE users ...)
if u.ID == userID { if u.ID == userID {
@@ -595,7 +632,7 @@ func (f *fakeRepo) UUIDInAllowlist(_ context.Context, n, uuid string) (bool, err
return f.allowUUID[n][uuid], nil return f.allowUUID[n][uuid], nil
} }
func (f *fakeRepo) UserByMCUUID(_ context.Context, uuid string) (string, error) { func (f *fakeRepo) UserByMCUUID(_ context.Context, uuid string) (string, error) {
if u, ok := f.links[uuid]; ok { if u, ok := f.links[uuid]; ok && !f.seededDead(u) {
return u, nil return u, nil
} }
return "", ErrNotFound return "", ErrNotFound
@@ -619,15 +656,15 @@ func (f *fakeRepo) IsUsernameBlacklisted(_ context.Context, mcUUID string) (bool
} }
// IsProtectedAdminLink mirrors PGRepo's JOIN of account_links to users: linked, // IsProtectedAdminLink mirrors PGRepo's JOIN of account_links to users: linked,
// auth_source 'thirdparty', and the linked user an admin — no password-hash test, so // auth_source 'thirdparty', and the linked user staff (admin OR owner) — no
// an SSO Operator (role='admin', with no password) is protected like any other. // password-hash test, so an SSO Operator or the Owner is protected like any other.
func (f *fakeRepo) IsProtectedAdminLink(_ context.Context, mcUUID string) (bool, error) { func (f *fakeRepo) IsProtectedAdminLink(_ context.Context, mcUUID string) (bool, error) {
userID, ok := f.links[mcUUID] userID, ok := f.links[mcUUID]
if !ok || f.linkAuthSource[mcUUID] != authSourceThirdParty { if !ok || f.linkAuthSource[mcUUID] != authSourceThirdParty {
return false, nil return false, nil
} }
for _, u := range f.staff { for _, u := range f.staff {
if u.ID == userID && u.Role == "admin" { if u.ID == userID && staffRole(u.Role) {
return true, nil return true, nil
} }
} }
@@ -638,6 +675,9 @@ func (f *fakeRepo) ClaimServer(_ context.Context, n, u string) (bool, error) {
if !present { if !present {
return false, ErrNotFound return false, ErrNotFound
} }
if ok && f.claimQuotaRefuse[n] {
return false, ErrQuotaExceeded // mirrors the atomic gate losing the race
}
return ok, nil return ok, nil
} }
func (f *fakeRepo) RecordJoin(_ context.Context, n, uuid string) error { func (f *fakeRepo) RecordJoin(_ context.Context, n, uuid string) error {
@@ -768,12 +808,18 @@ func (f *fakeRepo) CreateSession(_ context.Context, tokenHash, userID string, ex
return nil return nil
} }
func (f *fakeRepo) SessionUser(_ context.Context, tokenHash string, now time.Time) (*SessionedUser, error) { func (f *fakeRepo) SessionUser(_ context.Context, tokenHash string, now time.Time) (*SessionedUser, error) {
if f.failSessionUser != nil {
return nil, f.failSessionUser
}
s, ok := f.sessions[tokenHash] s, ok := f.sessions[tokenHash]
if !ok || s.revoked || !s.expiresAt.After(now) { if !ok || s.revoked || !s.expiresAt.After(now) {
return nil, ErrNotFound return nil, ErrNotFound
} }
for _, u := range f.staff { for _, u := range f.staff {
if u.ID == s.userID { if u.ID == s.userID {
if f.seededDead(u.ID) {
return nil, ErrNotFound
}
return &SessionedUser{ return &SessionedUser{
ID: u.ID, Email: u.Email, Role: u.Role, ID: u.ID, Email: u.Email, Role: u.Role,
}, nil }, nil
@@ -788,6 +834,9 @@ func (f *fakeRepo) RevokeSession(_ context.Context, tokenHash string) error {
return nil return nil
} }
func (f *fakeRepo) GetSetting(_ context.Context, key string) ([]byte, error) { func (f *fakeRepo) GetSetting(_ context.Context, key string) ([]byte, error) {
if f.failGetSetting != nil {
return nil, f.failGetSetting
}
if v, ok := f.settings[key]; ok { if v, ok := f.settings[key]; ok {
return v, nil return v, nil
} }
@@ -863,11 +912,29 @@ func (f *fakeRepo) ListUsers(_ context.Context, opts ListUsersOpts) ([]UserView,
} }
func (f *fakeRepo) UserDetail(_ context.Context, userID string) (*UserDetail, error) { func (f *fakeRepo) UserDetail(_ context.Context, userID string) (*UserDetail, error) {
deletedAt := time.Unix(1_700_000_000, 0)
for _, su := range f.seededUsers { for _, su := range f.seededUsers {
if su.view.ID == userID { if su.view.ID == userID {
return &su.detail, nil return &su.detail, nil
} }
} }
// Legacy fixtures seeded only into f.staff are live accounts (nothing marked
// them disabled or deleted), so detail reads must resolve them too — the
// liveness guards (discoverable login, owner protection) treat "unknown" as a
// fault, and these fixtures are known.
for _, u := range f.staff {
if u.ID == userID {
if f.deletedIDs[userID] {
// A soft-deleted account still HAS a detail row; it is flagged, not gone.
return &UserDetail{UserView: UserView{
ID: u.ID, Username: u.Username, Role: u.Role, Disabled: true,
}, DeletedAt: &deletedAt}, nil
}
return &UserDetail{UserView: UserView{
ID: u.ID, Username: u.Username, Email: u.Email, Role: u.Role,
}}, nil
}
}
return nil, ErrNotFound return nil, ErrNotFound
} }
@@ -904,6 +971,12 @@ func (f *fakeRepo) UpdateUser(_ context.Context, userID string, patch UpdateUser
f.seededUsers[i].detail.Username = *patch.Username f.seededUsers[i].detail.Username = *patch.Username
} }
if patch.Email != nil { if patch.Email != nil {
// Changing the address voids the proof of it, exactly like PGRepo:
// only VerifyEmailOTP may assert a verified address.
if *patch.Email != f.seededUsers[i].view.Email {
f.seededUsers[i].view.EmailVerified = false
f.seededUsers[i].detail.EmailVerified = false
}
f.seededUsers[i].view.Email = *patch.Email f.seededUsers[i].view.Email = *patch.Email
f.seededUsers[i].detail.Email = *patch.Email f.seededUsers[i].detail.Email = *patch.Email
} }
@@ -921,6 +994,20 @@ func (f *fakeRepo) DeleteUser(_ context.Context, userID, _ string) error {
for i, su := range f.seededUsers { for i, su := range f.seededUsers {
if su.view.ID == userID { if su.view.ID == userID {
f.seededUsers = append(f.seededUsers[:i], f.seededUsers[i+1:]...) f.seededUsers = append(f.seededUsers[:i], f.seededUsers[i+1:]...)
f.deletedIDs[userID] = true
// Mirror PGRepo: deletion severs the account's identity assets so the
// closed account keeps neither a login credential nor a MC-UUID claim.
for uuid, uid := range f.links {
if uid == userID {
delete(f.links, uuid)
delete(f.linkAuthSource, uuid)
}
}
for cid, cred := range f.passkeyCreds {
if cred.UserID == userID {
delete(f.passkeyCreds, cid)
}
}
return nil return nil
} }
} }
@@ -1068,7 +1155,52 @@ func (f *fakeRepo) RedeemMigration(_ context.Context, targetUserID, codeHash str
// ---- quota admin fakes ---- // ---- quota admin fakes ----
// liveUserExists mirrors PGRepo.requireLiveUser: the admin quota/link fakes
// only act on a live seeded row, so a unit test can drive the unknown-user 404
// the real FK would otherwise turn into a 500.
func (f *fakeRepo) liveUserExists(id string) bool {
for _, su := range f.seededUsers {
if su.view.ID == id && su.detail.DeletedAt == nil {
return true
}
}
return false
}
// seededDead mirrors PGRepo's liveness filters (audit #33): a seeded user that was
// disabled or soft-deleted is dead for the login doors and session validation. A
// fixture that was never seeded (legacy tests put it straight into f.staff) is
// treated as live, matching the fakes' pre-existing behavior.
func (f *fakeRepo) seededDead(id string) bool {
if f.deletedIDs[id] {
return true
}
for _, su := range f.seededUsers {
if su.view.ID == id {
return su.view.Disabled || su.detail.DeletedAt != nil
}
}
return false
}
// seededDeleted is the narrower liveness query: soft-deleted only (a disabled
// account still holds its identity, mirroring VerifyLinkCode's takeover rule).
func (f *fakeRepo) seededDeleted(id string) bool {
if f.deletedIDs[id] {
return true
}
for _, su := range f.seededUsers {
if su.view.ID == id {
return su.detail.DeletedAt != nil
}
}
return false
}
func (f *fakeRepo) GetQuotas(_ context.Context, userID string) (*QuotaView, error) { func (f *fakeRepo) GetQuotas(_ context.Context, userID string) (*QuotaView, error) {
if !f.liveUserExists(userID) {
return nil, ErrNotFound
}
v := &QuotaView{UserID: userID} v := &QuotaView{UserID: userID}
if f.fakeQuotas == nil { if f.fakeQuotas == nil {
return v, nil return v, nil
@@ -1083,6 +1215,9 @@ func (f *fakeRepo) GetQuotas(_ context.Context, userID string) (*QuotaView, erro
} }
func (f *fakeRepo) SetQuotas(_ context.Context, userID string, qi QuotaInput, _ string) (*QuotaView, error) { func (f *fakeRepo) SetQuotas(_ context.Context, userID string, qi QuotaInput, _ string) (*QuotaView, error) {
if !f.liveUserExists(userID) {
return nil, ErrNotFound
}
if f.fakeQuotas == nil { if f.fakeQuotas == nil {
f.fakeQuotas = map[string]*QuotaView{} f.fakeQuotas = map[string]*QuotaView{}
} }
@@ -1135,6 +1270,9 @@ func (f *fakeRepo) UnlinkAccount(_ context.Context, userID, mcUUID string) error
} }
func (f *fakeRepo) LinkAccount(_ context.Context, userID, mcUUID, authSource string) error { func (f *fakeRepo) LinkAccount(_ context.Context, userID, mcUUID, authSource string) error {
if !f.liveUserExists(userID) {
return ErrNotFound
}
if existing, ok := f.links[mcUUID]; ok && existing != userID { if existing, ok := f.links[mcUUID]; ok && existing != userID {
return ErrConflict return ErrConflict
} }
@@ -1153,7 +1291,7 @@ func (f *fakeRepo) LinkAccount(_ context.Context, userID, mcUUID, authSource str
// is indistinguishable from no account: both yield ErrNotFound. // is indistinguishable from no account: both yield ErrNotFound.
func (f *fakeRepo) UserByEmail(_ context.Context, email string) (*StaffUser, error) { func (f *fakeRepo) UserByEmail(_ context.Context, email string) (*StaffUser, error) {
for _, u := range f.staff { for _, u := range f.staff {
if u.EmailVerified && strings.EqualFold(u.Email, email) { if u.EmailVerified && strings.EqualFold(u.Email, email) && !f.seededDead(u.ID) {
su := *u su := *u
return &su, nil return &su, nil
} }
@@ -1313,6 +1451,7 @@ type fakeCluster struct {
desired map[string]v1alpha1.DesiredState desired map[string]v1alpha1.DesiredState
created map[string]CreateServerInput // name -> the validated input it was created from created map[string]CreateServerInput // name -> the validated input it was created from
patched map[string]ServerSpecPatch // name -> the validated spec patch it received patched map[string]ServerSpecPatch // name -> the validated spec patch it received
noWorld map[string]bool // server names modeled WITHOUT a world volume (never started / reaped)
createErr error createErr error
pingErr error pingErr error
} }
@@ -1320,7 +1459,7 @@ type fakeCluster struct {
func newFakeCluster() *fakeCluster { func newFakeCluster() *fakeCluster {
return &fakeCluster{byName: map[string]*ServerInfo{}, bySub: map[string]*ServerInfo{}, return &fakeCluster{byName: map[string]*ServerInfo{}, bySub: map[string]*ServerInfo{},
desired: map[string]v1alpha1.DesiredState{}, created: map[string]CreateServerInput{}, desired: map[string]v1alpha1.DesiredState{}, created: map[string]CreateServerInput{},
patched: map[string]ServerSpecPatch{}} patched: map[string]ServerSpecPatch{}, noWorld: map[string]bool{}}
} }
func (c *fakeCluster) GetServer(_ context.Context, n string) (*ServerInfo, error) { func (c *fakeCluster) GetServer(_ context.Context, n string) (*ServerInfo, error) {
if s, ok := c.byName[n]; ok { if s, ok := c.byName[n]; ok {
@@ -1335,7 +1474,14 @@ func (c *fakeCluster) GetBySubdomain(_ context.Context, s string) (*ServerInfo,
return nil, ErrNotFound return nil, ErrNotFound
} }
func (c *fakeCluster) ListServers(_ context.Context) ([]ServerInfo, error) { return c.list, nil } func (c *fakeCluster) ListServers(_ context.Context) ([]ServerInfo, error) { return c.list, nil }
func (c *fakeCluster) Ping(_ context.Context) error { return c.pingErr } func (c *fakeCluster) Ping(_ context.Context) error { return c.pingErr }
// WorldVolumeExists models the world PVC: present unless the test named the
// server in noWorld (never started / already reaped).
func (c *fakeCluster) WorldVolumeExists(_ context.Context, n string) (bool, error) {
return !c.noWorld[n], nil
}
func (c *fakeCluster) SetDesiredState(_ context.Context, n string, s v1alpha1.DesiredState) error { func (c *fakeCluster) SetDesiredState(_ context.Context, n string, s v1alpha1.DesiredState) error {
c.desired[n] = s c.desired[n] = s
return nil return nil
@@ -1671,6 +1817,44 @@ func TestFleetAdminRead(t *testing.T) {
} }
}) })
t.Run("system services are marked read-only", func(t *testing.T) {
// The login gate and the lobby carry reserved names, so every per-server
// route rejects them; the fleet row must say "system" so the cockpit
// renders them without actions that would 400.
sysCl := newFakeCluster()
sysCl.list = []ServerInfo{
{Name: "login", Phase: "Running", Ready: true},
{Name: "lobby", Phase: "Running", Ready: true},
{Name: "survival", Phase: "Stopped"},
}
api := newTestAPI(newFakeRepo(), sysCl)
api.External = staticExternal{p: &Principal{UserID: "a1", Email: "[email protected]",
Role: "admin", ViaAdminAccess: true}}
w := do(api.ExternalHandler(), "GET", "/api/v1/fleet", "", nil)
if w.Code != http.StatusOK {
t.Fatalf("code = %d, want 200 (%s)", w.Code, w.Body.String())
}
var got struct {
Servers []struct {
Name string `json:"name"`
System bool `json:"system"`
} `json:"servers"`
}
if err := json.Unmarshal(w.Body.Bytes(), &got); err != nil {
t.Fatalf("body not JSON: %v", err)
}
byName := map[string]bool{}
for _, r := range got.Servers {
byName[r.Name] = r.System
}
if !byName["login"] || !byName["lobby"] {
t.Errorf("system flags = %+v, want login+lobby marked", byName)
}
if byName["survival"] {
t.Errorf("survival marked system; only platform services are")
}
})
t.Run("owner merges for claimed, absent for unclaimed", func(t *testing.T) { t.Run("owner merges for claimed, absent for unclaimed", func(t *testing.T) {
repo := newFakeRepo() repo := newFakeRepo()
// Only "survival" is claimed; "creative"/"skyblock" stay unowned. // Only "survival" is claimed; "creative"/"skyblock" stay unowned.
@@ -1760,6 +1944,21 @@ func TestClaimStateMachine(t *testing.T) {
t.Fatalf("code = %d body %s", w.Code, w.Body.String()) t.Fatalf("code = %d body %s", w.Code, w.Body.String())
} }
}) })
t.Run("atomic gate refusal -> 403 quota_exceeded", func(t *testing.T) {
// The advisory pre-check passed, but ClaimServer's serialized re-check
// (audit #4) refuses: the caller must see the same 403, not a 500.
repo := newFakeRepo()
repo.linked["u1"] = true
repo.quota["u1"] = true
repo.claimOK["survival"] = true
repo.claimQuotaRefuse["survival"] = true
api := newTestAPI(repo, newFakeCluster())
api.External = staticExternal{p: user}
w := do(api.ExternalHandler(), "POST", "/api/v1/servers/survival/claim", "", nil)
if w.Code != http.StatusForbidden || decodeErr(t, w) != "quota_exceeded" {
t.Fatalf("code = %d body %s, want 403 quota_exceeded", w.Code, w.Body.String())
}
})
t.Run("already claimed -> 409", func(t *testing.T) { t.Run("already claimed -> 409", func(t *testing.T) {
repo := newFakeRepo() repo := newFakeRepo()
repo.linked["u1"] = true repo.linked["u1"] = true
@@ -2195,3 +2394,38 @@ func TestAccessVerifier(t *testing.T) {
} }
}) })
} }
// TestSessionAuthOutageIs503Not401: a session-store outage must surface as 503
// auth_unavailable, not a 401 that reads as "please log in again". Both failure
// points are covered — the local_auth_enabled read and the session row read —
// plus the regression that a genuinely missing session still answers 401.
func TestSessionAuthOutageIs503Not401(t *testing.T) {
apiWith := func(repo *fakeRepo) *API {
a := newTestAPI(repo, newFakeCluster())
a.External = SessionAuth{Repo: repo, RootDomain: testRoot, AdminHostname: "op.console." + testRoot}
return a
}
cookie := map[string]string{"Cookie": sessionCookieName + "=any"}
outage := errors.New("dial tcp 10.0.0.5:5432: connect: connection refused")
repo := newFakeRepo()
repo.settings[LocalAuthEnabledKey] = []byte("true")
repo.failGetSetting = outage
if w := do(apiWith(repo).ExternalHandler(), "GET", "/api/v1/me", "", cookie); w.Code != http.StatusServiceUnavailable || decodeErr(t, w) != "auth_unavailable" {
t.Fatalf("settings read outage = %d body %s, want 503 auth_unavailable", w.Code, w.Body.String())
}
repo = newFakeRepo()
repo.settings[LocalAuthEnabledKey] = []byte("true")
repo.failSessionUser = outage
if w := do(apiWith(repo).ExternalHandler(), "GET", "/api/v1/me", "", cookie); w.Code != http.StatusServiceUnavailable || decodeErr(t, w) != "auth_unavailable" {
t.Fatalf("session read outage = %d body %s, want 503 auth_unavailable", w.Code, w.Body.String())
}
// Regression: fail-closed auth (missing/invalid session) stays a 401.
repo = newFakeRepo()
repo.settings[LocalAuthEnabledKey] = []byte("true")
if w := do(apiWith(repo).ExternalHandler(), "GET", "/api/v1/me", "", cookie); w.Code != http.StatusUnauthorized || decodeErr(t, w) != "unauthorized" {
t.Fatalf("missing session = %d body %s, want 401 unauthorized", w.Code, w.Body.String())
}
}
+3 -3
View File
@@ -17,13 +17,13 @@ type Principal struct {
UserID string UserID string
// Email is the audited actor identity (spec §14: audit actor = Access email). // Email is the audited actor identity (spec §14: audit actor = Access email).
Email string Email string
// Role is "admin" or "user" (mirrors users.role). // Role is "owner", "admin", or "user" (mirrors users.role).
Role string Role string
// ViaAdminAccess is true only when the request arrived through an admin-graded // ViaAdminAccess is true only when the request arrived through an admin-graded
// path: the admin.* Zero-Trust hostname (Cloudflare Access, the remote face) OR // path: the admin.* Zero-Trust hostname (Cloudflare Access, the remote face) OR
// a local session presented on the op.console host (SessionAuth, the // a local session presented on the op.console host (SessionAuth, the
// passwordless face). Admin-tier operations require it in addition to // passwordless face). Admin-tier operations require it in addition to
// Role=="admin" (spec §14: ZT is graded by operation). A role=admin session // a staff role (spec §14: ZT is graded by operation). A staff session
// arriving on the player console (console.*) never sets it. // arriving on the player console (console.*) never sets it.
ViaAdminAccess bool ViaAdminAccess bool
// EmailVerified mirrors users.email_verified. The lockdown middleware gates // EmailVerified mirrors users.email_verified. The lockdown middleware gates
@@ -49,7 +49,7 @@ func staffRole(role string) bool {
} }
// IsAdmin reports whether the principal may perform admin-tier operations. // IsAdmin reports whether the principal may perform admin-tier operations.
// Both the role claim and the admin Access path are required: a role=admin // Both the role claim and the admin Access path are required: a staff
// session arriving on panel.* must not bypass the Zero-Trust boundary. // session arriving on panel.* must not bypass the Zero-Trust boundary.
// An owner implicitly passes this check (the owner role is a superset of admin). // An owner implicitly passes this check (the owner role is a superset of admin).
func (p *Principal) IsAdmin() bool { func (p *Principal) IsAdmin() bool {
+6
View File
@@ -81,6 +81,12 @@ type Cluster interface {
// GetServer reads one MinecraftServer's lifecycle view, or ErrNotFound. // GetServer reads one MinecraftServer's lifecycle view, or ErrNotFound.
GetServer(ctx context.Context, name string) (*ServerInfo, error) GetServer(ctx context.Context, name string) (*ServerInfo, error)
// WorldVolumeExists reports whether the server's world PVC exists in the
// server namespace. A server that never started — or whose world the
// retention reaper already archived and deleted — has no claim, and a
// backup/restore Job would hang Pending on the missing volume with nothing
// ever recorded, so both handlers refuse those up front.
WorldVolumeExists(ctx context.Context, name string) (bool, error)
// GetBySubdomain finds the MinecraftServer whose spec.subdomain matches, or // GetBySubdomain finds the MinecraftServer whose spec.subdomain matches, or
// ErrNotFound. // ErrNotFound.
GetBySubdomain(ctx context.Context, subdomain string) (*ServerInfo, error) GetBySubdomain(ctx context.Context, subdomain string) (*ServerInfo, error)
+28 -8
View File
@@ -15,6 +15,13 @@ var (
ErrNotFound = errors.New("not found") ErrNotFound = errors.New("not found")
// ErrConflict means an atomic precondition failed (e.g. claim lost the race). // ErrConflict means an atomic precondition failed (e.g. claim lost the race).
ErrConflict = errors.New("conflict") ErrConflict = errors.New("conflict")
// ErrQuotaExceeded means an ownership write would push the user over a quota
// cap (spec §9.3). ClaimServer — the atomic gate — returns it when a claim
// passes the handler's advisory pre-check but loses the serialized re-check
// (two concurrent claims by one user); handlers map it to a 403
// quota_exceeded, the same answer the pre-check gives, so the CONCURRENT case
// and the SEQUENTIAL case are indistinguishable to the caller.
ErrQuotaExceeded = errors.New("server quota exhausted")
// ErrLinkCodeInvalid means an account-link code is unknown or expired (spec // ErrLinkCodeInvalid means an account-link code is unknown or expired (spec
// §10). It is a client error (the verify endpoint exists; the code is bad), so // §10). It is a client error (the verify endpoint exists; the code is bad), so
// handlers map it to 400, not 404. // handlers map it to 400, not 404.
@@ -44,14 +51,22 @@ var (
// consumed, or expired) — so handlers map it to 400, not 404. // consumed, or expired) — so handlers map it to 400, not 404.
ErrPasskeyChallengeInvalid = errors.New("passkey challenge invalid or expired") ErrPasskeyChallengeInvalid = errors.New("passkey challenge invalid or expired")
// ErrPlayerBindForbidden means a public Bind-Code redemption resolved to a STAFF // ErrPlayerBindForbidden means a public Bind-Code redemption resolved to a STAFF
// account (role=admin), which the player-console bootstrap refuses (console-tier // account (admin or owner), which the player-console bootstrap refuses
// access model). Operators authenticate at op.console behind Zero Trust, never via // (console-tier access model). Staff authenticate at op.console behind Zero Trust,
// the account-less console.<root_domain> door, so the public bootstrap provably // never via the account-less console.<root_domain> door, so the public bootstrap
// never mints a session for an admin identity. It is distinct from ErrConflict so // provably never mints a session for a staff identity. It is distinct from
// the handler answers 403 (wrong door) rather than 409 (already linked). // ErrConflict so the handler answers 403 (wrong door) rather than 409.
ErrPlayerBindForbidden = errors.New("bind code belongs to a staff account") ErrPlayerBindForbidden = errors.New("bind code belongs to a staff account")
// ErrPlayerAccountRetired means a Bind-Code redemption resolved to an account the
// platform has closed: an owner soft-deleted it, or it is disabled (locked out).
// Reusing the row would mint a fresh session for a dead account — the same
// resurrection the login doors refuse by resolving only live accounts — so the
// redeemer gets an explicit 403 instead. The code is NOT consumed, so re-enabling
// the account and retrying still works within the code's TTL.
ErrPlayerAccountRetired = errors.New("player account is retired or disabled")
// ErrEmailTaken means a verified email would collide with another account's // ErrEmailTaken means a verified email would collide with another account's
// already-verified address (spec §B email-first login foundation, migration 0010). // already-verified address (spec §B email-first login foundation; the
// users_verified_email_unique index ships in migration 0020).
// VerifyEmailOTP returns it — WITHOUT consuming the code, since the address, not // VerifyEmailOTP returns it — WITHOUT consuming the code, since the address, not
// the code, is the problem — when a DIFFERENT user has already proven the same // the code, is the problem — when a DIFFERENT user has already proven the same
// address case-insensitively. It is the clean, application-level counterpart of // address case-insensitively. It is the clean, application-level counterpart of
@@ -89,8 +104,13 @@ func newError(status int, code, format string, a ...any) *apiError {
// Common errors reused across handlers. // Common errors reused across handlers.
var ( var (
errUnauthorized = newError(http.StatusUnauthorized, "unauthorized", "authentication required") errUnauthorized = newError(http.StatusUnauthorized, "unauthorized", "authentication required")
errForbidden = newError(http.StatusForbidden, "forbidden", "not permitted") // errAuthUnavailable answers when the session store itself is unreachable
errBadRequest = newError(http.StatusBadRequest, "bad_request", "invalid request") // (Postgres down): an outage is not a credential verdict, so the caller gets
// 503 "retry" instead of a 401 that reads as "log in again".
errAuthUnavailable = newError(http.StatusServiceUnavailable, "auth_unavailable",
"authentication is temporarily unavailable; retry shortly")
errForbidden = newError(http.StatusForbidden, "forbidden", "not permitted")
errBadRequest = newError(http.StatusBadRequest, "bad_request", "invalid request")
) )
// writeJSON writes v as an indented JSON body with the given status. // writeJSON writes v as an indented JSON body with the given status.
+1 -1
View File
@@ -416,7 +416,7 @@ type lpPermissionView struct {
// maxLPInfoPages bounds how many "permission info" pages the read projector // maxLPInfoPages bounds how many "permission info" pages the read projector
// chases per request. LuckPerms paginates its reply, so one command shows only // chases per request. LuckPerms paginates its reply, so one command shows only
// the first page; we follow the header's page count up to this cap. // the first page; we follow the header's page count up to this cap.
// ponytail: 10 pages ≈ 150 entries — raise if a real user outgrows it. // 10 pages ≈ 150 entries — raise if a real user outgrows it.
const maxLPInfoPages = 10 const maxLPInfoPages = 10
// handleAccessLuckPermsInfo is the read projector for a player's LuckPerms // handleAccessLuckPermsInfo is the read projector for a player's LuckPerms
+54
View File
@@ -1,7 +1,9 @@
package api package api
import ( import (
"context"
"encoding/json" "encoding/json"
"errors"
"net/http" "net/http"
"net/http/httptest" "net/http/httptest"
"strings" "strings"
@@ -252,6 +254,58 @@ func TestLinkVerifyIdempotent(t *testing.T) {
} }
} }
// A link whose account was SOFT-DELETED is unclaimed: a fresh in-game code lets a
// live account take it over (the migrated-source path — retire keeps the link but
// kills the account), while a merely disabled holder keeps its identity so the
// lockout cannot be re-linked away, and neither dead link has in-game standing
// (UserByMCUUID reads it exactly like an unlinked UUID). Audit #33's in-game half.
func TestLinkVerifyTakesOverDeletedLinkOnly(t *testing.T) {
ctx := context.Background()
const mcGone = "55555555-5555-5555-5555-555555555555"
const mcLocked = "66666666-6666-6666-6666-666666666666"
user := &Principal{UserID: "u-take", Email: "[email protected]", Role: "user"}
repo := newFakeRepo()
repo.seedUser(UserView{ID: "u-gone", Username: "gone", Email: "[email protected]", Role: "user"})
repo.seedUser(UserView{ID: "u-locked", Username: "locked", Email: "[email protected]", Role: "user"})
if err := repo.DeleteUser(ctx, "u-gone", "test"); err != nil {
t.Fatalf("DeleteUser: %v", err)
}
repo.links[mcGone] = "u-gone" // a retired source's link outlives the account
if _, err := repo.UserByMCUUID(ctx, mcGone); !errors.Is(err, ErrNotFound) {
t.Fatalf("UserByMCUUID(deleted link) = %v, want ErrNotFound (no in-game standing)", err)
}
repo.linkCodes["TAKEOVER"] = fakeLinkCode{mcUUID: mcGone, expiresAt: time.Unix(1_700_000_600, 0)}
api := newTestAPI(repo, newFakeCluster())
api.External = staticExternal{p: user}
if w := do(api.ExternalHandler(), "POST", "/api/v1/account/link/verify", `{"code":"TAKEOVER"}`, nil); w.Code != http.StatusOK {
t.Fatalf("takeover verify: code = %d, want 200 (%s)", w.Code, w.Body.String())
}
if repo.links[mcGone] != "u-take" {
t.Errorf("links[%s] = %q after takeover, want u-take", mcGone, repo.links[mcGone])
}
// A disabled (not deleted) holder keeps the identity: 409, code survives, link unmoved.
if err := repo.SetUserDisabled(ctx, "u-locked", true); err != nil {
t.Fatalf("disable: %v", err)
}
repo.links[mcLocked] = "u-locked"
repo.linkCodes["LOCKED12"] = fakeLinkCode{mcUUID: mcLocked, expiresAt: time.Unix(1_700_000_600, 0)}
if _, err := repo.UserByMCUUID(ctx, mcLocked); !errors.Is(err, ErrNotFound) {
t.Fatalf("UserByMCUUID(disabled link) = %v, want ErrNotFound", err)
}
if w := do(api.ExternalHandler(), "POST", "/api/v1/account/link/verify", `{"code":"LOCKED12"}`, nil); w.Code != http.StatusConflict {
t.Fatalf("takeover of a disabled holder: code = %d, want 409 (%s)", w.Code, w.Body.String())
}
if repo.links[mcLocked] != "u-locked" {
t.Errorf("disabled holder's link moved to %q", repo.links[mcLocked])
}
if _, ok := repo.linkCodes["LOCKED12"]; !ok {
t.Error("refused verify consumed the code")
}
}
// TestLinkAuthSourcePropagates proves auth_source survives the whole §10 flow: a // TestLinkAuthSourcePropagates proves auth_source survives the whole §10 flow: a
// thirdparty source captured in-game at mint reaches the durable link and the // thirdparty source captured in-game at mint reaches the durable link and the
// verify response — the value the web side can never originate itself. // verify response — the value the web side can never originate itself.
+3 -1
View File
@@ -241,7 +241,9 @@ func (a *API) handleLoginEmailVerify(w http.ResponseWriter, r *http.Request) {
// Code redeemed. Refuse staff here — never before the verify — so op.console keeps // Code redeemed. Refuse staff here — never before the verify — so op.console keeps
// its Zero-Trust + in-game-approval gates and this public door provably yields only // its Zero-Trust + in-game-approval gates and this public door provably yields only
// a role=user player session (mirrors handleBindRedeem's refuse-staff contract). // a role=user player session (mirrors handleBindRedeem's refuse-staff contract).
if u.Role == "admin" { // Staff means anything above role=user: an admin OR the role=owner identity. The
// player door must yield only player sessions.
if u.Role != "user" {
writeError(w, r, newError(http.StatusForbidden, "staff_account", writeError(w, r, newError(http.StatusForbidden, "staff_account",
"that account is staff; sign in at the operator console")) "that account is staff; sign in at the operator console"))
return return
Loaded 100 of 232 files, more files were not shown because too many files have changed in this diff. Show more