Commit Graph
586 Commits
Author SHA1 Message Date
Lemon-miaow ac403b9cd3 fix(build): let kaniko pull base images from the plain-HTTP registry — push-only insecure flags broke every FROM registry.felis.svc build (#70) 2026-09-24 02:56:56 +08:00
Lemon-miaow 7791ed74f2 docs(audit): batch-44 — modpack-scale (200 MiB) context drill + concurrent builds, zero defects 2026-09-24 02:50:31 +08:00
Lemon-miaow 6abce99b49 docs(audit): batch-43 addendum — #68 context-fetch retry, live-verified; long-poll closure 2026-09-24 02:44:46 +08:00
Lemon-miaow b76d0acec7 fix(build): retry the context fetch through a control-plane restart (#68) 2026-09-24 02:27:45 +08:00
Lemon-miaow 580032056e docs(audit): forty-third batch — build-job reaping (#67) + lifecycle/fault-injection soak
Six CR-level run/stop cycles with three chaos injections (operator pod
kill, postgres restart, api pod kill) all converged (Running 23-29s,
Stopped 3s, pods gone per cycle); PG outage keeps the 503-not-401
session semantics; 30-request burst all 200; ownerless wake correctly
403s.  The build lane's missing Job TTL (#67, fixed in 2755e41) was
verified live on auditfix62: a real build Job carries ttl=604800 and an
old Job patched to ttl=30s was reaped, pod and all, within 40s.
2026-09-24 02:23:18 +08:00
Lemon-miaow 2755e41ff3 fix(build): reap finished build Jobs — they accumulated forever
Every other Job family carries a TTLSecondsAfterFinished (fileedit 2m,
backup/restore 10m) but the build lane never set one: each build left a
completed Job + Pod in felis-build indefinitely (5 already on the drill
cluster, oldest 26h), growing etcd and — since completed pods count
against the node's pod budget (110 on stock k3s) — eventually blocking
new builds.  The code even anticipated GC it never got ("JobUnknown
means the Job was not found (e.g. GC'd)").

Set a deliberately long TTL (7 days): the kaniko log is the admin
failure-triage surface (GET /images/build/{id}/logs), so the window
keeps a week of logs while bounding steady-state pods; terminal builds
are idempotent under Sync, so a late log-404 is the only cost.
2026-09-24 02:03:20 +08:00
Lemon-miaow d5e623c7ba docs(audit): batch-42 addendum — static review of the never-run release pipeline
The release workflow has never executed (no tags/releases exist yet).
Its three load-bearing contracts were verified against their other
halves: the asset names vs bootstrap's download_release_binary (and its
'felis ${REF}' convergence check), the local-exporter path vs the
Dockerfile's final COPY, and the version assertion vs cmdVersion's first
line.  The one un-replicable check (file(1) machine type) was probed
with a cross-compiled linux/arm64 binary: 'ARM aarch64' matches.  No
pre-tag fixes needed.
2026-09-24 01:33:56 +08:00
Lemon-miaow aa82321907 docs(audit): forty-second batch — coverage sweep, README_EN (#66), §28 diagram alignment
Module-by-module coverage matrix against the ledger: repo hygiene clean,
no TODO/FIXME debt, CLI dispatch test-guarded, CI syntax-checks every
tracked shell script, all 22 internal packages accounted for (CRD types
and game-image dirs covered indirectly), panel's 17 page routes drilled
in earlier batches.  Fixes this batch: #66 README_EN was missing the
private-repo install workaround and the rerun/upgrade notes (an English
reader 404s on step one); the §28 claim/link sequence diagrams were
realigned with the audited implementations.  A local-only scripts/
sync.sh conflict (plugins/ exclusion vs go:embed) was reproduced and
fixed on the workstation — private per .git/info/exclude, not committed.
2026-09-24 01:27:42 +08:00
Lemon-miaow 62e8c87d0e docs(diagrams): align §28 claim/link sequences with the audited implementations
The claim-transaction note still described the pre-audit-#4 shape (SELECT
EXISTS plus a quota pre-check outside the transaction); the implemented
ClaimServer serializes on pg_advisory_xact_lock(user_id), takes the row
FOR UPDATE and re-runs the four-dimension gate inside the same
transaction — the diagram's QuotaAvailable step is only a fast path.
The link flow now selects the code FOR UPDATE, treats a same-user
re-verify as idempotent, and lets a live caller take over a retired
(soft-deleted) owner's link — the 409 is only for a different, live user.
Both re-read from internal/api/pgrepo.go and the handler mappings.
2026-09-24 01:26:32 +08:00
Lemon-miaow 62c5a8a408 docs(readme): port the private-repo install path and upgrade notes to the English README (#66)
README.md grew a required workaround — the repository is private, so the
plain raw.githubusercontent one-liner returns 404.  The credentialed form
(token handed to curl via --config - so it never touches argv, sudo -E so
the installer inherits it) and the rerun/upgrade notes (the rerun upgrades
felis-api; the channel is not inherited, so main-followers need
FELIS_VERSION_BOOTSTRAP=dev) never made it into README_EN.md, which is
linked as the English entry point.  An English reader following that page
could not install at all.  Bring it back in sync with the Chinese one.
2026-09-24 01:24:27 +08:00
Lemon-miaow 5f9380dfca docs(audit): forty-first batch — loader mod layer closed
The last module family nobody ever built (#64 gradlew exec bit, #65
license metadata), closed with three real dedicated-server boots
(fabric/forge/neoforge: mod loads, /link registers, console refused)
and a permanent JDK-17 gate in CI + release.  Reachability table now
#1–#65; conclusions item 20; remaining-queue note updated.  CI run
35892544563 = 5/5 with the new mods job green on its first run.
2026-09-24 01:14:49 +08:00
Lemon-miaow aa5abfa911 ci(plugins): gate the three loader mods (JDK 17 wrapper builds) in CI and release
Nothing ever compiled these modules — no CI job, no install path — which is
why the gradlew exec-bit bug (previous commit) shipped unnoticed.  Add
plugins/test-mods.sh: runs each module's vendored wrapper under JDK 17
(they target the Java-17 Minecraft lines; paper/limbo stay on the JDK 21
gate).  ci.yml gains a `mods` job (temurin 17 + setup-gradle, wrapper
pinned per module), release.yml gates both plugin gates before shipping.
README: status now records the server-boot verification, the Building
section pins limbo's `-PlimboVersion=<release>` (the `+` default is
unresolvable from the LOOHP repo) and documents both gates.
2026-09-24 00:59:24 +08:00
Lemon-miaow c2fe6a6bf4 fix(plugins): loader mod metadata — AGPL-3.0-only license, real issue tracker
The three loader mods declared `license = "MIT"` (and Forge/NeoForge a
placeholder `issueTrackerURL = https://example.invalid/felis`) while the
repository is AGPL-3.0-only (README.md:73, LICENSE).  The mods were added
2026-06-26, the LICENSE landed 2026-07-12 — stale leftovers that nothing
ever read back: no CI job, no install path.  Fabric loader prints the
license from fabric.mod.json at boot and both mods.toml files are parsed
by their loaders, so the wrong claim is user-visible.  Align both with
reality — boot-verified on real fabric/forge/neoforge dedicated servers
(AUDIT-2026-09-22.md, batch 41).
2026-09-24 00:59:19 +08:00
Lemon-miaow f6048f268b fix(plugins): commit the gradlew exec bit for the loader mods
All three vendored wrappers were tracked as 100644, so the README's documented
'plugins/<loader>/gradlew -p plugins/<loader> build' failed on every fresh
clone with 'Permission denied' (exit 126) — and since no install path or CI job
ever ran them, nothing caught it.

git update-index --chmod=+x for the three files; the VM then built all three
modules for the first time (results in the gate commit and the audit).
2026-09-24 00:47:32 +08:00
Lemon-miaow 2d04c448c2 docs(audit): fortieth batch — java plugin layer into CI, demo-up single origin, #62 screened out
#63 fixed in f5a76cf (demo-up delegates to bootstrap; T1/T2/T3 real-machine);
plugins/test.sh + the new CI plugins job verified on the VM and in CI
(c59b387, run 35888009965, 4/4 jobs); #62 re-checked against all three
install paths and screened out as unreachable/nonexistent; plugin-layer
real-machine E2E via the MC status/login probes recorded.
2026-09-24 00:25:43 +08:00
Lemon-miaow c59b38775b ci(plugins): run the plugin self-tests and build the shipped plugin jars
The three framework-free test mains under plugins/*/test were never run by
anything — not CI, not the plugin builds — and the velocity/paper/limbo jars
were only ever compiled by deploy/bootstrap.sh on a live host. CI gains a
'plugins' job (JDK 21 plus the Gradle 8.14 the plugin Dockerfiles pin) running
plugins/test.sh: the three mains (InviteCardTest's jars fetched from Maven
Central, pinned and digest-checked) and the three production builds.

The first real run surfaced and fixed two untested assumptions: InviteCardTest's
documented javac line omitted examination-api (adventure-api's Component
signatures reference Examinable, so javac needs it too), and limbo's '+' version
default cannot resolve — LOOHP's repository serves no maven-metadata — so the
script resolves the current release off the Limbo CI artifact name (the same
source bootstrap reads) and plugins/README.md stops advertising a bare
'gradle -p plugins/limbo build' that can never work.

Verified in gradle:8.14-jdk21 on the VM: mains OK (32/36/48 checks);
velocity/paper/limbo BUILD SUCCESSFUL.
2026-09-24 00:19:56 +08:00
Lemon-miaow f5a76cf88a fix(deploy): make demo-up a thin wrapper — bootstrap is the one origin for the game stack (#63)
demo-up.sh grew its own image lane in July, before bootstrap could build the
game stack. The copy has since drifted from the installer it duplicates: it
pinned Paper 1.21.8 while bootstrap derives one version from the Limbo login
gate (both hops of a login must speak one protocol), it never built
felis-velocity.jar (a proxy without it silently routes nothing), it imported
images under local tags the [velocity] wiring no longer points at, and it left
the docker daemon running.

It now runs deploy/bootstrap.sh (unless SKIP_BOOTSTRAP=1), checks that the
earlier run really left the full stack behind, and hands over to 'felis setup'
— so a half-installed base is a clear error instead of a proxy that accepts
logins and routes nowhere.

Verified on the VM: healthy base passes to the handoff (exit 0); a hidden
felis-velocity.jar dies naming the jar and the remedy (exit 1, jar restored).
2026-09-24 00:19:49 +08:00
Lemon-miaow 1184d6b787 docs(audit): thirty-ninth batch — reaper multi-node drill on a cloned second node 2026-09-23 23:58:58 +08:00
Lemon-miaow 811f2b8898 docs(audit): thirty-eighth batch — #56–#59 evidence archived, server detail and ops sweep clean 2026-09-23 23:50:47 +08:00
Lemon-miaow 4c842f5d70 docs(audit): thirty-seventh batch — admin interaction sweep clean, #61 LuckPerms write promise scoped 2026-09-23 23:29:13 +08:00
Lemon-miaow b4ef42d947 fix(panel): scope the LuckPerms write promise to names the server can resolve (#61) 2026-09-23 23:28:28 +08:00
Lemon-miaow ecd42ed484 docs(audit): thirty-sixth batch — panel error localization (#60), Run4c admin write-path review clean 2026-09-23 23:12:14 +08:00
Lemon-miaow c9af5481e7 fix(panel): localize every user-reachable API error code (#60) 2026-09-23 23:09:28 +08:00
Lemon-miaow a2ff2a1102 fix(panel): make the LuckPerms page honest about what the server said (#59)
Live with real LuckPerms 5.5.85: every lp command returns an empty RCON body
(list/plugins answer normally; a standalone RCON client sees the same, and
creategroup/permission-set still persist), so the read projection can never
populate and the page asserted "no parent groups / no explicit nodes" for a
state it could not actually read. The raw reply now rides the same disclosure
the rosters carry, a silent entry-less reply shows an explicit notice instead
of the false empty claims, and the write history's placeholder no longer
dresses up a fabricated "[RCON] ..." line as output.
2026-09-23 22:43:29 +08:00
Lemon-miaow 4d4cdd6ea7 fix(api,panel): refuse file operations on a server without a world volume (#58)
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
2026-09-23 22:17:08 +08:00
Lemon-miaow 70c988e702 fix(panel): render the console command's reply (#57)
sendCommand returns the RCON reply and the component's own contract says it
"displays its plain-text reply", but the reply was discarded and the pod log
does not echo command output, so a sent command produced no visible result at
all (verified live with 'list'). The last command echo + reply now render
above the prompt, terminal-style.
2026-09-23 22:08:19 +08:00
Lemon-miaow 0790f8dfd3 fix(panel): surface the RCON reply for player-access mutations (#56)
The whitelist/ban/kick mutations reported a canned success message and threw
away the server's reply, so a refused command still read as done: live, the
vanilla server answers "That player does not exist" for a name it has never
seen (any player who has not joined yet), while the panel said the player had
been whitelisted/banned. api.ts documents these replies as "surfaced verbatim
as confirmation"; now they are. The localized string stays as the fallback
for a silent server.
2026-09-23 22:08:19 +08:00
Lemon-miaow 9c3a1c5be1 docs(audit): thirty-fifth batch — NetworkPolicy live enforcement matrix green (whitelist closed loop), velocity refresh loop verified 2026-09-23 21:43:23 +08:00
Lemon-miaow cbb0f11288 docs(audit): thirty-fourth batch — felis nano sweep (red 13 -> green 60/60), #55 bad-source ladder stall 2026-09-23 21:30:53 +08:00
Lemon-miaow 9dad61f508 fix(nano): screen unusable profiles per source instead of stopping the ladder (#55)
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.

Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.

Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
2026-09-23 21:21:21 +08:00
Lemon-miaow 971ae01caf docs(audit): thirty-third batch — /updates window API+UI sweep green, #53 repo move, #54 update guidance red/green
Records: (1) the /updates maintenance-window API sweep (unset null, 400s/415 for the invalid set, write->read-back->survives API pod restart, 401/403 auth) and the CDP panel sweep (status transitions, validation copy, clear, audit x3, zero console errors); (2) #53 red/green with the live fix53 binary ('MliroLirrorsIngenuity/Felis' -> 'FelisMC/Felis' in the 404 line) plus the token'd check (old path 301 / new path 200) and the fact the new home has no stable release yet; (3) #54 red/green with the live fix54 binary (installer one-liner + single trailer replace 'run: sudo felis setup'), the setup-vs-installer evidence, and the doc/test synchronization. Reachability table gains #53 (2) and #54 (2, docs); stats 54 total; repro-entry notes updated to the new remote, host binary v0.0.0+fix54 and the batch's artifacts.
2026-09-23 21:04:02 +08:00
Lemon-miaow 397a400d57 fix(update): point the apply guidance at the installer, not felis setup (#54)
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.

Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
2026-09-23 21:01:32 +08:00
Lemon-miaow 2c6739ad76 fix(updater,install,docs): follow the move to FelisMC/Felis (#53)
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.

Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
2026-09-23 20:58:34 +08:00
Lemon-miaow cec9a98305 docs(audit): thirty-second batch — S3 storage wizard sweep, #52 mirror lag red/green, cleanup 2026-09-23 20:39:09 +08:00
Lemon-miaow de7fb2c936 fix(setup): converge the workload felis-config mirror on every apply path (#52)
felis setup's in-TUI applies (storage / connection / edge) refreshed only the
control-namespace felis-config Secret; the workload-namespace mirror kept the
render from the previous run's startup pass until the next setup or installer
run. Found live: after 's -> Local' the minecraft copy still carried
user_uploads_context = s3://felis-wizard-uploads while the control copy and
both tomls were local. The 'configure email' path already overwrote both
mirrors, so storage/connection were the odd ones out.

Move the mirror refresh into applyFelisConfigSecret — the single choke point
every apply path calls — best-effort with a warning, since a control-plane
default install may not have the workload namespace at all. The smtp helper
drops its now-duplicate felis-config block.
2026-09-23 20:31:48 +08:00
Lemon-miaow 01988305a8 docs(audit): thirty-first batch — first-install walkthrough, #51 replica refresh red/green 2026-09-23 20:15:46 +08:00
Lemon-miaow 328e570309 fix(setup,install): refresh the workload namespace's felis-config mirror (#51) 2026-09-23 20:07:20 +08:00
Lemon-miaow cf5a790ea8 docs(audit): thirtieth batch — #50 smtp carry hoard, healed by two live re-runs 2026-09-23 19:57:00 +08:00
Lemon-miaow 4d3c85fd06 fix(bootstrap): [smtp] carry stops hoarding the auth_source comment block (#50) 2026-09-23 19:45:26 +08:00
Lemon-miaow 28fe7c43a9 docs(audit): twenty-ninth batch — image durability live drills; #46–#49
- registry hosting + loopback pull path landed (a9b275a/13d64e0/fa0e8d7) and
  drilled live: three installer re-runs, then GC simulations on the control
  plane (rolled felis-api pulled back in 25ms) and a game image (lobby-0,
  182MB in 10ms).
- #46 registry OOM (475MB-layer push killed the 256Mi template; dmesg evidence)
  fixed and re-verified: oom-kill count unchanged across a full rebuild+push.
- #47 AppleDouble ._*.sql embedding broke felis migrate on a Mac-staged tree;
  .dockerignore fix probed live with a planted junk file.
- #48 per-image docker start/stop tripped systemd start-limit-hit mid-batch;
  one wrap per batch, re-run mirrors all four.
- #49 installer re-runs silently reverted operator [registry]/[archive] config;
  carry-forward landed + live-verified into host toml, pod toml and the Secret,
  and the carried pins drove a successful POST /images/build.
- reachability table extended to #49 (① 20 | ② 19 | ③ 3+ | ④ 4 | 决策 3).
2026-09-23 19:35:13 +08:00
Lemon-miaow 72553cb414 docs(troubleshooting): the installer leaves docker stopped — start it before manual pushes
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
2026-09-23 19:35:13 +08:00
Lemon-miaow 20a95da487 docs(update): the plugins note is a rebuild + registry re-mirror now, not a node re-import
The felis-paper/felis-limbo jars are baked into the lobby/limbo images; with
the images hosted in the in-cluster registry, the extra step is pushing the
rebuilt image there (which is also what survives an image GC), not a bare
containerd import. The installer re-run does both.
2026-09-23 19:33:04 +08:00
Lemon-miaow c7e585e21d fix(bootstrap): mirror the image batch under ONE docker start/stop
Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.

push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
2026-09-23 19:23:44 +08:00
Lemon-miaow 5fa8b7412e fix(build): keep macOS ._*/.DS_Store junk out of the image
A Mac-staged tree (BSD tar materializes extended attributes as ._<name>
sidecars) went through the docker build and one landed in
internal/store/migrations/ — //go:embed-ed into the binary, where every
`felis migrate` then died with 'migration "._0004..." has a non-numeric
version'. Observed live wiring up the auditfix42 image: the installer's own
run_migrations failed on it. Exclude the sidecars and .DS_Store from the
build context; deploy/*.yaml and plugins/ have the same exposure.
2026-09-23 19:19:04 +08:00
Lemon-miaow b8e554dac7 fix(bootstrap): carry the operator's [archive] keys across re-runs too
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.

persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
2026-09-23 19:09:57 +08:00
Lemon-miaow 765a8923a4 fix(bootstrap): re-runs keep the operator's [registry] overrides
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.

Found while landing the registry-hosting work, which depends on those same
keys surviving.

- persisted_registry_block carries the operator-owned [registry] keys and the
  [registry.s3] subtable forward, same first-readable-file rule as
  persisted_smtp_block; url/build_namespace stay installer-owned (they must
  match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
  as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
  idempotence (a second re-run writes a byte-identical file).
2026-09-23 19:05:54 +08:00
Lemon-miaow fa0e8d7d97 docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
2026-09-23 19:03:11 +08:00
Lemon-miaow 13d64e0000 feat(bootstrap): host every built image in the internal registry — GC-durable pulls
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:

- Every image the installer builds is now a registry ref
  (registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
  containerd under that exact name (first boot needs no registry round-trip)
  and mirrored into the registry after deploy_bundle (push_image_to_registry:
  push endpoint 127.0.0.1:5000, and only the path after the host matters to the
  registry — a push there lands where kubelet's mirrored pull looks). A ref
  outside the registry is warned about, not silently unmirrored.

- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
  registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
  registry Deployment binds (node containerd cannot dial the Service VIP — live
  drill: "Empty reply"). k3s regenerates containerd config only at agent start,
  so a CONTENT change restarts k3s and an identical file (every re-run)
  restarts nothing.

- import_registry_image caches registry:2 into containerd so the registry
  Deployment can start on a box that cannot reach Docker Hub.

- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
  'felis-paper:demo') at the registry refs — a user server created from those
  rows must not strand when GC collects the bare tag. Only recommended rows
  still holding the old seed are touched; enabled is preserved; a pre-existing
  target row wins over a duplicate.

bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
2026-09-23 19:02:58 +08:00
Lemon-miaow a9b275abbb fix(platform): registry OOM (audit #46) + loopback hostPort — the node-side pull path
Two changes to the registry Deployment, both prerequisite to GC-durable images:

- Dedicated resource template: the control plane's 256Mi memory limit was a
  live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
  mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
  same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.

- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
  reach the Service VIP (live stack: "Empty reply"), so the node-side pull
  path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
  http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
  the plain-HTTP registry off every other interface.

Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
2026-09-23 18:55:23 +08:00
Lemon-miaow 92c06ac8dc docs(audit): twenty-eighth batch — #45 blind-review fix live-verified (context download, byte-exact); reachability 1-45 2026-09-23 17:01:19 +08:00