Six CR-level run/stop cycles with three chaos injections (operator pod
kill, postgres restart, api pod kill) all converged (Running 23-29s,
Stopped 3s, pods gone per cycle); PG outage keeps the 503-not-401
session semantics; 30-request burst all 200; ownerless wake correctly
403s. The build lane's missing Job TTL (#67, fixed in 2755e41) was
verified live on auditfix62: a real build Job carries ttl=604800 and an
old Job patched to ttl=30s was reaped, pod and all, within 40s.
The release workflow has never executed (no tags/releases exist yet).
Its three load-bearing contracts were verified against their other
halves: the asset names vs bootstrap's download_release_binary (and its
'felis ${REF}' convergence check), the local-exporter path vs the
Dockerfile's final COPY, and the version assertion vs cmdVersion's first
line. The one un-replicable check (file(1) machine type) was probed
with a cross-compiled linux/arm64 binary: 'ARM aarch64' matches. No
pre-tag fixes needed.
Module-by-module coverage matrix against the ledger: repo hygiene clean,
no TODO/FIXME debt, CLI dispatch test-guarded, CI syntax-checks every
tracked shell script, all 22 internal packages accounted for (CRD types
and game-image dirs covered indirectly), panel's 17 page routes drilled
in earlier batches. Fixes this batch: #66 README_EN was missing the
private-repo install workaround and the rerun/upgrade notes (an English
reader 404s on step one); the §28 claim/link sequence diagrams were
realigned with the audited implementations. A local-only scripts/
sync.sh conflict (plugins/ exclusion vs go:embed) was reproduced and
fixed on the workstation — private per .git/info/exclude, not committed.
The last module family nobody ever built (#64 gradlew exec bit, #65
license metadata), closed with three real dedicated-server boots
(fabric/forge/neoforge: mod loads, /link registers, console refused)
and a permanent JDK-17 gate in CI + release. Reachability table now
#1–#65; conclusions item 20; remaining-queue note updated. CI run
35892544563 = 5/5 with the new mods job green on its first run.
#63 fixed in f5a76cf (demo-up delegates to bootstrap; T1/T2/T3 real-machine);
plugins/test.sh + the new CI plugins job verified on the VM and in CI
(c59b387, run 35888009965, 4/4 jobs); #62 re-checked against all three
install paths and screened out as unreachable/nonexistent; plugin-layer
real-machine E2E via the MC status/login probes recorded.
Records: (1) the /updates maintenance-window API sweep (unset null, 400s/415 for the invalid set, write->read-back->survives API pod restart, 401/403 auth) and the CDP panel sweep (status transitions, validation copy, clear, audit x3, zero console errors); (2) #53 red/green with the live fix53 binary ('MliroLirrorsIngenuity/Felis' -> 'FelisMC/Felis' in the 404 line) plus the token'd check (old path 301 / new path 200) and the fact the new home has no stable release yet; (3) #54 red/green with the live fix54 binary (installer one-liner + single trailer replace 'run: sudo felis setup'), the setup-vs-installer evidence, and the doc/test synchronization. Reachability table gains #53 (2) and #54 (2, docs); stats 54 total; repro-entry notes updated to the new remote, host binary v0.0.0+fix54 and the batch's artifacts.
- registry hosting + loopback pull path landed (a9b275a/13d64e0/fa0e8d7) and
drilled live: three installer re-runs, then GC simulations on the control
plane (rolled felis-api pulled back in 25ms) and a game image (lobby-0,
182MB in 10ms).
- #46 registry OOM (475MB-layer push killed the 256Mi template; dmesg evidence)
fixed and re-verified: oom-kill count unchanged across a full rebuild+push.
- #47 AppleDouble ._*.sql embedding broke felis migrate on a Mac-staged tree;
.dockerignore fix probed live with a planted junk file.
- #48 per-image docker start/stop tripped systemd start-limit-hit mid-batch;
one wrap per batch, re-run mirrors all four.
- #49 installer re-runs silently reverted operator [registry]/[archive] config;
carry-forward landed + live-verified into host toml, pod toml and the Secret,
and the carried pins drove a successful POST /images/build.
- reachability table extended to #49 (① 20 | ② 19 | ③ 3+ | ④ 4 | 决策 3).
Records the Access-policy upsert fix and the first end-to-end runs of
account/migrate (all four steps + negative matrix + retire assertions),
passkey credential management, the access player-management group, the
updates window, and fleet — plus the environment restore notes.
RCON-secret deletion lockup and the stale start anchor (with its permanent
Provisioned=False) get their full live evidence trail, plus the sts
accidental-deletion drill. Deployed image note bumped to auditfix24.
Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
The auditfix20 image (8e7c7bb) was exercised against a real SMTP sink with a
dedicated warntest server: delivery content, tier precedence, dedupe, retry
semantics (bad relay / unverified email / unowned), threshold boundaries, and
zero backup side effects. Ledger #24 records the defect and the evidence; the
deployed image note is bumped to auditfix20.
The deferred-seams entry that waited for a real-Postgres harness is struck:
ClaimServer owns the gate now, proven red-then-green by the pgint concurrency
test (two claims, one win, one 403), and the storage-cache zeroing found in the
same pass is recorded with its hermetic test. auditfix19 is live on the VM.
#20–#22 recorded with live evidence: the pgint harness caught the
never-shipped verified-email index on its first run; migration 0020 applied
and the 409 drill replayed live; the admin email edit now drops the stale
proof; the owner role is written by both provisioning paths and its staff
doors, guards, and reclaim protections were drilled end to end on auditfix18.
The remaining-work item "PG-level contract tests" is checked off.
- The control-plane probes (0c8e29b) and the user-modpack build lane
(f79e5eb + 02fd2de) get their live evidence recorded, including the three
drill-only defects the lane fixed (Job scheduling, Kaniko Dockerfile
ownership, Trivy DB egress).
- The stale 'unfixed' rows #7/#11-#15 are reconciled with their commits.
- Remaining-work list re-stated: image durability, PG contract tests, the
panel's /jobs block, multi-node reaper placement, alerting, and the newly
found player-visible build status gap.
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
build_mem_limit overrides; empty keeps the compiled-in defaults. An
air-gapped or mirrored install has no route to gcr.io/aquasec (the
build egress policy allows only DNS + registry + package mirrors), so
builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
impossible (PVCs cannot cross namespaces) and that the s3 lane also
lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
13b rewritten (verified eviction refusal, 5m pressure-transition,
image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
volume (worlds + config + plugins + cache) and a restore rolls all of
it back — OpenAPI/README wording updated to match (same-tag images are
still watched for regressions by the openapi parity gate).
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.
Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
6-way concurrent redeem of one code 500'd on users_username_key (each request
generated a fresh user id but the same uuid-derived username), plus the rarer
two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
locks its code row and handles the conflict.
- SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
- INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
cross-code racers converge on the winner's row (role checked, staff still
refused) instead of a unique-violation 500.
- account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.
Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
db clean; zero unmapped errors.
The interface doc promised 'each joined to its staff username', the fake and
the pending handler both project username and created_at, but the PG query
selected neither — live internal /op-login/pending returned username:"" and
created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
tests can't see PG-only regressions.