README.md grew a required workaround — the repository is private, so the
plain raw.githubusercontent one-liner returns 404. The credentialed form
(token handed to curl via --config - so it never touches argv, sudo -E so
the installer inherits it) and the rerun/upgrade notes (the rerun upgrades
felis-api; the channel is not inherited, so main-followers need
FELIS_VERSION_BOOTSTRAP=dev) never made it into README_EN.md, which is
linked as the English entry point. An English reader following that page
could not install at all. Bring it back in sync with the Chinese one.
The last module family nobody ever built (#64 gradlew exec bit, #65
license metadata), closed with three real dedicated-server boots
(fabric/forge/neoforge: mod loads, /link registers, console refused)
and a permanent JDK-17 gate in CI + release. Reachability table now
#1–#65; conclusions item 20; remaining-queue note updated. CI run
35892544563 = 5/5 with the new mods job green on its first run.
Nothing ever compiled these modules — no CI job, no install path — which is
why the gradlew exec-bit bug (previous commit) shipped unnoticed. Add
plugins/test-mods.sh: runs each module's vendored wrapper under JDK 17
(they target the Java-17 Minecraft lines; paper/limbo stay on the JDK 21
gate). ci.yml gains a `mods` job (temurin 17 + setup-gradle, wrapper
pinned per module), release.yml gates both plugin gates before shipping.
README: status now records the server-boot verification, the Building
section pins limbo's `-PlimboVersion=<release>` (the `+` default is
unresolvable from the LOOHP repo) and documents both gates.
The three loader mods declared `license = "MIT"` (and Forge/NeoForge a
placeholder `issueTrackerURL = https://example.invalid/felis`) while the
repository is AGPL-3.0-only (README.md:73, LICENSE). The mods were added
2026-06-26, the LICENSE landed 2026-07-12 — stale leftovers that nothing
ever read back: no CI job, no install path. Fabric loader prints the
license from fabric.mod.json at boot and both mods.toml files are parsed
by their loaders, so the wrong claim is user-visible. Align both with
reality — boot-verified on real fabric/forge/neoforge dedicated servers
(AUDIT-2026-09-22.md, batch 41).
All three vendored wrappers were tracked as 100644, so the README's documented
'plugins/<loader>/gradlew -p plugins/<loader> build' failed on every fresh
clone with 'Permission denied' (exit 126) — and since no install path or CI job
ever ran them, nothing caught it.
git update-index --chmod=+x for the three files; the VM then built all three
modules for the first time (results in the gate commit and the audit).
#63 fixed in f5a76cf (demo-up delegates to bootstrap; T1/T2/T3 real-machine);
plugins/test.sh + the new CI plugins job verified on the VM and in CI
(c59b387, run 35888009965, 4/4 jobs); #62 re-checked against all three
install paths and screened out as unreachable/nonexistent; plugin-layer
real-machine E2E via the MC status/login probes recorded.
The three framework-free test mains under plugins/*/test were never run by
anything — not CI, not the plugin builds — and the velocity/paper/limbo jars
were only ever compiled by deploy/bootstrap.sh on a live host. CI gains a
'plugins' job (JDK 21 plus the Gradle 8.14 the plugin Dockerfiles pin) running
plugins/test.sh: the three mains (InviteCardTest's jars fetched from Maven
Central, pinned and digest-checked) and the three production builds.
The first real run surfaced and fixed two untested assumptions: InviteCardTest's
documented javac line omitted examination-api (adventure-api's Component
signatures reference Examinable, so javac needs it too), and limbo's '+' version
default cannot resolve — LOOHP's repository serves no maven-metadata — so the
script resolves the current release off the Limbo CI artifact name (the same
source bootstrap reads) and plugins/README.md stops advertising a bare
'gradle -p plugins/limbo build' that can never work.
Verified in gradle:8.14-jdk21 on the VM: mains OK (32/36/48 checks);
velocity/paper/limbo BUILD SUCCESSFUL.
demo-up.sh grew its own image lane in July, before bootstrap could build the
game stack. The copy has since drifted from the installer it duplicates: it
pinned Paper 1.21.8 while bootstrap derives one version from the Limbo login
gate (both hops of a login must speak one protocol), it never built
felis-velocity.jar (a proxy without it silently routes nothing), it imported
images under local tags the [velocity] wiring no longer points at, and it left
the docker daemon running.
It now runs deploy/bootstrap.sh (unless SKIP_BOOTSTRAP=1), checks that the
earlier run really left the full stack behind, and hands over to 'felis setup'
— so a half-installed base is a clear error instead of a proxy that accepts
logins and routes nowhere.
Verified on the VM: healthy base passes to the handoff (exit 0); a hidden
felis-velocity.jar dies naming the jar and the remedy (exit 1, jar restored).
Live with real LuckPerms 5.5.85: every lp command returns an empty RCON body
(list/plugins answer normally; a standalone RCON client sees the same, and
creategroup/permission-set still persist), so the read projection can never
populate and the page asserted "no parent groups / no explicit nodes" for a
state it could not actually read. The raw reply now rides the same disclosure
the rosters carry, a silent entry-less reply shows an explicit notice instead
of the false empty claims, and the write history's placeholder no longer
dresses up a fabricated "[RCON] ..." line as output.
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
sendCommand returns the RCON reply and the component's own contract says it
"displays its plain-text reply", but the reply was discarded and the pod log
does not echo command output, so a sent command produced no visible result at
all (verified live with 'list'). The last command echo + reply now render
above the prompt, terminal-style.
The whitelist/ban/kick mutations reported a canned success message and threw
away the server's reply, so a refused command still read as done: live, the
vanilla server answers "That player does not exist" for a name it has never
seen (any player who has not joined yet), while the panel said the player had
been whitelisted/banned. api.ts documents these replies as "surfaced verbatim
as confirmation"; now they are. The localized string stays as the fallback
for a silent server.
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.
Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.
Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).
Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
Records: (1) the /updates maintenance-window API sweep (unset null, 400s/415 for the invalid set, write->read-back->survives API pod restart, 401/403 auth) and the CDP panel sweep (status transitions, validation copy, clear, audit x3, zero console errors); (2) #53 red/green with the live fix53 binary ('MliroLirrorsIngenuity/Felis' -> 'FelisMC/Felis' in the 404 line) plus the token'd check (old path 301 / new path 200) and the fact the new home has no stable release yet; (3) #54 red/green with the live fix54 binary (installer one-liner + single trailer replace 'run: sudo felis setup'), the setup-vs-installer evidence, and the doc/test synchronization. Reachability table gains #53 (2) and #54 (2, docs); stats 54 total; repro-entry notes updated to the new remote, host binary v0.0.0+fix54 and the batch's artifacts.
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.
Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.
Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.
Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).
Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
felis setup's in-TUI applies (storage / connection / edge) refreshed only the
control-namespace felis-config Secret; the workload-namespace mirror kept the
render from the previous run's startup pass until the next setup or installer
run. Found live: after 's -> Local' the minecraft copy still carried
user_uploads_context = s3://felis-wizard-uploads while the control copy and
both tomls were local. The 'configure email' path already overwrote both
mirrors, so storage/connection were the odd ones out.
Move the mirror refresh into applyFelisConfigSecret — the single choke point
every apply path calls — best-effort with a warning, since a control-plane
default install may not have the workload namespace at all. The smtp helper
drops its now-duplicate felis-config block.
- registry hosting + loopback pull path landed (a9b275a/13d64e0/fa0e8d7) and
drilled live: three installer re-runs, then GC simulations on the control
plane (rolled felis-api pulled back in 25ms) and a game image (lobby-0,
182MB in 10ms).
- #46 registry OOM (475MB-layer push killed the 256Mi template; dmesg evidence)
fixed and re-verified: oom-kill count unchanged across a full rebuild+push.
- #47 AppleDouble ._*.sql embedding broke felis migrate on a Mac-staged tree;
.dockerignore fix probed live with a planted junk file.
- #48 per-image docker start/stop tripped systemd start-limit-hit mid-batch;
one wrap per batch, re-run mirrors all four.
- #49 installer re-runs silently reverted operator [registry]/[archive] config;
carry-forward landed + live-verified into host toml, pod toml and the Secret,
and the carried pins drove a successful POST /images/build.
- reachability table extended to #49 (① 20 | ② 19 | ③ 3+ | ④ 4 | 决策 3).
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
The felis-paper/felis-limbo jars are baked into the lobby/limbo images; with
the images hosted in the in-cluster registry, the extra step is pushing the
rebuilt image there (which is also what survives an image GC), not a bare
containerd import. The installer re-run does both.
Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.
push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
A Mac-staged tree (BSD tar materializes extended attributes as ._<name>
sidecars) went through the docker build and one landed in
internal/store/migrations/ — //go:embed-ed into the binary, where every
`felis migrate` then died with 'migration "._0004..." has a non-numeric
version'. Observed live wiring up the auditfix42 image: the installer's own
run_migrations failed on it. Exclude the sidecars and .DS_Store from the
build context; deploy/*.yaml and plugins/ have the same exposure.
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.
persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.
Found while landing the registry-hosting work, which depends on those same
keys surviving.
- persisted_registry_block carries the operator-owned [registry] keys and the
[registry.s3] subtable forward, same first-readable-file rule as
persisted_smtp_block; url/build_namespace stay installer-owned (they must
match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
idempotence (a second re-run writes a byte-identical file).
- §8e: the executor-mirror recipe now pushes into the internal registry (the
node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
instead of advising bare node-containerd imports — GC collects those and an
air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
registries.yaml mirror); keeps the operator checks (registry pod, mirror
file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
point felis.toml at the registry ref.
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:
- Every image the installer builds is now a registry ref
(registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
containerd under that exact name (first boot needs no registry round-trip)
and mirrored into the registry after deploy_bundle (push_image_to_registry:
push endpoint 127.0.0.1:5000, and only the path after the host matters to the
registry — a push there lands where kubelet's mirrored pull looks). A ref
outside the registry is warned about, not silently unmirrored.
- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
registry Deployment binds (node containerd cannot dial the Service VIP — live
drill: "Empty reply"). k3s regenerates containerd config only at agent start,
so a CONTENT change restarts k3s and an identical file (every re-run)
restarts nothing.
- import_registry_image caches registry:2 into containerd so the registry
Deployment can start on a box that cannot reach Docker Hub.
- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
'felis-paper:demo') at the registry refs — a user server created from those
rows must not strand when GC collects the bare tag. Only recommended rows
still holding the old seed are touched; enabled is preserved; a pre-existing
target row wins over a duplicate.
bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
Two changes to the registry Deployment, both prerequisite to GC-durable images:
- Dedicated resource template: the control plane's 256Mi memory limit was a
live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.
- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
reach the Service VIP (live stack: "Empty reply"), so the node-side pull
path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
the plain-HTTP registry off every other interface.
Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
Two defects from the live Sync drill:
- The picker listed the system servers (login/lobby), which the backup API can
never accept (reserved names, no servers row): the pick died in name
validation with a raw "server name is reserved" error. backupPickable now
filters them out; the halt picker keeps them on purpose (break-glass retains
full power over system servers).
- backupErrorFromResponse mapped every 409 to the stopped gate, so the new
world-volume refusal would have displayed the wrong reason. The 409 arm now
keys on the body's error code; a code-less body still reads as the stopped
gate.
Live (auditfix38): the picker shows only user servers; a world-less pick shows
the API's own "no world volume yet — start it once" text; the not_stopped text
is unchanged.
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.
Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.
Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
Two defects live-drilled in the break-glass staff provisioning:
- An Owner reset that typed any username other than the occupied seat took
UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
existing seat — possibly the compromised account the reset was meant to
replace — live; every owner row is undeletable through the panel, so the tier
could never converge back to one. provisionOwner now refuses with
ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
form); bootstrap still mints, and the seat's own username still resets in
place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
console keys its rename prompt off api.ErrConflict — the "choose another
name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
the contract; PGRepo had drifted). Map the unique violation to ErrConflict
and pin it in pgint.
Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
The backend could list/read/write a stopped server's world volume since the
file-editor slice, but the panel had no entry, so the one repair path for a
server that will not boot (a wrong line in server.properties) was API-only.
New /servers/:name/files page: breadcrumb browser, editor dialog with the
base64 []byte codec, binary files open read-only, the stopped gate is owned
up front (with a stop action) instead of letting every call 409, and a
doorway card on the console. i18n files namespace + wire-shape tests.
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
DELETE /users/{id}/passkeys shipped as the owner-tier remediation for a
lost or compromised authenticator, but nothing in the panel reached it.
Add the danger-zone action with a confirm dialog; the account keeps its
other doors (email OTP, in-game op-login re-enrollment), so this severs
a credential without locking anyone out. Wire-shape test pins the call.
The backups page could list and restore archives but not create one,
and nothing surfaced backup/restore Job outcomes — a failed 202 was
visible only through kubectl. Add a Back up now action (enabled only on
a stopped server, the backend's own gate; a raced 409 is surfaced in
its words) and a Recent operations card fed by GET /servers/{name}/jobs
that shows running/succeeded/failed with the Job's failure message,
re-reads on an interval while a Job is running, and persists across
reloads. Wire-shape tests pin both endpoints.
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.
Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
UserByMCUUID now resolves only live accounts: claim, menu, wake
authorization, op-login vouch and the QR link-status poll treat a
disabled or soft-deleted link holder exactly like an unlinked UUID
instead of a retired identity. VerifyLinkCode lets a soft-deleted
link be taken over by a fresh in-game code (the deleted account is
gone, e.g. a migrated source), while a disabled holder still 409s so
the lockout is not bypassable; failed attempts still do not consume
the code. Fake repo and pgint coverage pin both branches.
'smtpSecretManifest' hardcoded namespace=felis, so the 'configure email'
refresh of the minecraft-namespace copies failed before it began: kubectl
refuses a manifest whose namespace conflicts with -n (found live: 'the
namespace from the provided object "felis" does not match the namespace
"minecraft"'), and the felis-config mirror never ran at all because the
smtp apply returned early. A later SMTP change could therefore never reach
the reaper's pre-reap warnings — the exact failure the refresh was added to
close.
Render the Secret with the caller's namespace (felis for the control-plane
apply, the workload namespace for the mirror). Regression test pins both.
Records the Access-policy upsert fix and the first end-to-end runs of
account/migrate (all four steps + negative matrix + retire assertions),
passkey credential management, the access player-management group, the
updates window, and fleet — plus the environment restore notes.
The comment described a deterministic-name collision that the unique random
suffix made near-impossible; align it with Backuper.Backup and jobspec's
contract (ErrAlreadyExists survives only as the defensive no-op).
CreateAccessPolicy treated a Cloudflare "policy_already_exists" as idempotent
success and kept whatever policy was there. On a re-run with a changed
identity — or against a hand-made broader policy — op.console would stay
guarded by something weaker than the fail-closed body this package builds and
guards, while Setup reported success. The fail-closed validation only ever ran
on the policy we built, never on the one that stayed live.
Now it upserts by name: lookup, PUT the guarded body over the existing policy,
POST only when absent (a racing POST re-looks up and PUTs). apiPost/apiPut
share one apiWrite; three httptest cases pin update-over-existing, create-when-
absent, and the race fallback.
RCON-secret deletion lockup and the stale start anchor (with its permanent
Provisioned=False) get their full live evidence trail, plus the sts
accidental-deletion drill. Deployed image note bumped to auditfix24.
Found live while validating the RCON-secret heal: a server that had already
recovered to Ready was marked Failed(StartupTimeout) minutes later, the moment
an unrelated pod rollout briefly dropped readyReplicas. The anchor
(status.startRequestedAt) was never cleared on success, so its 300s budget
kept ticking under a healthy server and any later blip spent it.
markRunningReady now clears the anchor: every start-or-recovery attempt gets
its own budget. It also flips ConditionProvisioned back to True — markFailed
sets it False and nothing ever reset it, leaving a permanent failure flag on
recovered servers that every conditions consumer would read.
Unit tests pin both: anchor cleared on Ready, Provisioned recovers from
Failed to Running.
Deleting the per-server RCON Secret used to leave a running pod authenticating
with the lost password while the operator re-minted a fresh one and probed
with it: the RCON gate failed forever (live: 96s+ of RconNotReachable, headed
for ReadinessTimeout) and nothing re-triggered a pod restart — the server only
came back when the pod was deleted by hand.
Two changes pair up:
- Owns(&corev1.Secret{}) so the deletion is noticed at all (a quiet Running
server emits no other events; the Secret is controller-owned, so the watch
maps it back to the CR).
- The pod template now carries a fingerprint of the current password
(RconSecretAnnotation). Re-creation changes the fingerprint, the
StatefulSet rolls, and the new pod picks the new password up; while the
Secret is untouched the value is stable so no spurious rolls.
Unit tests pin stability across reconciles and the change-on-recreation roll.
Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:
- The stop used a whole-object Update while the same reconcile loop writes
status; that risks clobbering a concurrent status write. Switch to the
reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
closed with 403 (visible in the operator log as 'cannot update resource
"minecraftservers"'). Grant patch and pin it in the RBAC scope test.
With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
EmptySince was stamped and then never revisited: player joins/leaves do not
touch the CRD, RCON is only probed inside Reconcile, and a steady Running
server produces no watch events (its status update goes out unchanged and is a
no-op). Live, an empty server with a 30s grace sat Running for minutes with
zero reconciles in the log — the auto-stop existed only on paper.
reconcileRunning now returns a RequeueAfter for idle-enabled servers: exactly
at the deadline while empty, or a 30s probe cadence while occupied so the
moment the last player leaves is noticed. New unit tests pin all three:
deadline requeue, occupied cadence, and no requeue when disabled.
The operator stamps EmptySince to time the idle auto-stop window, but the CRD's
status schema never declared it. Kubernetes pruned the field on every write
(apiserver warning: unknown field "status.emptySince"), so the timer reset to
nil on every read and idle auto-stop could never fire — a defect invisible to
the fake-client unit tests, which do not enforce the CRD schema. Found live:
the stamp was silently dropped the moment it was set.
Schema now declares emptySince (date-time) like its sibling timestamps.
The auditfix20 image (8e7c7bb) was exercised against a real SMTP sink with a
dedicated warntest server: delivery content, tier precedence, dedupe, retry
semantics (bad relay / unverified email / unowned), threshold boundaries, and
zero backup side effects. Ledger #24 records the defect and the evidence; the
deployed image note is bumped to auditfix20.
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:
- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
`warning suppressed — no warner wired` and does NOT stamp; a delivery error
logs and retries on the next daily run (bounded by the warning window). The
stamps are no longer burned by notices nobody received.
- A real channel: mail.SendNotice (the second and last message shape the mail
package sends) plus a mailWarner that resolves the owner's VERIFIED email
and mails the notice through the configured [smtp] relay. `felis reaper`
wires it when [smtp] is set (same password_ref convention as felis-api) and
prints exactly what happens when it is not.
Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.
Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
The deferred-seams entry that waited for a real-Postgres harness is struck:
ClaimServer owns the gate now, proven red-then-green by the pgint concurrency
test (two claims, one win, one 403), and the storage-cache zeroing found in the
same pass is recorded with its hermetic test. auditfix19 is live on the VM.
Two defects in the §9.3 quota path, both invisible to the hermetic suite:
- Audit #4's TOCTOU was real and documented: QuotaCheck and ClaimServer were
separate statements, so two concurrent claims by one user for two different
ownerless servers both read count < max_servers and both won. The gate now
lives inside ClaimServer, in the SAME transaction as the ownership write,
under pg_advisory_xact_lock(hashtext(user_id)) — the aggregate read, the
four-dimension re-check (shared with QuotaCheck via one helper so the two
cannot drift), and the UPDATE are one serialized decision. The loser gets
ErrQuotaExceeded, which both claim handlers map to the same 403 the
sequential path gives; the server row is additionally taken FOR UPDATE so
same-server races still resolve to exactly one winner.
- The server PATCH path called UpdateServerResources(..., 0) for storage even
though a resources patch cannot change storage. The cached columns are the
ONLY input to the quota aggregate, so every resource patch silently dropped
that server's storage contribution from its owner's cap. The handler now
reads the current spec and passes storage through.
Red-then-green: the new pgint test drives two real concurrent claims against
max_servers=1 (before: both win; now: exactly one win + one gated 403, and the
DB shows one owned row); the hermetic suite pins the 403 mapping and the
storage-preserving cache write.
#20–#22 recorded with live evidence: the pgint harness caught the
never-shipped verified-email index on its first run; migration 0020 applied
and the 409 drill replayed live; the admin email edit now drops the stale
proof; the owner role is written by both provisioning paths and its staff
doors, guards, and reclaim protections were drilled end to end on auditfix18.
The remaining-work item "PG-level contract tests" is checked off.
Found live while verifying the admin email-edit fix: the Owner account could
not load /api/v1/users at all. Root cause: migration 0011 adds the 'owner'
role and gates every user-administration route on it, but NOTHING ever wrote
it. break-glass (UpsertOwner), the setup MC-bind (CompleteOwnerSetup), and the
re-provision path all forced 'admin', so in a fresh install the entire
owner tier — list/create/edit/disable/delete users, quotas, sessions — was
unreachable. The role was a dead letter in the other direction too: staff
predicates that predate the role did not know it.
- UpsertOwner and CompleteOwnerSetup now write role='owner'; the username-
conflict arm re-asserts it, which is also the documented pre-0011 promotion
path ("re-provision via break-glass"). InsertOperator stays plain 'admin'.
- Staff doors learn the role: op-login start/finish admit the Owner; the
player email door refuses it like any staff account; the in-game approver
check already used staffRole.
- Reclaim protection: IsProtectedAdminLink (and the break-glass bootstrap
switch AdminExists) count admin OR owner — the Owner must never be displaced
by a Mojang-priority reclaim.
- Panel guards make migration 0011's claim true now that owner rows exist: an
owner can never be demoted, deleted, or disabled through the API (only the
local break-glass console resets the identity); username/email edits still
work.
Tests: pgint pins both provisioning paths, the protected-link predicate and
the reset/promote semantics; hermetic suites cover the owner-admitting staff
door, the owner-refusing player door, the three panel guards, and break-glass
attribution.
UpdateUser wrote a new address but kept email_verified, so patching a verified
account asserted a proof of an address nobody had proven — and the
pre-session login mails and resolves on exactly that flag, so a typo'd edit
could hand the account's sign-in codes to the wrong mailbox.
Changing the address now clears the flag in the same write; a no-op patch that
passes the same value keeps it. The fake mirrors the semantics, and the pgint
suite pins both halves (same value keeps proof, new value drops it).
The design has claimed since migration 0010 that at most one account can hold
a PROVEN email address, with ErrEmailTaken as the 409 a second verifier sees.
Neither half ever shipped: no migration created users_verified_email_unique,
and VerifyEmailOTP had no guard at all — the sentinel was defined but never
returned, so two accounts could both verify one address. The damage is not
cosmetic: the pre-session login resolves accounts BY verified email, so the
duplicate decided which identity a mailed sign-in code belonged to.
- Migration 0020 creates the partial unique index (lower(email) WHERE
email_verified) the comments have been citing — the database-level backstop.
- VerifyEmailOTP now refuses the take-over with ErrEmailTaken BEFORE consuming
the code (the address, not the code, is the problem), charges no attempt,
and maps a lost cross-user race (unique violation) to the same answer.
- The verify handler answers 409 email_taken instead of a generic 500.
Covered by the pgint suite (sequential double-verify refused with the code
still live, a direct duplicate write still loses to the index, the refused
account can still prove its own address) and a hermetic 409 case.
The hermetic suites encode the store contracts against fakes; PGRepo drifted
behind them three times (attempt accounting, a missing JOIN, a missing FOR
UPDATE) while every unit test stayed green. This harness replays the real
embedded migrations onto a throwaway database — its name must contain "pgint"
or the harness refuses to run — and exercises the SQL directly: sessions, the
onboarding email-OTP lifecycle (supersede/expiry/lockout), the pre-session
login consume, the op-login state machine, link and bind-code redemption,
submissions, and builds with the image admission round trip.
Run it after touching SQL under internal/api/pgrepo.go, internal/submit, or
internal/build; CONTRIBUTING.md carries the one-liner.
- The control-plane probes (0c8e29b) and the user-modpack build lane
(f79e5eb + 02fd2de) get their live evidence recorded, including the three
drill-only defects the lane fixed (Job scheduling, Kaniko Dockerfile
ownership, Trivy DB egress).
- The stale 'unfixed' rows #7/#11-#15 are reconciled with their commits.
- Remaining-work list re-stated: image durability, PG contract tests, the
panel's /jobs block, multi-node reaper placement, alerting, and the newly
found player-visible build status gap.
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:
- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
Pod never scheduled on the platform's own starter node: FailedScheduling /
Insufficient memory, Pending forever. Requests are now a small floor
(250m / 512Mi, never above a configured cap) while the limits stay the
safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
the source owner; a 65532-owned context (the distroless felis image uid)
fails that under the pod's dropped capabilities ('copying dockerfile:
chown /kaniko/Dockerfile: operation not permitted'). The fetch container
now extracts as root — the uid Kaniko already runs as — so the copy
succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
renders --db-repository, and docs/troubleshooting.md §8e now carries the
verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
internal registry; --insecure already covers its plain HTTP).
Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.