Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.
push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
A Mac-staged tree (BSD tar materializes extended attributes as ._<name>
sidecars) went through the docker build and one landed in
internal/store/migrations/ — //go:embed-ed into the binary, where every
`felis migrate` then died with 'migration "._0004..." has a non-numeric
version'. Observed live wiring up the auditfix42 image: the installer's own
run_migrations failed on it. Exclude the sidecars and .DS_Store from the
build context; deploy/*.yaml and plugins/ have the same exposure.
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.
persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.
Found while landing the registry-hosting work, which depends on those same
keys surviving.
- persisted_registry_block carries the operator-owned [registry] keys and the
[registry.s3] subtable forward, same first-readable-file rule as
persisted_smtp_block; url/build_namespace stay installer-owned (they must
match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
idempotence (a second re-run writes a byte-identical file).
- §8e: the executor-mirror recipe now pushes into the internal registry (the
node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
instead of advising bare node-containerd imports — GC collects those and an
air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
registries.yaml mirror); keeps the operator checks (registry pod, mirror
file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
point felis.toml at the registry ref.
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:
- Every image the installer builds is now a registry ref
(registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
containerd under that exact name (first boot needs no registry round-trip)
and mirrored into the registry after deploy_bundle (push_image_to_registry:
push endpoint 127.0.0.1:5000, and only the path after the host matters to the
registry — a push there lands where kubelet's mirrored pull looks). A ref
outside the registry is warned about, not silently unmirrored.
- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
registry Deployment binds (node containerd cannot dial the Service VIP — live
drill: "Empty reply"). k3s regenerates containerd config only at agent start,
so a CONTENT change restarts k3s and an identical file (every re-run)
restarts nothing.
- import_registry_image caches registry:2 into containerd so the registry
Deployment can start on a box that cannot reach Docker Hub.
- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
'felis-paper:demo') at the registry refs — a user server created from those
rows must not strand when GC collects the bare tag. Only recommended rows
still holding the old seed are touched; enabled is preserved; a pre-existing
target row wins over a duplicate.
bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
Two changes to the registry Deployment, both prerequisite to GC-durable images:
- Dedicated resource template: the control plane's 256Mi memory limit was a
live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.
- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
reach the Service VIP (live stack: "Empty reply"), so the node-side pull
path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
the plain-HTTP registry off every other interface.
Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
Two defects from the live Sync drill:
- The picker listed the system servers (login/lobby), which the backup API can
never accept (reserved names, no servers row): the pick died in name
validation with a raw "server name is reserved" error. backupPickable now
filters them out; the halt picker keeps them on purpose (break-glass retains
full power over system servers).
- backupErrorFromResponse mapped every 409 to the stopped gate, so the new
world-volume refusal would have displayed the wrong reason. The 409 arm now
keys on the body's error code; a code-less body still reads as the stopped
gate.
Live (auditfix38): the picker shows only user servers; a world-less pick shows
the API's own "no world volume yet — start it once" text; the not_stopped text
is unchanged.
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.
Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.
Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
Two defects live-drilled in the break-glass staff provisioning:
- An Owner reset that typed any username other than the occupied seat took
UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
existing seat — possibly the compromised account the reset was meant to
replace — live; every owner row is undeletable through the panel, so the tier
could never converge back to one. provisionOwner now refuses with
ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
form); bootstrap still mints, and the seat's own username still resets in
place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
console keys its rename prompt off api.ErrConflict — the "choose another
name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
the contract; PGRepo had drifted). Map the unique violation to ErrConflict
and pin it in pgint.
Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
The backend could list/read/write a stopped server's world volume since the
file-editor slice, but the panel had no entry, so the one repair path for a
server that will not boot (a wrong line in server.properties) was API-only.
New /servers/:name/files page: breadcrumb browser, editor dialog with the
base64 []byte codec, binary files open read-only, the stopped gate is owned
up front (with a stop action) instead of letting every call 409, and a
doorway card on the console. i18n files namespace + wire-shape tests.
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
DELETE /users/{id}/passkeys shipped as the owner-tier remediation for a
lost or compromised authenticator, but nothing in the panel reached it.
Add the danger-zone action with a confirm dialog; the account keeps its
other doors (email OTP, in-game op-login re-enrollment), so this severs
a credential without locking anyone out. Wire-shape test pins the call.
The backups page could list and restore archives but not create one,
and nothing surfaced backup/restore Job outcomes — a failed 202 was
visible only through kubectl. Add a Back up now action (enabled only on
a stopped server, the backend's own gate; a raced 409 is surfaced in
its words) and a Recent operations card fed by GET /servers/{name}/jobs
that shows running/succeeded/failed with the Job's failure message,
re-reads on an interval while a Job is running, and persists across
reloads. Wire-shape tests pin both endpoints.
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.
Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
UserByMCUUID now resolves only live accounts: claim, menu, wake
authorization, op-login vouch and the QR link-status poll treat a
disabled or soft-deleted link holder exactly like an unlinked UUID
instead of a retired identity. VerifyLinkCode lets a soft-deleted
link be taken over by a fresh in-game code (the deleted account is
gone, e.g. a migrated source), while a disabled holder still 409s so
the lockout is not bypassable; failed attempts still do not consume
the code. Fake repo and pgint coverage pin both branches.
'smtpSecretManifest' hardcoded namespace=felis, so the 'configure email'
refresh of the minecraft-namespace copies failed before it began: kubectl
refuses a manifest whose namespace conflicts with -n (found live: 'the
namespace from the provided object "felis" does not match the namespace
"minecraft"'), and the felis-config mirror never ran at all because the
smtp apply returned early. A later SMTP change could therefore never reach
the reaper's pre-reap warnings — the exact failure the refresh was added to
close.
Render the Secret with the caller's namespace (felis for the control-plane
apply, the workload namespace for the mirror). Regression test pins both.
Records the Access-policy upsert fix and the first end-to-end runs of
account/migrate (all four steps + negative matrix + retire assertions),
passkey credential management, the access player-management group, the
updates window, and fleet — plus the environment restore notes.
The comment described a deterministic-name collision that the unique random
suffix made near-impossible; align it with Backuper.Backup and jobspec's
contract (ErrAlreadyExists survives only as the defensive no-op).
CreateAccessPolicy treated a Cloudflare "policy_already_exists" as idempotent
success and kept whatever policy was there. On a re-run with a changed
identity — or against a hand-made broader policy — op.console would stay
guarded by something weaker than the fail-closed body this package builds and
guards, while Setup reported success. The fail-closed validation only ever ran
on the policy we built, never on the one that stayed live.
Now it upserts by name: lookup, PUT the guarded body over the existing policy,
POST only when absent (a racing POST re-looks up and PUTs). apiPost/apiPut
share one apiWrite; three httptest cases pin update-over-existing, create-when-
absent, and the race fallback.
RCON-secret deletion lockup and the stale start anchor (with its permanent
Provisioned=False) get their full live evidence trail, plus the sts
accidental-deletion drill. Deployed image note bumped to auditfix24.