Nothing ever removed a submission: users could not retract a pending row, no
route deleted blobs or rows, and the reaper never touches uploads — so every
upload accumulated on the 5 GiB PVC forever and the only cleanup was SQL or
kubectl against the store.
- Blobs.Delete on both transports (local: RemoveAll of the id-namespaced dir,
id re-validated at the boundary; S3: idempotent object DELETE).
- Store: DeleteSubmission (admin, any status) and DeletePendingSubmission
(owner+pending CAS — a reviewed row can never be withdrawn out from under
its build).
- Manager.Delete / Manager.Withdraw delete the ROW first (under the CAS for
withdraw) and the blob after, so a live row can never point at a reaped
blob; a cleanup failure names the orphan explicitly instead of failing mute.
- API: DELETE /me/submissions/{id} (withdraw, app tier) and
DELETE /api/v1/submissions/{id} (admin) both return the row as it was;
audit events submission.withdraw / submission.delete; openapi documents both
paths; admin route pinned in the admin-only table.
- Panel: two-step withdraw on a pending row (frees the pending slot and the
storage budget); two-step delete on every admin row; zh/en copy; wire tests.
Unit: submit (withdraw happy path / wrong owner / reviewed row / no transport /
blob-cleanup failure), local+S3 delete idempotence, api handlers (200/404/409/
503 + route tier); pgint: withdraw CAS + admin delete exactly-once.
go vet/go test/gofmt clean; panel vitest 120 + typecheck green.
A logged-in user could file submissions without bound and stream a 1 GiB
context per submission. The only limits were the single-blob size cap and the
5 GiB uploads PVC (platform/workloads.go); nothing counted a user's rows or
bytes, so one account could fill the volume and every other user's upload
would start failing.
- Create: per-user pending_review cap (default 5) — the review queue cannot
be parked full of one account's rows. Check-then-insert, documented soft.
- UploadContext: per-user stored-context budget (default 2 GiB) charged
against the blob store's REAL sizes (new Blobs.Size on local/S3 stores), so
the sum cannot drift from the volume; the write is capped at the remaining
budget, so the excess is refused before it is persisted, and a re-upload is
charged only for its new bytes.
- API: per-user create/upload throttles (30s/15s, cmd/felis-wired) on a
dedicated cooldown keyspace, reserve→release so a failed attempt never
burns the window and a burst collapses to one winner; ErrQuotaExceeded →
403 submission_quota_exceeded (distinct from the 400 an oversize blob
gets), 429 submission_cooldown for the throttles.
- Panel: zh/en copy for both codes; openapi documents 403/429 on the two
user routes; pgint covers the pending-queue count.
Unit tests: submit package (cap, budget boundary/exact-fit/replacement,
oversize-vs-quota split) and api handlers (quota 403 both paths, throttle
429 + recovery + failure-release). go vet/go test/gofmt clean; panel
vitest 118 + typecheck green.
Every other Job family carries a TTLSecondsAfterFinished (fileedit 2m,
backup/restore 10m) but the build lane never set one: each build left a
completed Job + Pod in felis-build indefinitely (5 already on the drill
cluster, oldest 26h), growing etcd and — since completed pods count
against the node's pod budget (110 on stock k3s) — eventually blocking
new builds. The code even anticipated GC it never got ("JobUnknown
means the Job was not found (e.g. GC'd)").
Set a deliberately long TTL (7 days): the kaniko log is the admin
failure-triage surface (GET /images/build/{id}/logs), so the window
keeps a week of logs while bounding steady-state pods; terminal builds
are idempotent under Sync, so a late log-404 is the only cost.
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.
Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.
Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).
Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.
Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).
Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:
- Every image the installer builds is now a registry ref
(registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
containerd under that exact name (first boot needs no registry round-trip)
and mirrored into the registry after deploy_bundle (push_image_to_registry:
push endpoint 127.0.0.1:5000, and only the path after the host matters to the
registry — a push there lands where kubelet's mirrored pull looks). A ref
outside the registry is warned about, not silently unmirrored.
- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
registry Deployment binds (node containerd cannot dial the Service VIP — live
drill: "Empty reply"). k3s regenerates containerd config only at agent start,
so a CONTENT change restarts k3s and an identical file (every re-run)
restarts nothing.
- import_registry_image caches registry:2 into containerd so the registry
Deployment can start on a box that cannot reach Docker Hub.
- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
'felis-paper:demo') at the registry refs — a user server created from those
rows must not strand when GC collects the bare tag. Only recommended rows
still holding the old seed are touched; enabled is preserved; a pre-existing
target row wins over a duplicate.
bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
Two changes to the registry Deployment, both prerequisite to GC-durable images:
- Dedicated resource template: the control plane's 256Mi memory limit was a
live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.
- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
reach the Service VIP (live stack: "Empty reply"), so the node-side pull
path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
the plain-HTTP registry off every other interface.
Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.
Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.
Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
Two defects live-drilled in the break-glass staff provisioning:
- An Owner reset that typed any username other than the occupied seat took
UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
existing seat — possibly the compromised account the reset was meant to
replace — live; every owner row is undeletable through the panel, so the tier
could never converge back to one. provisionOwner now refuses with
ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
form); bootstrap still mints, and the seat's own username still resets in
place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
console keys its rename prompt off api.ErrConflict — the "choose another
name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
the contract; PGRepo had drifted). Map the unique violation to ErrConflict
and pin it in pgint.
Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.
Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.