Manage Yggdrasil providers from the panel using durable platform settings, protected identity namespaces and atomic revisions. Apply changes to subsequent logins and profile lookups without restarting. Return operator-host logouts to the login method selection page.
Use shared action buttons and default to the login space. Allow staff to save startup-only experience settings while running and request a durable restart through the existing operator flow.
Grant the API PVC list permission needed to detect retained worlds before server creation, and show internal failures with a request ID. Cover permission, maintenance, restart and UI behavior with regression checks.
Configure connection and storage before creating or resuming the one-time Owner login. Remove Minecraft prerequisites from setup and preserve established login credentials.
Let staff preview and confirm roles from configured authentication sources using the existing account-link storage and game UUID mapping. Retain in-game code proof for players, add client-version and lobby guidance, and support NodePort passkey origins.
Nothing ever removed a submission: users could not retract a pending row, no
route deleted blobs or rows, and the reaper never touches uploads — so every
upload accumulated on the 5 GiB PVC forever and the only cleanup was SQL or
kubectl against the store.
- Blobs.Delete on both transports (local: RemoveAll of the id-namespaced dir,
id re-validated at the boundary; S3: idempotent object DELETE).
- Store: DeleteSubmission (admin, any status) and DeletePendingSubmission
(owner+pending CAS — a reviewed row can never be withdrawn out from under
its build).
- Manager.Delete / Manager.Withdraw delete the ROW first (under the CAS for
withdraw) and the blob after, so a live row can never point at a reaped
blob; a cleanup failure names the orphan explicitly instead of failing mute.
- API: DELETE /me/submissions/{id} (withdraw, app tier) and
DELETE /api/v1/submissions/{id} (admin) both return the row as it was;
audit events submission.withdraw / submission.delete; openapi documents both
paths; admin route pinned in the admin-only table.
- Panel: two-step withdraw on a pending row (frees the pending slot and the
storage budget); two-step delete on every admin row; zh/en copy; wire tests.
Unit: submit (withdraw happy path / wrong owner / reviewed row / no transport /
blob-cleanup failure), local+S3 delete idempotence, api handlers (200/404/409/
503 + route tier); pgint: withdraw CAS + admin delete exactly-once.
go vet/go test/gofmt clean; panel vitest 120 + typecheck green.
A logged-in user could file submissions without bound and stream a 1 GiB
context per submission. The only limits were the single-blob size cap and the
5 GiB uploads PVC (platform/workloads.go); nothing counted a user's rows or
bytes, so one account could fill the volume and every other user's upload
would start failing.
- Create: per-user pending_review cap (default 5) — the review queue cannot
be parked full of one account's rows. Check-then-insert, documented soft.
- UploadContext: per-user stored-context budget (default 2 GiB) charged
against the blob store's REAL sizes (new Blobs.Size on local/S3 stores), so
the sum cannot drift from the volume; the write is capped at the remaining
budget, so the excess is refused before it is persisted, and a re-upload is
charged only for its new bytes.
- API: per-user create/upload throttles (30s/15s, cmd/felis-wired) on a
dedicated cooldown keyspace, reserve→release so a failed attempt never
burns the window and a burst collapses to one winner; ErrQuotaExceeded →
403 submission_quota_exceeded (distinct from the 400 an oversize blob
gets), 429 submission_cooldown for the throttles.
- Panel: zh/en copy for both codes; openapi documents 403/429 on the two
user routes; pgint covers the pending-queue count.
Unit tests: submit package (cap, budget boundary/exact-fit/replacement,
oversize-vs-quota split) and api handlers (quota 403 both paths, throttle
429 + recovery + failure-release). go vet/go test/gofmt clean; panel
vitest 118 + typecheck green.
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:
- submit: derived context refs become the internal-face URL
/api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
image's new fetch-context entrypoint) that streams the blob with the
namespace-local service-token Secret — never mounted into Kaniko — and
extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
build namespace gets the token Secret through the existing replica mechanism
(bootstrap.sh + felis setup); the build egress lock opens exactly the control
namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
escapes/symlinks/non-gzip).
Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
The comments around hasJoined still described an authlib client that
is not in the path. Velocity reads -Dmojang.sessionserver and sends
the request itself, and it turns a 204 into its online-mode-only kick,
not authlib's "failed to verify username". The route comment in api.go
also offered "a thin login hook" as an alternative that does not
exist.
Other comments had drifted from the code:
- The [[auth_source]] doc said an empty list ships the multiplexer
off. Mojang is always prepended, so an empty list means Mojang is
the only source.
- The premium-name cache said Mojang does not recycle names. A name
frees up when its owner renames away. The day-long "taken" TTL still
holds, because a stale "taken" costs a third-party player only a
prefix.
- The cache bound claimed entries come only from players who
authenticated somewhere. Any third-party source that validates a
login adds one, so a hostile source can force the map to clear. That
costs repeat lookups, or a fail-closed prefix while Mojang is
unreachable, never an identity.
The rewrite rationale now states what it costs a backend operator. A
chat-session key that a third-party source signed over its native UUID
cannot verify against the canonical UUID, so chat from those players
can only be accepted unsigned.
In the tests, comments that repeated their subtest names are gone.
felis api wired the hasJoined multiplexer only when felis.toml had at
least one [[auth_source]]. With none, the source list stayed nil and
every hasJoined answer was a 204. That was harmless while nothing
pointed at the route, but the installer now starts Velocity with
-Dmojang.sessionserver aimed at felis-api unconditionally, and the
generated felis.toml tells the operator to delete the LittleSkin block
for a Mojang-only server. Doing exactly that turned every login away,
premium accounts included, and felis-api logged nothing about it.
Always build the list through authSourcesFromConfig, which prepends
Mojang in code, so an empty config is a Mojang-only relay. felis nano
already behaves this way with the same file.
The new test pins authSourcesFromConfig itself: Mojang first, the only
Identity source, and still present when nothing is configured. Marking
a configured source Identity makes it fail. The call site in cmdAPI is
now a single unconditional assignment and has no unit test of its own.
setup_required is what the SPA polls to decide whether the onboarding wall is
still owed, and it disagreed with the middleware that actually enforces the
wall. requireOnboarded lifts on a verified email OR an enrolled passkey;
setup_required answered `u.Email == "" || !hasPasskey`. A console-tier player
joins through the bind-code door with no email at all — by design, there is no
SMTP at that point — so the email term never clears and the SPA keeps them on
the setup screen forever, even after they enroll the passkey that already
unlocked the API for them.
The predicate now lives in one place (setupRequired) and both endpoints call
it, so the next edit to the unlock condition cannot drift them apart again.
Keying it on EmailVerified rather than email presence is the deliberate part:
presence is exactly the term that trapped the no-email player, and it was also
wrong on its own terms — an unverified address is not an authentication
factor, so it was never what the lockdown could safely lift on.
Also lands the regression test for the mechanism behind the live claim-403
report: /me/servers answers 200 for a bind-onboarded player (which is why the
dashboard renders the 认领 button at all) while claim, wake and status all
answer 403 with code "setup_required" — i.e. the refusal comes from
requireOnboarded before the handler, not from isOwnerOrAdmin inside it, which
would have said "forbidden". Enrolling a passkey and changing nothing else
lifts all three, which isolates the gate as the sole cause. The backend authz
is correct; the button that leads a locked-down player into a 403 is the
frontend's to hide.