A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.
Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.
Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
UserByMCUUID now resolves only live accounts: claim, menu, wake
authorization, op-login vouch and the QR link-status poll treat a
disabled or soft-deleted link holder exactly like an unlinked UUID
instead of a retired identity. VerifyLinkCode lets a soft-deleted
link be taken over by a fresh in-game code (the deleted account is
gone, e.g. a migrated source), while a disabled holder still 409s so
the lockout is not bypassable; failed attempts still do not consume
the code. Fake repo and pgint coverage pin both branches.
Two defects in the §9.3 quota path, both invisible to the hermetic suite:
- Audit #4's TOCTOU was real and documented: QuotaCheck and ClaimServer were
separate statements, so two concurrent claims by one user for two different
ownerless servers both read count < max_servers and both won. The gate now
lives inside ClaimServer, in the SAME transaction as the ownership write,
under pg_advisory_xact_lock(hashtext(user_id)) — the aggregate read, the
four-dimension re-check (shared with QuotaCheck via one helper so the two
cannot drift), and the UPDATE are one serialized decision. The loser gets
ErrQuotaExceeded, which both claim handlers map to the same 403 the
sequential path gives; the server row is additionally taken FOR UPDATE so
same-server races still resolve to exactly one winner.
- The server PATCH path called UpdateServerResources(..., 0) for storage even
though a resources patch cannot change storage. The cached columns are the
ONLY input to the quota aggregate, so every resource patch silently dropped
that server's storage contribution from its owner's cap. The handler now
reads the current spec and passes storage through.
Red-then-green: the new pgint test drives two real concurrent claims against
max_servers=1 (before: both win; now: exactly one win + one gated 403, and the
DB shows one owned row); the hermetic suite pins the 403 mapping and the
storage-preserving cache write.
Found live while verifying the admin email-edit fix: the Owner account could
not load /api/v1/users at all. Root cause: migration 0011 adds the 'owner'
role and gates every user-administration route on it, but NOTHING ever wrote
it. break-glass (UpsertOwner), the setup MC-bind (CompleteOwnerSetup), and the
re-provision path all forced 'admin', so in a fresh install the entire
owner tier — list/create/edit/disable/delete users, quotas, sessions — was
unreachable. The role was a dead letter in the other direction too: staff
predicates that predate the role did not know it.
- UpsertOwner and CompleteOwnerSetup now write role='owner'; the username-
conflict arm re-asserts it, which is also the documented pre-0011 promotion
path ("re-provision via break-glass"). InsertOperator stays plain 'admin'.
- Staff doors learn the role: op-login start/finish admit the Owner; the
player email door refuses it like any staff account; the in-game approver
check already used staffRole.
- Reclaim protection: IsProtectedAdminLink (and the break-glass bootstrap
switch AdminExists) count admin OR owner — the Owner must never be displaced
by a Mojang-priority reclaim.
- Panel guards make migration 0011's claim true now that owner rows exist: an
owner can never be demoted, deleted, or disabled through the API (only the
local break-glass console resets the identity); username/email edits still
work.
Tests: pgint pins both provisioning paths, the protected-link predicate and
the reset/promote semantics; hermetic suites cover the owner-admitting staff
door, the owner-refusing player door, the three panel guards, and break-glass
attribution.
UpdateUser wrote a new address but kept email_verified, so patching a verified
account asserted a proof of an address nobody had proven — and the
pre-session login mails and resolves on exactly that flag, so a typo'd edit
could hand the account's sign-in codes to the wrong mailbox.
Changing the address now clears the flag in the same write; a no-op patch that
passes the same value keeps it. The fake mirrors the semantics, and the pgint
suite pins both halves (same value keeps proof, new value drops it).
The design has claimed since migration 0010 that at most one account can hold
a PROVEN email address, with ErrEmailTaken as the 409 a second verifier sees.
Neither half ever shipped: no migration created users_verified_email_unique,
and VerifyEmailOTP had no guard at all — the sentinel was defined but never
returned, so two accounts could both verify one address. The damage is not
cosmetic: the pre-session login resolves accounts BY verified email, so the
duplicate decided which identity a mailed sign-in code belonged to.
- Migration 0020 creates the partial unique index (lower(email) WHERE
email_verified) the comments have been citing — the database-level backstop.
- VerifyEmailOTP now refuses the take-over with ErrEmailTaken BEFORE consuming
the code (the address, not the code, is the problem), charges no attempt,
and maps a lost cross-user race (unique violation) to the same answer.
- The verify handler answers 409 email_taken instead of a generic 500.
Covered by the pgint suite (sequential double-verify refused with the code
still live, a direct duplicate write still loses to the index, the refused
account can still prove its own address) and a hermetic 409 case.
Resolving a session cookie failed identically whether the credential was
missing or Postgres was unreachable: local_auth_enabled read errors fell into
the fail-closed 'disabled' branch and SessionUser errors into 'invalid
session', both surfacing as 401 'authentication required' — a lie that reads
as 'log in again' during an outage. Split the enabled-read into
(enabled, error), tag non-ErrNotFound store failures with errAuthBackend, and
map that to a new 503 auth_unavailable in requireExternal. Fail-closed is
unchanged: missing setting / bad value / missing session stay 401.
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.
Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
handlers; align tests (new account subtests, naming reserves
"console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
setup_panel.go, tui_root.go and tui_preflight.go instead of
hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.
Panel:
- Delete the ChangePassword page and all password UI; align
login/auth/api/types with the passwordless contract; add the
migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
ServerLuckPerms strings to translation keys; drop 72 orphan keys
per locale; unify the title as "Felis - Console".
Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
preference; unified /link lines in fabric/forge/neoforge; shared
link-client javadoc contract fixes.
Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.
BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
At bootstrap there is no SMTP, so the old /setup flow was unreachable: it
requested an emailed OTP that could never arrive. Setup now records the
Owner's email address unverified (no OTP round-trip) and requires a passkey,
deferring SMTP configuration to a later Settings page. Setup completes on
email-recorded + passkey-enrolled, and the lockdown lifts on the passkey, not
on email_verified: a passkey is the Owner's only pre-SMTP login credential
(email-OTP login refuses admin accounts).
The record-email endpoint (POST /account/email) now clears email_verified in
the same write. Only VerifyEmailOTP, which proves control of the address, may
set that flag; recording a fresh unproven address must never leave a stale
email_verified=true asserting a proof the user never gave. The change strictly
tightens the invariant, so no existing reader breaks.
Remove the dead ErrEmailTaken path and its documented 409: no migration puts a
unique index on users.email and the codebase does not enforce email
uniqueness, so the unique-violation branch was unreachable and the 409 an
impossible response.
The /setup route (Setup.tsx, setEmail helper, setup i18n copy) is rewritten to
match: record-email, mandatory passkey, no skip-for-now. The SMTP settings
page and post-setup configure-SMTP nudge are deferred.
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).
Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.
Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.
The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.
Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.
- 0015 migration: account_migrations state machine (initiated -> confirmed ->
code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem
Quota is grandfathered at redeem: no per-target quota re-check when servers move.
Previously /readyz only verified Repo != nil && Cluster != nil — a
process-liveness check, not a dependency-health check. The spec
requires the readyz probe to verify DB, K8s API, and CRD informer
are live before declaring the pod ready.
- Repo interface gains Ping(context.Context) error
- Cluster interface gains Ping(context.Context) error
- PGRepo.Ping delegates to sql.DB.PingContext
- K8sCluster.Ping lists MinecraftServer CRDs (Limit=1) in the
configured namespace, exercising both the API and CRD informer
- handleReadyz iterates ping checks; any failure returns 503 with
the failing dependency name in the error message
- fakeRepo and fakeCluster gain configurable pingErr for hermetic
test coverage of the failure paths
New test: TestReadyzPingsDependencies verifies 200 when healthy,
503 when DB or K8s API is down.