Add POST /api/v1/internal/servers/{name}/backup so the on-node break-glass
console can snapshot a stopped world while felis-api is alive. It goes through
the API (not direct-to-CRD like halt) because rendering the backup Job needs
deployment coordinates (FELIS_IMAGE, FELIS_BACKUP_PVC) only felis-api holds.
Service-token auth (no Principal); the middleware IS the authorization, since
the operator already has root on the node. Refactor the RWO stopped-gate,
optional-Backuper 503, async hand-off and audit+202 into a shared enqueueBackup
tail so the external (owner/admin) and internal (break-glass) faces cannot
diverge on the security-critical stopped-gate. The internal audit is attributed
to break-glass/internal so a console-initiated backup is distinguishable from an
owner self-service one.
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.
felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).
Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.
- 0015 migration: account_migrations state machine (initiated -> confirmed ->
code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem
Quota is grandfathered at redeem: no per-target quota re-check when servers move.
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.
Routes (both Public):
POST /api/v1/auth/passkey/login/discoverable/begin
POST /api/v1/auth/passkey/login/discoverable/finish
Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.
The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.
Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
The passwordless migration left ResetMailer (SendPasswordReset) and its API field with zero callers and no wiring; the web console authenticates via email-OTP and passkey only. Remove both, plus the now-orphaned context import that the interface was the last user of in handlers_users.go.
Reconcile the DeleteAllPasskeyCredentialsForUser docs in repo.go and pgrepo.go: they claimed there was no production caller, but 2f22027 wired the owner-tier DELETE /users/{id}/passkeys. Both now note that a complete authenticator remediation pairs the unbind with a session revoke (unbinding alone leaves the live hijacked session; revoking alone leaves a re-enrollable credential), and the OpenAPI operation carries the same guidance in a new description. Reword the stale local-password test-fake header, since the passwordless fakes carry no must_change_password field.
No behavior change. gofmt, build, and the full test tree are green; OpenAPI parity and passkey-unbind tests pass; a grep confirms ResetMailer/SendPasswordReset are gone from the Go tree.
Add DELETE /api/v1/users/{id}/passkeys (owner-only) to unbind every passkey a
target account holds — the authenticator remediation that stops a passkey planted
or retained via a transiently-hijacked session from surviving as a standing login
foothold. It wires the previously-uncalled DeleteAllPasskeyCredentialsForUser and
is deliberately not a lockout: the account re-enters via the email-OTP door
(players) or op-login's in-game approval (staff), then re-enrolls. Documented in
the OpenAPI, so the served/documented parity gate covers it.
Remove RevokeUserSessionsExcept: a change-password-era orphan with no callers
since the passwordless migration. Its keep-one ("log out my other devices")
semantics is inherently self-service, and no such slice is on the roadmap; the
admin remediation path already uses RevokeAllUserSessions.
The passwordless migration (b330d77) removed the password-login route, leaving
concurrencyLimiter — its bcrypt concurrency cap — with no caller, and scattered
stale "local-password" / "change-password" references through the surviving auth
code's comments.
- Remove the dead concurrencyLimiter (type + newConcurrencyLimiter + acquire):
no caller, no struct field, no test. Reword the one streamLimiter doc that
contrasted against it.
- Realign comments in repo.go, pgrepo.go, session.go, util.go to the passwordless
reality: staff lookups feed email-OTP / passkey / setup redeem, not a password
compare; RevokeUserSessionsExcept and DeleteAllPasskeyCredentialsForUser are
retained (uncalled) for the P5 account-remediation path (#78); "local sessions"
no longer implies a password.
Comments and dead code only; no behavior change. Full WSL test tree green.
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.
- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
methods an email can use. The single sanctioned existence oracle; methods
are computed with no role branch, so staff and player accounts in the same
credential state return byte-identical bodies (staffness invisible by
construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
/auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
(migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.
Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.
Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.
This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.
Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.
The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
Two admin-tier routes read and set a single platform-wide maintenance
window for the auto-update subsystem (decision core internal/updates):
GET /api/v1/updates/window
PUT /api/v1/updates/window
The window is stored as JSON {"start","end"} (RFC3339, or null when
unset) under the platform_settings key "update_window", reusing the
existing GetSetting/SetSetting KV seam -- no new Repo method, no
migration. Pointer times keep "unset" (null) distinct from a real
instant on both decode and encode; a never-set and an explicitly
cleared window both read back as {null,null}.
Validation mirrors the core's fail-closed Window: a window is either
fully set (both ends, end strictly after start) or fully cleared (both
null). A half-set, inverted, or empty-interval body is 400 and is never
persisted. Reads treat only a missing key as unset (ErrNotFound -> 200
nulls); any other store error 500s rather than fail open.
This is API + PERSISTENCE ONLY. Nothing consumes the stored window yet
-- the runner, the ReleaseSource/Notifier/Applier executors, and the
scheduler CronJob remain INTEGRATION-ONLY. Setting a window changes no
behavior until those land; it is the durable input they will read.
Nothing here force-updates ("不要强制自动更新").
Adds POST /api/v1/auth/bind, the one public pre-account entrypoint of the
player console (console.<root_domain>). An account-less player redeems the
one-time Bind Code minted in the in-game Login Lobby; in a single step the
platform creates a role=user player, links it to the verified in-game UUID,
and mints a host-only felis_session. Login is thus not forced at the edge
while operations stay app-authenticated.
The operator console (op.console.<root_domain>) is unaffected and stays
behind Zero Trust: a code whose UUID resolves to a staff (role=admin)
account is refused with 403 (ErrPlayerBindForbidden) without consuming the
code, so the public door provably never yields an admin principal — the
session it mints carries ViaAdminAccess=false and is host-only to console,
never sent to op.console.
Repo layer: new RedeemPlayerBindCode on the Repo interface, implemented on
PGRepo (single tx: resolve code, create-or-fetch the player, consume) and
the test fake. The returning-player branch is idempotent and is a deliberate
standing "log in via the game" door, not just first-time onboarding.
Honest labeling:
- ORACLE-VERIFIED (Go): account/session logic — role=user, refuse-staff,
idempotent create-or-fetch, single-use code, and the op.console redline
(player session rejected on admin routes). Covered by handlers_onboard_test
and the OpenAPI parity gate.
- INTEGRATION-dependent: the endpoint's security rests on the Bind Code having
been minted against an online-mode-Yggdrasil-authenticated UUID, a
precondition that lives in velocity/Java and is not verifiable from this
repo (CODE-ONLY). The Go layer proves the logic, not that identity guarantee.
- No app-level attempt cap: rate-limiting is deferred to the edge as for the
public /auth/login; the ~1e12 keyspace, single use and short TTL make a
blind app-level cap non-critical.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14), web app face.
An already-authenticated principal binds a passkey to their own account and
manages the credentials they have bound; email-OTP stays the fallback factor.
- four account routes: POST register/begin mints a credential-creation
challenge, POST register/finish verifies the attestation against the
server-stashed SessionData and binds the credential, GET/DELETE credentials
list and unbind the caller's OWN passkeys. App-tier, principal-scoped (the
body never names a user).
- PasskeyVerifier seam keeps go-webauthn out of this package: ceremony state
crosses as opaque bytes, attestation as an io.Reader, result as a plain
VerifiedCredential. A nil verifier makes begin/finish report 503 so the
authenticated boundary is exercised before the real verifier is wired in.
- the view never leaks the public key; credential_id collisions map to 409.
- OpenAPI: the four paths plus the PasskeyCredential schema, keeping the
served-routes parity gate green.
Scope: ENROLLMENT only. The passkey login/assertion path (proving a passkey
from an unauthenticated state) is deferred; every ceremony here rides on a
known principal.
Tests: handler + challenge state machine against a fake repo and a fake
verifier (no real attestation crypto, no SQL). The decisive assertion is the
session-data round-trip -- the finish body carries no challenge, so the only
path for the stashed blob into FinishRegistration is store-stash then consume,
proving the challenge is server-held and never client-echoed. Also covers
supersede-on-begin, single-use, expiry, 503-unavailable, 409-already-bound,
owner-scoped list/delete, and external-only face separation.
cooldownLimiter began as the wake-only throttle; the OTP-start hardening
reused it via the atomic reserve/release. Its type comment still called it
a per-server wake limiter and justified the per-replica behaviour as
"acceptable because the operator reconcile is idempotent" -- true for wake,
false for OTP, whose every admitted send is a non-idempotent email.
Rewrite the comment to describe the shared per-key limiter and record the
honest KNOWN-LIMITATION: the atomic reserve/release closes the
intra-replica concurrent burst, but the in-memory map throttles per
replica, so cross-replica bounding still needs a shared store. No
behaviour change.
The email-OTP resend cooldown checked the window with a peek (allowed)
and only recorded it after delivery. For OTP that throttle is the sole
defense and each admitted send is a real, non-idempotent email, so a
burst of truly concurrent starts all passed the peek before any recorded
and every one mailed: N concurrent starts bombed a mailbox with N codes.
Add an atomic reserve/release pair to cooldownLimiter: reserve checks and
records the window in one critical section under the mutex, so a
concurrent burst yields exactly one winner; release rolls a reservation
back only if it is still the current one, so a slow failing caller never
clobbers a newer holder. handleEmailOTPStart now reserves both the
principal and the recipient key up front and defers a rollback that frees
both windows on any mint, create, or delivery error — preserving the old
"a failed send does not consume the cooldown" property, now race-free.
The wake path keeps allowed→record: its real gate is the running cap and
its side effect (SetDesiredState) is idempotent, so the peek gap is
harmless there.
Tests: a frozen-clock gate-mailer fires 8 concurrent starts for one
victim from one principal and asserts exactly one mail and one 202; a
flaky-mailer test proves a failed delivery releases the window so an
immediate retry in the same instant is admitted.
handleEmailOTPStart minted and mailed a code on every call, so an
authenticated caller could drive unbounded mail to any address they
typed — an email-bomb primitive against arbitrary mailboxes.
Add a separate otpLimiter (its own sync.Once and map, distinct from the
wake limiter) and throttle each send on two keys before anything is
minted: the caller (user:<id>) and the recipient (email:<lower>). A
refused send mints no code and mails nothing; both cooldowns are
recorded only after delivery succeeds, mirroring the wake path so a
failed mint or delivery never consumes the throttle. The two-key design
stops both one account fanning out across addresses and many accounts
converging on one mailbox.
A wake refused by the §9.1 running-server cap returns 503, but the
per-server cooldown was recorded before the cap check ran. A player
held because the cluster was momentarily full would then also have to
wait out the wake cooldown once a slot freed, even though their refused
wake never actually flipped desiredState.
Split cooldownLimiter.allow into allowed (peek, no record) and record
(commit). Both wake paths now consult allowed for the 429, then call
record only after SetDesiredState succeeds — so neither a 503
at_capacity nor a SetDesiredState error consumes the cooldown. The
split is safe against the running cap, which counts CRD truth via
ListServers and is independent of the limiter.
QR scan-to-login is a device-code grant where the QR encodes the existing
short-lived account-link code (spec §B3 player game-login). velocity mints a
code in-game, renders it as a QR, the player scans it on a phone already signed
in to the panel, and that web session's verify writes the durable account_links
row bound to that user. The only new verifiable surface that flow needs is the
completion poll velocity calls to learn the link landed and admit the player.
Add GET /api/v1/internal/account/link/status/{mc_uuid}: a read-only, internal
handleLinkStatus keyed by the verified mc_uuid velocity already holds. It reuses
the existing UserByMCUUID, so it adds no migration and no mutation to the
load-bearing VerifyLinkCode; ErrNotFound maps to {linked:false} (pending /
not-yet-scanned), a hit to {linked:true, user_id}. Keying on the public UUID and
not the scanned code means the read carries no guessing surface and needs no
attempt cap — the internal face already gates it to service callers, and the poll
consumes nothing so a velocity restart re-polls safely.
QR render, limbo collision routing, in-game admit, and the reclaim
inherit-disambiguation stay CODE-ONLY (Java/Velocity) and are labeled as such;
this endpoint reports link completion only.
Document the route in openapi.yaml (x-felis-face internal, x-felis-tier service)
so the parity gate holds, and cover it with a hermetic vertical that proves the
poll reflects the durable link only after the external verify and binds the
verifier's id, plus unknown-uuid, idempotency, and internal-only face separation.
When the configured third-party Yggdrasil and the official Mojang service
issue the same username under different UUIDs, the non-genuine squatter is
displaced in favour of the real Mojang owner (正版优先). This adds the
Go-verifiable data layer of that flow on the internal (velocity) face.
- migration 0006: username_blacklist (barred squatter UUIDs) and
player_data_holds (the displaced account's 30-day data stash), both keyed
by mc_uuid so the genuine Mojang player — identical username, different
UUID — is never caught by the bar.
- POST /api/v1/internal/player/reclaim bars the squatter UUID and stashes
its data in one transaction (all-or-nothing). It is idempotent on a
retried callback and returns the hold's effective expiry — the first
reclaim's window, never a fresh now()+30d — so the rejected player is told
the truth about how long their data is kept.
- GET /api/v1/internal/player/blacklist/{mc_uuid} is the login-gate check
velocity calls to reject a barred squatter before admitting them.
Scope: velocity collision-routing, the limbo prompt, the authlib
dual-backend and the data-inherit flow are code-only (Java plus a QR-bound
device session a row cannot express) and are not part of this slice. Unit
tests cover the handlers and the in-memory repo contract; the Postgres SQL
path is exercised by integration only.
Forced web onboarding proves a player controls an email before it is
bound to their account. POST /api/v1/account/email/start mints a random
6-digit code, mails it (or logs it server-side when no Mailer is wired —
the demo has no SMTP), and POST /api/v1/account/email/verify redeems it,
flipping users.email_verified in the same transaction that consumes the
code.
Brute force is bounded two ways: a 10-minute TTL and a 5-attempt cap,
both enforced in the repo so the fake and Postgres agree. Only the
sha-256 of the code is stored; the digits live only in the email. Both
routes are app-tier external — verifying your own email is scoped to the
principal, never names another user.
Add username+password login for Owner/Operator staff accounts on
op.console, the primary web login when Zero Trust is not in front of the
API. Three handlers form the whole surface: login mints a server-side
session cookie, logout revokes it idempotently, and change-password
re-verifies the current password before rotating the hash and clearing
must_change_password.
- Session cookies are HttpOnly+Secure+SameSite=Lax, host-only, stored
server-side as a SHA-256 hash with a 12h TTL.
- Login is anti-enumeration: every failure runs a uniform bcrypt compare
against a dummy hash and returns the same vague error.
- Credential-bearing writes require Content-Type: application/json,
returning 415 otherwise, to close the cross-site form-POST forgery
vector as a belt to the SameSite cookie.
- Local auth fails closed: login is rejected unless local_auth_enabled
is set, so a Zero-Trust-only deployment never accepts a local password.
- Extend the users table with a nullable password_hash and
must_change_password; staff are role=admin rows with a hash, players
are role=user rows with hash NULL.
- /me now reports must_change_password so the panel can force a
first-login change.
Covered by Go unit tests (handlers, content-type guard, anti-enumeration,
forced-change lockdown) and the OpenAPI route-parity gate.
The dual-faced felis-api: internal (service) and external (public/app/admin) routes behind a Zero-Trust guard. Includes the access domain (whitelist, ban, and LuckPerms permission/group control over the owner-gated RCON path), the modpack submission endpoints, and the admin-tier SysAdmin fleet read. Structured access fields are charset-validated before assembly so no field can splice a second RCON command.