demo-up.sh collapses bootstrap -> build+import the limbo/lobby images -> wire [velocity] login_image/lobby_image into felis.host.toml -> felis setup into a single command, ending in the interactive Owner-creation TUI (the only step it cannot automate). Prefers prebuilt tars under deploy/images, else builds on the host, auto-resolving the LOOHP/Limbo CI jar and the latest stable Paper jar (all overridable by env); SKIP_BOOTSTRAP/SKIP_SETUP toggles for reruns.
The lobby image had never been built and two defects blocked it: the felis image .dockerignore excluded plugins/* and only re-included limbo/shared, so the lobby Dockerfile's COPY plugins/paper landed empty; and the plugin stage used eclipse-temurin:21-jdk, which ships no gradle (and the tree vendors no wrapper), failing with 'gradle: not found'. Re-include plugins/paper and build the paper plugin on gradle:8.14-jdk21, matching the limbo image. Verified: both images build and boot (limbo /healthz 200 on 25565; lobby reaches 'Done' with felis-paper enabled).
deploy/limbo assembles LOOHP/Limbo from its loose CI artifacts plus the felis-limbo plugin (and the shared link core), with an entrypoint that pins server-port to the operator's GamePort (25565) on every start. deploy/lobby carries the Paper + felis-paper hub image. .dockerignore re-includes plugins/limbo and plugins/shared so the plugin image build sees them.
The login limbo now performs the onboarding inside Limbo: on join it checks the collision blacklist, mints a bind code, opens a book linking the player to console.<root_domain> (guiding them to the system browser), polls link-status, and transfers to the lobby via BungeeCord Connect — fail-closed on blacklist, mint/transport error, or window elapse. FelisApiClient gains linkStatus/isBlacklisted on the existing internal transport.
WebAuthn is unusable inside the WeChat/QQ in-app WebViews, so a document navigation carrying those UAs is served a bilingual 'open in your system browser' interstitial instead of the passkey-centric SPA. API/config/health/asset requests pass through, and an ack cookie (ua_ack) lets a determined user or false-positive continue. Backend-only; the SPA is untouched.
setup builds the always-on, reaper-exempt login/lobby MinecraftServers (create-if-absent), bakes the login limbo's non-secret config (internal API URL, root domain, lobby name) into spec.env, and replicates the felis-service-token Secret from the control namespace into the minecraft namespace so the operator's namespace-local secretKeyRef on the login pod resolves.
InternalAPIBaseURL builds the felis-api internal-face DNS from SAAPI and the internal port for cross-namespace callers (the login limbo). The service-token Secret name/key now reference the shared naming constants so the Deployment wiring and the operator's login-pod injection cannot drift.
buildStatefulSet gates readiness on an HTTP probe when HealthHTTPPort is set (exposing it as a named container port). buildEnv injects FELIS_SERVICE_TOKEN into the login server only — keyed off the reserved name so it can never leak into a user pod — sourced from a Secret via secretKeyRef, never inlined into the CRD.
StartupSpec.HealthHTTPPort/Path switch pod readiness from plain-TCP to an HTTP GET for RCON-less loaders (LOOHP/Limbo) that report 'started' only after the first tick. User servers now default FallbackServer to the login gate, never the lobby, so a stopped/starting backend keeps authentication in front of a fresh connection.
SystemLoginServer/SystemLobbyServer plus ValidateSystemServerName (format rule without the reservation check) let the platform provision the reserved login/lobby names users can never claim. ServiceTokenSecretName/Key are the one source of truth for the internal-API credential Secret, shared by the platform renderer and the operator's login-pod injection.
setup provisions the always-on login/lobby system services only when these image refs are set; empty means skip-and-say-so (the same fail-loud stance manifests takes), since no official LOOHP/Limbo image exists and a deployment must build its own.
handleChangePassword revoked other sessions but never cleared webauthn_credentials, and enrollment needs no step-up. A passkey planted through a transiently-hijacked session needs no password, so it survived the reset + session-revoke as a standing login foothold. Add DeleteAllPasskeyCredentialsForUser and call it in the change-password remediation so every passkey is unbound alongside the session revoke. Removing zero rows is a successful no-op. Email-OTP remains the fallback factor, so this never locks anyone out; the user re-enrolls a passkey afterward if they want one.
Enrollment set no AuthenticatorSelection, so user verification defaulted to preferred (not enforced), and the UV/backup flags the ceremony reported were discarded. Set UserVerification=required so a bound passkey always proves possession AND user (a UV-incapable device falls back to email-OTP), and capture user_verified/backup_eligible/backup_state through VerifiedCredential -> PasskeyCredential -> webauthn_credentials (migration 0009) so a future login path can enforce UV per credential. Adds a negative test proving a presence-only authenticator is rejected, and asserts the roundtrip records UV=true.
webauthn_credentials.user_id and webauthn_challenges.user_id referenced users(id) with the default ON DELETE NO ACTION, so a future user-delete would either fail or leave orphaned auth material. Recreate both FKs ON DELETE CASCADE: a bound passkey and a pending challenge are ephemeral and must not outlive the account. Scoped to the passkey tables only, not blanket, so retention-bearing child data (world_backups) is not swept away with an account.
The supersede DELETE in CreatePasskeyChallenge filtered consumed_at IS NULL, so it only reaped the prior LIVE challenge; the row that each finish stamps consumed_at on was left behind. A begin->finish loop therefore accumulated one dead row per cycle, unbounded. Drop the consumed_at clause so a fresh begin reaps ALL prior rows for (user, purpose), bounding the table at one row per (user, purpose) with zero net growth per cycle. Deleting an already-consumed row is safe: it has been redeemed and nothing reads it. The fake mirrors the widened supersede.
handlePasskeyRegisterFinish logged an empty target for account.passkey.registered, while the delete half logs the credential id. An operator auditing the log could see that a passkey was bound but not which one. Pass cred.ID as the audit target so bind and unbind are symmetric, and tighten the enrollment test to assert both halves name the credential id.
The MyServers query lists both a user's own servers and unclaimed (owner_id IS NULL) servers, but computed owned as s.owner_id = $1. For an ownerless row that comparison is SQL NULL, which fails to scan into the Go bool and 500s the whole listing. Wrap it in COALESCE(..., false) so an ownerless row reports owned=false while still surfacing as claimable.
The per-write deadline that severs a stalled SSE reader was never cleared on
return. Server.WriteTimeout is deliberately unset -- a WriteTimeout would sever
a healthy long-lived stream -- and with it unset net/http never resets the
connection write deadline between keep-alive requests. So the deadline the last
writeChunk left set leaks onto the next request that reuses the pooled
connection and fails its first write for no reason. Clear it to the zero value
on return via a deferred rc.SetWriteDeadline; best-effort, a no-op on writers
without deadline support.
Also record honestly at the header flush that the connect-time stall stays
bounded only by the per-principal stream cap, not severed by this guard -- only
the mid-stream stall is closed. Adds a test pinning the clear (fails closed:
neutering the deferred clear leaves a +writeTimeout deadline set on return).
QuotaAvailable and ClaimServer run as two separate statements, so the
count read is not serialized against a concurrent claim's UPDATE: two
claims by one user for two different ownerless servers can both pass the
gate and both succeed, leaving the user one server over quota. It is low
severity — quota over-provisioning under a deliberate burst, not an
authorization, ownership, or isolation break, since each server is still
claimed atomically via UPDATE ... WHERE owner_id IS NULL.
Closing it requires Postgres transaction semantics (advisory-xact-lock on
the user, or SERIALIZABLE with retry) folding the gate into a single repo
method — verifiable only against a real Postgres, not the hermetic
fakeRepo suite. Documented at QuotaAvailable with back-references from the
two claim gates (handleClaim and the internal UUID claim) rather than
patched blind.
relayLogStream copied a pod-log follow to the client with a plain
flusher.Flush per event. On a client that stays connected but stops
reading (its TCP receive window shut), net/http buffers the small
"data:" line and only touches the socket at Flush, which then blocks
forever inside the write. The select's <-ctx.Done() branch is never
reached, because r.Context() cancels on an actual disconnect, not on a
stall, so the relay goroutine and its upstream apiserver follow leak for
the life of the process.
Route every event's write+flush through http.ResponseController with a
per-write deadline (writeTimeout, 30s): a stalled flush now returns
os.ErrDeadlineExceeded, the error plain http.Flusher.Flush swallows, and
the relay abandons the stream so the deferred cancel + src.Close release
the follow. SetWriteDeadline and rc.Flush are best-effort: a writer
without deadline support (httptest recorder; some HTTP/2 origins) ignores
the deadline and behaves exactly as before, so the guard degrades
gracefully.
This closes the leak the per-principal stream cap only bounded the blast
radius of. Verified by a deterministic test with a deadline-aware
ResponseWriter whose flush blocks until the deadline; the test times out
(fails closed) if the guard is removed.
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.
Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.
This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
The three felis-api http.Servers (internal, external, https) were built with
only Addr and Handler, leaving ReadHeaderTimeout, IdleTimeout, and ReadTimeout
at zero. A zero ReadHeaderTimeout is a Slowloris hole — a client trickling
header bytes pins a connection indefinitely — and a zero IdleTimeout lets
kept-alive connections accumulate (gosec G112).
Route all three listeners through a newAPIServer factory that sets a 10s
ReadHeaderTimeout and a 120s IdleTimeout. WriteTimeout and ReadTimeout are
left unset on purpose: the external and https faces stream Server-Sent Events
(console / build logs) for the lifetime of a client attachment, and a
WriteTimeout would sever a healthy long-lived stream. Slowloris is closed by
ReadHeaderTimeout, which bounds only the header phase.
withRequestID honored any inbound X-Request-Id verbatim, and that value is
echoed on the response, embedded in the error envelope, and persisted into
audit_logs.request_id. An unvalidated caller-supplied id is therefore an
audit-integrity vector: an arbitrarily long value bloats the audit row, and a
stray control byte (CR/LF) could smuggle a forged entry into a log sink.
Accept an inbound id only when it is well-formed — non-empty, at most 64
bytes, and restricted to a log-safe charset ([A-Za-z0-9._-]) — otherwise mint
a fresh server id. A rejected request loses its inbound trace link, which is
strictly better than storing attacker-controlled text in the audit trail.
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.
Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.
The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).