Commit Graph
79 Commits
Author SHA1 Message Date
flyemoji 0dbd557a7a fix(store): renumber discoverable-login migration 0013 -> 0014
A resource-cache migration (0013_resource_cache.sql) was merged onto main concurrently and also claimed version 0013. LoadMigrations rejects any duplicate migration version, so the app would refuse to boot with both files present.

Renumber the discoverable-login migration to 0014. The two migrations touch disjoint objects (0013 ALTERs servers to add cached_* columns; this one CREATEs webauthn_discoverable_challenges), so their relative order does not matter, and the table name is unchanged -- no Go reference moves.

Renaming a just-published migration is safe here because neither version has been applied to a persistent database yet: there is no schema_migrations row for version 13 to reconcile. This is a pre-application renumber, not a history rewrite of an already-applied migration.
2026-07-05 04:31:45 +09:00
flyemoji 154002edf3 docs(auth): cite MultiLogin reference for UUID-keyed reclaim split
Anchor the username-collision reclaim's UUID-keyed, proxy-detected design to
the multi-Yggdrasil reference: CaaMoe/MultiLogin v6 binds identity as
serviceId+online-UUID via "identity cards" that decouple the in-game name from
online identity — keyed by UUID, never by name. Note that §B3's Mojang-priority
reclaim goes beyond the common "protect the first-bound name" behavior by
evicting a squatter once the genuine Mojang owner appears and stashing the
squatter's data for the code-only inherit path.
2026-07-05 04:05:56 +09:00
flyemoji ec468baef9 feat(auth): add discoverable (usernameless) passkey login
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.

Routes (both Public):
  POST /api/v1/auth/passkey/login/discoverable/begin
  POST /api/v1/auth/passkey/login/discoverable/finish

Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.

The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.

Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
2026-07-05 04:05:56 +09:00
flyemoji 7db57b9fff feat(updater): add VersionGatherer extraction core and CLI gather seam
Give the Runner a way to read each component's CURRENT version so it can be
compared against the release sources already wired. Three pure extractors turn
raw system text into an updates.Version, each fail-closed:

  - versionFromCLI      — a `<tool> --version` banner   (k3s, cloudflared)
  - versionFromImageRef — a container image tag         (felis-api)
  - versionFromJarName  — a proxy jar filename          (velocity)

sysGatherer routes each Topology component to the right extractor over an
injected seam; every path is exercised with a fake runner, mirroring how the
release sources are proven against httptest.

The load-bearing case is k3s: its Git tag "v1.36.2+k3s1" parses stable, but a
registry cannot store '+', so the same build ships as image tag "v1.36.2-k3s1",
which parses as a prerelease unless repaired. versionFromImageRef normalizes
"-k3sN"/"-rke2rN" back to "+", so an image read and a CLI read agree instead of
the image masquerading as a prerelease and being barred from comparison.

Honest runtime state after this slice — a green suite is not "the updater runs
against real infra": only the CLI seam (execRunner) is wired, so of the four
tracked components just cloudflared is live end to end (gatherable AND
Scheduled/appliable). k3s is CLI-gatherable but Notify-only. felis-api and
velocity are NOT yet runtime-gatherable: their producing seams — a k8s read of
the control-plane Deployment image, and an off-cluster jar inspection — are left
nil, so both surface an explicit "gather seam not wired" error rather than a
wrong version. felis-api self-update is therefore not functional yet.

Remaining integration (tracked in doc.go): the two producing seams, the concrete
Notifier (SMTP + in-game), the Applier (image bump, cloudflared swap), the
`felis update` CLI + CronJob entry point, and the runtime append of the Pinned
Minecraft fleet.
2026-07-05 04:05:56 +09:00
Lemon-miaow e574749877 feat(api): enforce CPU/memory/storage quotas (spec §9.3, §22)
Add per-user resource quota enforcement across all four dimensions:
max_servers, max_cpu_milli, max_memory_mb, and max_storage_gb.

- Migration 0013: add cached_cpu_milli, cached_memory_mb, cached_storage_mb
  columns to servers table for pure-SQL per-owner aggregation
- SeedServer now writes resource cache alongside server row
- QuotaCheck replaces QuotaAvailable at claim time, checking all four caps
  against the owning user's cumulative usage
- handlePatchServer checks owner's quota before allowing memory/resource
  changes on owned servers; unowned servers skip the gate
- handleInternalClaim mirrors the full quota check
- UpdateServerResources keeps the cache in sync after spec mutations
- Reaper zeros resource cache on ReleaseWorld so released resources
  are not counted against a former owner
- quantityToMilli/quantityToMB helpers convert K8s quantities to
  quota-comparable integers

19 test packages pass.
2026-07-05 02:05:44 +08:00
Lemon-miaow 91bfa27e8c feat(operator): implement idle auto-stop (spec §8)
Adds empty-server auto-stop to the reconciler. When Idle.AutoStopEnabled
is true and the RCON player tally is zero for EmptySecondsBeforeStop seconds,
the operator flips desiredState to Stopped, which triggers the normal
graceful-shutdown path.

- Add EmptySince status field to track empty duration
- Clear EmptySince on stop and when players return
- Reuse existing RCON probe's player count (zero extra network cost)
- 4 new test cases covering timestamp, timeout, player-join reset, and
  disabled-by-default
2026-07-05 01:52:45 +08:00
flyemoji 7d27640c07 feat(updater): add GitHub Releases source and route felis-api/k3s/cloudflared
Give RoutingSource its second upstream so every non-pinned component now
resolves a real latest-stable: Velocity via PaperMC (already wired), and
felis-api, k3s and cloudflared via the GitHub REST API.

github.go queries /repos/{repo}/releases/latest (one request, rate-limit
friendly) and fails closed: a transport error, a non-200 status (404 = no
stable release), an undecodable body, a draft/prerelease flag, or an
unparseable / prerelease-parsing tag all return an error, never a zero
version. It sends the User-Agent GitHub requires (a UA-less request is
403'd) and tolerates the two live tag styles -- cloudflared's CalVer
"2026.6.1" and k3s's v-prefixed, build-tagged "v1.36.2+k3s1" -- while
String() keeps the raw tag for the report.

source.go routes sourceGitHub to it and drops the errGitHubNotWired stub;
velocity still routes to PaperMC.

Tests: github_test.go covers both tag styles, the User-Agent gate, and
fail-closed on 404 / prerelease-flag / unparseable tag, with fixtures
captured from api.github.com on 2026-07-05. runner_test.go now drives
PaperMC and GitHub through dual httptest servers end to end with no source
degrading to an error.

doc.go re-tiers the verification boundary: both release sources are now
built and live-grounded; the VersionGatherer's version-extraction core is
the next verifiable slice (logic over an exec seam, not pure I/O); the
genuine I/O remainder is the Notifier, Applier and felis update CLI/CronJob.
felis-api's coord is still a placeholder slug, so that component is dark at
runtime until a real repository is configured.
2026-07-05 01:03:32 +09:00
flyemoji 9896fe16c3 docs(updater): correct PaperMC UA/fixture overclaims, re-tier the boundary
An out-of-band curl of the live Fill v3 endpoint contradicted two claims the
previous commit shipped and surfaced a mis-tiering:

- User-Agent is NOT enforced: fill.papermc.io/v3/projects/velocity returned
  HTTP 200 to a bare curl UA. The comments claimed a generic UA "is refused"
  and the API "REQUIRES" a contact UA. Reword to what is true — PaperMC's usage
  policy asks for a descriptive UA and may block generic ones, but sending it is
  etiquette/defensive here, not a gate Felis depends on.
- The test fixture's shape was invented, not captured: the real "versions"
  object groups the entire 3.x line under a single key "3.0.0", not the
  per-minor keys the fixture used. Replace it with the real body (keys and
  version strings as returned). The key-agnostic parser already produced the
  right answer, and an independent max-stable check confirms 3.4.0.
- Re-tier doc.go: the GitHub Releases source is verifiable-here (the same
  httptest-testable shape as PaperMC), not integration remainder. It is why
  3 of 4 components report "latest unknown" today and is the next verifiable
  slice — the release-source work is only ~half done until it exists.

No production logic changed. WSL oracle: build + vet clean, internal/updater
10/10, full tree go test RC=0 (19 ok, 0 fail).
2026-07-05 00:28:36 +09:00
flyemoji 96b3cc901c feat(updater): wire updates.Run to a caller with PaperMC v3 release discovery
internal/updates is a pure, fakes-tested decision core with no production caller,
so nothing could produce its "版本号状态" report. Add internal/updater as that caller:

- topology: the fixed platform components and their user-set policies (felis-api
  and cloudflared Scheduled+manageable; k3s Notify, high-blast-radius single node;
  velocity Notify, off-cluster and unmanageable). Minecraft is pinned by ABSENCE,
  never force-tracked here, appended from the live fleet at runtime.
- PaperMC Fill v3 release source: the v2 API (api.papermc.io) was retired
  2026-07-01 and returns HTTP 410, so this targets fill.papermc.io/v3, sends the
  required non-generic User-Agent, and returns the newest STABLE version, filtering
  the -SNAPSHOT/rc prereleases the plan would otherwise suppress. Its test fixture
  is captured from the live v3 response shape (2026-07-04).
- RoutingSource: the single ReleaseSource updates.Run requires, dispatching
  velocity to PaperMC and returning errGitHubNotWired for the GitHub-backed
  components so they degrade to "latest unknown" honestly, never a fabricated one.
- Runner: gather current versions (seam) -> assemble Components -> updates.Run ->
  Report; report-only when notifier and applier are nil.

Verification boundary: the parse/plan/compose logic is unit-tested (httptest +
fakes, fixture grounded in the live v3 shape). Live network/TLS/User-Agent
enforcement, the GitHub Releases source, the concrete version gatherer, the
notifier and applier, and the felis update CLI/CronJob remain integration work,
enumerated in doc.go.
2026-07-04 23:41:46 +09:00
flyemoji c20b12c655 refactor(api): drop dead password-era ResetMailer, reconcile passkey-unbind docs
The passwordless migration left ResetMailer (SendPasswordReset) and its API field with zero callers and no wiring; the web console authenticates via email-OTP and passkey only. Remove both, plus the now-orphaned context import that the interface was the last user of in handlers_users.go.

Reconcile the DeleteAllPasskeyCredentialsForUser docs in repo.go and pgrepo.go: they claimed there was no production caller, but 2f22027 wired the owner-tier DELETE /users/{id}/passkeys. Both now note that a complete authenticator remediation pairs the unbind with a session revoke (unbinding alone leaves the live hijacked session; revoking alone leaves a re-enrollable credential), and the OpenAPI operation carries the same guidance in a new description. Reword the stale local-password test-fake header, since the passwordless fakes carry no must_change_password field.

No behavior change. gofmt, build, and the full test tree are green; OpenAPI parity and passkey-unbind tests pass; a grep confirms ResetMailer/SendPasswordReset are gone from the Go tree.
2026-07-04 21:47:13 +09:00
flyemoji 4f59d5128a feat(auth): add owner-tier passkey-unbind remediation endpoint
Add DELETE /api/v1/users/{id}/passkeys (owner-only) to unbind every passkey a
target account holds — the authenticator remediation that stops a passkey planted
or retained via a transiently-hijacked session from surviving as a standing login
foothold. It wires the previously-uncalled DeleteAllPasskeyCredentialsForUser and
is deliberately not a lockout: the account re-enters via the email-OTP door
(players) or op-login's in-game approval (staff), then re-enrolls. Documented in
the OpenAPI, so the served/documented parity gate covers it.

Remove RevokeUserSessionsExcept: a change-password-era orphan with no callers
since the passwordless migration. Its keep-one ("log out my other devices")
semantics is inherently self-service, and no such slice is on the roadmap; the
admin remediation path already uses RevokeAllUserSessions.
2026-07-04 21:47:12 +09:00
flyemoji 3b43f05a83 refactor(api): drop dead login concurrency limiter and reconcile passwordless comments
The passwordless migration (b330d77) removed the password-login route, leaving
concurrencyLimiter — its bcrypt concurrency cap — with no caller, and scattered
stale "local-password" / "change-password" references through the surviving auth
code's comments.

- Remove the dead concurrencyLimiter (type + newConcurrencyLimiter + acquire):
  no caller, no struct field, no test. Reword the one streamLimiter doc that
  contrasted against it.
- Realign comments in repo.go, pgrepo.go, session.go, util.go to the passwordless
  reality: staff lookups feed email-OTP / passkey / setup redeem, not a password
  compare; RevokeUserSessionsExcept and DeleteAllPasskeyCredentialsForUser are
  retained (uncalled) for the P5 account-remediation path (#78); "local sessions"
  no longer implies a password.

Comments and dead code only; no behavior change. Full WSL test tree green.
2026-07-04 21:47:12 +09:00
flyemoji 0c1cc598c1 feat(auth): migrate console login to passwordless
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.

- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
  login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
  methods an email can use. The single sanctioned existence oracle; methods
  are computed with no role branch, so staff and player accounts in the same
  credential state return byte-identical bodies (staffness invisible by
  construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
  /auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
  (migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
  tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.

Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
2026-07-04 21:47:12 +09:00
Lemon-miaow 3347cc05d5 feat(panel): implement user management administration panel with sessions and minecraft link support 2026-07-04 04:08:14 +08:00
Lemon-miaow 598f3d31f4 feat(submit): local + S3 backends for modpack upload contexts, installer-selectable 2026-07-02 23:38:46 +08:00
flyemoji a63f49dcb3 feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey
WebAuthn is unusable inside the WeChat/QQ in-app WebViews, so a document navigation carrying those UAs is served a bilingual 'open in your system browser' interstitial instead of the passkey-centric SPA. API/config/health/asset requests pass through, and an ack cookie (ua_ack) lets a determined user or false-positive continue. Backend-only; the SPA is untouched.
2026-07-02 19:38:38 +09:00
flyemoji 3fdb3d032e feat(platform): internal API base-URL helper and single-sourced token secret
InternalAPIBaseURL builds the felis-api internal-face DNS from SAAPI and the internal port for cross-namespace callers (the login limbo). The service-token Secret name/key now reference the shared naming constants so the Deployment wiring and the operator's login-pod injection cannot drift.
2026-07-02 19:38:37 +09:00
flyemoji dc23cb54d2 feat(operator): system-server pod readiness probe and login service-token env
buildStatefulSet gates readiness on an HTTP probe when HealthHTTPPort is set (exposing it as a named container port). buildEnv injects FELIS_SERVICE_TOKEN into the login server only — keyed off the reserved name so it can never leak into a user pod — sourced from a Secret via secretKeyRef, never inlined into the CRD.
2026-07-02 19:38:37 +09:00
flyemoji 159107b4e3 feat(api): HTTP readiness knob on MinecraftServer and login-gate fallback default
StartupSpec.HealthHTTPPort/Path switch pod readiness from plain-TCP to an HTTP GET for RCON-less loaders (LOOHP/Limbo) that report 'started' only after the first tick. User servers now default FallbackServer to the login gate, never the lobby, so a stopped/starting backend keeps authentication in front of a fresh connection.
2026-07-02 19:38:37 +09:00
flyemoji 9ef817f2b3 feat(naming): system-server names, validation, and service-token identifiers
SystemLoginServer/SystemLobbyServer plus ValidateSystemServerName (format rule without the reservation check) let the platform provision the reserved login/lobby names users can never claim. ServiceTokenSecretName/Key are the one source of truth for the internal-API credential Secret, shared by the platform renderer and the operator's login-pod injection.
2026-07-02 19:38:37 +09:00
flyemoji 9bed51b67f feat(config): add [velocity] login_image/lobby_image for system servers
setup provisions the always-on login/lobby system services only when these image refs are set; empty means skip-and-say-so (the same fail-loud stance manifests takes), since no official LOOHP/Limbo image exists and a deployment must build its own.
2026-07-02 19:38:37 +09:00
Lemon-miaow 8594622e23 feat(panel): implement email OTP verification and passkey registration management 2026-07-02 18:20:40 +08:00
flyemoji 54bc6ef211 fix(api): clear bound passkeys on password change to close a takeover foothold
handleChangePassword revoked other sessions but never cleared webauthn_credentials, and enrollment needs no step-up. A passkey planted through a transiently-hijacked session needs no password, so it survived the reset + session-revoke as a standing login foothold. Add DeleteAllPasskeyCredentialsForUser and call it in the change-password remediation so every passkey is unbound alongside the session revoke. Removing zero rows is a successful no-op. Email-OTP remains the fallback factor, so this never locks anyone out; the user re-enrolls a passkey afterward if they want one.
2026-07-02 06:55:21 +09:00
flyemoji 7278cd7c6a feat(passkey): require and record user verification at enrollment
Enrollment set no AuthenticatorSelection, so user verification defaulted to preferred (not enforced), and the UV/backup flags the ceremony reported were discarded. Set UserVerification=required so a bound passkey always proves possession AND user (a UV-incapable device falls back to email-OTP), and capture user_verified/backup_eligible/backup_state through VerifiedCredential -> PasskeyCredential -> webauthn_credentials (migration 0009) so a future login path can enforce UV per credential. Adds a negative test proving a presence-only authenticator is rejected, and asserts the roundtrip records UV=true.
2026-07-02 06:55:21 +09:00
flyemoji 20e31fb08f fix(store): cascade-delete passkeys and challenges on user removal
webauthn_credentials.user_id and webauthn_challenges.user_id referenced users(id) with the default ON DELETE NO ACTION, so a future user-delete would either fail or leave orphaned auth material. Recreate both FKs ON DELETE CASCADE: a bound passkey and a pending challenge are ephemeral and must not outlive the account. Scoped to the passkey tables only, not blanket, so retention-bearing child data (world_backups) is not swept away with an account.
2026-07-02 06:55:21 +09:00
flyemoji 99532759b2 fix(api): bound webauthn_challenges growth by superseding all prior rows
The supersede DELETE in CreatePasskeyChallenge filtered consumed_at IS NULL, so it only reaped the prior LIVE challenge; the row that each finish stamps consumed_at on was left behind. A begin->finish loop therefore accumulated one dead row per cycle, unbounded. Drop the consumed_at clause so a fresh begin reaps ALL prior rows for (user, purpose), bounding the table at one row per (user, purpose) with zero net growth per cycle. Deleting an already-consumed row is safe: it has been redeemed and nothing reads it. The fake mirrors the widened supersede.
2026-07-02 06:55:21 +09:00
flyemoji cdbb5abc35 fix(api): record credential id in passkey-register audit event
handlePasskeyRegisterFinish logged an empty target for account.passkey.registered, while the delete half logs the credential id. An operator auditing the log could see that a passkey was bound but not which one. Pass cred.ID as the audit target so bind and unbind are symmetric, and tighten the enrollment test to assert both halves name the credential id.
2026-07-02 06:55:21 +09:00
flyemoji 6368ab1914 fix(api): coalesce MyServers owned flag so ownerless rows do not 500
The MyServers query lists both a user's own servers and unclaimed (owner_id IS NULL) servers, but computed owned as s.owner_id = $1. For an ownerless row that comparison is SQL NULL, which fails to scan into the Go bool and 500s the whole listing. Wrap it in COALESCE(..., false) so an ownerless row reports owned=false while still surfacing as claimable.
2026-07-02 06:55:17 +09:00
Lemon-miaow c0d333bb98 feat(panel): support full server config edit dialog with status prefilling 2026-07-02 04:24:48 +08:00
Lemon-miaow 8ae65ae74a feat(panel): backup management 2026-07-02 03:30:50 +08:00
Lemon-miaow 15c58d982d feat(panel): player management 2026-07-02 03:30:50 +08:00
flyemoji 8f41a003b0 fix(api): clear the SSE write deadline on return so it can't leak to a reused connection
The per-write deadline that severs a stalled SSE reader was never cleared on
return. Server.WriteTimeout is deliberately unset -- a WriteTimeout would sever
a healthy long-lived stream -- and with it unset net/http never resets the
connection write deadline between keep-alive requests. So the deadline the last
writeChunk left set leaks onto the next request that reuses the pooled
connection and fails its first write for no reason. Clear it to the zero value
on return via a deferred rc.SetWriteDeadline; best-effort, a no-op on writers
without deadline support.

Also record honestly at the header flush that the connect-time stall stays
bounded only by the per-principal stream cap, not severed by this guard -- only
the mid-stream stall is closed. Adds a test pinning the clear (fails closed:
neutering the deferred clear leaves a +writeTimeout deadline set on return).
2026-07-01 23:06:26 +09:00
flyemoji 2c56d17989 docs(api): record the quota-claim TOCTOU as a KNOWN-LIMITATION (audit #4)
QuotaAvailable and ClaimServer run as two separate statements, so the
count read is not serialized against a concurrent claim's UPDATE: two
claims by one user for two different ownerless servers can both pass the
gate and both succeed, leaving the user one server over quota. It is low
severity — quota over-provisioning under a deliberate burst, not an
authorization, ownership, or isolation break, since each server is still
claimed atomically via UPDATE ... WHERE owner_id IS NULL.

Closing it requires Postgres transaction semantics (advisory-xact-lock on
the user, or SERIALIZABLE with retry) folding the gate into a single repo
method — verifiable only against a real Postgres, not the hermetic
fakeRepo suite. Documented at QuotaAvailable with back-references from the
two claim gates (handleClaim and the internal UUID claim) rather than
patched blind.
2026-07-01 22:37:36 +09:00
flyemoji d6e3189629 fix(api): bound SSE relay writes with a deadline to sever stalled readers
relayLogStream copied a pod-log follow to the client with a plain
flusher.Flush per event. On a client that stays connected but stops
reading (its TCP receive window shut), net/http buffers the small
"data:" line and only touches the socket at Flush, which then blocks
forever inside the write. The select's <-ctx.Done() branch is never
reached, because r.Context() cancels on an actual disconnect, not on a
stall, so the relay goroutine and its upstream apiserver follow leak for
the life of the process.

Route every event's write+flush through http.ResponseController with a
per-write deadline (writeTimeout, 30s): a stalled flush now returns
os.ErrDeadlineExceeded, the error plain http.Flusher.Flush swallows, and
the relay abandons the stream so the deferred cancel + src.Close release
the follow. SetWriteDeadline and rc.Flush are best-effort: a writer
without deadline support (httptest recorder; some HTTP/2 origins) ignores
the deadline and behaves exactly as before, so the guard degrades
gracefully.

This closes the leak the per-principal stream cap only bounded the blast
radius of. Verified by a deterministic test with a deadline-aware
ResponseWriter whose flush blocks until the deadline; the test times out
(fails closed) if the guard is removed.
2026-07-01 22:30:40 +09:00
flyemoji 3c1d64749f fix(api): cap concurrent SSE streams per principal
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.

Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.

This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
2026-07-01 22:04:19 +09:00
flyemoji 164ac447ef fix(api): validate inbound X-Request-Id before echo and audit persist
withRequestID honored any inbound X-Request-Id verbatim, and that value is
echoed on the response, embedded in the error envelope, and persisted into
audit_logs.request_id. An unvalidated caller-supplied id is therefore an
audit-integrity vector: an arbitrarily long value bloats the audit row, and a
stray control byte (CR/LF) could smuggle a forged entry into a log sink.

Accept an inbound id only when it is well-formed — non-empty, at most 64
bytes, and restricted to a log-safe charset ([A-Za-z0-9._-]) — otherwise mint
a fresh server id. A rejected request loses its inbound trace link, which is
strictly better than storing attacker-controlled text in the audit trail.
2026-07-01 21:07:08 +09:00
flyemoji 7a51c1d9c3 fix(api): bound concurrent login bcrypt to shed CPU-pin floods
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.

Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.

The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
2026-07-01 21:01:28 +09:00
flyemoji f34711c174 docs(api): record passkey login-handler deferral rationale
The passkey login/assertion HTTP handler stays deferred after its design
checkpoint; capture the reasoning in the handler header so the decision is
durable in the repo rather than only in task notes.

- RP boundary (resolved): felis-api is the app-login relying party (panel.*);
  the WebAuthn security gate lives at the Cloudflare Access edge. Spec §14 ties
  WebAuthn/posture to admin.* (Access) while panel.* is plain app login, so
  there is neither a spec-required assertion handler nor a backend step-up
  consumer for one.
- Identifier (blocking): a from-zero login needs a unique, human-typable handle
  to resolve an account, but users.email is nullable and non-unique and a
  player's username is their Minecraft uuid. Username-first assertion has
  nothing to key on; re-link stays the returning-player door.

Discoverable (usernameless) credentials are the future enabler; the adapter
crypto is already verified so that slice inherits correct crypto.
2026-07-01 19:15:44 +09:00
flyemoji e035142abc feat(passkey): add WebAuthn login/assertion crypto adapter
Build the assertion (login) half of the WebAuthn ceremony crypto in the
internal/passkey adapter, Oracle-verified against a virtual authenticator.

- BeginLogin/FinishLogin over go-webauthn BeginLogin/ValidateLogin,
  username-first (allowCredentials scoped to the known user's bound
  passkeys). Discoverable/usernameless login stays out of scope: the
  enrolled credentials are non-resident and the challenge store is
  user-keyed (migration 0007), so it would need a future migration.
- WebAuthnCredentials() now populates the stored COSE public key and
  signature counter (assertion validation needs both to verify the
  signature and detect clones); enrollment ignores them, so the change
  is backward-compatible and the enrollment tests guard it.
- VerifiedAssertion seam output: which credential signed plus the raw
  signature counter. Clone/regression policy is deliberately NOT here —
  the counter is a ceremony fact and the future handler, which holds the
  previously stored counter, decides reject/warn.

Scope: crypto adapter only. The login HTTP handlers, session minting,
and the panel.* relying-party boundary/tier decision remain a deferred
slice (no unauthenticated login route is added). BeginLogin/FinishLogin
live on the concrete adapter, not the api.PasskeyVerifier interface,
which grows only when a handler consumes them.

Tests (virtualwebauthn): a real enrollment chained into a real assertion
exercises the COSE public-key decode path and surfaces the advanced
signature counter, plus origin-mismatch and unbound-credential rejection.
2026-07-01 18:45:26 +09:00
flyemoji 7464fa700b fix(updates): tag Window JSON so the persisted maintenance window round-trips
The admin API persists the auto-update maintenance window as lowercase
JSON {"start","end"} (platform_settings key "update_window"), but
updates.Window had no json tags, so it marshaled/unmarshaled with
capitalized keys. The natural decode the update runner will use --
json.Unmarshal(stored, &updates.Window{}) -- would therefore miss every
key and silently yield the zero Window. That fails closed (a zero window
Contains nothing, so notify-only, never a rogue apply), so it is safe but
a latent silent-zero trap for the not-yet-built runner.

Add json:"start"/json:"end" to updates.Window so the obvious decode is
correct by construction; value time.Time treats a stored null as a no-op,
so a cleared/never-set window still decodes to the zero Window. Nothing
in the package serialized Window before, so this changes no existing
behavior.

Guarded by a cross-package contract test in internal/api that marshals
the real api.updateWindow DTO and unmarshals it into updates.Window --
asserting the interval survives (Contains(mid) is true) and that an empty
window decodes to the fail-closed zero Window -- so the two shapes cannot
drift apart silently.
2026-07-01 18:26:42 +09:00
flyemoji 3673af63c2 feat(api): add admin API for the SysAdmin-set auto-update maintenance window
Two admin-tier routes read and set a single platform-wide maintenance
window for the auto-update subsystem (decision core internal/updates):

  GET /api/v1/updates/window
  PUT /api/v1/updates/window

The window is stored as JSON {"start","end"} (RFC3339, or null when
unset) under the platform_settings key "update_window", reusing the
existing GetSetting/SetSetting KV seam -- no new Repo method, no
migration. Pointer times keep "unset" (null) distinct from a real
instant on both decode and encode; a never-set and an explicitly
cleared window both read back as {null,null}.

Validation mirrors the core's fail-closed Window: a window is either
fully set (both ends, end strictly after start) or fully cleared (both
null). A half-set, inverted, or empty-interval body is 400 and is never
persisted. Reads treat only a missing key as unset (ErrNotFound -> 200
nulls); any other store error 500s rather than fail open.

This is API + PERSISTENCE ONLY. Nothing consumes the stored window yet
-- the runner, the ReleaseSource/Notifier/Applier executors, and the
scheduler CronJob remain INTEGRATION-ONLY. Setting a window changes no
behavior until those land; it is the durable input they will read.
Nothing here force-updates ("不要强制自动更新").
2026-07-01 18:17:07 +09:00
flyemoji fe2ece08cc feat(api): add public Bind-Code onboarding for the player console
Adds POST /api/v1/auth/bind, the one public pre-account entrypoint of the
player console (console.<root_domain>). An account-less player redeems the
one-time Bind Code minted in the in-game Login Lobby; in a single step the
platform creates a role=user player, links it to the verified in-game UUID,
and mints a host-only felis_session. Login is thus not forced at the edge
while operations stay app-authenticated.

The operator console (op.console.<root_domain>) is unaffected and stays
behind Zero Trust: a code whose UUID resolves to a staff (role=admin)
account is refused with 403 (ErrPlayerBindForbidden) without consuming the
code, so the public door provably never yields an admin principal — the
session it mints carries ViaAdminAccess=false and is host-only to console,
never sent to op.console.

Repo layer: new RedeemPlayerBindCode on the Repo interface, implemented on
PGRepo (single tx: resolve code, create-or-fetch the player, consume) and
the test fake. The returning-player branch is idempotent and is a deliberate
standing "log in via the game" door, not just first-time onboarding.

Honest labeling:
- ORACLE-VERIFIED (Go): account/session logic — role=user, refuse-staff,
  idempotent create-or-fetch, single-use code, and the op.console redline
  (player session rejected on admin routes). Covered by handlers_onboard_test
  and the OpenAPI parity gate.
- INTEGRATION-dependent: the endpoint's security rests on the Bind Code having
  been minted against an online-mode-Yggdrasil-authenticated UUID, a
  precondition that lives in velocity/Java and is not verifiable from this
  repo (CODE-ONLY). The Go layer proves the logic, not that identity guarantee.
- No app-level attempt cap: rate-limiting is deferred to the edge as for the
  public /auth/login; the ~1e12 keyspace, single use and short TTL make a
  blind app-level cap non-critical.
2026-07-01 18:00:13 +09:00
flyemoji c01f133cd8 feat(updates): add pure decision core for component self-update
Introduce internal/updates: a pure, I/O-free engine that decides what
should happen to each tracked platform component (Felis control-plane,
k3s, cloudflared, Velocity) given its current version, the latest
discovered upstream, its policy, and the current time.

Updates are never force-applied. A component is Pinned (Minecraft, left
alone), Notify (a SysAdmin is told and applies out of band), or Scheduled
(Felis may apply, but only inside a maintenance window the SysAdmin set).
The load-bearing invariants are unit-tested: a pinned component never
changes, a downgrade is never proposed, a prerelease is never
auto-applied, and an apply happens only inside the window.

Version parsing tolerates the real feeds (leading v, k3s +k3s1 build
suffix, calendar versions, prerelease tails) and orders by SemVer
precedence. ReleaseSource/Notifier/Applier are declared as integration
seams and exercised via fakes; this package ships no network, SMTP, or
kubectl, and deliberately has no blind k3s-upgrade applier.
2026-07-01 15:59:33 +09:00
flyemoji a531f5e42a fix(cfsetup): keep connector install in the host apply layer only
Setup previously called runner.StartConnector (`cloudflared service install`)
while the TUI applyCloudflareEdge separately installs cloudflared-felis.service
for the same tunnel from the same config -- two managed services serving one
tunnel from one connector config.

Drop StartConnector from cfsetup: running a connector is a host-specific side
effect (systemd/launchd/Windows service) that belongs with the caller, not in
this host- and domain-agnostic package whose documented side effects are tunnel
creation, DNS routing, and the Access app/policy calls. installCloudflaredService
in the host layer stays the single connector installer, so the routed-but-dead
1033 is still closed; RouteDNS --overwrite-dns still closes the stale-DNS 1033.
2026-07-01 15:31:58 +09:00
flyemoji 7d3be64919 feat(cfsetup): start the tunnel connector as a setup step
Setup created the tunnel, routed DNS, and wrote config.yml, but nothing
installed or started a connector for it. A one-click run therefore left the
tunnel routed-but-dead: every web hostname returned Cloudflare error 1033
(tunnel has no connector) even though the config was correct on disk.

Add a StartConnector step to the Runner seam, invoked right after the config
is written (and gated on ConfigPath, so a caller wanting only the Access
config is not forced to install a service). The ExecRunner implementation
runs `cloudflared --config <path> service install`, which installs and starts
a managed system service (systemd/launchd/Windows), and is idempotent on an
already-installed service. The orchestration — connector started, and only
after its config exists — is unit-tested against the fake Runner; the actual
service install is INTEGRATION-ONLY.

Together with the RouteDNS --overwrite-dns fix, this closes both distinct
paths to a 1033 half-state from a fresh setup: a stale DNS binding and a
missing connector.
2026-07-01 15:11:35 +09:00
flyemoji 2810fe849c fix(cfsetup): repoint stale DNS record when routing a tunnel hostname
RouteDNS ran `cloudflared tunnel route dns` without --overwrite-dns and
swallowed the resulting "record already exists" error as success. When a
hostname already had a CNAME from an earlier tunnel that was deleted and
recreated, the record stayed bound to the dead tunnel: the setup reported
the hostname "routed" while it kept returning Cloudflare error 1033 (the
tunnel it pointed at has no connector).

Pass --overwrite-dns so the record is repointed at the tunnel just created,
making the route idempotent and correct on every re-run, and drop the
now-unnecessary "already exists" swallow. INTEGRATION-ONLY (ExecRunner
shells out to the real cloudflared binary).
2026-07-01 15:09:04 +09:00
flyemoji 0261204979 feat(passkey): add go-webauthn enrollment verifier adapter
Wrap github.com/go-webauthn/webauthn behind the api.PasskeyVerifier
seam so the api package stays free of go-webauthn types. The adapter
covers the credential-creation ceremony only (BeginRegistration /
CreateCredential); the login/assertion path is a deferred slice.

Ceremony state crosses the seam as opaque marshaled SessionData, the
attestation as an io.Reader, and the verified result as a plain
VerifiedCredential. SessionData carries no expiry so the challenge
row's TTL stays the single liveness authority. New rejects an empty
RP id or origin list so a misconfigured deployment fails at
construction rather than minting unverifiable challenges.

Tests drive a real relying party against a virtual authenticator
(descope/virtualwebauthn): a full creation round-trip plus adversarial
guards proving origin-mismatch and user-mismatch are rejected and
already-bound credentials are excluded.
2026-07-01 14:36:28 +09:00
Lemon-miaow d2de11af22 feat(panel): fleet 2026-07-01 03:46:07 +08:00
flyemoji 742f15f348 feat(api): add passkey enrollment endpoints
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14), web app face.
An already-authenticated principal binds a passkey to their own account and
manages the credentials they have bound; email-OTP stays the fallback factor.

- four account routes: POST register/begin mints a credential-creation
  challenge, POST register/finish verifies the attestation against the
  server-stashed SessionData and binds the credential, GET/DELETE credentials
  list and unbind the caller's OWN passkeys. App-tier, principal-scoped (the
  body never names a user).
- PasskeyVerifier seam keeps go-webauthn out of this package: ceremony state
  crosses as opaque bytes, attestation as an io.Reader, result as a plain
  VerifiedCredential. A nil verifier makes begin/finish report 503 so the
  authenticated boundary is exercised before the real verifier is wired in.
- the view never leaks the public key; credential_id collisions map to 409.
- OpenAPI: the four paths plus the PasskeyCredential schema, keeping the
  served-routes parity gate green.

Scope: ENROLLMENT only. The passkey login/assertion path (proving a passkey
from an unauthenticated state) is deferred; every ceremony here rides on a
known principal.

Tests: handler + challenge state machine against a fake repo and a fake
verifier (no real attestation crypto, no SQL). The decisive assertion is the
session-data round-trip -- the finish body carries no challenge, so the only
path for the stashed blob into FinishRegistration is store-stash then consume,
proving the challenge is server-held and never client-echoed. Also covers
supersede-on-begin, single-use, expiry, 503-unavailable, 409-already-bound,
owner-scoped list/delete, and external-only face separation.
2026-07-01 02:14:29 +09:00
flyemoji f2c916d378 feat(api): add passkey enrollment persistence layer
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14). Adds the data
layer an already-authenticated principal needs to bind and manage passkeys:

- migration 0007: webauthn_credentials (one bound passkey per row, public
  attestation material only) and webauthn_challenges (server-stashed ceremony
  state between begin and finish, single-use via consumed_at). Both rows are
  bound to a known user_id; there is no usernameless login lookup, since the
  assertion/login path is a deferred slice.
- PasskeyCredential type and five Repo methods (create/consume challenge,
  create/list/delete credential) with the PG semantics the handlers rely on:
  supersede-prior-live on begin, expiry-before-consume single-use on finish,
  credential_id UNIQUE -> ErrConflict, owner-scoped delete -> ErrNotFound.
- ErrPasskeyChallengeInvalid sentinel for a missing/expired/consumed ceremony.
2026-07-01 02:14:28 +09:00