Commit Graph
93 Commits
Author SHA1 Message Date
flyemoji 154002edf3 docs(auth): cite MultiLogin reference for UUID-keyed reclaim split
Anchor the username-collision reclaim's UUID-keyed, proxy-detected design to
the multi-Yggdrasil reference: CaaMoe/MultiLogin v6 binds identity as
serviceId+online-UUID via "identity cards" that decouple the in-game name from
online identity — keyed by UUID, never by name. Note that §B3's Mojang-priority
reclaim goes beyond the common "protect the first-bound name" behavior by
evicting a squatter once the genuine Mojang owner appears and stashing the
squatter's data for the code-only inherit path.
2026-07-05 04:05:56 +09:00
flyemoji ec468baef9 feat(auth): add discoverable (usernameless) passkey login
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.

Routes (both Public):
  POST /api/v1/auth/passkey/login/discoverable/begin
  POST /api/v1/auth/passkey/login/discoverable/finish

Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.

The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.

Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
2026-07-05 04:05:56 +09:00
flyemoji 7db57b9fff feat(updater): add VersionGatherer extraction core and CLI gather seam
Give the Runner a way to read each component's CURRENT version so it can be
compared against the release sources already wired. Three pure extractors turn
raw system text into an updates.Version, each fail-closed:

  - versionFromCLI      — a `<tool> --version` banner   (k3s, cloudflared)
  - versionFromImageRef — a container image tag         (felis-api)
  - versionFromJarName  — a proxy jar filename          (velocity)

sysGatherer routes each Topology component to the right extractor over an
injected seam; every path is exercised with a fake runner, mirroring how the
release sources are proven against httptest.

The load-bearing case is k3s: its Git tag "v1.36.2+k3s1" parses stable, but a
registry cannot store '+', so the same build ships as image tag "v1.36.2-k3s1",
which parses as a prerelease unless repaired. versionFromImageRef normalizes
"-k3sN"/"-rke2rN" back to "+", so an image read and a CLI read agree instead of
the image masquerading as a prerelease and being barred from comparison.

Honest runtime state after this slice — a green suite is not "the updater runs
against real infra": only the CLI seam (execRunner) is wired, so of the four
tracked components just cloudflared is live end to end (gatherable AND
Scheduled/appliable). k3s is CLI-gatherable but Notify-only. felis-api and
velocity are NOT yet runtime-gatherable: their producing seams — a k8s read of
the control-plane Deployment image, and an off-cluster jar inspection — are left
nil, so both surface an explicit "gather seam not wired" error rather than a
wrong version. felis-api self-update is therefore not functional yet.

Remaining integration (tracked in doc.go): the two producing seams, the concrete
Notifier (SMTP + in-game), the Applier (image bump, cloudflared swap), the
`felis update` CLI + CronJob entry point, and the runtime append of the Pinned
Minecraft fleet.
2026-07-05 04:05:56 +09:00
flyemoji 7d27640c07 feat(updater): add GitHub Releases source and route felis-api/k3s/cloudflared
Give RoutingSource its second upstream so every non-pinned component now
resolves a real latest-stable: Velocity via PaperMC (already wired), and
felis-api, k3s and cloudflared via the GitHub REST API.

github.go queries /repos/{repo}/releases/latest (one request, rate-limit
friendly) and fails closed: a transport error, a non-200 status (404 = no
stable release), an undecodable body, a draft/prerelease flag, or an
unparseable / prerelease-parsing tag all return an error, never a zero
version. It sends the User-Agent GitHub requires (a UA-less request is
403'd) and tolerates the two live tag styles -- cloudflared's CalVer
"2026.6.1" and k3s's v-prefixed, build-tagged "v1.36.2+k3s1" -- while
String() keeps the raw tag for the report.

source.go routes sourceGitHub to it and drops the errGitHubNotWired stub;
velocity still routes to PaperMC.

Tests: github_test.go covers both tag styles, the User-Agent gate, and
fail-closed on 404 / prerelease-flag / unparseable tag, with fixtures
captured from api.github.com on 2026-07-05. runner_test.go now drives
PaperMC and GitHub through dual httptest servers end to end with no source
degrading to an error.

doc.go re-tiers the verification boundary: both release sources are now
built and live-grounded; the VersionGatherer's version-extraction core is
the next verifiable slice (logic over an exec seam, not pure I/O); the
genuine I/O remainder is the Notifier, Applier and felis update CLI/CronJob.
felis-api's coord is still a placeholder slug, so that component is dark at
runtime until a real repository is configured.
2026-07-05 01:03:32 +09:00
flyemoji 9896fe16c3 docs(updater): correct PaperMC UA/fixture overclaims, re-tier the boundary
An out-of-band curl of the live Fill v3 endpoint contradicted two claims the
previous commit shipped and surfaced a mis-tiering:

- User-Agent is NOT enforced: fill.papermc.io/v3/projects/velocity returned
  HTTP 200 to a bare curl UA. The comments claimed a generic UA "is refused"
  and the API "REQUIRES" a contact UA. Reword to what is true — PaperMC's usage
  policy asks for a descriptive UA and may block generic ones, but sending it is
  etiquette/defensive here, not a gate Felis depends on.
- The test fixture's shape was invented, not captured: the real "versions"
  object groups the entire 3.x line under a single key "3.0.0", not the
  per-minor keys the fixture used. Replace it with the real body (keys and
  version strings as returned). The key-agnostic parser already produced the
  right answer, and an independent max-stable check confirms 3.4.0.
- Re-tier doc.go: the GitHub Releases source is verifiable-here (the same
  httptest-testable shape as PaperMC), not integration remainder. It is why
  3 of 4 components report "latest unknown" today and is the next verifiable
  slice — the release-source work is only ~half done until it exists.

No production logic changed. WSL oracle: build + vet clean, internal/updater
10/10, full tree go test RC=0 (19 ok, 0 fail).
2026-07-05 00:28:36 +09:00
flyemoji 96b3cc901c feat(updater): wire updates.Run to a caller with PaperMC v3 release discovery
internal/updates is a pure, fakes-tested decision core with no production caller,
so nothing could produce its "版本号状态" report. Add internal/updater as that caller:

- topology: the fixed platform components and their user-set policies (felis-api
  and cloudflared Scheduled+manageable; k3s Notify, high-blast-radius single node;
  velocity Notify, off-cluster and unmanageable). Minecraft is pinned by ABSENCE,
  never force-tracked here, appended from the live fleet at runtime.
- PaperMC Fill v3 release source: the v2 API (api.papermc.io) was retired
  2026-07-01 and returns HTTP 410, so this targets fill.papermc.io/v3, sends the
  required non-generic User-Agent, and returns the newest STABLE version, filtering
  the -SNAPSHOT/rc prereleases the plan would otherwise suppress. Its test fixture
  is captured from the live v3 response shape (2026-07-04).
- RoutingSource: the single ReleaseSource updates.Run requires, dispatching
  velocity to PaperMC and returning errGitHubNotWired for the GitHub-backed
  components so they degrade to "latest unknown" honestly, never a fabricated one.
- Runner: gather current versions (seam) -> assemble Components -> updates.Run ->
  Report; report-only when notifier and applier are nil.

Verification boundary: the parse/plan/compose logic is unit-tested (httptest +
fakes, fixture grounded in the live v3 shape). Live network/TLS/User-Agent
enforcement, the GitHub Releases source, the concrete version gatherer, the
notifier and applier, and the felis update CLI/CronJob remain integration work,
enumerated in doc.go.
2026-07-04 23:41:46 +09:00
flyemoji 0a2accd245 chore: stop tracking Autohand-generated AGENTS.md
Added by the tooling in commit 0c1cc59, not authored guidance. Untrack and gitignore it: the file claims precedence over CLAUDE.md and tells agents to run go fmt, which rewrites the CRLF working tree. The file stays on disk (git rm --cached) so local tooling keeps it, but it is no longer tracked or committed.
2026-07-04 23:11:25 +09:00
flyemoji c20b12c655 refactor(api): drop dead password-era ResetMailer, reconcile passkey-unbind docs
The passwordless migration left ResetMailer (SendPasswordReset) and its API field with zero callers and no wiring; the web console authenticates via email-OTP and passkey only. Remove both, plus the now-orphaned context import that the interface was the last user of in handlers_users.go.

Reconcile the DeleteAllPasskeyCredentialsForUser docs in repo.go and pgrepo.go: they claimed there was no production caller, but 2f22027 wired the owner-tier DELETE /users/{id}/passkeys. Both now note that a complete authenticator remediation pairs the unbind with a session revoke (unbinding alone leaves the live hijacked session; revoking alone leaves a re-enrollable credential), and the OpenAPI operation carries the same guidance in a new description. Reword the stale local-password test-fake header, since the passwordless fakes carry no must_change_password field.

No behavior change. gofmt, build, and the full test tree are green; OpenAPI parity and passkey-unbind tests pass; a grep confirms ResetMailer/SendPasswordReset are gone from the Go tree.
2026-07-04 21:47:13 +09:00
flyemoji 4f59d5128a feat(auth): add owner-tier passkey-unbind remediation endpoint
Add DELETE /api/v1/users/{id}/passkeys (owner-only) to unbind every passkey a
target account holds — the authenticator remediation that stops a passkey planted
or retained via a transiently-hijacked session from surviving as a standing login
foothold. It wires the previously-uncalled DeleteAllPasskeyCredentialsForUser and
is deliberately not a lockout: the account re-enters via the email-OTP door
(players) or op-login's in-game approval (staff), then re-enrolls. Documented in
the OpenAPI, so the served/documented parity gate covers it.

Remove RevokeUserSessionsExcept: a change-password-era orphan with no callers
since the passwordless migration. Its keep-one ("log out my other devices")
semantics is inherently self-service, and no such slice is on the roadmap; the
admin remediation path already uses RevokeAllUserSessions.
2026-07-04 21:47:12 +09:00
flyemoji 3b43f05a83 refactor(api): drop dead login concurrency limiter and reconcile passwordless comments
The passwordless migration (b330d77) removed the password-login route, leaving
concurrencyLimiter — its bcrypt concurrency cap — with no caller, and scattered
stale "local-password" / "change-password" references through the surviving auth
code's comments.

- Remove the dead concurrencyLimiter (type + newConcurrencyLimiter + acquire):
  no caller, no struct field, no test. Reword the one streamLimiter doc that
  contrasted against it.
- Realign comments in repo.go, pgrepo.go, session.go, util.go to the passwordless
  reality: staff lookups feed email-OTP / passkey / setup redeem, not a password
  compare; RevokeUserSessionsExcept and DeleteAllPasskeyCredentialsForUser are
  retained (uncalled) for the P5 account-remediation path (#78); "local sessions"
  no longer implies a password.

Comments and dead code only; no behavior change. Full WSL test tree green.
2026-07-04 21:47:12 +09:00
flyemoji 0c1cc598c1 feat(auth): migrate console login to passwordless
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.

- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
  login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
  methods an email can use. The single sanctioned existence oracle; methods
  are computed with no role branch, so staff and player accounts in the same
  credential state return byte-identical bodies (staffness invisible by
  construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
  /auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
  (migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
  tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.

Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
2026-07-04 21:47:12 +09:00
flyemoji b84debf872 feat(deploy): one-shot demo bring-up wrapper
demo-up.sh collapses bootstrap -> build+import the limbo/lobby images -> wire [velocity] login_image/lobby_image into felis.host.toml -> felis setup into a single command, ending in the interactive Owner-creation TUI (the only step it cannot automate). Prefers prebuilt tars under deploy/images, else builds on the host, auto-resolving the LOOHP/Limbo CI jar and the latest stable Paper jar (all overridable by env); SKIP_BOOTSTRAP/SKIP_SETUP toggles for reruns.
2026-07-03 01:26:15 +09:00
flyemoji d9e866fcdb fix(deploy): make the lobby image actually build
The lobby image had never been built and two defects blocked it: the felis image .dockerignore excluded plugins/* and only re-included limbo/shared, so the lobby Dockerfile's COPY plugins/paper landed empty; and the plugin stage used eclipse-temurin:21-jdk, which ships no gradle (and the tree vendors no wrapper), failing with 'gradle: not found'. Re-include plugins/paper and build the paper plugin on gradle:8.14-jdk21, matching the limbo image. Verified: both images build and boot (limbo /healthz 200 on 25565; lobby reaches 'Done' with felis-paper enabled).
2026-07-03 01:26:15 +09:00
flyemoji c7315e44b6 feat(deploy): login-limbo and lobby images with game-port pinning
deploy/limbo assembles LOOHP/Limbo from its loose CI artifacts plus the felis-limbo plugin (and the shared link core), with an entrypoint that pins server-port to the operator's GamePort (25565) on every start. deploy/lobby carries the Paper + felis-paper hub image. .dockerignore re-includes plugins/limbo and plugins/shared so the plugin image build sees them.
2026-07-02 19:38:38 +09:00
flyemoji 241fe21f8a feat(limbo): felis-limbo in-game login flow over the shared account-link client
The login limbo now performs the onboarding inside Limbo: on join it checks the collision blacklist, mints a bind code, opens a book linking the player to console.<root_domain> (guiding them to the system browser), polls link-status, and transfers to the lobby via BungeeCord Connect — fail-closed on blacklist, mint/transport error, or window elapse. FelisApiClient gains linkStatus/isBlacklisted on the existing internal transport.
2026-07-02 19:38:38 +09:00
flyemoji a63f49dcb3 feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey
WebAuthn is unusable inside the WeChat/QQ in-app WebViews, so a document navigation carrying those UAs is served a bilingual 'open in your system browser' interstitial instead of the passkey-centric SPA. API/config/health/asset requests pass through, and an ack cookie (ua_ack) lets a determined user or false-positive continue. Backend-only; the SPA is untouched.
2026-07-02 19:38:38 +09:00
flyemoji f554d525d4 feat(cli): provision login/lobby system servers with login env and token replica
setup builds the always-on, reaper-exempt login/lobby MinecraftServers (create-if-absent), bakes the login limbo's non-secret config (internal API URL, root domain, lobby name) into spec.env, and replicates the felis-service-token Secret from the control namespace into the minecraft namespace so the operator's namespace-local secretKeyRef on the login pod resolves.
2026-07-02 19:38:37 +09:00
flyemoji 3fdb3d032e feat(platform): internal API base-URL helper and single-sourced token secret
InternalAPIBaseURL builds the felis-api internal-face DNS from SAAPI and the internal port for cross-namespace callers (the login limbo). The service-token Secret name/key now reference the shared naming constants so the Deployment wiring and the operator's login-pod injection cannot drift.
2026-07-02 19:38:37 +09:00
flyemoji dc23cb54d2 feat(operator): system-server pod readiness probe and login service-token env
buildStatefulSet gates readiness on an HTTP probe when HealthHTTPPort is set (exposing it as a named container port). buildEnv injects FELIS_SERVICE_TOKEN into the login server only — keyed off the reserved name so it can never leak into a user pod — sourced from a Secret via secretKeyRef, never inlined into the CRD.
2026-07-02 19:38:37 +09:00
flyemoji 159107b4e3 feat(api): HTTP readiness knob on MinecraftServer and login-gate fallback default
StartupSpec.HealthHTTPPort/Path switch pod readiness from plain-TCP to an HTTP GET for RCON-less loaders (LOOHP/Limbo) that report 'started' only after the first tick. User servers now default FallbackServer to the login gate, never the lobby, so a stopped/starting backend keeps authentication in front of a fresh connection.
2026-07-02 19:38:37 +09:00
flyemoji 9ef817f2b3 feat(naming): system-server names, validation, and service-token identifiers
SystemLoginServer/SystemLobbyServer plus ValidateSystemServerName (format rule without the reservation check) let the platform provision the reserved login/lobby names users can never claim. ServiceTokenSecretName/Key are the one source of truth for the internal-API credential Secret, shared by the platform renderer and the operator's login-pod injection.
2026-07-02 19:38:37 +09:00
flyemoji 9bed51b67f feat(config): add [velocity] login_image/lobby_image for system servers
setup provisions the always-on login/lobby system services only when these image refs are set; empty means skip-and-say-so (the same fail-loud stance manifests takes), since no official LOOHP/Limbo image exists and a deployment must build its own.
2026-07-02 19:38:37 +09:00
flyemoji 54bc6ef211 fix(api): clear bound passkeys on password change to close a takeover foothold
handleChangePassword revoked other sessions but never cleared webauthn_credentials, and enrollment needs no step-up. A passkey planted through a transiently-hijacked session needs no password, so it survived the reset + session-revoke as a standing login foothold. Add DeleteAllPasskeyCredentialsForUser and call it in the change-password remediation so every passkey is unbound alongside the session revoke. Removing zero rows is a successful no-op. Email-OTP remains the fallback factor, so this never locks anyone out; the user re-enrolls a passkey afterward if they want one.
2026-07-02 06:55:21 +09:00
flyemoji 7278cd7c6a feat(passkey): require and record user verification at enrollment
Enrollment set no AuthenticatorSelection, so user verification defaulted to preferred (not enforced), and the UV/backup flags the ceremony reported were discarded. Set UserVerification=required so a bound passkey always proves possession AND user (a UV-incapable device falls back to email-OTP), and capture user_verified/backup_eligible/backup_state through VerifiedCredential -> PasskeyCredential -> webauthn_credentials (migration 0009) so a future login path can enforce UV per credential. Adds a negative test proving a presence-only authenticator is rejected, and asserts the roundtrip records UV=true.
2026-07-02 06:55:21 +09:00
flyemoji 20e31fb08f fix(store): cascade-delete passkeys and challenges on user removal
webauthn_credentials.user_id and webauthn_challenges.user_id referenced users(id) with the default ON DELETE NO ACTION, so a future user-delete would either fail or leave orphaned auth material. Recreate both FKs ON DELETE CASCADE: a bound passkey and a pending challenge are ephemeral and must not outlive the account. Scoped to the passkey tables only, not blanket, so retention-bearing child data (world_backups) is not swept away with an account.
2026-07-02 06:55:21 +09:00
flyemoji 99532759b2 fix(api): bound webauthn_challenges growth by superseding all prior rows
The supersede DELETE in CreatePasskeyChallenge filtered consumed_at IS NULL, so it only reaped the prior LIVE challenge; the row that each finish stamps consumed_at on was left behind. A begin->finish loop therefore accumulated one dead row per cycle, unbounded. Drop the consumed_at clause so a fresh begin reaps ALL prior rows for (user, purpose), bounding the table at one row per (user, purpose) with zero net growth per cycle. Deleting an already-consumed row is safe: it has been redeemed and nothing reads it. The fake mirrors the widened supersede.
2026-07-02 06:55:21 +09:00
flyemoji cdbb5abc35 fix(api): record credential id in passkey-register audit event
handlePasskeyRegisterFinish logged an empty target for account.passkey.registered, while the delete half logs the credential id. An operator auditing the log could see that a passkey was bound but not which one. Pass cred.ID as the audit target so bind and unbind are symmetric, and tighten the enrollment test to assert both halves name the credential id.
2026-07-02 06:55:21 +09:00
flyemoji 6368ab1914 fix(api): coalesce MyServers owned flag so ownerless rows do not 500
The MyServers query lists both a user's own servers and unclaimed (owner_id IS NULL) servers, but computed owned as s.owner_id = $1. For an ownerless row that comparison is SQL NULL, which fails to scan into the Go bool and 500s the whole listing. Wrap it in COALESCE(..., false) so an ownerless row reports owned=false while still surfacing as claimable.
2026-07-02 06:55:17 +09:00
flyemoji 8f41a003b0 fix(api): clear the SSE write deadline on return so it can't leak to a reused connection
The per-write deadline that severs a stalled SSE reader was never cleared on
return. Server.WriteTimeout is deliberately unset -- a WriteTimeout would sever
a healthy long-lived stream -- and with it unset net/http never resets the
connection write deadline between keep-alive requests. So the deadline the last
writeChunk left set leaks onto the next request that reuses the pooled
connection and fails its first write for no reason. Clear it to the zero value
on return via a deferred rc.SetWriteDeadline; best-effort, a no-op on writers
without deadline support.

Also record honestly at the header flush that the connect-time stall stays
bounded only by the per-principal stream cap, not severed by this guard -- only
the mid-stream stall is closed. Adds a test pinning the clear (fails closed:
neutering the deferred clear leaves a +writeTimeout deadline set on return).
2026-07-01 23:06:26 +09:00
flyemoji 2c56d17989 docs(api): record the quota-claim TOCTOU as a KNOWN-LIMITATION (audit #4)
QuotaAvailable and ClaimServer run as two separate statements, so the
count read is not serialized against a concurrent claim's UPDATE: two
claims by one user for two different ownerless servers can both pass the
gate and both succeed, leaving the user one server over quota. It is low
severity — quota over-provisioning under a deliberate burst, not an
authorization, ownership, or isolation break, since each server is still
claimed atomically via UPDATE ... WHERE owner_id IS NULL.

Closing it requires Postgres transaction semantics (advisory-xact-lock on
the user, or SERIALIZABLE with retry) folding the gate into a single repo
method — verifiable only against a real Postgres, not the hermetic
fakeRepo suite. Documented at QuotaAvailable with back-references from the
two claim gates (handleClaim and the internal UUID claim) rather than
patched blind.
2026-07-01 22:37:36 +09:00
flyemoji d6e3189629 fix(api): bound SSE relay writes with a deadline to sever stalled readers
relayLogStream copied a pod-log follow to the client with a plain
flusher.Flush per event. On a client that stays connected but stops
reading (its TCP receive window shut), net/http buffers the small
"data:" line and only touches the socket at Flush, which then blocks
forever inside the write. The select's <-ctx.Done() branch is never
reached, because r.Context() cancels on an actual disconnect, not on a
stall, so the relay goroutine and its upstream apiserver follow leak for
the life of the process.

Route every event's write+flush through http.ResponseController with a
per-write deadline (writeTimeout, 30s): a stalled flush now returns
os.ErrDeadlineExceeded, the error plain http.Flusher.Flush swallows, and
the relay abandons the stream so the deferred cancel + src.Close release
the follow. SetWriteDeadline and rc.Flush are best-effort: a writer
without deadline support (httptest recorder; some HTTP/2 origins) ignores
the deadline and behaves exactly as before, so the guard degrades
gracefully.

This closes the leak the per-principal stream cap only bounded the blast
radius of. Verified by a deterministic test with a deadline-aware
ResponseWriter whose flush blocks until the deadline; the test times out
(fails closed) if the guard is removed.
2026-07-01 22:30:40 +09:00
flyemoji 3c1d64749f fix(api): cap concurrent SSE streams per principal
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.

Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.

This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
2026-07-01 22:04:19 +09:00
flyemoji c6c0772a7a fix(api): set read/idle timeouts on the felis-api listeners
The three felis-api http.Servers (internal, external, https) were built with
only Addr and Handler, leaving ReadHeaderTimeout, IdleTimeout, and ReadTimeout
at zero. A zero ReadHeaderTimeout is a Slowloris hole — a client trickling
header bytes pins a connection indefinitely — and a zero IdleTimeout lets
kept-alive connections accumulate (gosec G112).

Route all three listeners through a newAPIServer factory that sets a 10s
ReadHeaderTimeout and a 120s IdleTimeout. WriteTimeout and ReadTimeout are
left unset on purpose: the external and https faces stream Server-Sent Events
(console / build logs) for the lifetime of a client attachment, and a
WriteTimeout would sever a healthy long-lived stream. Slowloris is closed by
ReadHeaderTimeout, which bounds only the header phase.
2026-07-01 21:12:03 +09:00
flyemoji 164ac447ef fix(api): validate inbound X-Request-Id before echo and audit persist
withRequestID honored any inbound X-Request-Id verbatim, and that value is
echoed on the response, embedded in the error envelope, and persisted into
audit_logs.request_id. An unvalidated caller-supplied id is therefore an
audit-integrity vector: an arbitrarily long value bloats the audit row, and a
stray control byte (CR/LF) could smuggle a forged entry into a log sink.

Accept an inbound id only when it is well-formed — non-empty, at most 64
bytes, and restricted to a log-safe charset ([A-Za-z0-9._-]) — otherwise mint
a fresh server id. A rejected request loses its inbound trace link, which is
strictly better than storing attacker-controlled text in the audit trail.
2026-07-01 21:07:08 +09:00
flyemoji 7a51c1d9c3 fix(api): bound concurrent login bcrypt to shed CPU-pin floods
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.

Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.

The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
2026-07-01 21:01:28 +09:00
flyemoji f34711c174 docs(api): record passkey login-handler deferral rationale
The passkey login/assertion HTTP handler stays deferred after its design
checkpoint; capture the reasoning in the handler header so the decision is
durable in the repo rather than only in task notes.

- RP boundary (resolved): felis-api is the app-login relying party (panel.*);
  the WebAuthn security gate lives at the Cloudflare Access edge. Spec §14 ties
  WebAuthn/posture to admin.* (Access) while panel.* is plain app login, so
  there is neither a spec-required assertion handler nor a backend step-up
  consumer for one.
- Identifier (blocking): a from-zero login needs a unique, human-typable handle
  to resolve an account, but users.email is nullable and non-unique and a
  player's username is their Minecraft uuid. Username-first assertion has
  nothing to key on; re-link stays the returning-player door.

Discoverable (usernameless) credentials are the future enabler; the adapter
crypto is already verified so that slice inherits correct crypto.
2026-07-01 19:15:44 +09:00
flyemoji e035142abc feat(passkey): add WebAuthn login/assertion crypto adapter
Build the assertion (login) half of the WebAuthn ceremony crypto in the
internal/passkey adapter, Oracle-verified against a virtual authenticator.

- BeginLogin/FinishLogin over go-webauthn BeginLogin/ValidateLogin,
  username-first (allowCredentials scoped to the known user's bound
  passkeys). Discoverable/usernameless login stays out of scope: the
  enrolled credentials are non-resident and the challenge store is
  user-keyed (migration 0007), so it would need a future migration.
- WebAuthnCredentials() now populates the stored COSE public key and
  signature counter (assertion validation needs both to verify the
  signature and detect clones); enrollment ignores them, so the change
  is backward-compatible and the enrollment tests guard it.
- VerifiedAssertion seam output: which credential signed plus the raw
  signature counter. Clone/regression policy is deliberately NOT here —
  the counter is a ceremony fact and the future handler, which holds the
  previously stored counter, decides reject/warn.

Scope: crypto adapter only. The login HTTP handlers, session minting,
and the panel.* relying-party boundary/tier decision remain a deferred
slice (no unauthenticated login route is added). BeginLogin/FinishLogin
live on the concrete adapter, not the api.PasskeyVerifier interface,
which grows only when a handler consumes them.

Tests (virtualwebauthn): a real enrollment chained into a real assertion
exercises the COSE public-key decode path and surfaces the advanced
signature counter, plus origin-mismatch and unbound-credential rejection.
2026-07-01 18:45:26 +09:00
flyemoji 7464fa700b fix(updates): tag Window JSON so the persisted maintenance window round-trips
The admin API persists the auto-update maintenance window as lowercase
JSON {"start","end"} (platform_settings key "update_window"), but
updates.Window had no json tags, so it marshaled/unmarshaled with
capitalized keys. The natural decode the update runner will use --
json.Unmarshal(stored, &updates.Window{}) -- would therefore miss every
key and silently yield the zero Window. That fails closed (a zero window
Contains nothing, so notify-only, never a rogue apply), so it is safe but
a latent silent-zero trap for the not-yet-built runner.

Add json:"start"/json:"end" to updates.Window so the obvious decode is
correct by construction; value time.Time treats a stored null as a no-op,
so a cleared/never-set window still decodes to the zero Window. Nothing
in the package serialized Window before, so this changes no existing
behavior.

Guarded by a cross-package contract test in internal/api that marshals
the real api.updateWindow DTO and unmarshals it into updates.Window --
asserting the interval survives (Contains(mid) is true) and that an empty
window decodes to the fail-closed zero Window -- so the two shapes cannot
drift apart silently.
2026-07-01 18:26:42 +09:00
flyemoji 3673af63c2 feat(api): add admin API for the SysAdmin-set auto-update maintenance window
Two admin-tier routes read and set a single platform-wide maintenance
window for the auto-update subsystem (decision core internal/updates):

  GET /api/v1/updates/window
  PUT /api/v1/updates/window

The window is stored as JSON {"start","end"} (RFC3339, or null when
unset) under the platform_settings key "update_window", reusing the
existing GetSetting/SetSetting KV seam -- no new Repo method, no
migration. Pointer times keep "unset" (null) distinct from a real
instant on both decode and encode; a never-set and an explicitly
cleared window both read back as {null,null}.

Validation mirrors the core's fail-closed Window: a window is either
fully set (both ends, end strictly after start) or fully cleared (both
null). A half-set, inverted, or empty-interval body is 400 and is never
persisted. Reads treat only a missing key as unset (ErrNotFound -> 200
nulls); any other store error 500s rather than fail open.

This is API + PERSISTENCE ONLY. Nothing consumes the stored window yet
-- the runner, the ReleaseSource/Notifier/Applier executors, and the
scheduler CronJob remain INTEGRATION-ONLY. Setting a window changes no
behavior until those land; it is the durable input they will read.
Nothing here force-updates ("不要强制自动更新").
2026-07-01 18:17:07 +09:00
flyemoji fe2ece08cc feat(api): add public Bind-Code onboarding for the player console
Adds POST /api/v1/auth/bind, the one public pre-account entrypoint of the
player console (console.<root_domain>). An account-less player redeems the
one-time Bind Code minted in the in-game Login Lobby; in a single step the
platform creates a role=user player, links it to the verified in-game UUID,
and mints a host-only felis_session. Login is thus not forced at the edge
while operations stay app-authenticated.

The operator console (op.console.<root_domain>) is unaffected and stays
behind Zero Trust: a code whose UUID resolves to a staff (role=admin)
account is refused with 403 (ErrPlayerBindForbidden) without consuming the
code, so the public door provably never yields an admin principal — the
session it mints carries ViaAdminAccess=false and is host-only to console,
never sent to op.console.

Repo layer: new RedeemPlayerBindCode on the Repo interface, implemented on
PGRepo (single tx: resolve code, create-or-fetch the player, consume) and
the test fake. The returning-player branch is idempotent and is a deliberate
standing "log in via the game" door, not just first-time onboarding.

Honest labeling:
- ORACLE-VERIFIED (Go): account/session logic — role=user, refuse-staff,
  idempotent create-or-fetch, single-use code, and the op.console redline
  (player session rejected on admin routes). Covered by handlers_onboard_test
  and the OpenAPI parity gate.
- INTEGRATION-dependent: the endpoint's security rests on the Bind Code having
  been minted against an online-mode-Yggdrasil-authenticated UUID, a
  precondition that lives in velocity/Java and is not verifiable from this
  repo (CODE-ONLY). The Go layer proves the logic, not that identity guarantee.
- No app-level attempt cap: rate-limiting is deferred to the edge as for the
  public /auth/login; the ~1e12 keyspace, single use and short TTL make a
  blind app-level cap non-critical.
2026-07-01 18:00:13 +09:00
flyemoji c01f133cd8 feat(updates): add pure decision core for component self-update
Introduce internal/updates: a pure, I/O-free engine that decides what
should happen to each tracked platform component (Felis control-plane,
k3s, cloudflared, Velocity) given its current version, the latest
discovered upstream, its policy, and the current time.

Updates are never force-applied. A component is Pinned (Minecraft, left
alone), Notify (a SysAdmin is told and applies out of band), or Scheduled
(Felis may apply, but only inside a maintenance window the SysAdmin set).
The load-bearing invariants are unit-tested: a pinned component never
changes, a downgrade is never proposed, a prerelease is never
auto-applied, and an apply happens only inside the window.

Version parsing tolerates the real feeds (leading v, k3s +k3s1 build
suffix, calendar versions, prerelease tails) and orders by SemVer
precedence. ReleaseSource/Notifier/Applier are declared as integration
seams and exercised via fakes; this package ships no network, SMTP, or
kubectl, and deliberately has no blind k3s-upgrade applier.
2026-07-01 15:59:33 +09:00
flyemoji e058a64a9b feat(edge): close the panel NodePort to the public after the tunnel is up
After the Cloudflare tunnel connector is installed and the origin has rolled
out, applyCloudflareEdge now fences the panel NodePort so the origin is
reachable only over loopback -- the hop the host-side connector uses -- and
never from a public interface. This closes the Access-bypass hole where a direct
https://<node-ip>:<nodeport>/ with the right Host header reached the origin
behind Cloudflare Access.

The fence is an nftables table hooked at prerouting priority -300 (raw), before
kube-proxy's NodePort DNAT (dstnat, -100), so it catches the packet on its
original destination port; a filter/INPUT rule would miss the DNAT'd, then
FORWARDed NodePort packet. Loopback is accepted first, so the connector origin
hop is untouched; the inet family fences a public IPv6 NodePort too.

It is gated on the connector actually serving (verifyConnectorServing polls
`cloudflared tunnel info`): fencing a dead tunnel would sever the only web path
to a still-up origin. If serving cannot be confirmed the port is left open (its
pre-tunnel state) and the failure is surfaced loudly. unfenceOriginNodePort is
the on-host break-glass reversal. The nft/cloudflared calls are INTEGRATION-ONLY;
the ruleset shape and the conn-count gate are pure and unit-tested.

KNOWN-LIMITATION: targets nftables; firewalld-native coordination is not yet
handled (a firewalld reload can flush the standalone table).
2026-07-01 15:39:05 +09:00
flyemoji a531f5e42a fix(cfsetup): keep connector install in the host apply layer only
Setup previously called runner.StartConnector (`cloudflared service install`)
while the TUI applyCloudflareEdge separately installs cloudflared-felis.service
for the same tunnel from the same config -- two managed services serving one
tunnel from one connector config.

Drop StartConnector from cfsetup: running a connector is a host-specific side
effect (systemd/launchd/Windows service) that belongs with the caller, not in
this host- and domain-agnostic package whose documented side effects are tunnel
creation, DNS routing, and the Access app/policy calls. installCloudflaredService
in the host layer stays the single connector installer, so the routed-but-dead
1033 is still closed; RouteDNS --overwrite-dns still closes the stale-DNS 1033.
2026-07-01 15:31:58 +09:00
flyemoji 7d3be64919 feat(cfsetup): start the tunnel connector as a setup step
Setup created the tunnel, routed DNS, and wrote config.yml, but nothing
installed or started a connector for it. A one-click run therefore left the
tunnel routed-but-dead: every web hostname returned Cloudflare error 1033
(tunnel has no connector) even though the config was correct on disk.

Add a StartConnector step to the Runner seam, invoked right after the config
is written (and gated on ConfigPath, so a caller wanting only the Access
config is not forced to install a service). The ExecRunner implementation
runs `cloudflared --config <path> service install`, which installs and starts
a managed system service (systemd/launchd/Windows), and is idempotent on an
already-installed service. The orchestration — connector started, and only
after its config exists — is unit-tested against the fake Runner; the actual
service install is INTEGRATION-ONLY.

Together with the RouteDNS --overwrite-dns fix, this closes both distinct
paths to a 1033 half-state from a fresh setup: a stale DNS binding and a
missing connector.
2026-07-01 15:11:35 +09:00
flyemoji 2810fe849c fix(cfsetup): repoint stale DNS record when routing a tunnel hostname
RouteDNS ran `cloudflared tunnel route dns` without --overwrite-dns and
swallowed the resulting "record already exists" error as success. When a
hostname already had a CNAME from an earlier tunnel that was deleted and
recreated, the record stayed bound to the dead tunnel: the setup reported
the hostname "routed" while it kept returning Cloudflare error 1033 (the
tunnel it pointed at has no connector).

Pass --overwrite-dns so the record is repointed at the tunnel just created,
making the route idempotent and correct on every re-run, and drop the
now-unnecessary "already exists" swallow. INTEGRATION-ONLY (ExecRunner
shells out to the real cloudflared binary).
2026-07-01 15:09:04 +09:00
flyemoji fce0fceac4 feat(passkey): wire enrollment verifier into felis-api
Construct the go-webauthn verifier at the composition root and attach
it to the API when auth.panel_hostname is configured (RP id = panel
hostname, origin = https://<panel hostname>, display name Felis). When
the hostname is unset or the verifier fails to build it stays nil and
the passkey ceremony routes report 503, matching the existing
nil-when-unconfigured subsystem pattern. An admin passkey, if ever
added, is a separate relying party on the admin host and is
intentionally not wired here.
2026-07-01 14:36:28 +09:00
flyemoji 0261204979 feat(passkey): add go-webauthn enrollment verifier adapter
Wrap github.com/go-webauthn/webauthn behind the api.PasskeyVerifier
seam so the api package stays free of go-webauthn types. The adapter
covers the credential-creation ceremony only (BeginRegistration /
CreateCredential); the login/assertion path is a deferred slice.

Ceremony state crosses the seam as opaque marshaled SessionData, the
attestation as an io.Reader, and the verified result as a plain
VerifiedCredential. SessionData carries no expiry so the challenge
row's TTL stays the single liveness authority. New rejects an empty
RP id or origin list so a misconfigured deployment fails at
construction rather than minting unverifiable challenges.

Tests drive a real relying party against a virtual authenticator
(descope/virtualwebauthn): a full creation round-trip plus adversarial
guards proving origin-mismatch and user-mismatch are rejected and
already-bound credentials are excluded.
2026-07-01 14:36:28 +09:00
flyemoji 742f15f348 feat(api): add passkey enrollment endpoints
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14), web app face.
An already-authenticated principal binds a passkey to their own account and
manages the credentials they have bound; email-OTP stays the fallback factor.

- four account routes: POST register/begin mints a credential-creation
  challenge, POST register/finish verifies the attestation against the
  server-stashed SessionData and binds the credential, GET/DELETE credentials
  list and unbind the caller's OWN passkeys. App-tier, principal-scoped (the
  body never names a user).
- PasskeyVerifier seam keeps go-webauthn out of this package: ceremony state
  crosses as opaque bytes, attestation as an io.Reader, result as a plain
  VerifiedCredential. A nil verifier makes begin/finish report 503 so the
  authenticated boundary is exercised before the real verifier is wired in.
- the view never leaks the public key; credential_id collisions map to 409.
- OpenAPI: the four paths plus the PasskeyCredential schema, keeping the
  served-routes parity gate green.

Scope: ENROLLMENT only. The passkey login/assertion path (proving a passkey
from an unauthenticated state) is deferred; every ceremony here rides on a
known principal.

Tests: handler + challenge state machine against a fake repo and a fake
verifier (no real attestation crypto, no SQL). The decisive assertion is the
session-data round-trip -- the finish body carries no challenge, so the only
path for the stashed blob into FinishRegistration is store-stash then consume,
proving the challenge is server-held and never client-echoed. Also covers
supersede-on-begin, single-use, expiry, 503-unavailable, 409-already-bound,
owner-scoped list/delete, and external-only face separation.
2026-07-01 02:14:29 +09:00
flyemoji f2c916d378 feat(api): add passkey enrollment persistence layer
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14). Adds the data
layer an already-authenticated principal needs to bind and manage passkeys:

- migration 0007: webauthn_credentials (one bound passkey per row, public
  attestation material only) and webauthn_challenges (server-stashed ceremony
  state between begin and finish, single-use via consumed_at). Both rows are
  bound to a known user_id; there is no usernameless login lookup, since the
  assertion/login path is a deferred slice.
- PasskeyCredential type and five Repo methods (create/consume challenge,
  create/list/delete credential) with the PG semantics the handlers rely on:
  supersede-prior-live on begin, expiry-before-consume single-use on finish,
  credential_id UNIQUE -> ErrConflict, owner-scoped delete -> ErrNotFound.
- ErrPasskeyChallengeInvalid sentinel for a missing/expired/consumed ceremony.
2026-07-01 02:14:28 +09:00
flyemoji 29f5341cad docs(api): correct cooldownLimiter doc for its OTP reuse
cooldownLimiter began as the wake-only throttle; the OTP-start hardening
reused it via the atomic reserve/release. Its type comment still called it
a per-server wake limiter and justified the per-replica behaviour as
"acceptable because the operator reconcile is idempotent" -- true for wake,
false for OTP, whose every admitted send is a non-idempotent email.

Rewrite the comment to describe the shared per-key limiter and record the
honest KNOWN-LIMITATION: the atomic reserve/release closes the
intra-replica concurrent burst, but the in-memory map throttles per
replica, so cross-replica bounding still needs a shared store. No
behaviour change.
2026-07-01 00:17:35 +09:00
flyemoji 879b1777f5 fix(api): make OTP-start throttle atomic to close concurrent-burst bypass
The email-OTP resend cooldown checked the window with a peek (allowed)
and only recorded it after delivery. For OTP that throttle is the sole
defense and each admitted send is a real, non-idempotent email, so a
burst of truly concurrent starts all passed the peek before any recorded
and every one mailed: N concurrent starts bombed a mailbox with N codes.

Add an atomic reserve/release pair to cooldownLimiter: reserve checks and
records the window in one critical section under the mutex, so a
concurrent burst yields exactly one winner; release rolls a reservation
back only if it is still the current one, so a slow failing caller never
clobbers a newer holder. handleEmailOTPStart now reserves both the
principal and the recipient key up front and defers a rollback that frees
both windows on any mint, create, or delivery error — preserving the old
"a failed send does not consume the cooldown" property, now race-free.

The wake path keeps allowed→record: its real gate is the running cap and
its side effect (SetDesiredState) is idempotent, so the peek gap is
harmless there.

Tests: a frozen-clock gate-mailer fires 8 concurrent starts for one
victim from one principal and asserts exactly one mail and one 202; a
flaky-mailer test proves a failed delivery releases the window so an
immediate retry in the same instant is admitted.
2026-06-30 23:04:07 +09:00
flyemoji 98739044a5 fix(operator): populate Status.Players from an RCON list probe
A Running, ready server always reported 0/0 players: markRunningReady
never wrote Status.Players, and markStopped only cleared it. The panel
therefore showed an empty tally for live servers.

Extend the readiness probe to also sample the player count. Prober.Probe
now returns a PlayerCount{Online, Max}: RconProber still gates readiness
on Dial+auth, then runs a best-effort `list` and parses the vanilla
reply ("There are N of a max of M players online"). A failed or
unparseable tally is swallowed (0/0) so it never blocks readiness. The
reconciler threads the count into markRunningReady, which writes
Status.Players; markStopped still resets it to zero.
2026-06-30 20:04:46 +09:00
flyemoji 6c3999a067 fix(api): rate-limit email-OTP sends to close the email-bomb vector
handleEmailOTPStart minted and mailed a code on every call, so an
authenticated caller could drive unbounded mail to any address they
typed — an email-bomb primitive against arbitrary mailboxes.

Add a separate otpLimiter (its own sync.Once and map, distinct from the
wake limiter) and throttle each send on two keys before anything is
minted: the caller (user:<id>) and the recipient (email:<lower>). A
refused send mints no code and mails nothing; both cooldowns are
recorded only after delivery succeeds, mirroring the wake path so a
failed mint or delivery never consumes the throttle. The two-key design
stops both one account fanning out across addresses and many accounts
converging on one mailbox.
2026-06-30 20:04:23 +09:00
flyemoji 2a4a81b9b2 fix(api): don't burn wake cooldown when refused at capacity
A wake refused by the §9.1 running-server cap returns 503, but the
per-server cooldown was recorded before the cap check ran. A player
held because the cluster was momentarily full would then also have to
wait out the wake cooldown once a slot freed, even though their refused
wake never actually flipped desiredState.

Split cooldownLimiter.allow into allowed (peek, no record) and record
(commit). Both wake paths now consult allowed for the 429, then call
record only after SetDesiredState succeeds — so neither a 503
at_capacity nor a SetDesiredState error consumes the cooldown. The
split is safe against the running cap, which counts CRD truth via
ListServers and is independent of the limiter.
2026-06-30 20:03:55 +09:00
flyemoji c14ed170cc fix(docker): keep embedded panel/ and deploy/ in the image build context
`docker build` failed twice over because .dockerignore excluded two trees the
image actually needs. The panel stage's `COPY panel/ ./` hit `"/panel": not
found`, and even past that the Go build would fail: the root felis package
//go:embeds deploy/bootstrap.sh and deploy/crd/*.yaml, which `COPY . .` dropped
along with the excluded deploy/.

The stale header comment claimed only internal/store/migrations was embedded,
which is what licensed the over-broad exclusions. Rewrite it to name all three
embedded trees (migrations, deploy assets, panel static) and warn against
re-adding panel/, deploy/, or internal/ without re-checking the go:embed list.

Tighten node_modules -> **/node_modules so a working-tree build no longer drags
panel/node_modules over the Linux modules npm ci installs in the panel stage.
2026-06-30 18:23:29 +09:00
flyemoji eb5875a699 feat(felis): add Operator break-glass op behind an operation menu
When a staff account already exists, the break-glass console now opens on a
thin top-level menu (menuModel) where account operations are peers rather than
tails of one wizard: provision/reset the Owner, or add an Operator. A fresh
machine with no Owner skips the menu and goes straight to Owner bootstrap, since
minting an Operator first would create a staff account the login gate rejects.

The Operator path reuses ownerModel via a bgOperation discriminator. It is
insert-only (performAddOperator -> InsertOperator), wraps a duplicate username as
api.ErrConflict and routes back to the provision form for a retry rather than
tearing down, and deliberately never flips the global local_auth toggle the way
the Owner thread does. The post-exit summary and audit trail distinguish the two
outcomes (isOperator); only the Owner provision claims local-password login was
enabled.

Tests cover the operator-model defaults, path selection (insert vs upsert and
the local-auth gate), conflict-retry versus generic teardown, isOperator
propagation, and the root menu routing for both fresh and admin-present
machines.
2026-06-30 15:40:24 +09:00
flyemoji ac02c69612 docs(troubleshooting): add operator failure-mode checklist
Add docs/troubleshooting.md covering the common failure modes the spec
implies, grounded in the actual control-plane code paths:

- Stuck Starting (PodNotReady / RconSecretUnavailable / RconNotReachable)
  and the deliberate absence of a Starting->Failed timeout.
- Failed reachable only via InvalidSpec on a malformed spec.storage.size,
  plus the stale status.endpoint=direct caveat after a failure.
- Routing via status.endpoint direct/fallback and the empty fallbackServer
  pitfall; wake 403/429/503 gate order.
- online-mode coupling and Velocity's offline-mode routing refusal.
- Cloudflare Access 401/403, nil-Keyfunc fail-closed, audience checks,
  the absence of an issuer check, and local-session gating.
- Internal service-token (FELIS_SERVICE_TOKEN) rejection path.
- link/claim error codes, Kaniko build denials (SA-by-absence RBAC,
  default-deny egress, internal-registry push gate), and the registry
  DNS contract.
- Reaper backup-before-delete invariant and false-delete vectors.
- Unimplemented idle auto-stop, permanently-zero players.online, the
  inert CRD fields, and the always-survives world PVC behaviour.

Each item is labelled with its evidence grade (GO-TESTED / CODE-ONLY /
INTEGRATION-ONLY / INERT) so operators know what is verified versus
asserted.
2026-06-30 13:38:37 +09:00
flyemoji 676407d036 docs(diagrams): align §28 sequence diagrams with implemented routes
Adversarial cross-check of the three §28 diagrams against the wake, claim
and link code paths surfaced two fidelity drifts:

- The wake/status/join-event lanes used abbreviated /internal/... paths;
  the registered internal-face routes carry the /api/v1 prefix (api.go),
  matching the convention the claim and link diagrams already use.
- The /link diagram showed the game posting {mc_uuid, auth_source}, but no
  shipped in-game caller sends auth_source — LinkClient posts {mc_uuid}
  and the API defaults auth_source to mojang server-side.

The claim diagram already matched the code (verbatim atomic UPDATE,
404/409/200 mapping) and is unchanged.
2026-06-30 13:18:47 +09:00
flyemoji 8ac5e64d8a feat(metrics): observe felis_start_duration_seconds across the start lifecycle
Wire the fourth mandated §23 metric to a real producer. The histogram
spans two reconcile passes, so anchor and observation must persist in
status:

- Add status.startRequestedAt, set once on the first Starting reconcile
  of a start attempt and cleared on Stopped so the next start re-anchors.
- Observe felis_start_duration_seconds exactly when readiness is first
  reached (ReadySignalAt - StartRequestedAt), guarded so a server that
  reaches ready without a Starting pass records nothing.
- Mirror the field into the deepcopy and the structural CRD schema so the
  apiserver does not prune it on patchStatus round-trips.
- Promote prometheus/client_golang and client_model to direct deps now
  that the operator and its tests import them.

Tests drive a step clock through Starting -> Running asserting the exact
observed duration, and through Running -> Stopped asserting the metric is
observed once and the anchor clears.
2026-06-30 13:11:49 +09:00
flyemoji 79eae7f669 feat(metrics): publish felis_servers_total from a fleet snapshot
A per-object reconcile cannot maintain felis_servers_total (spec §23): it
sees one server per call, so it could never Set a correct fleet-wide gauge
and inc/dec on transitions would drift on any missed event. Add a snapshot
producer instead.

metrics.SyncServerGauge Resets the GaugeVec then Sets one child per state,
so a state that drains to zero reports 0 rather than a stale last value.
operator.GaugeSyncer is a manager.Runnable that periodically Lists the
fleet and republishes from it, defaulting an unset desiredState to Stopped.

SyncOnce is exercised end-to-end against a fake client (List, default,
republish); the ticker loop in Start is the only untested I/O edge.
2026-06-30 12:52:53 +09:00
flyemoji 2a93a9e0cb feat(metrics): record felis_image_build_failures_total on failed builds
Wire the build subsystem to the felis_image_build_failures_total counter
(spec §23). It advances at the two terminal-failure producers: finishAt
(the Sync JobFailed/JobUnknown verdict — a kaniko failure or a CRITICAL
CVE from trivy's --exit-code 1) and Submit's job-creation bypass path,
which records its failure directly without going through finishAt.
Cancellations and successful builds are deliberately not counted.

A delta-asserting test exercises both Inc sites plus a successful-build
negative control that proves the StatusFailed guard discriminates rather
than firing on every terminal write, all over the existing in-memory
Store/Jobs fakes.
2026-06-30 12:47:49 +09:00
flyemoji 75642d90cf feat(metrics): add named felis_* Prometheus collectors
Introduce internal/metrics exposing the four metric families spec §23
mandates at minimum: felis_servers_total (gauge by desired state),
felis_start_duration_seconds (histogram with Minecraft cold-start
buckets), felis_image_build_failures_total and
felis_reaper_worlds_deleted_total (counters). Collectors are
package-level vars so any subsystem records without an import cycle;
Register wires them into a prometheus.Registerer and is idempotent.

Wire registration into the operator against controller-runtime's global
Registry, so /metrics on the manager's existing metrics endpoint carries
the felis_* families. Instrument the reaper to increment
felis_reaper_worlds_deleted_total in lockstep with Summary.WorldsReaped,
at the one point a world's PVC has actually been deleted.
2026-06-30 12:38:55 +09:00
flyemoji 50b8487ff5 feat(panel): wire role-switcher into the app shell
Build the React consumer over the fail-closed view-mode logic so an admin
can view the app as each of the three homes (User/Admin/SysAdmin) and step
down to a plain User-Side home.

- ViewModeProvider holds the raw requested home (seeded from localStorage,
  shape-checked only) and resolves it live against is_admin on every render,
  so a demotion or transient /me failure collapses to the User home with no
  flash, while an unentitled value is never stored or applied.
- RoleSwitcher renders only for admins (availableViewModes > 1); switching
  re-gates the choice and navigates to the chosen home's root.
- AppShell drives its sidebar from sectionsForView(view, isAdmin), which only
  ever narrows visibleSections — an admin viewing as a user sees a plain
  user's sidebar and lands on the Dashboard at /.
- landingPathForView / viewModeLabelKey added to the logic layer (tested);
  view_* and view_switch_label i18n keys added for en-US and zh-CN.

Placement note: the spec calls for a top-right avatar control, but the panel
has no desktop top bar, so the switcher lives in the sidebar foot beside the
user strip. Functionally complete; placement is not yet spec-parity.
2026-06-30 12:18:51 +09:00
flyemoji 563041ada3 feat(panel): add fail-closed role-switcher view-mode logic
Pure logic layer for the top-right avatar role-switcher: derive the home
a principal is in and may switch into, mirroring nav.ts/auth.ts so the
decision is unit-tested without a React renderer.

- ViewMode is derived from NavSection["id"], so the three switchable
  homes (User/Admin/SysAdmin) can never drift from the nav sections.
- availableViewModes / effectiveViewMode resolve a requested view against
  the live is_admin flag, failing closed: a non-admin or a demoted admin
  always collapses to the User home.
- restoreViewMode re-gates a persisted (localStorage) choice on every
  read, never trusting the stored value over the live flag, closing the
  one escalation vector a client-side persona could open.
- sectionsForView composes the view ceiling on top of visibleSections, so
  the switcher only ever narrows the sidebar, never widens access.

The avatar dropdown UI that consumes this lands as a separate increment.
2026-06-30 12:01:24 +09:00
flyemoji a94b0015c4 feat(deploy): add break-glass Operator account provisioning
Add the insert-only Operator-creation path to the break-glass console
(felis breakGlass). An Operator is an additional staff admin: role=admin
with must_change_password=true, identical in shape to the Owner, since
Felis has no separate operator DB role (migration 0003).

Unlike the Owner upsert, provisioning is insert-only -- a username already
taken returns ErrConflict (ON CONFLICT DO NOTHING + zero RowsAffected)
rather than silently resetting a live account, so adding an Operator can
never clobber the Owner's or another Operator's credential. A typed
password is used as-is; an empty one is replaced with a generated
one-time credential returned for display. Operator-add does not touch
local_auth_enabled -- that global gate belongs to the Owner thread alone.
Accountability is recorded best-effort under a break_glass.operator_create
audit action, written only after a successful provision.

The TUI menu router that reaches this path is deferred; this lands the
fully unit-testable logic layer (provisionOperator, performAddOperator,
auditAddOperator) with the PGRepo insert kept integration-only.
2026-06-30 04:38:40 +09:00
flyemoji 116595f3ee feat(api): add QR scan-login completion poll on the internal face
QR scan-to-login is a device-code grant where the QR encodes the existing
short-lived account-link code (spec §B3 player game-login). velocity mints a
code in-game, renders it as a QR, the player scans it on a phone already signed
in to the panel, and that web session's verify writes the durable account_links
row bound to that user. The only new verifiable surface that flow needs is the
completion poll velocity calls to learn the link landed and admit the player.

Add GET /api/v1/internal/account/link/status/{mc_uuid}: a read-only, internal
handleLinkStatus keyed by the verified mc_uuid velocity already holds. It reuses
the existing UserByMCUUID, so it adds no migration and no mutation to the
load-bearing VerifyLinkCode; ErrNotFound maps to {linked:false} (pending /
not-yet-scanned), a hit to {linked:true, user_id}. Keying on the public UUID and
not the scanned code means the read carries no guessing surface and needs no
attempt cap — the internal face already gates it to service callers, and the poll
consumes nothing so a velocity restart re-polls safely.

QR render, limbo collision routing, in-game admit, and the reclaim
inherit-disambiguation stay CODE-ONLY (Java/Velocity) and are labeled as such;
this endpoint reports link completion only.

Document the route in openapi.yaml (x-felis-face internal, x-felis-tier service)
so the parity gate holds, and cover it with a hermetic vertical that proves the
poll reflects the durable link only after the external verify and binds the
verifier's id, plus unknown-uuid, idempotency, and internal-only face separation.
2026-06-30 04:08:01 +09:00
flyemoji 5450c268f4 chore: normalize line endings and apply formatting
- Convert CRLF to LF across Go, panel, and plugin files
- Add Cloudflare API token template URL to breakGlass TUI edge intro
- Verify API token in cfsetup before creating tunnel, DNS, or Access app
2026-06-28 16:43:10 +09:00
flyemoji 9c46632929 feat(cli): add felis setup first-run console with reclaim protection and cfsetup idempotency
- Add `felis setup` TUI for initial Owner provisioning and optional Cloudflare edge
- Refactor breakGlass to share console TUI model (runConsoleTUI) with setup mode
- Session auth respects configured [auth].admin_hostname; fallback to op.console.<root>
- Protect linked Yggdrasil admins from Mojang-priority reclaim (spec §B3)
- cfsetup: idempotent Access app/policy creation, better 401/403 errors, GET + lookup
- Bootstrap: auto-install cloudflared, symlink /etc/felis/felis.toml
- Add sequence diagrams for ping-to-join, claim, and link flows
2026-06-28 16:41:37 +09:00
flyemoji b81b33453a Merge branch 'main' of https://github.com/MliroLirrorsIngenuity/Felis 2026-06-27 20:19:06 +09:00
flyemoji ba13839341 feat(breakglass): optional Cloudflare Tunnel + Access setup in the TUI
Add an optional edge-setup flow to the `felis breakGlass` sudo TUI,
reachable as an independent peer of Owner provisioning through a new
top-level menu (so reaching it never forces an Owner password reset).

The flow drives the operator's own Cloudflare consent (interactive
`cloudflared tunnel login`, suspending the alt-screen, plus an API
token) and then calls cfsetup to stand up a Tunnel routing the
configured admin and panel hosts and a fail-closed Access application.

It stays gated shut unless an admin hostname is configured and the
operator is logged in (edgeReady), and refuses empty or bare-domain
credentials before any side effect. On success the TUI surfaces the
issued Access aud and an explicit ACTION REQUIRED note; it never edits
felis.toml. The live cloudflared and Cloudflare API calls are
integration-only and exercised against a real account.
2026-06-27 15:03:48 +09:00
flyemoji 53a76640a4 feat(cfsetup): recommended Cloudflare Tunnel + Access edge setup
Add internal/cfsetup, the verifiable core of an optional one-click
Cloudflare Tunnel + Access provisioning flow for the SysAdmin edge
(spec §14). It is domain-agnostic (every FQDN is composed from the
configured root_domain) and IdP-agnostic (any valid Access JWT aud is
accepted, whichever IdP fronts it), so a SysAdmin who brings their own
domain or Zero-Trust scheme stays fully supported.

The load-bearing safety property is a fail-closed guard on the
recommended Access policy. validateFailClosed is an allowlist that
refuses any policy that could be public: a bypass/non-allow decision, an
empty include, an "everyone" include not narrowed by a constraining
require (include rules are OR, so "everyone" beside an identity is still
public), or any include rule it cannot positively recognize as a scoped
identity. Setup runs the guard before any side effect, so a public
policy aborts the run with nothing created.

The tunnel ingress routes only the web hostnames to the local panel
origin and terminates in the mandatory fail-shut 404 catch-all; the raw
game host is never proxied. Gating preconditions (cloudflared present,
tunnel login completed, API token) are hard checks with no side effects
on failure.

The actual cloudflared exec, DNS routing, and Access API calls live in
runner.go and are integration-only: they require the operator's own live
Cloudflare account and interactive browser consent, which cannot be
unit-tested. The policy guard, ingress generation, request bodies, and
gating are unit-tested.
2026-06-27 14:18:49 +09:00
flyemoji a29571de39 feat(api): reclaim squatted usernames for Mojang-priority players (spec §B3)
When the configured third-party Yggdrasil and the official Mojang service
issue the same username under different UUIDs, the non-genuine squatter is
displaced in favour of the real Mojang owner (正版优先). This adds the
Go-verifiable data layer of that flow on the internal (velocity) face.

- migration 0006: username_blacklist (barred squatter UUIDs) and
  player_data_holds (the displaced account's 30-day data stash), both keyed
  by mc_uuid so the genuine Mojang player — identical username, different
  UUID — is never caught by the bar.
- POST /api/v1/internal/player/reclaim bars the squatter UUID and stashes
  its data in one transaction (all-or-nothing). It is idempotent on a
  retried callback and returns the hold's effective expiry — the first
  reclaim's window, never a fresh now()+30d — so the rejected player is told
  the truth about how long their data is kept.
- GET /api/v1/internal/player/blacklist/{mc_uuid} is the login-gate check
  velocity calls to reject a barred squatter before admitting them.

Scope: velocity collision-routing, the limbo prompt, the authlib
dual-backend and the data-inherit flow are code-only (Java plus a QR-bound
device session a row cannot express) and are not part of this slice. Unit
tests cover the handlers and the in-memory repo contract; the Postgres SQL
path is exercised by integration only.
2026-06-27 13:41:19 +09:00
flyemoji 1f8b9bb5d0 feat(api): record account-link auth source (mojang|thirdparty)
Capture which Yggdrasil authenticated an in-game UUID when a link code is
minted (spec §10 dual-Yggdrasil) and copy it onto the durable account_links
row at verify. The value originates in-game — the web verify side never sees
the authentication — so it threads through account_link_codes, mirroring how
mc_uuid (not user_id) lives on a code.

- migration 0005: add link_auth_source enum + auth_source column on both
  account_link_codes and account_links; DEFAULT 'mojang' backfills existing
  rows and sets the Mojang-priority default for a mint that omits the field
- mint validates an explicit auth_source (unknown value -> 400); verify
  surfaces it in the 200 body and refreshes it on idempotent re-verify
2026-06-27 13:01:19 +09:00
flyemoji dbe34a175f feat(api): add player email OTP verification (spec §B2 onboarding)
Forced web onboarding proves a player controls an email before it is
bound to their account. POST /api/v1/account/email/start mints a random
6-digit code, mails it (or logs it server-side when no Mailer is wired —
the demo has no SMTP), and POST /api/v1/account/email/verify redeems it,
flipping users.email_verified in the same transaction that consumes the
code.

Brute force is bounded two ways: a 10-minute TTL and a 5-attempt cap,
both enforced in the repo so the fake and Postgres agree. Only the
sha-256 of the code is stored; the digits live only in the email. Both
routes are app-tier external — verifying your own email is scoped to the
principal, never names another user.
2026-06-27 12:29:52 +09:00
flyemoji 2d0bbb0c37 feat(cli): attribute break-glass recovery to the SysAdmin who runs it
Root is machine authority, not a human identity, so `felis breakGlass`
now also records WHICH SysAdmin broke the glass. Even under
`sudo felis breakGlass` an account and password are entered in the TUI;
the root gate is necessary but no longer sufficient for accountability.

The console resolves one of three modes up front and audits the
difference:

- bootstrap (no staff account exists yet): the typed credential mints
  the first Owner; the act is attributed to the OS user ($SUDO_USER,
  else root) and recorded verified:false.
- recovery (an admin already exists): the operator authenticates as an
  existing admin via bcrypt; the verified identity is the accountable
  actor and the row is recorded verified:true.
- root override (the typed credential did not verify): a deliberate
  OVERRIDE token proceeds under local-root authority, attributed to the
  OS user and recorded verified:false. Break-glass never refuses -
  recovering when no admin password can be produced is its whole job.

Attribution is best-effort, not proof (whoever runs this is root and can
edit Postgres directly); the audit row is honest about which it is.

- internal/api: AuditEntry gains an optional jsonb Payload (nil maps to
  SQL NULL, so existing callers are unaffected); PGRepo.Audit writes it
  and a new PGRepo.AdminExists drives the bootstrap-vs-recovery switch.
- the accountability row is written the instant the credential changes,
  before local auth is enabled, so a failed toggle write can never leave
  a reset credential with no "who did it" record.
- local_auth_enabled is now one exported api.LocalAuthEnabledKey shared
  by the break-glass writer and the per-request reader, replacing two
  drifting copies of the literal.
- break-glass password entry reuses the panel's 8-72-byte rule so a
  credential set here is never later rejected by web change-password.

Covered by Go unit tests over a fake owner store: auth match/non-match,
the three audit modes and their payloads, that a dead audit sink does
not fail the recovery, that the audit precedes the toggle write, and a
headless drive of the TUI state machine asserting no credential reaches
provisioning without a verified admin or an explicit OVERRIDE.
2026-06-27 11:48:41 +09:00
flyemoji 885c4a9bd8 feat(panel): local-password login and forced password change
Add the op.console login and forced first-login password-change flow to
the panel. RequireAuth bounces an unauthenticated visitor to /login;
both /login and /change-password render outside the app shell with their
own centered chrome.

- TierProvider now derives auth state (deriveAuth) and exposes refresh()
  so a successful login re-fetches identity without a full reload; only a
  genuine 401 marks the session unauthenticated, so a transient /me
  failure keeps a healthy Zero-Trust principal in the app.
- api.login/logout/changePassword send Content-Type: application/json on
  bodied requests to satisfy the backend guard; humanizeError maps the
  auth error codes to stable copy.

Covered by vitest unit tests for deriveAuth branch coverage and the
login/change-password wire-shape contracts.
2026-06-27 04:23:32 +09:00
flyemoji e108a3709a feat(cli): break-glass emergency console TUI
Add `felis breakGlass`, a root-only interactive TUI that provisions or
resets the Owner account directly against Postgres and enables local
password login. It is the local-root recovery path that bypasses web
Zero Trust by design - used to bootstrap the first Owner credential and
to recover when the web login is unreachable.

- Bare `felis` prints CLI usage only; breakGlass is the sole subcommand
  that enters a TUI rather than running as a CLI.
- Refuses to run unless euid is 0 (try: sudo felis breakGlass); on
  non-Unix platforms the euid check also refuses.
- Generates a one-time Owner password, sets must_change_password, and
  prints a durable summary (username, one-time password, op.console
  login URL derived from the configured root domain) after the
  alt-screen TUI is torn down.

Covered by Go unit tests over a fake owner store.
2026-06-27 04:23:21 +09:00
flyemoji af14f02f38 feat(api): local-password authentication backend
Add username+password login for Owner/Operator staff accounts on
op.console, the primary web login when Zero Trust is not in front of the
API. Three handlers form the whole surface: login mints a server-side
session cookie, logout revokes it idempotently, and change-password
re-verifies the current password before rotating the hash and clearing
must_change_password.

- Session cookies are HttpOnly+Secure+SameSite=Lax, host-only, stored
  server-side as a SHA-256 hash with a 12h TTL.
- Login is anti-enumeration: every failure runs a uniform bcrypt compare
  against a dummy hash and returns the same vague error.
- Credential-bearing writes require Content-Type: application/json,
  returning 415 otherwise, to close the cross-site form-POST forgery
  vector as a belt to the SameSite cookie.
- Local auth fails closed: login is rejected unless local_auth_enabled
  is set, so a Zero-Trust-only deployment never accepts a local password.
- Extend the users table with a nullable password_hash and
  must_change_password; staff are role=admin rows with a hash, players
  are role=user rows with hash NULL.
- /me now reports must_change_password so the panel can force a
  first-login change.

Covered by Go unit tests (handlers, content-type guard, anti-enumeration,
forced-change lockdown) and the OpenAPI route-parity gate.
2026-06-27 04:22:45 +09:00
flyemoji 58fa4b0af8 feat(deploy): add one-line bootstrap installer and container image
bootstrap.sh auto-detects the host package manager (apt/dnf) and installs whatever is missing: Docker, k3s, and PostgreSQL. It builds and imports the felis image, opens pg_hba to the pod CIDR, runs migrations, and applies the rendered control-plane bundle, leaving Web disabled pending 'felis setup'. The Dockerfile builds the distroless felis image; deploy/crd holds the MinecraftServer CRD.
2026-06-27 00:40:39 +09:00
flyemoji 99de43f74f chore: ignore plugin build artifacts and editor config 2026-06-27 00:40:39 +09:00
flyemoji 7d913737af fix(migrate): honor -config flag placed after the up verb
Go's flag.Parse stops at the first non-flag token, so a -config given as 'felis migrate up -config path' was silently dropped and the default path used instead. Pull the up verb off the front, then parse the remaining flags so the configured path is honored.
2026-06-27 00:40:38 +09:00
flyemoji ce0ba76a3e chore(api): add kubebuilder object-generation markers to v1alpha1 2026-06-27 00:40:26 +09:00
flyemoji 93f143f5b6 feat(plugins): add Velocity proxy and Fabric/Forge/NeoForge/Paper integration mods
Server-side integration plugins: the Velocity proxy plugin plus Fabric, Forge, NeoForge, and Paper mods with a shared module. Gradle build output is not tracked.
2026-06-26 23:32:40 +09:00
flyemoji eee00c2772 feat(panel): add three-sided web console (User, Admin, SysAdmin)
Vite + TypeScript + Tailwind single-page console presenting the three operator tiers and consuming the external felis-api face. Build output and design notes are not tracked.
2026-06-26 23:32:39 +09:00
flyemoji 47fcd90f75 feat(platform): add node orchestration and the felis entrypoint
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them.
2026-06-26 23:32:38 +09:00
flyemoji b508fccc6f feat(api): add felis-api service with permissions, modpack lane, and fleet read
The dual-faced felis-api: internal (service) and external (public/app/admin) routes behind a Zero-Trust guard. Includes the access domain (whitelist, ban, and LuckPerms permission/group control over the owner-gated RCON path), the modpack submission endpoints, and the admin-tier SysAdmin fleet read. Structured access fields are charset-validated before assembly so no field can splice a second RCON command.
2026-06-26 23:32:38 +09:00
flyemoji d39605e05e feat(submit): add user modpack build and approval pipeline
A user-directed extension over the build subsystem: an uploaded modpack stays in pending_review and is never built until an admin approves. Approval is a single-winner compare-and-swap that hands off to the image-build Job, keeping the mandatory vulnerability scan in front of any push.
2026-06-26 23:32:38 +09:00
flyemoji 78b8cf6ded feat(operator): add MinecraftServer controller and reconcilers
The Kubernetes controller that drives MinecraftServer resources through their lifecycle and issues RCON where readiness requires it.
2026-06-26 23:31:58 +09:00
flyemoji 43ab92151f feat(backup): add backup, restore, and reaper subsystems
Archive-based world backup and restore, plus the reaper that enforces retention and reclaims idle servers.
2026-06-26 23:31:58 +09:00
flyemoji 708cdfc5b8 feat(core): add naming, RCON, store, config, and image-build libraries
Foundational libraries: deterministic resource naming, the RCON client, the Postgres store with embedded SQL migrations, configuration loading, and container image-build helpers.
2026-06-26 23:31:58 +09:00
flyemoji 7fbebfe843 feat(apis): add MinecraftServer CRD types (v1alpha1)
Kubernetes API types for the MinecraftServer custom resource, the lifecycle source of truth (spec §1). Leaf package with no internal dependencies.
2026-06-26 23:31:57 +09:00
flyemoji 5a30aa5073 docs: add OpenAPI 3.1 served-route contract
Machine-readable contract for the felis-api faces. The parity test (internal/api/openapi_test.go) checks every served route against this document's x-felis-face and x-felis-tier, so served and documented routes cannot drift.
2026-06-26 23:31:57 +09:00
flyemoji 5b7b38d8bb chore: add Go module manifest and ignore rules
Go 1.26 module felis.lolicon.best. Ignore build artifacts (node_modules,
gradle/plugin build output, panel dist), secrets (keys/env), and local agent
state; keep docs/openapi.yaml (the served-route contract) tracked.
2026-06-26 23:31:20 +09:00