Commit Graph
66 Commits
Author SHA1 Message Date
flyemoji fe4c92c1c5 feat(files): add the server file editor
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.

felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.

What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.

Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.

The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.

Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:

  * A write accepts arbitrary bytes at any path, so an owner can place a
    loadable plugin jar. This is deliberate — it is what a hosting panel's
    file manager does, scoped to a server the caller already owns and
    already drives through /command — but it is the one owner-tier route
    that lands executable code in a backend pod, since images are
    admin-only and modpack submissions need an admin verdict.
  * config/paper-global.yml is refused on read. felis-lobby's entrypoint
    writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
    identical on every backend, so reading it from a server you own would
    hand you the handshake key for everyone else's. It is the only path in
    the mount that is not the caller's own data, and therefore the only
    denial. The comparison is on the cleaned path, or ./config/... would
    walk straight through it.

Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.

The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
2026-07-20 14:33:29 +09:00
flyemoji d26acc20ae feat(api): let in-game staff manage any server without claiming it
The web face has always granted staff the run of the fleet (isOwnerOrAdmin
passes an admin for stop/command/console/access on any node), but the
internal face explicitly had "no admin tier": a linked administrator in game
could only wake servers they owned or that autostartPolicy permitted. The
only way to manage another player's (or an unclaimed) server from inside the
game was to claim it — seizing ownership and burning the admin's own quota.

Give authorizeWakeByUUID the admin tier on the same trust anchor the op-login
approve already uses: verified online-mode UUID -> account link -> stored
role. A linked staff member now wakes ANY node under any policy (so
`/felis go` works fleet-wide without claiming); the owner bypass and the
policy gates are unchanged, and an unlinked UUID still fails safe.

Centralize the staff-role rule while at it: staffRole(role) in auth.go
(admin, plus owner as its superset) now backs Principal.IsAdmin, the session
ViaAdminAccess grading, the op-login approve gate and the new wake tier.
That also fixes a real hole in the approve gate, which required role=admin
exactly: an Owner manually promoted to role='owner' per migration 0011's
upgrade note would have been refused by their own in-game approval door.

The lobby menu still renders "Claim & Start" on ownerless tiles — claiming
becomes optional for staff rather than the only entry — so the velocity
plugin needs no change.
2026-07-20 11:10:23 +09:00
flyemoji f0b79e9edd feat(mail): deliver email one-time codes over SMTP and add the setup email screen
Felis never actually sent mail: OTP codes for onboarding, email login and
op-login were only written to the felis-api log behind a "demo has no SMTP"
limitation, and the Settings/SMTP flow those comments promised was never
built. Combined with the bootstrap Owner's address being recorded unverified
(87279a1), op-login start always took the anti-enumeration neutral branch and
minted a fake request_id, so the in-game approve inevitably answered "No
pending operator sign-in with that code".

Give the codes a real delivery path, configured in felis.toml rather than a
web settings page so config keeps a single source of truth:

- config: new [smtp] table (host, port defaulting to 587, from, username,
  password_ref). Validation requires a plausible from address and a sane
  port; the password itself never enters the config file.
- internal/mail (new): stdlib net/smtp mailer implementing the api.OTPMailer
  seam. Port 465 dials implicit TLS, other ports upgrade via STARTTLS when
  advertised; AUTH only when a username is configured (PlainAuth itself
  refuses plaintext, so the password cannot leak to a TLS-less relay).
  Ping() proves reachability and credentials without sending mail. The
  message shape (CRLF, Q-encoded bilingual subject) is pinned by test.
- platform: felis-smtp Secret constants and an optional FELIS_SMTP_PASSWORD
  env var on the felis-api Deployment, mirroring felis-uploads-s3.
- cmd/felis api: construct the real mailer when [smtp] is configured; keep
  the log fallback otherwise and say so at startup. Warn when a username is
  set but the credentials env is empty.
- setup TUI: "e" on the summary/status screen opens the email form (host,
  port, from, optional auth). Apply order: Ping preflight, [smtp] into both
  host and pod config files, felis-smtp Secret piped to kubectl via stdin,
  config Secret, felis-api rollout. A failed preflight leaves the install
  untouched. SMTP is deliberately not a wizard rail step: first-run stays
  mail-less by design, and the passkey minted at onboarding is the pre-SMTP
  owner credential.

Also make PGRepo.UserByEmail match case-insensitively (lower(email) =
lower($1)), honoring the interface contract and the users_verified_email_
unique partial index; the fake repo already matched with EqualFold.

Existing installs need the felis-api Deployment manifest re-applied (e.g. a
bootstrap re-run) before the new env var exists; a rollout restart alone
cannot add it.
2026-07-20 10:24:52 +09:00
flyemoji 7860152f57 feat(auth)!: go fully passwordless and fix cross-check review findings
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.

Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
  handlers; align tests (new account subtests, naming reserves
  "console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
  setup_panel.go, tui_root.go and tui_preflight.go instead of
  hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
  into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.

Panel:
- Delete the ChangePassword page and all password UI; align
  login/auth/api/types with the passwordless contract; add the
  migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
  ServerLuckPerms strings to translation keys; drop 72 orphan keys
  per locale; unify the title as "Felis - Console".

Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
  MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
  preference; unified /link lines in fabric/forge/neoforge; shared
  link-client javadoc contract fixes.

Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.

BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
2026-07-20 04:47:32 +09:00
flyemoji 87279a1366 fix(setup): record email unverified so onboarding works without SMTP
At bootstrap there is no SMTP, so the old /setup flow was unreachable: it
requested an emailed OTP that could never arrive. Setup now records the
Owner's email address unverified (no OTP round-trip) and requires a passkey,
deferring SMTP configuration to a later Settings page. Setup completes on
email-recorded + passkey-enrolled, and the lockdown lifts on the passkey, not
on email_verified: a passkey is the Owner's only pre-SMTP login credential
(email-OTP login refuses admin accounts).

The record-email endpoint (POST /account/email) now clears email_verified in
the same write. Only VerifyEmailOTP, which proves control of the address, may
set that flag; recording a fresh unproven address must never leave a stale
email_verified=true asserting a proof the user never gave. The change strictly
tightens the invariant, so no existing reader breaks.

Remove the dead ErrEmailTaken path and its documented 409: no migration puts a
unique index on users.email and the codebase does not enforce email
uniqueness, so the unique-violation branch was unreachable and the 409 an
impossible response.

The /setup route (Setup.tsx, setEmail helper, setup i18n copy) is rewritten to
match: record-email, mandatory passkey, no skip-for-now. The SMTP settings
page and post-setup configure-SMTP nudge are deferred.
2026-07-17 01:41:07 +09:00
flyemoji 60732a6283 feat(operator): gate op.console to staff and land owner setup there
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).

Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.

Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.

The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.

Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
2026-07-16 18:03:30 +09:00
flyemoji 9ea35304e3 fix(setup): source the console host from the panel hostname, not op.console
The owner setup URL and the limbo login link were built from the admin host
(op.console.<root>, with an op.console.localhost fallback) and a hardcoded
console.<root>, so an operator who set a custom panel_hostname got an unreachable setup
link and a wrong login target. Thread the resolved panel host (defaultPanelHostname)
through performSetupMCBind, the MC-bind TUI, and the login system-server env
(new FELIS_PANEL_HOSTNAME); the limbo plugin prefers it and keeps console.<root> only as
the fallback for an older operator whose env predates it. This also matters for security:
the only wired WebAuthn verifier is scoped to the panel host, so passkey enrollment must
land on the panel face, never op.console.

While here, the limbo login handler checks link status before minting a bind code: an
already-linked player is sent straight to the lobby instead of being shown a useless code.
2026-07-16 13:27:02 +09:00
flyemoji b5cd4501e5 fix(passkey): serve flat WebAuthn options to register and username login
go-webauthn marshals CredentialCreation/CredentialAssertion as {"publicKey": {...}},
but the panel's register (Account.tsx) and username-first login (Login.tsx) read the
options flat (options.challenge, options.user.id), so base64urlToBytes(undefined) threw
"Cannot read properties of undefined (reading 'replace')" and neither ceremony could
start. Strip the envelope in the register-begin and username-login-begin handlers via a
small unwrapPublicKey helper; discoverable login keeps the envelope because it reads
options.publicKey.* plus a top-level options.login_id. The begin tests now feed a wrapped
body and assert the handlers return it flat, so they genuinely exercise the unwrap.
2026-07-16 13:26:56 +09:00
flyemoji dab8fc214b feat(setup): bind owner through login gate 2026-07-14 02:57:15 +09:00
flyemoji fd062882ed feat(nano): give a Mojang player's name back to them, by prefixing the squatter
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.

The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.

A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.

The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.

prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.

Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.

Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:

  premium FLYEMOJ1     -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
  LittleSkin FLYEMOJ1  -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
  both online at once, zero "already connected" rejections
  LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3

The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.

Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.

BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
2026-07-13 13:00:54 +09:00
flyemoji d177428fb0 feat(felis): add felis nano — Yggdrasil hasJoined multiplexer without a control plane
`felis nano` serves the vanilla sessionserver protocol
(GET /session/minecraft/hasJoined) as a federating multiplexer over
Mojang plus any number of third-party Yggdrasil roots, with no k3s,
Postgres, or panel — a MultiLogin-style auth front-end delivered as a
subcommand of the single felis binary rather than a separate build.

- config.LoadNano reads only [[auth_source]] blocks; it skips the
  database.url / root_domain / archive requirements the full server
  needs. Zero sources is valid (Mojang-only).
- Mojang is prepended in code (Identity:true), never from config, so it
  is always the sole identity root. Third-party profiles are rewritten
  to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID).
- validateAuthSources rejects unknown keys, duplicate tags, and
  scheme-less URLs — a malformed nano config fails loud at load.
- Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend);
  a rejected login is a 204, matching the vanilla sessionserver.
- nano.go binds the -listen flag and ignores [server] listen in config.

Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime
smoke against the template config returns 204 on a miss and logs
"Mojang + 0 third-party source(s)"; a duplicate-tag config exits
non-zero citing "unique".
2026-07-13 02:39:41 +09:00
flyemoji ff550c41ef feat(nano): federating hasJoined multiplexer with per-source UUID namespacing 2026-07-12 02:09:47 +09:00
flyemoji 85b8a92a0e test(api): pin restore's owner gate against a superseded former owner
handleRestoreBackup's owner-or-admin gate was not pinned by any test: the former-owner gate backstopped every non-owner case the suite exercised, so a broken owner gate would not redden. Add the mirror of the former-owner test — a released former owner (still the backup's former_owner, no longer the current owner) must get 403 — the sole subtest that fails when the owner gate is disabled. Found by the round-2 backup/restore mutation audit; production code unchanged.
2026-07-08 05:58:00 +09:00
flyemoji fc748d3462 feat(breakglass): add "back up a world now" console peer (§B4 Sync)
Adds a break-glass console operation that snapshots a stopped world by
calling the felis-api internal face while the API is alive, rather than
rendering the backup Job locally: the Job needs felis-api deployment
coordinates the console does not hold.

The peer resolves the felis-api-internal ClusterIP Service + service
token from the control namespace, POSTs the internal backup endpoint
with the operator os_user for audit attribution, and maps 409/503/404
to friendly outcome cards. Core decision logic lives in backupnow.go
(unit-tested against a fake client + httptest); tui_backupnow.go is the
untested bubbletea glue mirroring tui_halt.go.
2026-07-07 19:32:13 +09:00
flyemoji f2fc57cad9 feat(api): add internal-face break-glass world backup endpoint (§B4 Sync)
Add POST /api/v1/internal/servers/{name}/backup so the on-node break-glass
console can snapshot a stopped world while felis-api is alive. It goes through
the API (not direct-to-CRD like halt) because rendering the backup Job needs
deployment coordinates (FELIS_IMAGE, FELIS_BACKUP_PVC) only felis-api holds.

Service-token auth (no Principal); the middleware IS the authorization, since
the operator already has root on the node. Refactor the RWO stopped-gate,
optional-Backuper 503, async hand-off and audit+202 into a shared enqueueBackup
tail so the external (owner/admin) and internal (break-glass) faces cannot
diverge on the security-critical stopped-gate. The internal audit is attributed
to break-glass/internal so a console-initiated backup is distinguishable from an
owner self-service one.
2026-07-07 10:48:01 +09:00
flyemoji 7a7c0d53ab feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
2026-07-07 10:04:30 +09:00
flyemoji fdb6efbd88 feat(account): migrate a live account's owned servers to a new account (§B3 inherit)
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.

- 0015 migration: account_migrations state machine (initiated -> confirmed ->
  code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem

Quota is grandfathered at redeem: no per-target quota re-check when servers move.
2026-07-05 20:50:48 +09:00
Lemon-miaow 7becb38488 fix(api): implement /readyz with real DB + K8s API + CRD checks (§7)
Previously /readyz only verified Repo != nil && Cluster != nil — a
process-liveness check, not a dependency-health check. The spec
requires the readyz probe to verify DB, K8s API, and CRD informer
are live before declaring the pod ready.

- Repo interface gains Ping(context.Context) error
- Cluster interface gains Ping(context.Context) error
- PGRepo.Ping delegates to sql.DB.PingContext
- K8sCluster.Ping lists MinecraftServer CRDs (Limit=1) in the
  configured namespace, exercising both the API and CRD informer
- handleReadyz iterates ping checks; any failure returns 503 with
  the failing dependency name in the error message
- fakeRepo and fakeCluster gain configurable pingErr for hermetic
  test coverage of the failure paths

New test: TestReadyzPingsDependencies verifies 200 when healthy,
503 when DB or K8s API is down.
2026-07-05 16:22:08 +08:00
flyemoji 9e1df12975 feat(passkey): advance sign_count, reject clone-warned assertions
Both login doors (username-first and discoverable) now run a shared applyAssertionCounter after a verified assertion. A signature-counter regression — go-webauthn's CloneWarning, the possible-cloned-authenticator signal — is refused fail-closed with the same opaque passkey_login_invalid envelope any other finish failure returns (no clone oracle to a prober) and audited distinctly as auth.passkey_clone_rejected under the resolved account. A clean assertion advances the stored sign_count to the asserted value and stamps last_used_at, before any session is minted.

Counter-less/synced authenticators report 0 and never warn, so they pass through and simply re-stamp 0; the check gates only counter-keeping hardware authenticators, where a rollback is the meaningful signal. Email-OTP and username-first passkey remain fallbacks, so a rejected clone is never bricked.

Adds Repo.AdvanceCredentialSignCount (pgrepo UPDATE by credential_id) and surfaces CloneWarning from the internal/passkey adapter's FinishLogin/FinishDiscoverableLogin. Proven by real-crypto adapter tests (a counter regression still verifies but flags CloneWarning), handler tests (advance-and-stamp on success, fail-closed on clone), and a symmetric test on each door so both call sites of the shared helper are covered.
2026-07-05 16:02:30 +09:00
flyemoji 154002edf3 docs(auth): cite MultiLogin reference for UUID-keyed reclaim split
Anchor the username-collision reclaim's UUID-keyed, proxy-detected design to
the multi-Yggdrasil reference: CaaMoe/MultiLogin v6 binds identity as
serviceId+online-UUID via "identity cards" that decouple the in-game name from
online identity — keyed by UUID, never by name. Note that §B3's Mojang-priority
reclaim goes beyond the common "protect the first-bound name" behavior by
evicting a squatter once the genuine Mojang owner appears and stashing the
squatter's data for the code-only inherit path.
2026-07-05 04:05:56 +09:00
flyemoji ec468baef9 feat(auth): add discoverable (usernameless) passkey login
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.

Routes (both Public):
  POST /api/v1/auth/passkey/login/discoverable/begin
  POST /api/v1/auth/passkey/login/discoverable/finish

Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.

The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.

Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
2026-07-05 04:05:56 +09:00
Lemon-miaow e574749877 feat(api): enforce CPU/memory/storage quotas (spec §9.3, §22)
Add per-user resource quota enforcement across all four dimensions:
max_servers, max_cpu_milli, max_memory_mb, and max_storage_gb.

- Migration 0013: add cached_cpu_milli, cached_memory_mb, cached_storage_mb
  columns to servers table for pure-SQL per-owner aggregation
- SeedServer now writes resource cache alongside server row
- QuotaCheck replaces QuotaAvailable at claim time, checking all four caps
  against the owning user's cumulative usage
- handlePatchServer checks owner's quota before allowing memory/resource
  changes on owned servers; unowned servers skip the gate
- handleInternalClaim mirrors the full quota check
- UpdateServerResources keeps the cache in sync after spec mutations
- Reaper zeros resource cache on ReleaseWorld so released resources
  are not counted against a former owner
- quantityToMilli/quantityToMB helpers convert K8s quantities to
  quota-comparable integers

19 test packages pass.
2026-07-05 02:05:44 +08:00
flyemoji c20b12c655 refactor(api): drop dead password-era ResetMailer, reconcile passkey-unbind docs
The passwordless migration left ResetMailer (SendPasswordReset) and its API field with zero callers and no wiring; the web console authenticates via email-OTP and passkey only. Remove both, plus the now-orphaned context import that the interface was the last user of in handlers_users.go.

Reconcile the DeleteAllPasskeyCredentialsForUser docs in repo.go and pgrepo.go: they claimed there was no production caller, but 2f22027 wired the owner-tier DELETE /users/{id}/passkeys. Both now note that a complete authenticator remediation pairs the unbind with a session revoke (unbinding alone leaves the live hijacked session; revoking alone leaves a re-enrollable credential), and the OpenAPI operation carries the same guidance in a new description. Reword the stale local-password test-fake header, since the passwordless fakes carry no must_change_password field.

No behavior change. gofmt, build, and the full test tree are green; OpenAPI parity and passkey-unbind tests pass; a grep confirms ResetMailer/SendPasswordReset are gone from the Go tree.
2026-07-04 21:47:13 +09:00
flyemoji 4f59d5128a feat(auth): add owner-tier passkey-unbind remediation endpoint
Add DELETE /api/v1/users/{id}/passkeys (owner-only) to unbind every passkey a
target account holds — the authenticator remediation that stops a passkey planted
or retained via a transiently-hijacked session from surviving as a standing login
foothold. It wires the previously-uncalled DeleteAllPasskeyCredentialsForUser and
is deliberately not a lockout: the account re-enters via the email-OTP door
(players) or op-login's in-game approval (staff), then re-enrolls. Documented in
the OpenAPI, so the served/documented parity gate covers it.

Remove RevokeUserSessionsExcept: a change-password-era orphan with no callers
since the passwordless migration. Its keep-one ("log out my other devices")
semantics is inherently self-service, and no such slice is on the roadmap; the
admin remediation path already uses RevokeAllUserSessions.
2026-07-04 21:47:12 +09:00
flyemoji 3b43f05a83 refactor(api): drop dead login concurrency limiter and reconcile passwordless comments
The passwordless migration (b330d77) removed the password-login route, leaving
concurrencyLimiter — its bcrypt concurrency cap — with no caller, and scattered
stale "local-password" / "change-password" references through the surviving auth
code's comments.

- Remove the dead concurrencyLimiter (type + newConcurrencyLimiter + acquire):
  no caller, no struct field, no test. Reword the one streamLimiter doc that
  contrasted against it.
- Realign comments in repo.go, pgrepo.go, session.go, util.go to the passwordless
  reality: staff lookups feed email-OTP / passkey / setup redeem, not a password
  compare; RevokeUserSessionsExcept and DeleteAllPasskeyCredentialsForUser are
  retained (uncalled) for the P5 account-remediation path (#78); "local sessions"
  no longer implies a password.

Comments and dead code only; no behavior change. Full WSL test tree green.
2026-07-04 21:47:12 +09:00
flyemoji 0c1cc598c1 feat(auth): migrate console login to passwordless
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.

- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
  login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
  methods an email can use. The single sanctioned existence oracle; methods
  are computed with no role branch, so staff and player accounts in the same
  credential state return byte-identical bodies (staffness invisible by
  construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
  /auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
  (migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
  tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.

Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
2026-07-04 21:47:12 +09:00
Lemon-miaow 3347cc05d5 feat(panel): implement user management administration panel with sessions and minecraft link support 2026-07-04 04:08:14 +08:00
Lemon-miaow 598f3d31f4 feat(submit): local + S3 backends for modpack upload contexts, installer-selectable 2026-07-02 23:38:46 +08:00
flyemoji 159107b4e3 feat(api): HTTP readiness knob on MinecraftServer and login-gate fallback default
StartupSpec.HealthHTTPPort/Path switch pod readiness from plain-TCP to an HTTP GET for RCON-less loaders (LOOHP/Limbo) that report 'started' only after the first tick. User servers now default FallbackServer to the login gate, never the lobby, so a stopped/starting backend keeps authentication in front of a fresh connection.
2026-07-02 19:38:37 +09:00
Lemon-miaow 8594622e23 feat(panel): implement email OTP verification and passkey registration management 2026-07-02 18:20:40 +08:00
flyemoji 54bc6ef211 fix(api): clear bound passkeys on password change to close a takeover foothold
handleChangePassword revoked other sessions but never cleared webauthn_credentials, and enrollment needs no step-up. A passkey planted through a transiently-hijacked session needs no password, so it survived the reset + session-revoke as a standing login foothold. Add DeleteAllPasskeyCredentialsForUser and call it in the change-password remediation so every passkey is unbound alongside the session revoke. Removing zero rows is a successful no-op. Email-OTP remains the fallback factor, so this never locks anyone out; the user re-enrolls a passkey afterward if they want one.
2026-07-02 06:55:21 +09:00
flyemoji 7278cd7c6a feat(passkey): require and record user verification at enrollment
Enrollment set no AuthenticatorSelection, so user verification defaulted to preferred (not enforced), and the UV/backup flags the ceremony reported were discarded. Set UserVerification=required so a bound passkey always proves possession AND user (a UV-incapable device falls back to email-OTP), and capture user_verified/backup_eligible/backup_state through VerifiedCredential -> PasskeyCredential -> webauthn_credentials (migration 0009) so a future login path can enforce UV per credential. Adds a negative test proving a presence-only authenticator is rejected, and asserts the roundtrip records UV=true.
2026-07-02 06:55:21 +09:00
flyemoji 99532759b2 fix(api): bound webauthn_challenges growth by superseding all prior rows
The supersede DELETE in CreatePasskeyChallenge filtered consumed_at IS NULL, so it only reaped the prior LIVE challenge; the row that each finish stamps consumed_at on was left behind. A begin->finish loop therefore accumulated one dead row per cycle, unbounded. Drop the consumed_at clause so a fresh begin reaps ALL prior rows for (user, purpose), bounding the table at one row per (user, purpose) with zero net growth per cycle. Deleting an already-consumed row is safe: it has been redeemed and nothing reads it. The fake mirrors the widened supersede.
2026-07-02 06:55:21 +09:00
flyemoji cdbb5abc35 fix(api): record credential id in passkey-register audit event
handlePasskeyRegisterFinish logged an empty target for account.passkey.registered, while the delete half logs the credential id. An operator auditing the log could see that a passkey was bound but not which one. Pass cred.ID as the audit target so bind and unbind are symmetric, and tighten the enrollment test to assert both halves name the credential id.
2026-07-02 06:55:21 +09:00
flyemoji 6368ab1914 fix(api): coalesce MyServers owned flag so ownerless rows do not 500
The MyServers query lists both a user's own servers and unclaimed (owner_id IS NULL) servers, but computed owned as s.owner_id = $1. For an ownerless row that comparison is SQL NULL, which fails to scan into the Go bool and 500s the whole listing. Wrap it in COALESCE(..., false) so an ownerless row reports owned=false while still surfacing as claimable.
2026-07-02 06:55:17 +09:00
Lemon-miaow c0d333bb98 feat(panel): support full server config edit dialog with status prefilling 2026-07-02 04:24:48 +08:00
Lemon-miaow 8ae65ae74a feat(panel): backup management 2026-07-02 03:30:50 +08:00
Lemon-miaow 15c58d982d feat(panel): player management 2026-07-02 03:30:50 +08:00
flyemoji 8f41a003b0 fix(api): clear the SSE write deadline on return so it can't leak to a reused connection
The per-write deadline that severs a stalled SSE reader was never cleared on
return. Server.WriteTimeout is deliberately unset -- a WriteTimeout would sever
a healthy long-lived stream -- and with it unset net/http never resets the
connection write deadline between keep-alive requests. So the deadline the last
writeChunk left set leaks onto the next request that reuses the pooled
connection and fails its first write for no reason. Clear it to the zero value
on return via a deferred rc.SetWriteDeadline; best-effort, a no-op on writers
without deadline support.

Also record honestly at the header flush that the connect-time stall stays
bounded only by the per-principal stream cap, not severed by this guard -- only
the mid-stream stall is closed. Adds a test pinning the clear (fails closed:
neutering the deferred clear leaves a +writeTimeout deadline set on return).
2026-07-01 23:06:26 +09:00
flyemoji 2c56d17989 docs(api): record the quota-claim TOCTOU as a KNOWN-LIMITATION (audit #4)
QuotaAvailable and ClaimServer run as two separate statements, so the
count read is not serialized against a concurrent claim's UPDATE: two
claims by one user for two different ownerless servers can both pass the
gate and both succeed, leaving the user one server over quota. It is low
severity — quota over-provisioning under a deliberate burst, not an
authorization, ownership, or isolation break, since each server is still
claimed atomically via UPDATE ... WHERE owner_id IS NULL.

Closing it requires Postgres transaction semantics (advisory-xact-lock on
the user, or SERIALIZABLE with retry) folding the gate into a single repo
method — verifiable only against a real Postgres, not the hermetic
fakeRepo suite. Documented at QuotaAvailable with back-references from the
two claim gates (handleClaim and the internal UUID claim) rather than
patched blind.
2026-07-01 22:37:36 +09:00
flyemoji d6e3189629 fix(api): bound SSE relay writes with a deadline to sever stalled readers
relayLogStream copied a pod-log follow to the client with a plain
flusher.Flush per event. On a client that stays connected but stops
reading (its TCP receive window shut), net/http buffers the small
"data:" line and only touches the socket at Flush, which then blocks
forever inside the write. The select's <-ctx.Done() branch is never
reached, because r.Context() cancels on an actual disconnect, not on a
stall, so the relay goroutine and its upstream apiserver follow leak for
the life of the process.

Route every event's write+flush through http.ResponseController with a
per-write deadline (writeTimeout, 30s): a stalled flush now returns
os.ErrDeadlineExceeded, the error plain http.Flusher.Flush swallows, and
the relay abandons the stream so the deferred cancel + src.Close release
the follow. SetWriteDeadline and rc.Flush are best-effort: a writer
without deadline support (httptest recorder; some HTTP/2 origins) ignores
the deadline and behaves exactly as before, so the guard degrades
gracefully.

This closes the leak the per-principal stream cap only bounded the blast
radius of. Verified by a deterministic test with a deadline-aware
ResponseWriter whose flush blocks until the deadline; the test times out
(fails closed) if the guard is removed.
2026-07-01 22:30:40 +09:00
flyemoji 3c1d64749f fix(api): cap concurrent SSE streams per principal
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.

Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.

This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
2026-07-01 22:04:19 +09:00
flyemoji 164ac447ef fix(api): validate inbound X-Request-Id before echo and audit persist
withRequestID honored any inbound X-Request-Id verbatim, and that value is
echoed on the response, embedded in the error envelope, and persisted into
audit_logs.request_id. An unvalidated caller-supplied id is therefore an
audit-integrity vector: an arbitrarily long value bloats the audit row, and a
stray control byte (CR/LF) could smuggle a forged entry into a log sink.

Accept an inbound id only when it is well-formed — non-empty, at most 64
bytes, and restricted to a log-safe charset ([A-Za-z0-9._-]) — otherwise mint
a fresh server id. A rejected request loses its inbound trace link, which is
strictly better than storing attacker-controlled text in the audit trail.
2026-07-01 21:07:08 +09:00
flyemoji 7a51c1d9c3 fix(api): bound concurrent login bcrypt to shed CPU-pin floods
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.

Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.

The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
2026-07-01 21:01:28 +09:00
flyemoji f34711c174 docs(api): record passkey login-handler deferral rationale
The passkey login/assertion HTTP handler stays deferred after its design
checkpoint; capture the reasoning in the handler header so the decision is
durable in the repo rather than only in task notes.

- RP boundary (resolved): felis-api is the app-login relying party (panel.*);
  the WebAuthn security gate lives at the Cloudflare Access edge. Spec §14 ties
  WebAuthn/posture to admin.* (Access) while panel.* is plain app login, so
  there is neither a spec-required assertion handler nor a backend step-up
  consumer for one.
- Identifier (blocking): a from-zero login needs a unique, human-typable handle
  to resolve an account, but users.email is nullable and non-unique and a
  player's username is their Minecraft uuid. Username-first assertion has
  nothing to key on; re-link stays the returning-player door.

Discoverable (usernameless) credentials are the future enabler; the adapter
crypto is already verified so that slice inherits correct crypto.
2026-07-01 19:15:44 +09:00
flyemoji e035142abc feat(passkey): add WebAuthn login/assertion crypto adapter
Build the assertion (login) half of the WebAuthn ceremony crypto in the
internal/passkey adapter, Oracle-verified against a virtual authenticator.

- BeginLogin/FinishLogin over go-webauthn BeginLogin/ValidateLogin,
  username-first (allowCredentials scoped to the known user's bound
  passkeys). Discoverable/usernameless login stays out of scope: the
  enrolled credentials are non-resident and the challenge store is
  user-keyed (migration 0007), so it would need a future migration.
- WebAuthnCredentials() now populates the stored COSE public key and
  signature counter (assertion validation needs both to verify the
  signature and detect clones); enrollment ignores them, so the change
  is backward-compatible and the enrollment tests guard it.
- VerifiedAssertion seam output: which credential signed plus the raw
  signature counter. Clone/regression policy is deliberately NOT here —
  the counter is a ceremony fact and the future handler, which holds the
  previously stored counter, decides reject/warn.

Scope: crypto adapter only. The login HTTP handlers, session minting,
and the panel.* relying-party boundary/tier decision remain a deferred
slice (no unauthenticated login route is added). BeginLogin/FinishLogin
live on the concrete adapter, not the api.PasskeyVerifier interface,
which grows only when a handler consumes them.

Tests (virtualwebauthn): a real enrollment chained into a real assertion
exercises the COSE public-key decode path and surfaces the advanced
signature counter, plus origin-mismatch and unbound-credential rejection.
2026-07-01 18:45:26 +09:00
flyemoji 7464fa700b fix(updates): tag Window JSON so the persisted maintenance window round-trips
The admin API persists the auto-update maintenance window as lowercase
JSON {"start","end"} (platform_settings key "update_window"), but
updates.Window had no json tags, so it marshaled/unmarshaled with
capitalized keys. The natural decode the update runner will use --
json.Unmarshal(stored, &updates.Window{}) -- would therefore miss every
key and silently yield the zero Window. That fails closed (a zero window
Contains nothing, so notify-only, never a rogue apply), so it is safe but
a latent silent-zero trap for the not-yet-built runner.

Add json:"start"/json:"end" to updates.Window so the obvious decode is
correct by construction; value time.Time treats a stored null as a no-op,
so a cleared/never-set window still decodes to the zero Window. Nothing
in the package serialized Window before, so this changes no existing
behavior.

Guarded by a cross-package contract test in internal/api that marshals
the real api.updateWindow DTO and unmarshals it into updates.Window --
asserting the interval survives (Contains(mid) is true) and that an empty
window decodes to the fail-closed zero Window -- so the two shapes cannot
drift apart silently.
2026-07-01 18:26:42 +09:00
flyemoji 3673af63c2 feat(api): add admin API for the SysAdmin-set auto-update maintenance window
Two admin-tier routes read and set a single platform-wide maintenance
window for the auto-update subsystem (decision core internal/updates):

  GET /api/v1/updates/window
  PUT /api/v1/updates/window

The window is stored as JSON {"start","end"} (RFC3339, or null when
unset) under the platform_settings key "update_window", reusing the
existing GetSetting/SetSetting KV seam -- no new Repo method, no
migration. Pointer times keep "unset" (null) distinct from a real
instant on both decode and encode; a never-set and an explicitly
cleared window both read back as {null,null}.

Validation mirrors the core's fail-closed Window: a window is either
fully set (both ends, end strictly after start) or fully cleared (both
null). A half-set, inverted, or empty-interval body is 400 and is never
persisted. Reads treat only a missing key as unset (ErrNotFound -> 200
nulls); any other store error 500s rather than fail open.

This is API + PERSISTENCE ONLY. Nothing consumes the stored window yet
-- the runner, the ReleaseSource/Notifier/Applier executors, and the
scheduler CronJob remain INTEGRATION-ONLY. Setting a window changes no
behavior until those land; it is the durable input they will read.
Nothing here force-updates ("不要强制自动更新").
2026-07-01 18:17:07 +09:00
flyemoji fe2ece08cc feat(api): add public Bind-Code onboarding for the player console
Adds POST /api/v1/auth/bind, the one public pre-account entrypoint of the
player console (console.<root_domain>). An account-less player redeems the
one-time Bind Code minted in the in-game Login Lobby; in a single step the
platform creates a role=user player, links it to the verified in-game UUID,
and mints a host-only felis_session. Login is thus not forced at the edge
while operations stay app-authenticated.

The operator console (op.console.<root_domain>) is unaffected and stays
behind Zero Trust: a code whose UUID resolves to a staff (role=admin)
account is refused with 403 (ErrPlayerBindForbidden) without consuming the
code, so the public door provably never yields an admin principal — the
session it mints carries ViaAdminAccess=false and is host-only to console,
never sent to op.console.

Repo layer: new RedeemPlayerBindCode on the Repo interface, implemented on
PGRepo (single tx: resolve code, create-or-fetch the player, consume) and
the test fake. The returning-player branch is idempotent and is a deliberate
standing "log in via the game" door, not just first-time onboarding.

Honest labeling:
- ORACLE-VERIFIED (Go): account/session logic — role=user, refuse-staff,
  idempotent create-or-fetch, single-use code, and the op.console redline
  (player session rejected on admin routes). Covered by handlers_onboard_test
  and the OpenAPI parity gate.
- INTEGRATION-dependent: the endpoint's security rests on the Bind Code having
  been minted against an online-mode-Yggdrasil-authenticated UUID, a
  precondition that lives in velocity/Java and is not verifiable from this
  repo (CODE-ONLY). The Go layer proves the logic, not that identity guarantee.
- No app-level attempt cap: rate-limiting is deferred to the edge as for the
  public /auth/login; the ~1e12 keyspace, single use and short TTL make a
  blind app-level cap non-critical.
2026-07-01 18:00:13 +09:00
Lemon-miaow d2de11af22 feat(panel): fleet 2026-07-01 03:46:07 +08:00