Commit Graph
504 Commits
Author SHA1 Message Date
flyemoji be0c4c41f4 feat(bootstrap): install authenticated game stack 2026-07-14 02:58:21 +09:00
flyemoji dab8fc214b feat(setup): bind owner through login gate 2026-07-14 02:57:15 +09:00
flyemoji 5dc8eb92a8 feat(operator): secure system server workloads 2026-07-14 02:56:19 +09:00
flyemoji 688c86da18 feat(proxy): enforce login-first routing 2026-07-14 02:55:11 +09:00
flyemoji 16b7ad2756 feat(runtime): add authenticated system backends 2026-07-14 02:54:13 +09:00
flyemoji fd062882ed feat(nano): give a Mojang player's name back to them, by prefixing the squatter
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.

The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.

A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.

The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.

prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.

Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.

Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:

  premium FLYEMOJ1     -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
  LittleSkin FLYEMOJ1  -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
  both online at once, zero "already connected" rejections
  LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3

The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.

Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.

BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
2026-07-13 13:00:54 +09:00
flyemoji b323975ddb fix(nano): -Dmojang.sessionserver takes the full hasJoined URL, not the base
d417efc got this backwards, in both the code comment and the installer summary.
It claimed authlib appends /session/minecraft/hasJoined itself, so the property
should be given the base URL only. Velocity does not work that way, and a real
login says so: pointed at http://127.0.0.1:8081, a Mojang login arrives at nano
as

    GET /?username=FLYEMOJ1&serverId=-23ae0b50...

with no path at all. Velocity appends the query string to the property verbatim
and issues the request itself; authlib is not in the loop. nano has no route on
/, so it answers 404 and Velocity kicks the player with authservers_down.
Velocity's own default for the property is the full URL,
https://sessionserver.mojang.com/session/minecraft/hasJoined, which is the same
thing said another way.

With the full endpoint URL the same account logs straight in, so both the
nano.go header and summary_nano now print

    -Dmojang.sessionserver=http://127.0.0.1:8081/session/minecraft/hasJoined

and note that the flag belongs between `java` and `-jar`.

Verified against Velocity 3.5.1 + Paper 26.2 on the deploy host: a Mojang login
reaches the backend with its real Mojang UUID unchanged, and a LittleSkin login
under the same username reaches it as UUIDv3(felisAuthNS, "littleskin:"+id) --
two different players on the backend, which is the point.
2026-07-13 11:05:45 +09:00
flyemoji d417efc8bf fix(nano): make the installer work on EL10 and stop serving hasJoined to the world
Deploying `felis nano` to a real Rocky Linux 10 host surfaced four defects that
no local check could see. Fixed together because they all sit on the same path
from `curl|bash` to a running felis-nano.service.

* docker killed the nano install on EL10. `acquire_nano_binary` pulled in
  docker purely to build the binary; on Rocky 10.2 the docker-ce el10 rpms
  install but dockerd refuses to start, so the install died at
  `systemctl enable --now docker`. nano needs one static binary, not an image,
  so the docker dependency is gone: fetch_source -> install_go_toolchain
  (pinned FELIS_GO_VERSION, default 1.26.4, amd64/arm64) -> build_nano_binary.

* the built binary could not be exec'd by systemd (203/EXEC). The Go linker
  renames its output out of $TMPDIR, and a same-filesystem rename carries the
  source SELinux label, so `go build -o /usr/local/bin/felis` produced a binary
  labelled user_tmp_t rather than bin_t. root is unconfined and could run it by
  hand, which is what made this look fine, but the DynamicUser service could
  not. build_nano_binary now stages the output and installs it as a fresh file
  so the policy type transition labels it bin_t, with restorecon as a belt.

* re-running the installer did not converge. `systemctl enable --now` is a
  no-op on an already-active unit, so a rebuilt binary was installed while the
  old process kept running. Now enable + restart.

* the Velocity wiring comment in cmd/felis/nano.go was wrong. authlib appends
  /session/minecraft/hasJoined itself, so -Dmojang.sessionserver takes the base
  URL only, as the installer has always printed.

Also bind to loopback by default. hasJoined is unauthenticated by protocol --
authlib speaks the vanilla sessionserver dialect and sends no token -- so a
public bind is an open auth relay: anyone can point their own proxy at it and
spend this host's egress IP on Mojang until Mojang rate-limits it and the
operator's own players stop getting in. It is not an identity bypass (a caller
still needs a serverId hash bound to their own server key, which the upstream
Yggdrasil validates), but it is someone else's traffic on your address.
FELIS_NANO_LISTEN and the -listen flag now default to 127.0.0.1:8081, which a
same-host Velocity reaches unchanged; serving an off-host proxy is an explicit
opt-in. configure_nano_firewall no longer opens a port for a loopback bind, and
summary_nano prints the real bind address plus the relay warning.

Verified on the target host: installs with no docker present, service active,
binary labelled bin_t, `ss` shows LISTEN 127.0.0.1:8081, an external request is
unreachable, and an in-host request returns 204 with the login logged.

BREAKING CHANGE: felis nano defaults to 127.0.0.1:8081 instead of 0.0.0.0:8081.
A Velocity proxy on another machine must now set FELIS_NANO_LISTEN (or -listen)
to a reachable address, and should allow that port only from the proxy's IP.
2026-07-13 10:30:18 +09:00
flyemoji 87aa9ea625 feat(deploy): bootstrap can install Felis-nano only, chosen at the start prompt
bootstrap.sh now asks up front whether to install the full Felis control
plane or only Felis-nano, and grows a parallel install path for the
nano-only case.

- prompt_install_mode() runs right after OS detection and reads /dev/tty
  (so it works under `curl ... | sudo bash`) offering [1] Felis / [2]
  Felis-nano, default full. FELIS_INSTALL_MODE=full|nano skips the prompt
  for non-interactive runs; no tty falls back to full.
- main_nano() installs only what nano needs: the felis binary (reusing
  the embedded-binary / docker-build acquisition), a template felis.toml
  carrying a commented [[auth_source]] example (Mojang-only until edited),
  a felis-nano.service unit running `felis nano -config ... -listen ...`
  under DynamicUser hardening, and a firewalld port-open for the listen
  port. None of the k3s / Postgres / migrate / bundle steps run.
- write_nano_config is idempotent (leaves any existing config untouched)
  and its template is valid as-is. summary_nano prints the hasJoined
  endpoint and the Velocity -Dmojang.sessionserver flag, offering the
  127.0.0.1 form when the proxy is on the same host.

Verified on WSL: `bash -n` clean; the emitted template loads via
config.LoadNano and the exact systemd ExecStart command serves 204 on a
miss ("Mojang + 0 third-party source(s)"); a duplicate-tag config still
exits non-zero citing "unique". Not exercised: a full main_nano run,
systemd activation of the unit, and shellcheck (unavailable in this env).
2026-07-13 02:39:50 +09:00
flyemoji d177428fb0 feat(felis): add felis nano — Yggdrasil hasJoined multiplexer without a control plane
`felis nano` serves the vanilla sessionserver protocol
(GET /session/minecraft/hasJoined) as a federating multiplexer over
Mojang plus any number of third-party Yggdrasil roots, with no k3s,
Postgres, or panel — a MultiLogin-style auth front-end delivered as a
subcommand of the single felis binary rather than a separate build.

- config.LoadNano reads only [[auth_source]] blocks; it skips the
  database.url / root_domain / archive requirements the full server
  needs. Zero sources is valid (Mojang-only).
- Mojang is prepended in code (Identity:true), never from config, so it
  is always the sole identity root. Third-party profiles are rewritten
  to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID).
- validateAuthSources rejects unknown keys, duplicate tags, and
  scheme-less URLs — a malformed nano config fails loud at load.
- Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend);
  a rejected login is a 204, matching the vanilla sessionserver.
- nano.go binds the -listen flag and ignores [server] listen in config.

Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime
smoke against the template config returns 204 on a miss and logs
"Mojang + 0 third-party source(s)"; a duplicate-tag config exits
non-zero citing "unique".
2026-07-13 02:39:41 +09:00
Lemon-miaow cc0c945811 docs: Update README.md 2026-07-12 04:37:11 +08:00
Lemon-miaow 241374db58 docs: Update README.md 2026-07-12 03:41:52 +08:00
Lemon-miaow 037eb248d2 docs: Update LICENSE 2026-07-12 01:42:07 +08:00
flyemoji ecea20ee7c feat(nano): configure hasJoined auth sources via [[auth_source]], Mojang-anchored
Step 2 of Felis-nano: a [[auth_source]] array-of-tables (tag + full hasJoined
url, config order = priority) supplies the multiplexer's third-party Yggdrasil
roots; cmd/felis prepends Mojang as the sole code-owned identity anchor and
wires them into API.AuthSources. With no sources configured the endpoint stays
inert (204s), unchanged from step 1.

The config deliberately has no identity/trusted field: Mojang is the only source
whose self-asserted UUIDs are trusted verbatim, so no misconfiguration can
reopen the impersonation hole the per-source UUID rewrite closes. An identity=
key is an unknown key and Load rejects it. Validate adds two fail-fast guards:
unique tags (namespace collision) and a scheme-qualified url (else the source is
silently dead, never validating any login).
2026-07-12 02:27:02 +09:00
flyemoji ed8fa0cdbf docs(changes): record the Felis-nano hasJoined resolver and reconcile the audit ledger row 2026-07-12 02:10:49 +09:00
flyemoji ff550c41ef feat(nano): federating hasJoined multiplexer with per-source UUID namespacing 2026-07-12 02:09:47 +09:00
flyemoji 667c6d33d2 docs(changes): record the adversarial input-validation audit (sink-first negative-path) 2026-07-08 12:38:40 +09:00
flyemoji b7b4a3be45 docs(changes): record the round-2 backup/restore mutation audit and index the owner-gate test
Round-2 mutation audit of the backup/restore data-safety surface (7 fail-open gates pinned, 1 coverage gap found and closed by the owner-gate test in 85b8a92). Reverts the Pending section and indexes c67a4d3 + 85b8a92 into the committed ledger.
2026-07-08 05:59:13 +09:00
flyemoji 85b8a92a0e test(api): pin restore's owner gate against a superseded former owner
handleRestoreBackup's owner-or-admin gate was not pinned by any test: the former-owner gate backstopped every non-owner case the suite exercised, so a broken owner gate would not redden. Add the mirror of the former-owner test — a released former owner (still the backup's former_owner, no longer the current owner) must get 403 — the sole subtest that fails when the owner gate is disabled. Found by the round-2 backup/restore mutation audit; production code unchanged.
2026-07-08 05:58:00 +09:00
flyemoji c67a4d3f35 docs(changes): close §B4 with the S3 archive backend deferred by design
The operational break-glass ops (Owner provision, OP create, halt, Sync backup) are built and oracle-verified. The fourth §B4 line item — the tarS3 archive backend — is recorded as a deliberate deferral, not a silent gap: config.Load fail-closes store=tarS3 (frozen by TestLoadRejectsUnimplementedArchiveStore), tarLocal is the tested baseline every backup/restore path uses today, and offsite/cross-cluster DR is opt-in future work. Supersedes the "S3 remains open" note in the phase-2b peer doc.
2026-07-08 05:56:59 +09:00
flyemoji 729bd7ba0f docs(changes): index the mutation audit and two lagging ledger rows 2026-07-07 23:45:50 +09:00
flyemoji 4626ab580e docs(changes): mutation-audit the ledger's "unit-tested" safety claims
Break each load-bearing safety gate the change ledger names as "unit-tested" and
confirm the specific test turns red — passing proves GREEN, not that the test would
catch a regression. All 18 fail-open crown-jewel gates across the subsystems (auto-update
pin/no-downgrade/no-prerelease/window, passkey clone-refuse, modpack CAS + Trivy scan-gate,
cfsetup fail-closed + NodePort fence conn-count, SSE cap, OTP atomic reserve, idle stop,
startup/readiness timeout, /readyz deps, naming reservation, service-token login-pod-only)
are mutation-proven to pin behaviour. Every documented "unit-tested" claim is reconciled
one-for-one; the fence conn-count gate, previously mis-classified as integration-only, is
corrected and verified. No code changed — read-and-verify only.
2026-07-07 23:45:12 +09:00
flyemoji 5a7cd5abe0 docs(changes): fold 346ec68 cloudflare-edge walkthrough into its detail doc
A completeness re-check of the ledger backfill (in-scope pre-fad48ff
feat/fix/refactor commits vs SHAs actually cited in detail docs, excluding
the auto-generated ledger table) surfaced one backend functional commit with
no detail-doc home: 346ec68 refactor(deploy) — an 868-insertion rework of the
cloudflare-edge TUI walkthrough plus a tested cfsetup integration-runner path.
The "improved cloudflare walkthrough" subject undersold a behaviour change, so
it is folded into the existing cloudflare-tunnel-access-edge detail doc and its
INDEX row rather than left orphaned. Backend detail-doc coverage of the
pre-ledger history is now complete (0 backend orphans).
2026-07-07 21:05:24 +09:00
flyemoji 096d59716e docs(changes): backfill detail docs for pre-ledger functional commits
Retroactively author 15 grouped detail docs covering the backend
functional (feat/fix) commits made before the change ledger was
established (fad48ff), closing the ledger's detail-doc axis for the
pre-convention history. Each doc groups a feature's constituent commits,
lists their SHAs with subjects, and carries a backfill note stating it
was reconstructed from git history on 2026-07-07 and not independently
re-verified (current tree green at 9911b8c).

Add a Detail docs section to INDEX.md linking every detail doc (the 6
existing + 15 backfill) to the commit(s) it covers, so a doc is findable
from the index without a column on the auto-generated ledger table. Catch
the table up with the missing 9911b8c row.

Scope: backend (Go/Java/K8s) only, per the ledger's stated convention
that frontend/panel commits are the collaborator's UI work; non-functional
commits (docs/style/chore/refactor) keep their table row without a
dedicated detail doc.
2026-07-07 20:55:51 +09:00
flyemoji 9911b8cd41 docs(changes): record the break-glass backup console peer (§B4 Sync phase 2b) 2026-07-07 19:34:00 +09:00
flyemoji fc748d3462 feat(breakglass): add "back up a world now" console peer (§B4 Sync)
Adds a break-glass console operation that snapshots a stopped world by
calling the felis-api internal face while the API is alive, rather than
rendering the backup Job locally: the Job needs felis-api deployment
coordinates the console does not hold.

The peer resolves the felis-api-internal ClusterIP Service + service
token from the control namespace, POSTs the internal backup endpoint
with the operator os_user for audit attribution, and maps 409/503/404
to friendly outcome cards. Core decision logic lives in backupnow.go
(unit-tested against a fake client + httptest); tui_backupnow.go is the
untested bubbletea glue mirroring tui_halt.go.
2026-07-07 19:32:13 +09:00
flyemoji 4918d93412 docs(changes): record the felis-api internal-face ClusterIP Service fix
Detail doc + ledger row for 2ba9948: the separate felis-api-internal ClusterIP
Service that gives the login pod (and the break-glass console) a routable 8081.
2026-07-07 11:24:38 +09:00
flyemoji 2ba994889e fix(platform): front the felis-api internal face on its own ClusterIP Service
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.

Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).

Manifest-level fix; the live packet path is pending real-cluster verification.
2026-07-07 11:24:09 +09:00
flyemoji 73195ca48f docs(changes): record the internal-face backup endpoint (§B4 Sync phase 2a)
Detail doc + ledger row for f2fc57c: the internal (service-token) twin of the
on-demand world backup endpoint that the break-glass console peer will call.
2026-07-07 10:48:51 +09:00
flyemoji f2fc57cad9 feat(api): add internal-face break-glass world backup endpoint (§B4 Sync)
Add POST /api/v1/internal/servers/{name}/backup so the on-node break-glass
console can snapshot a stopped world while felis-api is alive. It goes through
the API (not direct-to-CRD like halt) because rendering the backup Job needs
deployment coordinates (FELIS_IMAGE, FELIS_BACKUP_PVC) only felis-api holds.

Service-token auth (no Principal); the middleware IS the authorization, since
the operator already has root on the node. Refactor the RWO stopped-gate,
optional-Backuper 503, async hand-off and audit+202 into a shared enqueueBackup
tail so the external (owner/admin) and internal (break-glass) faces cannot
diverge on the security-critical stopped-gate. The internal audit is attributed
to break-glass/internal so a console-initiated backup is distinguishable from an
owner self-service one.
2026-07-07 10:48:01 +09:00
flyemoji c3b4d7b957 docs(changes): backfill 7a7c0d5 into the change ledger
Move the on-demand world backup entry from Pending into the committed
ledger and record its SHA in the detail doc.
2026-07-07 10:05:58 +09:00
flyemoji 7a7c0d53ab feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
2026-07-07 10:04:30 +09:00
flyemoji fad48ff21d docs(changes): establish the change ledger for functional changes
Add docs/changes/ — a durable, in-repo map of every functional change and
the commit that records it, independent of git log. INDEX.md carries the
convention (each functional change gets a dated detail doc plus a ledger
row) and the full oldest-first ledger, regenerable losslessly from git.
Seed detail docs for the two changes just landed: the break-glass halt op
(c2ee21a) and the /felis migrate command (c1aa38b).
2026-07-06 23:52:12 +09:00
flyemoji c1aa38bac1 feat(velocity): add /felis migrate to open an account migration (§B3 inherit)
Add the in-game /felis migrate command that a player runs to open an
account migration, the entry point for handing their owned servers to
another account (§B3 inherit, scenario A). The command posts the player's
Mojang-verified UUID to the existing handleMigrateStart backend, which puts
the account into migrate mode; the player then finishes on the web console
(prove identity, name the receiving account, redeem a one-time code).

Mirrors the existing /felis claim path: requires a real player past login
limbo, acts on the caller's account rather than the current server, expects
201 Created affirming started=true (a 201 without it is a contract breach,
not a refusal), and maps the handler refusals (404 not_linked, 409
account_retired) to player-facing guidance. On success it points the player
at https://console.<root_domain>, derived from config, never hardcoded.
Compile-verified against velocity-api:3.3.0-SNAPSHOT via the podman gradle
toolchain. Closes the code-only gap named in handlers_account_migrate.go.
2026-07-06 23:50:57 +09:00
flyemoji c2ee21ae05 feat(breakglass): add halt-a-server op to the recovery console (§B4)
Add a root-gated "Halt a running server" operation to the break-glass
console. The operator picks a server from the live fleet and the console
flips that MinecraftServer CRD's spec.desiredState to Stopped via a
spec-only merge patch, disjoint from the operator's status writes, so it
cannot race or clobber reconciliation. It is the panel-independent
emergency stop for when the box still has root plus a kubeconfig.

System servers (login/lobby) are allowed but flagged: a system tag in the
picker and an explicit WARNING in the post-exit summary, since halting
login takes the shared auth front door down with no fallback. Audit is
best-effort so a halt still works with the audit sink down. Already-stopped
is a distinct no-op. The core (halt.go) is unit-tested against a real
controller-runtime fake client that applies the patch.
2026-07-06 23:50:09 +09:00
Lemon-miaow abad137a01 style(panel): unify vertical spacing below PageHeader across pages 2026-07-06 02:04:54 +08:00
flyemoji fdb6efbd88 feat(account): migrate a live account's owned servers to a new account (§B3 inherit)
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.

- 0015 migration: account_migrations state machine (initiated -> confirmed ->
  code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem

Quota is grandfathered at redeem: no per-target quota re-check when servers move.
2026-07-05 20:50:48 +09:00
Lemon-miaow bbcfaeb6c4 refactor(panel): remove redundant voxel network topology description subtitle 2026-07-05 16:51:12 +08:00
Lemon-miaow cfe68aedfc fix(panel): align status distribution order to put Stopped at the end 2026-07-05 16:50:16 +08:00
Lemon-miaow 9079a2cd67 feat(panel): implement email otp and passkey login interface 2026-07-05 16:48:03 +08:00
Lemon-miaow 7becb38488 fix(api): implement /readyz with real DB + K8s API + CRD checks (§7)
Previously /readyz only verified Repo != nil && Cluster != nil — a
process-liveness check, not a dependency-health check. The spec
requires the readyz probe to verify DB, K8s API, and CRD informer
are live before declaring the pod ready.

- Repo interface gains Ping(context.Context) error
- Cluster interface gains Ping(context.Context) error
- PGRepo.Ping delegates to sql.DB.PingContext
- K8sCluster.Ping lists MinecraftServer CRDs (Limit=1) in the
  configured namespace, exercising both the API and CRD informer
- handleReadyz iterates ping checks; any failure returns 503 with
  the failing dependency name in the error message
- fakeRepo and fakeCluster gain configurable pingErr for hermetic
  test coverage of the failure paths

New test: TestReadyzPingsDependencies verifies 200 when healthy,
503 when DB or K8s API is down.
2026-07-05 16:22:08 +08:00
Lemon-miaow 7f7e459746 fix(operator): enforce startup and readiness timeouts (§5, §8)
Previously a server whose pod was ready but RCON probe kept failing
would stay in Starting phase forever. The CRD defines TimeoutSeconds
and ReadinessTimeoutSeconds but the reconciler never checked them.

- PodNotReady path: if the pod stays not-ready past timeoutSeconds
  (default 300s), transition to Failed
- RCON unreachable path: if RCON stays unreachable past
  readinessTimeoutSeconds (default 300s), transition to Failed
- Helper methods startupTimedOut/readinessTimedOut compare
  StartRequestedAt against the respective timeout, falling back to
  300s defaults when unset
- 2 new tests: ReadinessTimeoutConvertsToFailed, StartupTimeoutConvertsToFailed

Fixes the scenario where a broken backend (bad jar, crash-looping
process) would permanently occupy a Starting server slot.
2026-07-05 16:22:08 +08:00
flyemoji 9e1df12975 feat(passkey): advance sign_count, reject clone-warned assertions
Both login doors (username-first and discoverable) now run a shared applyAssertionCounter after a verified assertion. A signature-counter regression — go-webauthn's CloneWarning, the possible-cloned-authenticator signal — is refused fail-closed with the same opaque passkey_login_invalid envelope any other finish failure returns (no clone oracle to a prober) and audited distinctly as auth.passkey_clone_rejected under the resolved account. A clean assertion advances the stored sign_count to the asserted value and stamps last_used_at, before any session is minted.

Counter-less/synced authenticators report 0 and never warn, so they pass through and simply re-stamp 0; the check gates only counter-keeping hardware authenticators, where a rollback is the meaningful signal. Email-OTP and username-first passkey remain fallbacks, so a rejected clone is never bricked.

Adds Repo.AdvanceCredentialSignCount (pgrepo UPDATE by credential_id) and surfaces CloneWarning from the internal/passkey adapter's FinishLogin/FinishDiscoverableLogin. Proven by real-crypto adapter tests (a counter regression still verifies but flags CloneWarning), handler tests (advance-and-stamp on success, fail-closed on clone), and a symmetric test on each door so both call sites of the shared helper are covered.
2026-07-05 16:02:30 +09:00
flyemoji 0dbd557a7a fix(store): renumber discoverable-login migration 0013 -> 0014
A resource-cache migration (0013_resource_cache.sql) was merged onto main concurrently and also claimed version 0013. LoadMigrations rejects any duplicate migration version, so the app would refuse to boot with both files present.

Renumber the discoverable-login migration to 0014. The two migrations touch disjoint objects (0013 ALTERs servers to add cached_* columns; this one CREATEs webauthn_discoverable_challenges), so their relative order does not matter, and the table name is unchanged -- no Go reference moves.

Renaming a just-published migration is safe here because neither version has been applied to a persistent database yet: there is no schema_migrations row for version 13 to reconcile. This is a pre-application renumber, not a history rewrite of an already-applied migration.
2026-07-05 04:31:45 +09:00
flyemoji 154002edf3 docs(auth): cite MultiLogin reference for UUID-keyed reclaim split
Anchor the username-collision reclaim's UUID-keyed, proxy-detected design to
the multi-Yggdrasil reference: CaaMoe/MultiLogin v6 binds identity as
serviceId+online-UUID via "identity cards" that decouple the in-game name from
online identity — keyed by UUID, never by name. Note that §B3's Mojang-priority
reclaim goes beyond the common "protect the first-bound name" behavior by
evicting a squatter once the genuine Mojang owner appears and stashing the
squatter's data for the code-only inherit path.
2026-07-05 04:05:56 +09:00
flyemoji ec468baef9 feat(auth): add discoverable (usernameless) passkey login
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.

Routes (both Public):
  POST /api/v1/auth/passkey/login/discoverable/begin
  POST /api/v1/auth/passkey/login/discoverable/finish

Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.

The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.

Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
2026-07-05 04:05:56 +09:00
flyemoji 7db57b9fff feat(updater): add VersionGatherer extraction core and CLI gather seam
Give the Runner a way to read each component's CURRENT version so it can be
compared against the release sources already wired. Three pure extractors turn
raw system text into an updates.Version, each fail-closed:

  - versionFromCLI      — a `<tool> --version` banner   (k3s, cloudflared)
  - versionFromImageRef — a container image tag         (felis-api)
  - versionFromJarName  — a proxy jar filename          (velocity)

sysGatherer routes each Topology component to the right extractor over an
injected seam; every path is exercised with a fake runner, mirroring how the
release sources are proven against httptest.

The load-bearing case is k3s: its Git tag "v1.36.2+k3s1" parses stable, but a
registry cannot store '+', so the same build ships as image tag "v1.36.2-k3s1",
which parses as a prerelease unless repaired. versionFromImageRef normalizes
"-k3sN"/"-rke2rN" back to "+", so an image read and a CLI read agree instead of
the image masquerading as a prerelease and being barred from comparison.

Honest runtime state after this slice — a green suite is not "the updater runs
against real infra": only the CLI seam (execRunner) is wired, so of the four
tracked components just cloudflared is live end to end (gatherable AND
Scheduled/appliable). k3s is CLI-gatherable but Notify-only. felis-api and
velocity are NOT yet runtime-gatherable: their producing seams — a k8s read of
the control-plane Deployment image, and an off-cluster jar inspection — are left
nil, so both surface an explicit "gather seam not wired" error rather than a
wrong version. felis-api self-update is therefore not functional yet.

Remaining integration (tracked in doc.go): the two producing seams, the concrete
Notifier (SMTP + in-game), the Applier (image bump, cloudflared swap), the
`felis update` CLI + CronJob entry point, and the runtime append of the Pinned
Minecraft fleet.
2026-07-05 04:05:56 +09:00
Lemon-miaow e574749877 feat(api): enforce CPU/memory/storage quotas (spec §9.3, §22)
Add per-user resource quota enforcement across all four dimensions:
max_servers, max_cpu_milli, max_memory_mb, and max_storage_gb.

- Migration 0013: add cached_cpu_milli, cached_memory_mb, cached_storage_mb
  columns to servers table for pure-SQL per-owner aggregation
- SeedServer now writes resource cache alongside server row
- QuotaCheck replaces QuotaAvailable at claim time, checking all four caps
  against the owning user's cumulative usage
- handlePatchServer checks owner's quota before allowing memory/resource
  changes on owned servers; unowned servers skip the gate
- handleInternalClaim mirrors the full quota check
- UpdateServerResources keeps the cache in sync after spec mutations
- Reaper zeros resource cache on ReleaseWorld so released resources
  are not counted against a former owner
- quantityToMilli/quantityToMB helpers convert K8s quantities to
  quota-comparable integers

19 test packages pass.
2026-07-05 02:05:44 +08:00
Lemon-miaow 91bfa27e8c feat(operator): implement idle auto-stop (spec §8)
Adds empty-server auto-stop to the reconciler. When Idle.AutoStopEnabled
is true and the RCON player tally is zero for EmptySecondsBeforeStop seconds,
the operator flips desiredState to Stopped, which triggers the normal
graceful-shutdown path.

- Add EmptySince status field to track empty duration
- Clear EmptySince on stop and when players return
- Reuse existing RCON probe's player count (zero extra network cost)
- 4 new test cases covering timestamp, timeout, player-join reset, and
  disabled-by-default
2026-07-05 01:52:45 +08:00
Lemon-miaow bd49313e7e refactor(panel): 抽取 10 个公共组件,消除 ~150 处重复代码
新增组件:
- PageHeader: 统一页面头部(14 个页面)
- StatCard: 统一数值统计卡片(6 个页面)
- MessageLine: 统一错误/成功提示条(16 处)
- ConfirmFooter: 统一 Dialog 底部按钮(8 个文件)
- SearchInput: 统一搜索框(5 个页面)
- SubmissionStatusBadge: 统一提交状态徽章
- RoleBadge / UserStatusBadge: 统一用户角色/状态徽章
- BackLink: 统一返回导航链接(5 个页面)
- InlineConfirm: 统一二次确认交互(3 个 section)

修复:
- Pagination 图标尺寸统一 h-4 w-4
- 右上角按钮统一 size=sm gap-1.5
- StatCard 默认值恢复 text-2xl(24px)

States.tsx 新增 NotYours / NotRunning 通用状态组件

已有文件 -659 行,新组件 +353 行,净减 306 行
2026-07-05 01:15:03 +08:00