Commit Graph
95 Commits
Author SHA1 Message Date
Lemon-miaow b6ef27cd2d fix(auth): enforce the verified-email uniqueness that email login assumes
The design has claimed since migration 0010 that at most one account can hold
a PROVEN email address, with ErrEmailTaken as the 409 a second verifier sees.
Neither half ever shipped: no migration created users_verified_email_unique,
and VerifyEmailOTP had no guard at all — the sentinel was defined but never
returned, so two accounts could both verify one address. The damage is not
cosmetic: the pre-session login resolves accounts BY verified email, so the
duplicate decided which identity a mailed sign-in code belonged to.

- Migration 0020 creates the partial unique index (lower(email) WHERE
  email_verified) the comments have been citing — the database-level backstop.
- VerifyEmailOTP now refuses the take-over with ErrEmailTaken BEFORE consuming
  the code (the address, not the code, is the problem), charges no attempt,
  and maps a lost cross-user race (unique violation) to the same answer.
- The verify handler answers 409 email_taken instead of a generic 500.

Covered by the pgint suite (sequential double-verify refused with the code
still live, a direct duplicate write still loses to the index, the refused
account can still prove its own address) and a hermetic 409 case.
2026-09-23 03:19:35 +08:00
Lemon-miaow f79e5ebb5e feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:

- submit: derived context refs become the internal-face URL
  /api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
  gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
  route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
  image's new fetch-context entrypoint) that streams the blob with the
  namespace-local service-token Secret — never mounted into Kaniko — and
  extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
  reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
  build namespace gets the token Secret through the existing replica mechanism
  (bootstrap.sh + felis setup); the build egress lock opens exactly the control
  namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
  escapes/symlinks/non-gzip).

Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
2026-09-22 22:45:09 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow 2a8f897e61 fix(api): a session-store outage answers 503, not 401
Resolving a session cookie failed identically whether the credential was
missing or Postgres was unreachable: local_auth_enabled read errors fell into
the fail-closed 'disabled' branch and SessionUser errors into 'invalid
session', both surfacing as 401 'authentication required' — a lie that reads
as 'log in again' during an outage. Split the enabled-read into
(enabled, error), tag non-ErrNotFound store failures with errAuthBackend, and
map that to a new 503 auth_unavailable in requireExternal. Fail-closed is
unchanged: missing setting / bad value / missing session stay 401.
2026-09-22 20:30:43 +08:00
Lemon-miaow 0414913bc7 fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
6-way concurrent redeem of one code 500'd on users_username_key (each request
generated a fresh user id but the same uuid-derived username), plus the rarer
two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
locks its code row and handles the conflict.

- SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
  as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
- INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
  cross-code racers converge on the winner's row (role checked, staff still
  refused) instead of a unique-violation 500.
- account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.

Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
db clean; zero unmapped errors.
2026-09-22 20:08:38 +08:00
Lemon-miaow dcc3b7403e fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake)
The interface doc promised 'each joined to its staff username', the fake and
the pending handler both project username and created_at, but the PG query
selected neither — live internal /op-login/pending returned username:"" and
created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
tests can't see PG-only regressions.
2026-09-22 20:04:32 +08:00
Lemon-miaow 52549f7b3a fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
The PG implementation was a single UPDATE ... WHERE code_hash that returned
ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
the pre-session email-login door (and the op-login finish / migration confirm
doors) fell through to writeError's unmapped-error 500, and attempts were never
charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
intended contract; only the PG side had drifted.

Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
UPDATE the newest live row, expiry + attempt cap before the hash compare,
mismatch charges one attempt and returns ErrOTPInvalid without consuming,
match consumes and commits. Verified live on the VM: 5 wrong guesses return
400 and stop at attempts=5 (correct code then also refused, unconsumed);
fresh code redeems; replay returns 400.
2026-09-22 19:51:04 +08:00
Lemon-miaow 9309ff5a7f chore: apply the missed S1016 conversions in handlers_users
The gofmt/staticcheck commit staged handlers_user.go (singular) for the
formatting fix but missed this sibling for its two struct-literal-to-
conversion cleanups.
2026-09-22 17:57:03 +08:00
Lemon-miaow a05edc934c chore: gofmt the tree, clear staticcheck, add a CI gofmt gate
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
2026-09-22 17:49:54 +08:00
flyemoji e9f74f3f0f fix(nano): stop trusting an expired free name while mojang is failing
When the premium-name lookup failed, isPremiumName fell back to any
cached answer, however old. An expired "free" is exactly the answer
that may have stopped being true: someone can buy the name after it
was last seen free. For as long as api.mojang.com kept failing (429,
5xx, a timeout), a third-party player holding that name kept it on
every reconnect, and the Velocity registry, keyed on the name, turned
its new owner away as already connected. A hostile source could drive
the host into Mojang's rate limit on purpose to hold names that way.

A failed lookup now always counts as taken, so the player is renamed
with the source's prefix. An expired "taken" already gave that answer,
so only the stale "free" case changes. The cost is cosmetic: during an
outage an ordinary third-party player may get a prefix they do not
need, and their data follows the UUID, not the name.

A new test gives the cache a free entry past its TTL and has Mojang
answer 429. It fails on the old fallback. The two comments that
described the fallback now describe the fail-closed rule.
2026-09-22 14:17:53 +09:00
flyemoji c2a5645c55 fix: keep internal section numbers out of runtime messages
Four messages that reach an operator or an API client cited sections
of a specification nobody outside the project can read:

- the unimplemented archive store error from config load
- the running-server cap refusal, from both the user wake and the
  internal wake
- the missing memory ceiling guard, in the API and in felis apply

The references are gone and the wording is otherwise unchanged. Each
message still says what went wrong and, where there is one, what to
do about it. The test for the archive store message checks for the
tarLocal remediation, which is still there.
2026-09-22 13:44:38 +09:00
flyemoji 1ebd73a309 docs(nano): describe the hasjoined path as it works
The comments around hasJoined still described an authlib client that
is not in the path. Velocity reads -Dmojang.sessionserver and sends
the request itself, and it turns a 204 into its online-mode-only kick,
not authlib's "failed to verify username". The route comment in api.go
also offered "a thin login hook" as an alternative that does not
exist.

Other comments had drifted from the code:

- The [[auth_source]] doc said an empty list ships the multiplexer
  off. Mojang is always prepended, so an empty list means Mojang is
  the only source.
- The premium-name cache said Mojang does not recycle names. A name
  frees up when its owner renames away. The day-long "taken" TTL still
  holds, because a stale "taken" costs a third-party player only a
  prefix.
- The cache bound claimed entries come only from players who
  authenticated somewhere. Any third-party source that validates a
  login adds one, so a hostile source can force the map to clear. That
  costs repeat lookups, or a fail-closed prefix while Mojang is
  unreachable, never an identity.

The rewrite rationale now states what it costs a backend operator. A
chat-session key that a third-party source signed over its native UUID
cannot verify against the canonical UUID, so chat from those players
can only be accepted unsigned.

In the tests, comments that repeated their subtest names are gone.
2026-09-22 13:42:03 +09:00
flyemoji 30b4e1dfb2 test(nano): cover the bar-list error, bad identity id and ip relay
Three paths in handleHasJoined had no test that fails when they break:

- A bar-list lookup error answers 500. Logging it and carrying on would
  admit a reclaimed squatter during a database outage.
- An identity (Mojang) id that does not parse answers 204. Ignoring the
  parse error would emit the nil UUID for every such login, so they all
  share one player's data.
- The ip parameter is relayed to each source. Dropping it turns off the
  sources' check that the session is used from the player's own address.

One subtest each. Mutants that ignore the bar-list error, ignore the id
parse error, or stop appending ip each fail their subtest.
2026-09-22 13:35:50 +09:00
flyemoji 942e9a5ff8 test(nano): cover the premium-name cache rules
isPremiumName decides on every third-party login whether the player
keeps their name, and none of its rules had a test that fails when the
rule breaks: treating a 429 or 5xx from api.mojang.com as "free",
swapping the free and taken TTLs, flipping the freshness comparison,
answering "free" from an expired taken entry during an outage, or
dropping the clear-at-4096 bound. Each of those leaves a squatter
holding a name its owner has bought, or grows the cache without limit,
with CI green.

TestPremiumNameCache drives isPremiumName against a stub that answers
with a fixed status and counts lookups, and seeds cache entries at chosen
ages. Five mutants of handlers_hasjoined.go, one per rule above, each
fail at least one subtest. It does not test an expired "free" entry
during an outage; what that case should return is still open.
2026-09-22 13:34:43 +09:00
flyemoji 3f7274d29f test(nano): pin the auth namespace and one rewritten uuid as literals
The rewrite test computed its expected UUID from felisAuthNS itself, so
a change to the namespace seed moved both sides together and still
passed. Such a change gives every third-party player a new UUID on next
login, orphaning their playerdata and account links and letting any
squatter barred by the old UUID back in.

The test now also compares felisAuthNS and the rewrite of
littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e-
9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were
computed independently with Python's uuid5/uuid3, not read back from
the code. Prefixing the seed with https:// fails the test.
2026-09-22 13:32:45 +09:00
flyemoji a7fe525bfc test(api): keep the package's tests off the live mojang profile api
mojangProfileAPI defaults to https://api.mojang.com, and only the tests
that call stubMojangNames or setProfileAPI swap it out. A new test that
reaches a third-party login without doing so would query the real
service: its result then depends on network access and on whether
someone owns the name that day, and the shared premium cache can carry
that answer into later tests.

A TestMain now points the lookup at an address nothing listens on
before any test runs, so a forgotten stub always takes the same
fail-closed path. Tests that stub it restore this address, not the live
one, when they finish.
2026-09-22 13:31:53 +09:00
flyemoji 3338d6f0fe fix(nano): drop oversized hasJoined parameters before asking sources
username, serverId and ip were forwarded to every configured source at
whatever length the caller sent, up to the megabyte net/http allows in a
request line. Velocity never sends more than a 16-character name, a
41-character signed SHA-1 serverId and a textual IP address, so only a
direct caller reaches those sizes, and each such request cost one
oversized upstream call per source.

Any of the three over 64 bytes is now answered 204 before a source is
asked, the same as a missing username or serverId. 64 bytes still
leaves room for a 16-character name in multi-byte UTF-8.

The subtest behind this points a source that validates anything at the
handler and sends missing and oversized fields, expecting 204 and zero
upstream requests, then a well-formed login that gets 200. It replaces
the old missing-username case, whose only source was unreachable, so
the test passed even with the guard removed. Dropping the length check
now fails it on the long username; dropping the whole guard fails it on
the first missing field.
2026-09-22 13:21:43 +09:00
flyemoji a4779186a4 fix(nano): refuse hasJoined requests that declare a body
A GET to hasJoined with a Content-Length and no body held its
connection indefinitely. The handler returned, but net/http tries to
drain an unread body before it writes the answer, and nothing bounds
that wait: ReadHeaderTimeout ends with the headers. One such request
per socket pins a goroutine and a descriptor on nano or on felis-api's
internal face.

Velocity never sends a body, so any request that declares one, including
a chunked one, now gets a 400 with Connection: close, which skips the
drain and releases the connection once the answer is written.

The new subtest writes that request over a raw socket and waits three
seconds for an answer. Before the change it times out with no response
at all; now it reads a 400 marked close.
2026-09-22 13:19:55 +09:00
flyemoji ff81295aa9 fix(nano): report failing sources instead of treating them as a no
A source that timed out, answered 5xx or 429, redirected, or sent a 200
without a usable profile was skipped exactly like one that answered 204.
With nobody else validating, the login got a 204 and Velocity told the
player their account is offline-mode. Nothing was logged, so a dead or
mistyped source URL, or an http:// root that now redirects to https since
redirects stopped being followed, failed every one of its players with
no trace.

Each such failure now logs the source tag and the cause; for a 3xx it
names the Location to configure instead. When no source validates and at
least one failed, the answer is 503, which Velocity reports as the auth
servers being down and logs with the status. A source answering 204 is
still a plain no, and a validating source still wins regardless of
failures before it.

The new subtest puts a 503 source, a redirecting source and an
unreachable one each behind a Mojang that answers 204, and expects 503.
Against the previous handler every case returns 204.
2026-09-22 13:18:33 +09:00
flyemoji 1dd62a9bdc fix(nano): cap upstream response headers at 16 KiB
The hasJoined and name-lookup clients limited the body to 64 KiB but
left headers at the transport default of 1 MiB. A configured root could
answer with a megabyte of headers and stall the body, holding a few MiB
of heap per in-flight login for the full five seconds; enough parallel
logins take down the host, and every source's logins with it.

Both clients now share a transport with MaxResponseHeaderBytes set to
16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends
338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over
the cap fails the request and the resolver moves on to the next one.

The new subtest puts a source with 64 KiB of headers and a valid profile
ahead of an honest one and expects the honest player. Without the cap
the padded source wins.
2026-09-22 13:02:48 +09:00
flyemoji 2180e77cf5 chore: drop tool-name markers from source comments
Seventeen comments opened with a tag naming the tool that wrote them.
The tag goes and each comment keeps its reasoning, now starting as a
plain sentence. None of the reasoning changes.

The AGENTS.md note in .gitignore drops the story of how the file got
into the tree and keeps the one fact a reader needs: its advice to run
go fmt is destructive on this CRLF working tree.

Comments only; no code, build or test changes.
2026-09-22 12:57:19 +09:00
flyemoji 1976fca809 fix(nano): always relay properties as an array
sessionProfile tagged properties with omitempty, so an upstream answer
of "properties": [] (or null, or no key at all) reached Velocity with
no properties key. A Yggdrasil root may legitimately answer that way for
a player without a skin. Velocity 3.5.1's GameProfile deserializer
passes the missing key on as null and ImmutableList.copyOf throws, so
that player hangs at login with nothing logged, even though the same
answer sent straight to Velocity is accepted. Mojang always sends
textures, which is why the premium path and the hardware runs never hit
it.

Drop omitempty and replace a nil slice with an empty one before the
response is written. Removing omitempty alone is not enough: a nil
slice marshals as null, which Velocity rejects the same way.

The new subtest feeds the relay [], null and a missing key and expects
"properties":[] every time. The previous handler fails all three.
2026-09-22 12:52:04 +09:00
flyemoji 28d3638952 fix(nano): stop following redirects from upstream Yggdrasil roots
authHTTPClient kept net/http's default redirect policy, so a configured
third-party root that answered hasJoined with a 3xx made this host fetch
whatever URL it named, up to ten hops. That is a blind SSRF into
anything the host can reach, and it includes the multiplexer's own
listener: a root that redirects back to /session/minecraft/hasJoined
re-enters the handler, which queries Mojang and every source again and
gets redirected again, until the outer 5s client timeout fires. With a
50ms Mojang stub, one login produced 86 nested handler calls and 86
Mojang requests from this host's egress IP. The loopback default does
not help, because the redirect target is resolved from this host.

Return the 3xx as the response instead. resolveHasJoined already skips
any non-200 answer and closes its body, so a redirecting source is now
treated like one that is down, and the next source gets its turn. The
same probe now makes one handler call and one Mojang request.
Neither Mojang's nor LittleSkin's hasJoined redirects.

The new subtest puts a redirecting root ahead of an honest one and
checks that the redirect target is never contacted and the honest
source's player is returned. The pre-fix handler fails it.
2026-09-22 12:50:56 +09:00
flyemoji 8fe255e38f fix(api): relay Mojang logins when no auth source is configured
felis api wired the hasJoined multiplexer only when felis.toml had at
least one [[auth_source]]. With none, the source list stayed nil and
every hasJoined answer was a 204. That was harmless while nothing
pointed at the route, but the installer now starts Velocity with
-Dmojang.sessionserver aimed at felis-api unconditionally, and the
generated felis.toml tells the operator to delete the LittleSkin block
for a Mojang-only server. Doing exactly that turned every login away,
premium accounts included, and felis-api logged nothing about it.

Always build the list through authSourcesFromConfig, which prepends
Mojang in code, so an empty config is a Mojang-only relay. felis nano
already behaves this way with the same file.

The new test pins authSourcesFromConfig itself: Mojang first, the only
Identity source, and still present when nothing is configured. Marking
a configured source Identity makes it fail. The call site in cmdAPI is
now a single unconditional assignment and has no unit test of its own.
2026-09-22 12:46:00 +09:00
flyemoji 800a9042a1 test: use placeholder domains in setup and system-server tests
Three tests carried the maintainer's production root domain, a personal
mailbox and the public IP of a live demo host as fixture values. None of
them needs the value to be real: the re-domain test only needs two
different roots, and the setup flow only needs a well-formed address.

Swap them for the placeholders the rest of the suite already uses
(mc.example.net, [email protected]), and move the "before" root in the
re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737
documentation range, so the stale install the test models still has an
IP-derived hostname, which is the case the refresh exists for.
2026-09-22 12:43:54 +09:00
flyemoji a8c4077202 fix(api): log the panic value and stack behind the opaque 500
withRecover turned a panicking handler into a 500 envelope and threw the panic
away. The client is meant to get an opaque "internal error" — that part is
right — but nothing was written server-side, so a recovered panic was an
untraceable 500: an operator holding "internal error" has no message, no
stack and no request to grep for, and diagnosis degrades into guessing against
a live install. That is what it cost during the email-OTP report.

The panic value, a stack and the method+path are now logged first, keyed by
the same request_id writeError already stamps on unmapped errors, so the
client envelope and the server log can be joined. The test pins all three
markers plus the unchanged 500/"panic" response, because a silent recover
looks exactly like a working one from the outside.
2026-07-22 14:40:42 +09:00
flyemoji c4c964578e fix(setup): stop the forced-onboarding gate trapping players who have no email
setup_required is what the SPA polls to decide whether the onboarding wall is
still owed, and it disagreed with the middleware that actually enforces the
wall. requireOnboarded lifts on a verified email OR an enrolled passkey;
setup_required answered `u.Email == "" || !hasPasskey`. A console-tier player
joins through the bind-code door with no email at all — by design, there is no
SMTP at that point — so the email term never clears and the SPA keeps them on
the setup screen forever, even after they enroll the passkey that already
unlocked the API for them.

The predicate now lives in one place (setupRequired) and both endpoints call
it, so the next edit to the unlock condition cannot drift them apart again.
Keying it on EmailVerified rather than email presence is the deliberate part:
presence is exactly the term that trapped the no-email player, and it was also
wrong on its own terms — an unverified address is not an authentication
factor, so it was never what the lockdown could safely lift on.

Also lands the regression test for the mechanism behind the live claim-403
report: /me/servers answers 200 for a bind-onboarded player (which is why the
dashboard renders the 认领 button at all) while claim, wake and status all
answer 403 with code "setup_required" — i.e. the refusal comes from
requireOnboarded before the handler, not from isOwnerOrAdmin inside it, which
would have said "forbidden". Enrolling a passkey and changing nothing else
lifts all three, which isolates the gate as the sole cause. The backend authz
is correct; the button that leads a locked-down player into a 403 is the
frontend's to hide.
2026-07-22 14:40:28 +09:00
flyemoji 694e3cb800 feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.

The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.

Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.

RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.

Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.

That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.

Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.

Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.

Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
2026-07-21 00:12:37 +09:00
flyemoji b3989fa4af fix(mail): prove SMTP deliverability before saving, and stop losing the relay
A live install passed the SMTP setup screen and then failed every one-time
code with a bare `internal error`. Four separate defects had to line up for
that, and each is fixed here.

The relay was configured with `from = noreply@<domain-A>` on an account
authenticated as `<user>@<domain-B>`. Providers that validate sender identity
— Fastmail among them — answer MAIL FROM with an unconditional 250 and only
refuse at end-of-DATA. Ping stopped at NOOP, so it never saw the refusal: the
wizard reported success, wrote the config, rolled felis-api, and every OTP
afterwards died at w.Close().

Ping now runs the same transaction a real code takes — connect, (STARTTLS,)
AUTH, MAIL FROM, RCPT TO, DATA — delivering one self-test message to the From
address, and SendOTP and Ping share deliver() so the check cannot drift from
the thing it checks. The self-test recipient cannot cause a false negative:
an authenticated submission relay accepts RCPT for any destination by
definition, while the sender identity it does validate is exactly what we
want tested. The setup screen now says a message will be sent, names the
address it went to, and warns that From must be an address the account is
allowed to send as.

A relay refusal also answered 500 `internal`, which reads as a broken panel
and sends the operator hunting through handler code instead of their [smtp]
block. It is now 502 `mail_undeliverable`, mapped inside deliverOTP so all
four doors that mail a code (onboarding, email login, op-login, migrate
step-up) answer alike. The relay's own text stays out of the response — it
can name the SMTP account, and these routes are reachable by any signed-in
player — and goes to the log instead.

writeError logged nothing when it collapsed an unmapped error to 500, so an
operator holding an `internal error` had nothing to grep for and diagnosis
degraded into guessing against a live install. It now logs the method, path,
wrapped chain and the same request_id the caller is shown.

Finally, write_felis_toml regenerated the config wholesale and never emitted
[smtp], so re-running the installer — the documented way to update felis-api —
silently erased a working relay and reverted OTP delivery to the no-Mailer
path, logging codes instead of sending them. It now carries the block forward,
cached on first read because the host toml is clobbered before the pod toml is
written. Same defect family as the root_domain loss fixed in ecbeb20: a
generated file holding a hand-set value with no carry-forward.

Tests cover the case a MAIL FROM probe cannot see: a fake relay that answers
250 to MAIL FROM and 550 at end-of-DATA must fail both Ping and SendOTP, and
the 502 must carry a distinct machine code without leaking the relay's text.
2026-07-21 00:11:40 +09:00
flyemoji fe4c92c1c5 feat(files): add the server file editor
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.

felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.

What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.

Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.

The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.

Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:

  * A write accepts arbitrary bytes at any path, so an owner can place a
    loadable plugin jar. This is deliberate — it is what a hosting panel's
    file manager does, scoped to a server the caller already owns and
    already drives through /command — but it is the one owner-tier route
    that lands executable code in a backend pod, since images are
    admin-only and modpack submissions need an admin verdict.
  * config/paper-global.yml is refused on read. felis-lobby's entrypoint
    writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
    identical on every backend, so reading it from a server you own would
    hand you the handshake key for everyone else's. It is the only path in
    the mount that is not the caller's own data, and therefore the only
    denial. The comparison is on the cleaned path, or ./config/... would
    walk straight through it.

Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.

The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
2026-07-20 14:33:29 +09:00
flyemoji d26acc20ae feat(api): let in-game staff manage any server without claiming it
The web face has always granted staff the run of the fleet (isOwnerOrAdmin
passes an admin for stop/command/console/access on any node), but the
internal face explicitly had "no admin tier": a linked administrator in game
could only wake servers they owned or that autostartPolicy permitted. The
only way to manage another player's (or an unclaimed) server from inside the
game was to claim it — seizing ownership and burning the admin's own quota.

Give authorizeWakeByUUID the admin tier on the same trust anchor the op-login
approve already uses: verified online-mode UUID -> account link -> stored
role. A linked staff member now wakes ANY node under any policy (so
`/felis go` works fleet-wide without claiming); the owner bypass and the
policy gates are unchanged, and an unlinked UUID still fails safe.

Centralize the staff-role rule while at it: staffRole(role) in auth.go
(admin, plus owner as its superset) now backs Principal.IsAdmin, the session
ViaAdminAccess grading, the op-login approve gate and the new wake tier.
That also fixes a real hole in the approve gate, which required role=admin
exactly: an Owner manually promoted to role='owner' per migration 0011's
upgrade note would have been refused by their own in-game approval door.

The lobby menu still renders "Claim & Start" on ownerless tiles — claiming
becomes optional for staff rather than the only entry — so the velocity
plugin needs no change.
2026-07-20 11:10:23 +09:00
flyemoji f0b79e9edd feat(mail): deliver email one-time codes over SMTP and add the setup email screen
Felis never actually sent mail: OTP codes for onboarding, email login and
op-login were only written to the felis-api log behind a "demo has no SMTP"
limitation, and the Settings/SMTP flow those comments promised was never
built. Combined with the bootstrap Owner's address being recorded unverified
(87279a1), op-login start always took the anti-enumeration neutral branch and
minted a fake request_id, so the in-game approve inevitably answered "No
pending operator sign-in with that code".

Give the codes a real delivery path, configured in felis.toml rather than a
web settings page so config keeps a single source of truth:

- config: new [smtp] table (host, port defaulting to 587, from, username,
  password_ref). Validation requires a plausible from address and a sane
  port; the password itself never enters the config file.
- internal/mail (new): stdlib net/smtp mailer implementing the api.OTPMailer
  seam. Port 465 dials implicit TLS, other ports upgrade via STARTTLS when
  advertised; AUTH only when a username is configured (PlainAuth itself
  refuses plaintext, so the password cannot leak to a TLS-less relay).
  Ping() proves reachability and credentials without sending mail. The
  message shape (CRLF, Q-encoded bilingual subject) is pinned by test.
- platform: felis-smtp Secret constants and an optional FELIS_SMTP_PASSWORD
  env var on the felis-api Deployment, mirroring felis-uploads-s3.
- cmd/felis api: construct the real mailer when [smtp] is configured; keep
  the log fallback otherwise and say so at startup. Warn when a username is
  set but the credentials env is empty.
- setup TUI: "e" on the summary/status screen opens the email form (host,
  port, from, optional auth). Apply order: Ping preflight, [smtp] into both
  host and pod config files, felis-smtp Secret piped to kubectl via stdin,
  config Secret, felis-api rollout. A failed preflight leaves the install
  untouched. SMTP is deliberately not a wizard rail step: first-run stays
  mail-less by design, and the passkey minted at onboarding is the pre-SMTP
  owner credential.

Also make PGRepo.UserByEmail match case-insensitively (lower(email) =
lower($1)), honoring the interface contract and the users_verified_email_
unique partial index; the fake repo already matched with EqualFold.

Existing installs need the felis-api Deployment manifest re-applied (e.g. a
bootstrap re-run) before the new env var exists; a rollout restart alone
cannot add it.
2026-07-20 10:24:52 +09:00
flyemoji 7860152f57 feat(auth)!: go fully passwordless and fix cross-check review findings
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.

Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
  handlers; align tests (new account subtests, naming reserves
  "console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
  setup_panel.go, tui_root.go and tui_preflight.go instead of
  hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
  into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.

Panel:
- Delete the ChangePassword page and all password UI; align
  login/auth/api/types with the passwordless contract; add the
  migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
  ServerLuckPerms strings to translation keys; drop 72 orphan keys
  per locale; unify the title as "Felis - Console".

Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
  MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
  preference; unified /link lines in fabric/forge/neoforge; shared
  link-client javadoc contract fixes.

Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.

BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
2026-07-20 04:47:32 +09:00
flyemoji 87279a1366 fix(setup): record email unverified so onboarding works without SMTP
At bootstrap there is no SMTP, so the old /setup flow was unreachable: it
requested an emailed OTP that could never arrive. Setup now records the
Owner's email address unverified (no OTP round-trip) and requires a passkey,
deferring SMTP configuration to a later Settings page. Setup completes on
email-recorded + passkey-enrolled, and the lockdown lifts on the passkey, not
on email_verified: a passkey is the Owner's only pre-SMTP login credential
(email-OTP login refuses admin accounts).

The record-email endpoint (POST /account/email) now clears email_verified in
the same write. Only VerifyEmailOTP, which proves control of the address, may
set that flag; recording a fresh unproven address must never leave a stale
email_verified=true asserting a proof the user never gave. The change strictly
tightens the invariant, so no existing reader breaks.

Remove the dead ErrEmailTaken path and its documented 409: no migration puts a
unique index on users.email and the codebase does not enforce email
uniqueness, so the unique-violation branch was unreachable and the 409 an
impossible response.

The /setup route (Setup.tsx, setEmail helper, setup i18n copy) is rewritten to
match: record-email, mandatory passkey, no skip-for-now. The SMTP settings
page and post-setup configure-SMTP nudge are deferred.
2026-07-17 01:41:07 +09:00
flyemoji 60732a6283 feat(operator): gate op.console to staff and land owner setup there
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).

Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.

Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.

The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.

Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
2026-07-16 18:03:30 +09:00
flyemoji 9ea35304e3 fix(setup): source the console host from the panel hostname, not op.console
The owner setup URL and the limbo login link were built from the admin host
(op.console.<root>, with an op.console.localhost fallback) and a hardcoded
console.<root>, so an operator who set a custom panel_hostname got an unreachable setup
link and a wrong login target. Thread the resolved panel host (defaultPanelHostname)
through performSetupMCBind, the MC-bind TUI, and the login system-server env
(new FELIS_PANEL_HOSTNAME); the limbo plugin prefers it and keeps console.<root> only as
the fallback for an older operator whose env predates it. This also matters for security:
the only wired WebAuthn verifier is scoped to the panel host, so passkey enrollment must
land on the panel face, never op.console.

While here, the limbo login handler checks link status before minting a bind code: an
already-linked player is sent straight to the lobby instead of being shown a useless code.
2026-07-16 13:27:02 +09:00
flyemoji b5cd4501e5 fix(passkey): serve flat WebAuthn options to register and username login
go-webauthn marshals CredentialCreation/CredentialAssertion as {"publicKey": {...}},
but the panel's register (Account.tsx) and username-first login (Login.tsx) read the
options flat (options.challenge, options.user.id), so base64urlToBytes(undefined) threw
"Cannot read properties of undefined (reading 'replace')" and neither ceremony could
start. Strip the envelope in the register-begin and username-login-begin handlers via a
small unwrapPublicKey helper; discoverable login keeps the envelope because it reads
options.publicKey.* plus a top-level options.login_id. The begin tests now feed a wrapped
body and assert the handlers return it flat, so they genuinely exercise the unwrap.
2026-07-16 13:26:56 +09:00
flyemoji dab8fc214b feat(setup): bind owner through login gate 2026-07-14 02:57:15 +09:00
flyemoji fd062882ed feat(nano): give a Mojang player's name back to them, by prefixing the squatter
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.

The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.

A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.

The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.

prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.

Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.

Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:

  premium FLYEMOJ1     -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
  LittleSkin FLYEMOJ1  -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
  both online at once, zero "already connected" rejections
  LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3

The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.

Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.

BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
2026-07-13 13:00:54 +09:00
flyemoji d177428fb0 feat(felis): add felis nano — Yggdrasil hasJoined multiplexer without a control plane
`felis nano` serves the vanilla sessionserver protocol
(GET /session/minecraft/hasJoined) as a federating multiplexer over
Mojang plus any number of third-party Yggdrasil roots, with no k3s,
Postgres, or panel — a MultiLogin-style auth front-end delivered as a
subcommand of the single felis binary rather than a separate build.

- config.LoadNano reads only [[auth_source]] blocks; it skips the
  database.url / root_domain / archive requirements the full server
  needs. Zero sources is valid (Mojang-only).
- Mojang is prepended in code (Identity:true), never from config, so it
  is always the sole identity root. Third-party profiles are rewritten
  to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID).
- validateAuthSources rejects unknown keys, duplicate tags, and
  scheme-less URLs — a malformed nano config fails loud at load.
- Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend);
  a rejected login is a 204, matching the vanilla sessionserver.
- nano.go binds the -listen flag and ignores [server] listen in config.

Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime
smoke against the template config returns 204 on a miss and logs
"Mojang + 0 third-party source(s)"; a duplicate-tag config exits
non-zero citing "unique".
2026-07-13 02:39:41 +09:00
flyemoji ff550c41ef feat(nano): federating hasJoined multiplexer with per-source UUID namespacing 2026-07-12 02:09:47 +09:00
flyemoji 85b8a92a0e test(api): pin restore's owner gate against a superseded former owner
handleRestoreBackup's owner-or-admin gate was not pinned by any test: the former-owner gate backstopped every non-owner case the suite exercised, so a broken owner gate would not redden. Add the mirror of the former-owner test — a released former owner (still the backup's former_owner, no longer the current owner) must get 403 — the sole subtest that fails when the owner gate is disabled. Found by the round-2 backup/restore mutation audit; production code unchanged.
2026-07-08 05:58:00 +09:00
flyemoji fc748d3462 feat(breakglass): add "back up a world now" console peer (§B4 Sync)
Adds a break-glass console operation that snapshots a stopped world by
calling the felis-api internal face while the API is alive, rather than
rendering the backup Job locally: the Job needs felis-api deployment
coordinates the console does not hold.

The peer resolves the felis-api-internal ClusterIP Service + service
token from the control namespace, POSTs the internal backup endpoint
with the operator os_user for audit attribution, and maps 409/503/404
to friendly outcome cards. Core decision logic lives in backupnow.go
(unit-tested against a fake client + httptest); tui_backupnow.go is the
untested bubbletea glue mirroring tui_halt.go.
2026-07-07 19:32:13 +09:00
flyemoji f2fc57cad9 feat(api): add internal-face break-glass world backup endpoint (§B4 Sync)
Add POST /api/v1/internal/servers/{name}/backup so the on-node break-glass
console can snapshot a stopped world while felis-api is alive. It goes through
the API (not direct-to-CRD like halt) because rendering the backup Job needs
deployment coordinates (FELIS_IMAGE, FELIS_BACKUP_PVC) only felis-api holds.

Service-token auth (no Principal); the middleware IS the authorization, since
the operator already has root on the node. Refactor the RWO stopped-gate,
optional-Backuper 503, async hand-off and audit+202 into a shared enqueueBackup
tail so the external (owner/admin) and internal (break-glass) faces cannot
diverge on the security-critical stopped-gate. The internal audit is attributed
to break-glass/internal so a console-initiated backup is distinguishable from an
owner self-service one.
2026-07-07 10:48:01 +09:00
flyemoji 7a7c0d53ab feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
2026-07-07 10:04:30 +09:00
flyemoji fdb6efbd88 feat(account): migrate a live account's owned servers to a new account (§B3 inherit)
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.

- 0015 migration: account_migrations state machine (initiated -> confirmed ->
  code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem

Quota is grandfathered at redeem: no per-target quota re-check when servers move.
2026-07-05 20:50:48 +09:00
Lemon-miaow 7becb38488 fix(api): implement /readyz with real DB + K8s API + CRD checks (§7)
Previously /readyz only verified Repo != nil && Cluster != nil — a
process-liveness check, not a dependency-health check. The spec
requires the readyz probe to verify DB, K8s API, and CRD informer
are live before declaring the pod ready.

- Repo interface gains Ping(context.Context) error
- Cluster interface gains Ping(context.Context) error
- PGRepo.Ping delegates to sql.DB.PingContext
- K8sCluster.Ping lists MinecraftServer CRDs (Limit=1) in the
  configured namespace, exercising both the API and CRD informer
- handleReadyz iterates ping checks; any failure returns 503 with
  the failing dependency name in the error message
- fakeRepo and fakeCluster gain configurable pingErr for hermetic
  test coverage of the failure paths

New test: TestReadyzPingsDependencies verifies 200 when healthy,
503 when DB or K8s API is down.
2026-07-05 16:22:08 +08:00
flyemoji 9e1df12975 feat(passkey): advance sign_count, reject clone-warned assertions
Both login doors (username-first and discoverable) now run a shared applyAssertionCounter after a verified assertion. A signature-counter regression — go-webauthn's CloneWarning, the possible-cloned-authenticator signal — is refused fail-closed with the same opaque passkey_login_invalid envelope any other finish failure returns (no clone oracle to a prober) and audited distinctly as auth.passkey_clone_rejected under the resolved account. A clean assertion advances the stored sign_count to the asserted value and stamps last_used_at, before any session is minted.

Counter-less/synced authenticators report 0 and never warn, so they pass through and simply re-stamp 0; the check gates only counter-keeping hardware authenticators, where a rollback is the meaningful signal. Email-OTP and username-first passkey remain fallbacks, so a rejected clone is never bricked.

Adds Repo.AdvanceCredentialSignCount (pgrepo UPDATE by credential_id) and surfaces CloneWarning from the internal/passkey adapter's FinishLogin/FinishDiscoverableLogin. Proven by real-crypto adapter tests (a counter regression still verifies but flags CloneWarning), handler tests (advance-and-stamp on success, fail-closed on clone), and a symmetric test on each door so both call sites of the shared helper are covered.
2026-07-05 16:02:30 +09:00
flyemoji 154002edf3 docs(auth): cite MultiLogin reference for UUID-keyed reclaim split
Anchor the username-collision reclaim's UUID-keyed, proxy-detected design to
the multi-Yggdrasil reference: CaaMoe/MultiLogin v6 binds identity as
serviceId+online-UUID via "identity cards" that decouple the in-game name from
online identity — keyed by UUID, never by name. Note that §B3's Mojang-priority
reclaim goes beyond the common "protect the first-bound name" behavior by
evicting a squatter once the genuine Mojang owner appears and stashing the
squatter's data for the code-only inherit path.
2026-07-05 04:05:56 +09:00
flyemoji ec468baef9 feat(auth): add discoverable (usernameless) passkey login
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.

Routes (both Public):
  POST /api/v1/auth/passkey/login/discoverable/begin
  POST /api/v1/auth/passkey/login/discoverable/finish

Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.

The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.

Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
2026-07-05 04:05:56 +09:00