06d5e652c68b95c2c3ab053658117a73ad00c3b1
248
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c839454a1f |
fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from another namespace either (live drill: 'error looking up service account minecraft/felis-reaper: serviceaccount not found'). Move the SA and its RoleBinding subject to the Minecraft namespace alongside the CronJob. |
||
|
|
e4f2cff532 |
fix(manifests): render the retention reaper CronJob into the Minecraft namespace
A Pod can only mount PVCs from its own namespace; the CronJob referenced the minecraft-namespace backup PVC while being rendered under ControlNamespace, so it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim felis-backups not found'. The reaper Role/RoleBinding were already minecraft-scoped (the objects it touches live there), so the CronJob was the odd one out. The minecraft felis-config replica (felis setup, backup Job fix) supplies its config mount. |
||
|
|
0414913bc7 |
fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
6-way concurrent redeem of one code 500'd on users_username_key (each request generated a fresh user id but the same uuid-derived username), plus the rarer two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already locks its code row and handles the conflict. - SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted. - INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username: cross-code racers converge on the winner's row (role checked, staff still refused) instead of a unique-violation 500. - account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race. Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user; db clean; zero unmapped errors. |
||
|
|
dcc3b7403e |
fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake)
The interface doc promised 'each joined to its staff username', the fake and the pending handler both project username and created_at, but the PG query selected neither — live internal /op-login/pending returned username:"" and created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based tests can't see PG-only regressions. |
||
|
|
52549f7b3a |
fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
The PG implementation was a single UPDATE ... WHERE code_hash that returned
ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
the pre-session email-login door (and the op-login finish / migration confirm
doors) fell through to writeError's unmapped-error 500, and attempts were never
charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
intended contract; only the PG side had drifted.
Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
UPDATE the newest live row, expiry + attempt cap before the hash compare,
mismatch charges one attempt and returns ErrOTPInvalid without consuming,
match consumes and commits. Verified live on the VM: 5 wrong guesses return
400 and stop at attempts=5 (correct code then also refused, unconsumed);
fresh code redeems; replay returns 400.
|
||
|
|
9309ff5a7f |
chore: apply the missed S1016 conversions in handlers_users
The gofmt/staticcheck commit staged handlers_user.go (singular) for the formatting fix but missed this sibling for its two struct-literal-to- conversion cleanups. |
||
|
|
e690b058db |
fix(restore): wait for the tracking finalizer before recreating
Live verification of the previous commit showed the immediate retry STILL stranded: deleting a finished Job leaves it terminating (job-tracking finalizer), so the re-Create collided with the dying object and was mapped to ErrAlreadyExists a second time. Poll until the name actually frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck finalizer ever outlives the budget. Fake-client tests pin both the replace-finished and coalesce-in-flight branches. |
||
|
|
90ccbfede4 |
fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its deterministic Job name for the rest of the 10-minute TTL, so the next restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was treated as success unconditionally). K8sJobs now inspects the colliding Job: in-flight still coalesces, finished (succeeded or failed) is deleted and replaced. The minecraft-namespace Role gains jobs:get/delete for exactly that replacement. The same audit found the backup Job mounts the felis-config Secret but the installer only provisions it in the control namespace, so every backup Job stranded on FailedMount. felis setup now replicates it into the minecraft namespace beside the service-token and forwarding secrets. |
||
|
|
a05edc934c |
chore: gofmt the tree, clear staticcheck, add a CI gofmt gate
Nine files had drifted from gofmt and nothing checked; nine staticcheck findings were live (three dead symbols, capitalization, a redundant Sprintf, two literal-to-conversion sites, a nil test context). Fix all of them and make CI fail on unformatted Go so this cannot re-drift. |
||
|
|
99c31c1d4e |
fix(config): refuse plaintext auth-source urls to public hosts
An auth_source url could be http:// to any host. Anyone on the path to a public root, or anyone who can spoof its DNS name, can then answer hasJoined with a 200 and log in as any player of that source, including a third-party account linked to staff. The player's IP also travels in cleartext. Mojang logins are unaffected, since that source is built in over https. Config load now refuses http:// unless the host is localhost or a loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12, 192.168/16, fc00::/7), so a root on the same host or the LAN still works without TLS. The decision is made on the literal host because nothing is resolved at load time, so a LAN root named by hostname needs its IP address or https. The error says what to change. The new test covers public names and addresses, link-local, 0.0.0.0 and the first address past 172.16/12 (all refused over http, all accepted over https), and the loopback and private forms that stay allowed. It fails on the old check. |
||
|
|
e9f74f3f0f |
fix(nano): stop trusting an expired free name while mojang is failing
When the premium-name lookup failed, isPremiumName fell back to any cached answer, however old. An expired "free" is exactly the answer that may have stopped being true: someone can buy the name after it was last seen free. For as long as api.mojang.com kept failing (429, 5xx, a timeout), a third-party player holding that name kept it on every reconnect, and the Velocity registry, keyed on the name, turned its new owner away as already connected. A hostile source could drive the host into Mojang's rate limit on purpose to hold names that way. A failed lookup now always counts as taken, so the player is renamed with the source's prefix. An expired "taken" already gave that answer, so only the stale "free" case changes. The cost is cosmetic: during an outage an ordinary third-party player may get a prefix they do not need, and their data follows the UUID, not the name. A new test gives the cache a free entry past its TTL and has Mojang answer 429. It fails on the old fallback. The two comments that described the fallback now describe the fail-closed rule. |
||
|
|
c2a5645c55 |
fix: keep internal section numbers out of runtime messages
Four messages that reach an operator or an API client cited sections of a specification nobody outside the project can read: - the unimplemented archive store error from config load - the running-server cap refusal, from both the user wake and the internal wake - the missing memory ceiling guard, in the API and in felis apply The references are gone and the wording is otherwise unchanged. Each message still says what went wrong and, where there is one, what to do about it. The test for the archive store message checks for the tarLocal remediation, which is still there. |
||
|
|
1ebd73a309 |
docs(nano): describe the hasjoined path as it works
The comments around hasJoined still described an authlib client that is not in the path. Velocity reads -Dmojang.sessionserver and sends the request itself, and it turns a 204 into its online-mode-only kick, not authlib's "failed to verify username". The route comment in api.go also offered "a thin login hook" as an alternative that does not exist. Other comments had drifted from the code: - The [[auth_source]] doc said an empty list ships the multiplexer off. Mojang is always prepended, so an empty list means Mojang is the only source. - The premium-name cache said Mojang does not recycle names. A name frees up when its owner renames away. The day-long "taken" TTL still holds, because a stale "taken" costs a third-party player only a prefix. - The cache bound claimed entries come only from players who authenticated somewhere. Any third-party source that validates a login adds one, so a hostile source can force the map to clear. That costs repeat lookups, or a fail-closed prefix while Mojang is unreachable, never an identity. The rewrite rationale now states what it costs a backend operator. A chat-session key that a third-party source signed over its native UUID cannot verify against the canonical UUID, so chat from those players can only be accepted unsigned. In the tests, comments that repeated their subtest names are gone. |
||
|
|
e0ad78af98 |
test(config): make the identity-key test fail when the key is accepted
TestLoadRejectsAuthSourceIdentityKey is the guard against a config line identity = true making a third-party source's UUIDs trusted as-is. Its fixture had no prefix, so Load failed on the prefix rule and the test passed on that error. With the unknown-key check in decodeConfig disabled, the test still passed. The fixture now carries a valid prefix, the error must mention unknown keys and identity, and LoadNano is checked alongside Load. With the unknown-key check disabled, both loaders now fail the test; the old version of the test passes against the same change. |
||
|
|
30b4e1dfb2 |
test(nano): cover the bar-list error, bad identity id and ip relay
Three paths in handleHasJoined had no test that fails when they break: - A bar-list lookup error answers 500. Logging it and carrying on would admit a reclaimed squatter during a database outage. - An identity (Mojang) id that does not parse answers 204. Ignoring the parse error would emit the nil UUID for every such login, so they all share one player's data. - The ip parameter is relayed to each source. Dropping it turns off the sources' check that the session is used from the player's own address. One subtest each. Mutants that ignore the bar-list error, ignore the id parse error, or stop appending ip each fail their subtest. |
||
|
|
942e9a5ff8 |
test(nano): cover the premium-name cache rules
isPremiumName decides on every third-party login whether the player keeps their name, and none of its rules had a test that fails when the rule breaks: treating a 429 or 5xx from api.mojang.com as "free", swapping the free and taken TTLs, flipping the freshness comparison, answering "free" from an expired taken entry during an outage, or dropping the clear-at-4096 bound. Each of those leaves a squatter holding a name its owner has bought, or grows the cache without limit, with CI green. TestPremiumNameCache drives isPremiumName against a stub that answers with a fixed status and counts lookups, and seeds cache entries at chosen ages. Five mutants of handlers_hasjoined.go, one per rule above, each fail at least one subtest. It does not test an expired "free" entry during an outage; what that case should return is still open. |
||
|
|
3f7274d29f |
test(nano): pin the auth namespace and one rewritten uuid as literals
The rewrite test computed its expected UUID from felisAuthNS itself, so a change to the namespace seed moved both sides together and still passed. Such a change gives every third-party player a new UUID on next login, orphaning their playerdata and account links and letting any squatter barred by the old UUID back in. The test now also compares felisAuthNS and the rewrite of littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e- 9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were computed independently with Python's uuid5/uuid3, not read back from the code. Prefixing the seed with https:// fails the test. |
||
|
|
a7fe525bfc |
test(api): keep the package's tests off the live mojang profile api
mojangProfileAPI defaults to https://api.mojang.com, and only the tests that call stubMojangNames or setProfileAPI swap it out. A new test that reaches a third-party login without doing so would query the real service: its result then depends on network access and on whether someone owns the name that day, and the shared premium cache can carry that answer into later tests. A TestMain now points the lookup at an address nothing listens on before any test runs, so a forgotten stub always takes the same fail-closed path. Tests that stub it restore this address, not the live one, when they finish. |
||
|
|
8e9c8ca4e6 |
fix(nano): say that [server] listen is ignored instead of defaulting it
LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset, and a test pinned that value, but felis nano never reads it: it binds the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An operator moving nano off loopback by writing [server] listen in its config got connection refused from the proxy and no hint that the key did nothing. LoadNano no longer sets the default, and nano prints a line naming the ignored value and the address it actually binds whenever the key is set. It is a warning rather than a load error so a full felis.toml copied onto a nano host keeps starting. The assertion that pinned the unused default is removed along with it. The new test runs cmdNano against a config that sets [server] listen and one that does not, with an unbindable -listen so it returns after loading. The first must warn and the second must not; with the old default restored, the second prints a warning about 0.0.0.0:8080. |
||
|
|
1905cac950 |
fix(config): refuse auth-source tags padded with whitespace
A third-party player's UUID is hashed from the source tag byte for byte, so the tag is a permanent namespace: change it and every player of that source comes back as someone new, with their playerdata, permissions, account links and reclaim bans left behind. Nothing said so, and a tag with a stray leading or trailing space, which nobody can see in the file, loaded as a brand new namespace. Such a tag is now rejected at load, and the AuthSourceConfig doc states that the tag is permanent, case included. The charset stays otherwise open: tightening it would force existing installs to rename, which is the very thing that rekeys their players. The new test loads a tag with a trailing space, a leading space and a trailing tab through LoadNano; all three loaded before this change. |
||
|
|
72a2750461 |
fix(config): refuse mojang as an auth-source tag
Mojang is prepended in code as the first, identity source, and the config templates say not to list it. Nothing enforced that. A listed tag = "mojang" loaded, and nano's startup list printed it as if Mojang had been pointed at that url, while the real Mojang was still asked first. The listed entry was a separate third-party source: asked again on every login that got past Mojang, adding up to five seconds when its url was Mojang's own and it answered 204 each time. Any case of "mojang" is now rejected at load with a message saying Mojang is built in and must not be listed. The duplicate-tag check could not catch this because the built-in source never passes through it. The new test loads "mojang" and "Mojang" through LoadNano; both loaded before this change. |
||
|
|
2c74080b78 |
fix(config): reject auth-source urls the resolver cannot query
The url check only looked for an http:// or https:// prefix. Several
shapes passed it and then left the source dead at login time: no host
("https://"), a bad port, surrounding whitespace (sent as %20 and
answered 404), and any query or fragment. The resolver appends
"?username=…&serverId=…" to the url as a string, so an existing query
swallows those parameters and a fragment hides them from the request
entirely. Each loaded green, and every login from that source failed.
The url is now parsed and must be http or https with a host, no query,
no fragment and no surrounding whitespace. Load and LoadNano share the
check. The shipped LittleSkin default and plain http:// endpoints, such
as a same-host root on loopback, still load.
The new test feeds each rejected shape to LoadNano. Against the previous
prefix check, six of the seven load; only ftp:// was refused.
|
||
|
|
3338d6f0fe |
fix(nano): drop oversized hasJoined parameters before asking sources
username, serverId and ip were forwarded to every configured source at whatever length the caller sent, up to the megabyte net/http allows in a request line. Velocity never sends more than a 16-character name, a 41-character signed SHA-1 serverId and a textual IP address, so only a direct caller reaches those sizes, and each such request cost one oversized upstream call per source. Any of the three over 64 bytes is now answered 204 before a source is asked, the same as a missing username or serverId. 64 bytes still leaves room for a 16-character name in multi-byte UTF-8. The subtest behind this points a source that validates anything at the handler and sends missing and oversized fields, expecting 204 and zero upstream requests, then a well-formed login that gets 200. It replaces the old missing-username case, whose only source was unreachable, so the test passed even with the guard removed. Dropping the length check now fails it on the long username; dropping the whole guard fails it on the first missing field. |
||
|
|
a4779186a4 |
fix(nano): refuse hasJoined requests that declare a body
A GET to hasJoined with a Content-Length and no body held its connection indefinitely. The handler returned, but net/http tries to drain an unread body before it writes the answer, and nothing bounds that wait: ReadHeaderTimeout ends with the headers. One such request per socket pins a goroutine and a descriptor on nano or on felis-api's internal face. Velocity never sends a body, so any request that declares one, including a chunked one, now gets a 400 with Connection: close, which skips the drain and releases the connection once the answer is written. The new subtest writes that request over a raw socket and waits three seconds for an answer. Before the change it times out with no response at all; now it reads a 400 marked close. |
||
|
|
ff81295aa9 |
fix(nano): report failing sources instead of treating them as a no
A source that timed out, answered 5xx or 429, redirected, or sent a 200 without a usable profile was skipped exactly like one that answered 204. With nobody else validating, the login got a 204 and Velocity told the player their account is offline-mode. Nothing was logged, so a dead or mistyped source URL, or an http:// root that now redirects to https since redirects stopped being followed, failed every one of its players with no trace. Each such failure now logs the source tag and the cause; for a 3xx it names the Location to configure instead. When no source validates and at least one failed, the answer is 503, which Velocity reports as the auth servers being down and logs with the status. A source answering 204 is still a plain no, and a validating source still wins regardless of failures before it. The new subtest puts a 503 source, a redirecting source and an unreachable one each behind a Mojang that answers 204, and expects 503. Against the previous handler every case returns 204. |
||
|
|
1dd62a9bdc |
fix(nano): cap upstream response headers at 16 KiB
The hasJoined and name-lookup clients limited the body to 64 KiB but left headers at the transport default of 1 MiB. A configured root could answer with a megabyte of headers and stall the body, holding a few MiB of heap per in-flight login for the full five seconds; enough parallel logins take down the host, and every source's logins with it. Both clients now share a transport with MaxResponseHeaderBytes set to 16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends 338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over the cap fails the request and the resolver moves on to the next one. The new subtest puts a source with 64 KiB of headers and a valid profile ahead of an honest one and expects the honest player. Without the cap the padded source wins. |
||
|
|
2180e77cf5 |
chore: drop tool-name markers from source comments
Seventeen comments opened with a tag naming the tool that wrote them. The tag goes and each comment keeps its reasoning, now starting as a plain sentence. None of the reasoning changes. The AGENTS.md note in .gitignore drops the story of how the file got into the tree and keeps the one fact a reader needs: its advice to run go fmt is destructive on this CRLF working tree. Comments only; no code, build or test changes. |
||
|
|
4e98ae6e56 |
fix(config): reject auth-source tags that contain a colon
A third-party player's canonical UUID is UUIDv3 over tag+":"+nativeID, and the native id is whatever the source answers. Tags were only checked for being non-empty and unique, so both "guild" and "guild:eu" could be configured. The "guild" root could then answer hasJoined with id "eu:X" and receive exactly the UUID of "guild:eu"'s player X, along with their playerdata, permissions and account links. Real native ids are 32 hex digits, so only the shorter tag's source can do this, and only when the operator has configured such a pair; when they have, it is a full impersonation. Reject a ':' in a tag at load. With colon-free tags the join is unambiguous: two different (tag, id) pairs can no longer produce the same input, since equal inputs force equal tags and duplicate tags are already refused. The tag is deliberately not narrowed any further. It is a permanent UUID namespace, and forcing an operator to rename a tag that has no colon would move every one of its players to a new UUID. The hash input and the native id are left exactly as they were, so no existing player's UUID changes. Load and LoadNano share validateAuthSources; the new test runs both against the guild / guild:eu pair and fails on the previous config.go. |
||
|
|
1976fca809 |
fix(nano): always relay properties as an array
sessionProfile tagged properties with omitempty, so an upstream answer of "properties": [] (or null, or no key at all) reached Velocity with no properties key. A Yggdrasil root may legitimately answer that way for a player without a skin. Velocity 3.5.1's GameProfile deserializer passes the missing key on as null and ImmutableList.copyOf throws, so that player hangs at login with nothing logged, even though the same answer sent straight to Velocity is accepted. Mojang always sends textures, which is why the premium path and the hardware runs never hit it. Drop omitempty and replace a nil slice with an empty one before the response is written. Removing omitempty alone is not enough: a nil slice marshals as null, which Velocity rejects the same way. The new subtest feeds the relay [], null and a missing key and expects "properties":[] every time. The previous handler fails all three. |
||
|
|
28d3638952 |
fix(nano): stop following redirects from upstream Yggdrasil roots
authHTTPClient kept net/http's default redirect policy, so a configured third-party root that answered hasJoined with a 3xx made this host fetch whatever URL it named, up to ten hops. That is a blind SSRF into anything the host can reach, and it includes the multiplexer's own listener: a root that redirects back to /session/minecraft/hasJoined re-enters the handler, which queries Mojang and every source again and gets redirected again, until the outer 5s client timeout fires. With a 50ms Mojang stub, one login produced 86 nested handler calls and 86 Mojang requests from this host's egress IP. The loopback default does not help, because the redirect target is resolved from this host. Return the 3xx as the response instead. resolveHasJoined already skips any non-200 answer and closes its body, so a redirecting source is now treated like one that is down, and the next source gets its turn. The same probe now makes one handler call and one Mojang request. Neither Mojang's nor LittleSkin's hasJoined redirects. The new subtest puts a redirecting root ahead of an honest one and checks that the redirect target is never contacted and the honest source's player is returned. The pre-fix handler fails it. |
||
|
|
8fe255e38f |
fix(api): relay Mojang logins when no auth source is configured
felis api wired the hasJoined multiplexer only when felis.toml had at least one [[auth_source]]. With none, the source list stayed nil and every hasJoined answer was a 204. That was harmless while nothing pointed at the route, but the installer now starts Velocity with -Dmojang.sessionserver aimed at felis-api unconditionally, and the generated felis.toml tells the operator to delete the LittleSkin block for a Mojang-only server. Doing exactly that turned every login away, premium accounts included, and felis-api logged nothing about it. Always build the list through authSourcesFromConfig, which prepends Mojang in code, so an empty config is a Mojang-only relay. felis nano already behaves this way with the same file. The new test pins authSourcesFromConfig itself: Mojang first, the only Identity source, and still present when nothing is configured. Marking a configured source Identity makes it fail. The call site in cmdAPI is now a single unconditional assignment and has no unit test of its own. |
||
|
|
800a9042a1 |
test: use placeholder domains in setup and system-server tests
Three tests carried the maintainer's production root domain, a personal mailbox and the public IP of a live demo host as fixture values. None of them needs the value to be real: the re-domain test only needs two different roots, and the setup flow only needs a well-formed address. Swap them for the placeholders the rest of the suite already uses (mc.example.net, [email protected]), and move the "before" root in the re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737 documentation range, so the stale install the test models still has an IP-derived hostname, which is the case the refresh exists for. |
||
|
|
afdbfac7a8 |
docs: index the deferred integration seams and correct two stale markers
INTEGRATION-ONLY and KNOWN-LIMITATION are grep-able, but the grep answers the wrong question. Thirty-four Go sites share the two markers and they carry four different meanings: "declared, nothing implements it" reads exactly like "implemented, only its I/O is unreachable from here", and neither reads differently from a limitation that was accepted on purpose and is not coming back. docs/deferred-seams.md sorts them, following the bucketed shape internal/updater/doc.go already uses for its own package rather than starting a second convention. Sorting them turned up two markers that had outlived the condition they describe. config.go called the modpack upload transport a deferred integration after both backends had shipped -- LocalContextStore and S3ContextStore, selected in cmd/felis by the shape of user_uploads_context, with the uploads PVC mounted and the felis-uploads-s3 Secret rendered. What is still deferred is the far end: Kaniko reading that context from inside the build Pod. updater/doc.go listed the `felis update` CLI and the off-cluster Velocity jar read under REMAINING INTEGRATION. Both exist -- cmd/felis/update.go, and gatherer_host.go, which answers Velocity from the installed jar's manifest and felis-api from the running binary's build stamp. The two nil seams that bullet also names are real, but they belong to the in-cluster gatherer only, so the bullet now says which caller has what and which is still empty. The index also records the collision that makes a naive grep misleading: docs/troubleshooting.md uses [INTEGRATION-ONLY] for something else, defined in its own opening at :19 -- the symptom is produced by the kubelet, kaniko or a live handshake, so it cannot be reproduced from the repository. Those twelve marks say where a failure comes from, not that something is unbuilt, and are excluded. Both code changes are comments. Every file:line the index cites was checked against the line it points at. |
||
|
|
23792d6251 |
fix(crd): remove spec.storage.retainOnDelete rather than leave it inert
The field validated, shipped in the CRD, and reached no controller. The world PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate, StatefulSet deletion does not cascade to template PVCs, and no finalizer exists anywhere in the operator. So setting it true described what already happened, and setting it false did nothing at all. False is the worse half: it reads as a request to delete a world, and was silently ignored. This departs from spec v4.1 §5, which asks for "删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half was ever built. Restoring that line means adding a finalizer whose other listed duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs, so the only work it would newly do is delete worlds, on a path that does not pass the reaper's verified-backup check. The reaper is the one thing in the system allowed to destroy a world and it earns that by proving a backup first. A second door without that check is not an improvement. The spec is a frozen versioned document, so it is left alone and the departure is recorded in troubleshooting.md §13, beside the behaviour it explains. §12 loses its inert row and its opening sentence, which existed to introduce this one field: every field in that table is now read by a controller. Deployed installs need nothing. A CR still carrying retainOnDelete keeps working, because a v1 CRD prunes unknown keys on the next write and the behaviour the field claimed to control was never conditional. go build, go vet and go test ./... pass on Linux with zero failures; the CRD still parses and storage keeps size and storageClassName. |
||
|
|
5d4f3063a9 |
feat(operator): make any Paper image joinable behind the forwarding proxy
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed handshake does not degrade -- it rejects every login the proxy forwards. Until now the only backends that could verify it were the two images Felis builds itself (deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own entrypoints. An arbitrary Paper image a user brings does not, so it passed admission, started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the lobby image as a base for a user's own world (0018_recommended_images.sql), which was never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining player away, the exact opposite of a server you mean to stay on. The fix configures forwarding from OUTSIDE the image instead of requiring it inside. The operator now injects a root `felis init-forwarding` initContainer into every user server; it writes the proxies.velocity block into config/paper-global.yml and forces online-mode=false in server.properties on the /data PVC before the main container starts. The image needs no forwarding logic of its own, so the joinable set stops being "images that self-configure forwarding" and becomes every Paper-family image the platform runs. buildStatefulSet gates the injection on the ABSENCE of the system-role label: the Felis-built system servers already consume the secret in their entrypoints and the login gate is a limbo, not Paper. It is also gated on a non-empty felis image name -- the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it skips the injection rather than failing, because a cluster whose proxy is not in modern mode has nothing to configure. The init runs as root deliberately. The world volume's ownership comes from the storage provisioner and the main container runs as whatever UID its image declares, so root is the only UID that can reliably write these files; it then chmods them 0666/0777 so that non-root main container can rewrite them on boot. The privilege is bounded -- the init exits before the server container starts and the server container keeps its own UID. The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the init ever stops running as root. The writer merges rather than overwrites, both because Paper expands paper-global.yml to its full default tree on first boot and because the panel file editor may edit either file between boots. It sets proxies.velocity.* and the single online-mode key and leaves every other setting alone. It is a no-op on an empty secret, for the same reason the env var is optional: a proxy that is not in modern mode provisions no Secret, and wedging every server's init on a missing optional value would be worse than the status quo. felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and 0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the online-player list and permission commands work out of the box. 0018's row is left in place -- an admin who kept it can keep it; this only adds the better default beside it. Three fixes ride along, each of which the 1.8 path hit in practice. bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only builds its block-connection provider when Via's lowest supported protocol is below 1.13; under modern forwarding the Velocity injector reports 393, so init() returns early, blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining. Every call site is behind isServersideBlockConnections(), so switching it off skips all of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected. ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with this one key suffices -- Config#loadConfig parses the bundled default as the base map and merges the on-disk file over it, so every other option stays current across version bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this worked: init() gates on the protocol version too, and that half fails on its own, so the line is missing either way. The config value is the only evidence, which is what the test asserts. The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13 and has nowhere to go. Only the handshake address field survives Via, so the Felis fork forwards the named servers BungeeCord-style while every other backend keeps modern+secret untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer CRs is the upgrade path. deploy/lobby's set_prop escapes the value before substituting it. The RCON password is operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare `sed s|...|...|` and silently kills the key -- taking the console, the online-player list and permission commands with it. deploy/paper was written with the escaping, so the lobby gets the same rather than leaving the sibling caller broken. Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping. The new tests cover the initContainer's image, root UID, world mount and secret env; the merge preserving unrelated config trees; the properties upsert including the commented-key case; and the bootstrap script both writing the Via key and still calling the function that writes it. Not verified: the initContainer has never run in a real cluster, and the felis-paper image is code-only here as the other game-stack images are -- no Go CI builds them. The ViaVersion pin is the one piece with live evidence, and that evidence is what it was written from. Before it, a client was cut within a second of "logged in with entity id" on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on 2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes. Neither session says which client version it was. The proxy never logged a protocol number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16 players and below" warning that fired for that player on the lobby leg, and no lower -- Via floors every handshake to the proxy's 393, so anything from 47 up is admissible. Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the NPE proves is that the pin was load-bearing, not who was holding the mouse. That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8 behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the fault above; off, holds. No automated test in THIS repository exercises the leg. |
||
|
|
37ab87f07c |
docs(mail): say plainly what a green SMTP self-test does not prove
The Ping comment claimed the self-test could not produce a false negative and left the impression it therefore proved deliverability. It does not, and the distinction is the whole trap: a relay that gates sender identity at end-of-DATA gates it on the way OUT. Fastmail answers 250 for [email protected] addressed to the account's own mailbox and 551 5.7.1 "Not authorised to send from this header address" for that same From addressed to anyone else — and only the account's exact authorized identity passes the second one, another local-part on the same domain is refused too. So a green Ping means connect, TLS, AUTH and message shape are good, and nothing more; the operator still has to have authorized From as a sending identity with their provider, and the first real OTP is what proves they did. Addressing the self-test somewhere external would not fix it either — the only mailbox an operator can reliably check is usually inside the same account — so the honest move is to scope the claim rather than buy false confidence with a bigger probe. |
||
|
|
a8c4077202 |
fix(api): log the panic value and stack behind the opaque 500
withRecover turned a panicking handler into a 500 envelope and threw the panic away. The client is meant to get an opaque "internal error" — that part is right — but nothing was written server-side, so a recovered panic was an untraceable 500: an operator holding "internal error" has no message, no stack and no request to grep for, and diagnosis degrades into guessing against a live install. That is what it cost during the email-OTP report. The panic value, a stack and the method+path are now logged first, keyed by the same request_id writeError already stamps on unmapped errors, so the client envelope and the server log can be joined. The test pins all three markers plus the unchanged 500/"panic" response, because a silent recover looks exactly like a working one from the outside. |
||
|
|
c4c964578e |
fix(setup): stop the forced-onboarding gate trapping players who have no email
setup_required is what the SPA polls to decide whether the onboarding wall is still owed, and it disagreed with the middleware that actually enforces the wall. requireOnboarded lifts on a verified email OR an enrolled passkey; setup_required answered `u.Email == "" || !hasPasskey`. A console-tier player joins through the bind-code door with no email at all — by design, there is no SMTP at that point — so the email term never clears and the SPA keeps them on the setup screen forever, even after they enroll the passkey that already unlocked the API for them. The predicate now lives in one place (setupRequired) and both endpoints call it, so the next edit to the unlock condition cannot drift them apart again. Keying it on EmailVerified rather than email presence is the deliberate part: presence is exactly the term that trapped the no-email player, and it was also wrong on its own terms — an unverified address is not an authentication factor, so it was never what the lockdown could safely lift on. Also lands the regression test for the mechanism behind the live claim-403 report: /me/servers answers 200 for a bind-onboarded player (which is why the dashboard renders the 认领 button at all) while claim, wake and status all answer 403 with code "setup_required" — i.e. the refusal comes from requireOnboarded before the handler, not from isOwnerOrAdmin inside it, which would have said "forbidden". Enrolling a passkey and changing nothing else lifts all three, which isolates the gate as the sole cause. The backend authz is correct; the button that leads a locked-down player into a 403 is the frontend's to hide. |
||
|
|
694e3cb800 |
feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.
The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.
Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.
RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.
Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.
That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.
Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.
Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.
Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
|
||
|
|
b3989fa4af |
fix(mail): prove SMTP deliverability before saving, and stop losing the relay
A live install passed the SMTP setup screen and then failed every one-time
code with a bare `internal error`. Four separate defects had to line up for
that, and each is fixed here.
The relay was configured with `from = noreply@<domain-A>` on an account
authenticated as `<user>@<domain-B>`. Providers that validate sender identity
— Fastmail among them — answer MAIL FROM with an unconditional 250 and only
refuse at end-of-DATA. Ping stopped at NOOP, so it never saw the refusal: the
wizard reported success, wrote the config, rolled felis-api, and every OTP
afterwards died at w.Close().
Ping now runs the same transaction a real code takes — connect, (STARTTLS,)
AUTH, MAIL FROM, RCPT TO, DATA — delivering one self-test message to the From
address, and SendOTP and Ping share deliver() so the check cannot drift from
the thing it checks. The self-test recipient cannot cause a false negative:
an authenticated submission relay accepts RCPT for any destination by
definition, while the sender identity it does validate is exactly what we
want tested. The setup screen now says a message will be sent, names the
address it went to, and warns that From must be an address the account is
allowed to send as.
A relay refusal also answered 500 `internal`, which reads as a broken panel
and sends the operator hunting through handler code instead of their [smtp]
block. It is now 502 `mail_undeliverable`, mapped inside deliverOTP so all
four doors that mail a code (onboarding, email login, op-login, migrate
step-up) answer alike. The relay's own text stays out of the response — it
can name the SMTP account, and these routes are reachable by any signed-in
player — and goes to the log instead.
writeError logged nothing when it collapsed an unmapped error to 500, so an
operator holding an `internal error` had nothing to grep for and diagnosis
degraded into guessing against a live install. It now logs the method, path,
wrapped chain and the same request_id the caller is shown.
Finally, write_felis_toml regenerated the config wholesale and never emitted
[smtp], so re-running the installer — the documented way to update felis-api —
silently erased a working relay and reverted OTP delivery to the no-Mailer
path, logging codes instead of sending them. It now carries the block forward,
cached on first read because the host toml is clobbered before the pod toml is
written. Same defect family as the root_domain loss fixed in
|
||
|
|
646d514a65 |
test(updates): pin the ordering of a dev build's own version stamp
|
||
|
|
c4300cb005 |
feat(updater): authenticate GitHub polling and track the real release repo
felis-api's coord was the placeholder "felis/felis", which resolves against nothing on real GitHub. It is now MliroLirrorsIngenuity/Felis — the same slug deploy/bootstrap.sh clones from — so update reporting for the control plane itself is live rather than parked. That repo is private today, so the github source gained an optional token, read from FELIS_GITHUB_TOKEN: the variable bootstrap already needs, so an operator sets one value once. It comes from the environment and is never compiled in. A constant would be committed to the very repository it protects, ship inside every felis binary where strings(1) recovers it, reach every node the image is imported onto, and need a rebuild and a redeploy to rotate. Empty stays the correct posture for the other tracked components — k3s and cloudflared are public — and an empty token sends no Authorization header at all rather than an empty one. GitHub answers 404, not 401 or 403, for a private repo the caller cannot see, so "no token" and "no stable release published yet" arrive as the same status. On an unauthenticated 404 the error now names both causes and the variable that fixes the actionable one. With a token already set that hint would be wrong, so it is suppressed. Tests pin both halves: the Bearer header is sent only when the token is set, and the diagnostic names the variable only when it is not. doc.go's CAVEATS bullet still described the coord as a placeholder and the component as "dark at runtime". Both were true only until this change; it now records the real condition, which is that the component resolves like the others but needs a credential while the repo is private. |
||
|
|
659c8e5e9f |
feat(bootstrap): install the published release build instead of compiling on the host
deploy/bootstrap.sh now resolves the newest published GitHub release, downloads
the binary CI built for that tag, and builds a thin image around it. Compiling
on the target host becomes the fallback and the opt-in, not the default.
The panel is not a separate artifact. The Dockerfile copies panel/dist into
internal/panel/static before the go build, so the control plane — panel and
backend — ships as ONE file. The release channel therefore downloads exactly
one asset, felis-linux-<arch>, and needs no registry, no Go toolchain and no
checkout on the host.
The Minecraft game stack (limbo, lobby, the Velocity plugin) is still always
built locally. game_stack_source now keys on HAVE_PREBUILT_BINARY — the same
flag build_image uses — so on any prebuilt path it unpacks the tar embedded in
that binary instead of trusting a checkout an earlier install left behind.
Trusting the checkout would build the plugin from an old commit against a
freshly downloaded control plane: a silent version skew across the plugin/API
boundary.
Channels:
(default) newest published release, downloaded
FELIS_VERSION_BOOTSTRAP=dev clone main and compile
FELIS_REF=<ref> pins the tree, forces the source path
The download is best-effort. A tag whose assets are not uploaded yet, an
architecture with no published asset, or an asset that fails validation each
warn and fall back to compiling THE SAME TAG from source — never a different
commit.
The ref is resolved right after install_base, the first point curl exists and
well before docker and k3s, so a missing FELIS_GITHUB_TOKEN or an unpublished
release costs the operator seconds instead of a k3s install they then have to
unwind. It is skipped on exactly the paths that never consume the result: the
TUI, which rebuilds the binary it is already running, and FELIS_SKIP_FETCH,
which builds whatever is staged. Resolving anyway would set FELIS_VERSION to
the newest tag and stamp a staged tree as that release.
The asset is staged next to HOST_BIN rather than in TMPDIR. Validation EXECUTES
it, and /tmp is noexec on CIS-hardened images, where the exec dies 126, the
check reads it as a bad asset, and every such host silently falls back to the
full on-host compile this path exists to avoid. It also keeps a private-repo
artifact out of a world-readable 1777 directory.
git_auth, which supplies the token to git for a private-repo clone, passes an EMPTY
credential.helper before the inline one. credential.helper is multi-valued: a bare
`-c credential.helper=...` APPENDS to whatever the host has configured rather than
replacing it, and an empty value is git's documented list reset. Without it, on a host
with a persistent helper (Git for Windows ships `manager` at SYSTEM scope) two things
go wrong. Git runs `credential approve` automatically after a successful clone and
feeds every helper in the list, so a `store` helper writes the PAT to
~/.git-credentials in cleartext — the token outlives the install, in a file bootstrap
never created and never cleans up. And because the inline helper is LAST, a
pre-existing helper answers `fill` first, so a stale cached credential can win and the
clone authenticates as the wrong account — surfacing as exactly the 404-on-private-repo
the surrounding code works hard to explain. Reproduced both against a real clone, and
confirmed the reset closes both.
internal/panel parses the new stamp. The dev channel now emits "<tag>+g<sha>", which
matched neither describeSuffix ("-N-g<sha>") nor releaseTag, so a dev build fell through
to the default case and the version badge rendered the entire stamp as the release with
no commit. A devSuffix case handles it; the git-describe case stays for hand-rolled
`-ldflags "-X main.version=$(git describe)"` builds. Table test covers both forms plus
the release, dirty and unstamped cases.
CRD application no longer branches on the install path: it is always
`felis bootstrap-assets crd`. That output is byte-identical to deploy/crd/ —
bootstrap_asset.go embeds that very file — and needs no checkout, so one source
replaces a branch whose two arms had to be kept in agreement by hand.
Dockerfile gains a FELIS_VERSION build arg wired into -X main.version, declared
after `go mod download` so a version bump does not invalidate that layer. Both
build stages are pinned to $BUILDPLATFORM so a multi-platform buildx run never
emulates them: the panel's output is architecture-independent and the Go stage
cross-compiles via TARGETARCH. The final stage stays on the target platform and
is COPY-only, which BuildKit performs without QEMU.
.github/workflows/release.yml publishes on a vX.Y.Z tag: vet, tests, then one
buildx run producing both architectures through the repo Dockerfile. Not a bare
`go build` — internal/panel/static holds a tracked placeholder index.html so the
//go:embed compiles without node, which means a direct build succeeds and
quietly ships a release whose panel is that placeholder.
The stamp is asserted end to end, because it fails silently: an unstamped binary
reports "dev", which the updater refuses to compare, disabling update reporting
for every install built from that release. The arm64 artifact is checked by ELF
machine type rather than by running it — runners have binfmt registered, so
executing an amd64 binary misnamed arm64 would succeed.
Prerelease tags are flagged explicitly. The trigger glob is v*, gh does not read
semver out of a tag name, and an RC published as a full release becomes
/releases/latest — the single endpoint the default channel installs from and
`felis update` polls.
No SHA256SUMS. A checksum fetched over the same TLS session, with the same
credential, from the same host as the binary adds no trust root; signing is the
real answer and is a separate decision.
Not verified: the download -> validate -> image -> k3s path has never run on a
host against a real published release, because no tag exists yet. The shell
logic around it is verified out of tree; the network and exec behaviour is not.
|
||
|
|
fe4c92c1c5 |
feat(files): add the server file editor
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.
felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.
What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.
Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.
The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.
Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:
* A write accepts arbitrary bytes at any path, so an owner can place a
loadable plugin jar. This is deliberate — it is what a hosting panel's
file manager does, scoped to a server the caller already owns and
already drives through /command — but it is the one owner-tier route
that lands executable code in a backend pod, since images are
admin-only and modpack submissions need an admin verdict.
* config/paper-global.yml is refused on read. felis-lobby's entrypoint
writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
identical on every backend, so reading it from a server you own would
hand you the handshake key for everyone else's. It is the only path in
the mount that is not the caller's own data, and therefore the only
denial. The comparison is on the cleaned path, or ./config/... would
walk straight through it.
Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.
The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
|
||
|
|
05cb8f6320 |
feat(cli): report component updates and make the router a data table
Add `felis update`, which reports which platform components have newer versions available, and route `felis version`, which shipped implemented but unreachable. That bug is why the subcommand router is now a map rather than a switch. cmdVersion existed with nothing dispatching to it and no usage line, so `felis version` fell through to "unknown command" and no test noticed — a switch offers no way to enumerate what it routes, so the usage text and the router could not be compared. As data, they can: a test now walks the Commands: block and the table in both directions, failing an entry added to one without the other. bootstrap-assets stays deliberately undocumented and is listed as such, which makes its absence a decision rather than an oversight. The host gatherer answers the two seams NewSysGatherer leaves nil, for the one caller that can satisfy them without a cluster client. felis-api is answered from the running binary's own build stamp rather than the Deployment's image tag: deploy/bootstrap.sh builds the image from the same checkout it installs /usr/local/bin/felis from and stamps both with one git describe, so it is the same artifact, and it is the identity `felis version` reports. Reading the Deployment answers a slightly different question — what is rolled out — and stays the right seam for the in-cluster path. Velocity is read from the jar's own META-INF/MANIFEST.MF Implementation-Version, which is what the proxy reports about itself at runtime, because bootstrap installs the jar under a fixed name with no version in it. The filename extractor remains only as a fallback for a hand-placed velocity-3.5.1.jar. An unstamped local build reports "dev" and is refused with an actionable message rather than being treated as 0.0.0, which would make every release upstream look like an upgrade. The panel and the plugin jars have no version of their own on purpose: they are embedded in or built alongside the felis binary, so the felis version is theirs. |
||
|
|
f36d5b87f6 |
feat(images): mark platform-curated images and seed the lobby
The create-server form has no way to tell a user which of the whitelisted images is a sensible starting point. Add 'recommended' as a third image_whitelist.source alongside 'built' and 'external', and seed it with the one image that has earned it. The marker is presentation only. ImageAdmitted still turns solely on enabled, so a recommended row is admitted by exactly the rule that governs every other row and carries no extra privilege; a test pins both halves, because the failure modes are silent and opposite — make admission source-aware and the curated images vanish from the form, or let curation bypass the disable switch and an admin who pulled an image finds it still creatable. Only one image is seeded, and the restraint is the point. Velocity runs proxy-wide modern forwarding, so a backend that cannot verify the signed handshake rejects every login the proxy sends it. The operator injects FELIS_FORWARDING_SECRET into every backend but cannot make an image consume it. An arbitrary public Minecraft image therefore passes admission, builds, schedules, reports Ready — and then refuses every join, with nothing in the server's status explaining why. Exactly two images read that variable, deploy/limbo and deploy/lobby; limbo is the login gate and is nonsense as a base for a user's server, which leaves lobby. The list grows when Felis ships another forwarding-aware image, not before. AdmitBuiltImage now preserves a 'recommended' source through its ON CONFLICT path. Rebuilding a curated tag is the expected way to patch it, and that rebuild arrives through this exact path, so a blind SET source = 'built' would demote the curation on the first rebuild with nothing in the request saying so. AddExternalImage deliberately does not preserve it: an admin POSTing the ref is an explicit, named re-admission, and the 201 body reports the Image it constructed without re-reading the row, so a sticky source there would report a value the database does not hold. The migration is idempotent via ON CONFLICT DO NOTHING, so an admin who disabled or re-pointed the row does not have that decision undone on the next apply. |
||
|
|
d26acc20ae |
feat(api): let in-game staff manage any server without claiming it
The web face has always granted staff the run of the fleet (isOwnerOrAdmin passes an admin for stop/command/console/access on any node), but the internal face explicitly had "no admin tier": a linked administrator in game could only wake servers they owned or that autostartPolicy permitted. The only way to manage another player's (or an unclaimed) server from inside the game was to claim it — seizing ownership and burning the admin's own quota. Give authorizeWakeByUUID the admin tier on the same trust anchor the op-login approve already uses: verified online-mode UUID -> account link -> stored role. A linked staff member now wakes ANY node under any policy (so `/felis go` works fleet-wide without claiming); the owner bypass and the policy gates are unchanged, and an unlinked UUID still fails safe. Centralize the staff-role rule while at it: staffRole(role) in auth.go (admin, plus owner as its superset) now backs Principal.IsAdmin, the session ViaAdminAccess grading, the op-login approve gate and the new wake tier. That also fixes a real hole in the approve gate, which required role=admin exactly: an Owner manually promoted to role='owner' per migration 0011's upgrade note would have been refused by their own in-game approval door. The lobby menu still renders "Claim & Start" on ownerless tiles — claiming becomes optional for staff rather than the only entry — so the velocity plugin needs no change. |
||
|
|
f0b79e9edd |
feat(mail): deliver email one-time codes over SMTP and add the setup email screen
Felis never actually sent mail: OTP codes for onboarding, email login and
op-login were only written to the felis-api log behind a "demo has no SMTP"
limitation, and the Settings/SMTP flow those comments promised was never
built. Combined with the bootstrap Owner's address being recorded unverified
(
|
||
|
|
7860152f57 |
feat(auth)!: go fully passwordless and fix cross-check review findings
Remove password authentication everywhere; the only session doors are passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and op-login vouching. Remediates the 33-finding cross-check review across backend, CLI, panel, plugins, and docs. Backend/CLI: - Drop password routes and fields from account/user/onboard/auth handlers; align tests (new account subtests, naming reserves "console", op-login/onboard/qr-login test updates). - Add migrations 0016_op_login.sql and 0017_drop_password.sql. - Thread panel/admin hostnames from hostcfg through api.go, setup_panel.go, tui_root.go and tui_preflight.go instead of hardcoding; bootstrap.sh writes panel-hostname/admin-hostname into felis.toml. - Reword breakglass and TUI copy for passwordless flows. Panel: - Delete the ChangePassword page and all password UI; align login/auth/api/types with the passwordless contract; add the migration and op-login approval flows. - i18n: convert ImageBuildPage durations/status badges and ServerLuckPerms strings to translation keys; drop 72 orphan keys per locale; unify the title as "Felis - Console". Plugins (all six rebuilt): - Velocity waiting router returns 503 at_capacity during wake; MOTD/control-channel copy and config comments. - Paper zh menu title; Limbo bind-code TTL 600s with panel_url preference; unified /link lines in fabric/forge/neoforge; shared link-client javadoc contract fixes. Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and plugins/README.md aligned with the implementation. BREAKING CHANGE: migration 0017 irreversibly drops users.password_hash and users.must_change_password; password login cannot be restored after migrating. |
||
|
|
87279a1366 |
fix(setup): record email unverified so onboarding works without SMTP
At bootstrap there is no SMTP, so the old /setup flow was unreachable: it requested an emailed OTP that could never arrive. Setup now records the Owner's email address unverified (no OTP round-trip) and requires a passkey, deferring SMTP configuration to a later Settings page. Setup completes on email-recorded + passkey-enrolled, and the lockdown lifts on the passkey, not on email_verified: a passkey is the Owner's only pre-SMTP login credential (email-OTP login refuses admin accounts). The record-email endpoint (POST /account/email) now clears email_verified in the same write. Only VerifyEmailOTP, which proves control of the address, may set that flag; recording a fresh unproven address must never leave a stale email_verified=true asserting a proof the user never gave. The change strictly tightens the invariant, so no existing reader breaks. Remove the dead ErrEmailTaken path and its documented 409: no migration puts a unique index on users.email and the codebase does not enforce email uniqueness, so the unique-violation branch was unreachable and the 409 an impossible response. The /setup route (Setup.tsx, setEmail helper, setup i18n copy) is rewritten to match: record-email, mandatory passkey, no skip-for-now. The SMTP settings page and post-setup configure-SMTP nudge are deferred. |