Loading
Commits on Source 61
-
Minseong Choi authored
chore: work the tracker items that need no cluster
-
Minseong Choi authored
-
Minseong Choi authored
-
Minseong Choi authored
Three tests carried the maintainer's production root domain, a personal mailbox and the public IP of a live demo host as fixture values. None of them needs the value to be real: the re-domain test only needs two different roots, and the setup flow only needs a well-formed address. Swap them for the placeholders the rest of the suite already uses (mc.example.net, [email protected]), and move the "before" root in the re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737 documentation range, so the stale install the test models still has an IP-derived hostname, which is the case the refresh exists for.
-
Minseong Choi authored
felis api wired the hasJoined multiplexer only when felis.toml had at least one [[auth_source]]. With none, the source list stayed nil and every hasJoined answer was a 204. That was harmless while nothing pointed at the route, but the installer now starts Velocity with -Dmojang.sessionserver aimed at felis-api unconditionally, and the generated felis.toml tells the operator to delete the LittleSkin block for a Mojang-only server. Doing exactly that turned every login away, premium accounts included, and felis-api logged nothing about it. Always build the list through authSourcesFromConfig, which prepends Mojang in code, so an empty config is a Mojang-only relay. felis nano already behaves this way with the same file. The new test pins authSourcesFromConfig itself: Mojang first, the only Identity source, and still present when nothing is configured. Marking a configured source Identity makes it fail. The call site in cmdAPI is now a single unconditional assignment and has no unit test of its own.
-
Minseong Choi authored
write_felis_toml regenerates felis.host.toml and felis.pod.toml with a wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin block in that heredoc. Re-running the installer, which is also what `felis setup` does, threw away any edit to the list: a root the operator added stopped admitting logins, and a root they removed came back. The generated comment invited exactly that edit. Carry the tables forward the way [smtp] already is: read every [[auth_source]] table from the existing felis.host.toml (falling back to felis.pod.toml) and emit the LittleSkin default only when there is no earlier file at all. An earlier file with no tables stays empty, because that is a Mojang-only server rather than a missing value; felis-api now treats an empty list that way. The file header and the comment above the list now say what survives a re-run, and point at felis.host.toml, which is what the next run reads. bootstrap_test.sh extracts the new function from bootstrap.sh and checks the fresh-install default, an operator's own table carried without the default or the following section, an empty list staying empty, and the indented form the setup TUI writes. It passes under dash with gawk and with mawk; forcing the function to always return the default fails five of the new cases.
-
Minseong Choi authored
authHTTPClient kept net/http's default redirect policy, so a configured third-party root that answered hasJoined with a 3xx made this host fetch whatever URL it named, up to ten hops. That is a blind SSRF into anything the host can reach, and it includes the multiplexer's own listener: a root that redirects back to /session/minecraft/hasJoined re-enters the handler, which queries Mojang and every source again and gets redirected again, until the outer 5s client timeout fires. With a 50ms Mojang stub, one login produced 86 nested handler calls and 86 Mojang requests from this host's egress IP. The loopback default does not help, because the redirect target is resolved from this host. Return the 3xx as the response instead. resolveHasJoined already skips any non-200 answer and closes its body, so a redirecting source is now treated like one that is down, and the next source gets its turn. The same probe now makes one handler call and one Mojang request. Neither Mojang's nor LittleSkin's hasJoined redirects. The new subtest puts a redirecting root ahead of an honest one and checks that the redirect target is never contacted and the honest source's player is returned. The pre-fix handler fails it.
-
Minseong Choi authored
sessionProfile tagged properties with omitempty, so an upstream answer of "properties": [] (or null, or no key at all) reached Velocity with no properties key. A Yggdrasil root may legitimately answer that way for a player without a skin. Velocity 3.5.1's GameProfile deserializer passes the missing key on as null and ImmutableList.copyOf throws, so that player hangs at login with nothing logged, even though the same answer sent straight to Velocity is accepted. Mojang always sends textures, which is why the premium path and the hardware runs never hit it. Drop omitempty and replace a nil slice with an empty one before the response is written. Removing omitempty alone is not enough: a nil slice marshals as null, which Velocity rejects the same way. The new subtest feeds the relay [], null and a missing key and expects "properties":[] every time. The previous handler fails all three.
-
Minseong Choi authored
A third-party player's canonical UUID is UUIDv3 over tag+":"+nativeID, and the native id is whatever the source answers. Tags were only checked for being non-empty and unique, so both "guild" and "guild:eu" could be configured. The "guild" root could then answer hasJoined with id "eu:X" and receive exactly the UUID of "guild:eu"'s player X, along with their playerdata, permissions and account links. Real native ids are 32 hex digits, so only the shorter tag's source can do this, and only when the operator has configured such a pair; when they have, it is a full impersonation. Reject a ':' in a tag at load. With colon-free tags the join is unambiguous: two different (tag, id) pairs can no longer produce the same input, since equal inputs force equal tags and duplicate tags are already refused. The tag is deliberately not narrowed any further. It is a permanent UUID namespace, and forcing an operator to rename a tag that has no colon would move every one of its players to a new UUID. The hash input and the native id are left exactly as they were, so no existing player's UUID changes. Load and LoadNano share validateAuthSources; the new test runs both against the guild / guild:eu pair and fails on the previous config.go.
-
Minseong Choi authored
Seventeen comments opened with a tag naming the tool that wrote them. The tag goes and each comment keeps its reasoning, now starting as a plain sentence. None of the reasoning changes. The AGENTS.md note in .gitignore drops the story of how the file got into the tree and keeps the one fact a reader needs: its advice to run go fmt is destructive on this CRLF working tree. Comments only; no code, build or test changes.
-
Minseong Choi authored
The hasJoined and name-lookup clients limited the body to 64 KiB but left headers at the transport default of 1 MiB. A configured root could answer with a megabyte of headers and stall the body, holding a few MiB of heap per in-flight login for the full five seconds; enough parallel logins take down the host, and every source's logins with it. Both clients now share a transport with MaxResponseHeaderBytes set to 16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends 338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over the cap fails the request and the resolver moves on to the next one. The new subtest puts a source with 64 KiB of headers and a valid profile ahead of an honest one and expects the honest player. Without the cap the padded source wins.
-
Minseong Choi authored
felis-nano runs as a systemd DynamicUser, so it can read /etc/felis/felis.toml only if others may search /etc/felis. write_nano_config made the directory with a bare mkdir -p, which takes its mode from root's umask. On a host hardened to umask 027 that is 0750: nano exits on "permission denied", the unit restarts every five seconds, and no login gets through. A missing directory is now created 0755 explicitly. An existing one keeps its mode, because the full install sets it to 0700 to protect its secrets and widening that from the nano path would expose them. A nano unit locked out that way is left for the install to report. The harness runs the extracted function under umask 027 and checks both cases. Reverting to the bare mkdir fails the first; an unconditional chmod 0755 fails the second. The mode checks skip on filesystems that ignore chmod, such as Git Bash on NTFS.
-
Minseong Choi authored
install_nano_service printed "enabled and started" straight after systemctl restart, which returns as soon as the process is forked. An upgrade that keeps an old felis.toml the new binary rejects (an [[auth_source]] without a prefix, say) left the unit crash-looping in auto-restart while the installer reported success, and every login through the proxy failed. The install now waits two seconds and asks systemctl is-active. A unit that exited is in "activating (auto-restart)", which is-active does not count as active; on real systemd a unit whose process exits 1 under Restart=on-failure reads activating/auto-restart and is-active returns non-zero, while a running one reads active/running and returns 0. On failure the install prints the unit's last 20 journal lines and stops. This also surfaces a nano unit locked out of an existing 0700 /etc/felis. The harness runs the extracted function with systemctl stubbed both ways. Without the check, the dead-unit cases fail.
-
Minseong Choi authored
A source that timed out, answered 5xx or 429, redirected, or sent a 200 without a usable profile was skipped exactly like one that answered 204. With nobody else validating, the login got a 204 and Velocity told the player their account is offline-mode. Nothing was logged, so a dead or mistyped source URL, or an http:// root that now redirects to https since redirects stopped being followed, failed every one of its players with no trace. Each such failure now logs the source tag and the cause; for a 3xx it names the Location to configure instead. When no source validates and at least one failed, the answer is 503, which Velocity reports as the auth servers being down and logs with the status. A source answering 204 is still a plain no, and a validating source still wins regardless of failures before it. The new subtest puts a 503 source, a redirecting source and an unreachable one each behind a Mojang that answers 204, and expects 503. Against the previous handler every case returns 204.
-
Minseong Choi authored
A GET to hasJoined with a Content-Length and no body held its connection indefinitely. The handler returned, but net/http tries to drain an unread body before it writes the answer, and nothing bounds that wait: ReadHeaderTimeout ends with the headers. One such request per socket pins a goroutine and a descriptor on nano or on felis-api's internal face. Velocity never sends a body, so any request that declares one, including a chunked one, now gets a 400 with Connection: close, which skips the drain and releases the connection once the answer is written. The new subtest writes that request over a raw socket and waits three seconds for an answer. Before the change it times out with no response at all; now it reads a 400 marked close.
-
Minseong Choi authored
username, serverId and ip were forwarded to every configured source at whatever length the caller sent, up to the megabyte net/http allows in a request line. Velocity never sends more than a 16-character name, a 41-character signed SHA-1 serverId and a textual IP address, so only a direct caller reaches those sizes, and each such request cost one oversized upstream call per source. Any of the three over 64 bytes is now answered 204 before a source is asked, the same as a missing username or serverId. 64 bytes still leaves room for a 16-character name in multi-byte UTF-8. The subtest behind this points a source that validates anything at the handler and sends missing and oversized fields, expecting 204 and zero upstream requests, then a well-formed login that gets 200. It replaces the old missing-username case, whose only source was unreachable, so the test passed even with the guard removed. Dropping the length check now fails it on the long username; dropping the whole guard fails it on the first missing field.
-
Minseong Choi authored
The url check only looked for an http:// or https:// prefix. Several shapes passed it and then left the source dead at login time: no host ("https://"), a bad port, surrounding whitespace (sent as %20 and answered 404), and any query or fragment. The resolver appends "?username=…&serverId=…" to the url as a string, so an existing query swallows those parameters and a fragment hides them from the request entirely. Each loaded green, and every login from that source failed. The url is now parsed and must be http or https with a host, no query, no fragment and no surrounding whitespace. Load and LoadNano share the check. The shipped LittleSkin default and plain http:// endpoints, such as a same-host root on loopback, still load. The new test feeds each rejected shape to LoadNano. Against the previous prefix check, six of the seven load; only ftp:// was refused. -
Minseong Choi authored
Mojang is prepended in code as the first, identity source, and the config templates say not to list it. Nothing enforced that. A listed tag = "mojang" loaded, and nano's startup list printed it as if Mojang had been pointed at that url, while the real Mojang was still asked first. The listed entry was a separate third-party source: asked again on every login that got past Mojang, adding up to five seconds when its url was Mojang's own and it answered 204 each time. Any case of "mojang" is now rejected at load with a message saying Mojang is built in and must not be listed. The duplicate-tag check could not catch this because the built-in source never passes through it. The new test loads "mojang" and "Mojang" through LoadNano; both loaded before this change.
-
Minseong Choi authored
A third-party player's UUID is hashed from the source tag byte for byte, so the tag is a permanent namespace: change it and every player of that source comes back as someone new, with their playerdata, permissions, account links and reclaim bans left behind. Nothing said so, and a tag with a stray leading or trailing space, which nobody can see in the file, loaded as a brand new namespace. Such a tag is now rejected at load, and the AuthSourceConfig doc states that the tag is permanent, case included. The charset stays otherwise open: tightening it would force existing installs to rename, which is the very thing that rekeys their players. The new test loads a tag with a trailing space, a leading space and a trailing tab through LoadNano; all three loaded before this change.
-
Minseong Choi authored
felis nano logged every request with the raw RequestURI and %s. That text is the caller's: a right-to-left override reordered the line as displayed, an invalid UTF-8 byte made journald store the entry as a binary blob that journalctl -f shows as "[N blob data]", and a query near net/http's one-megabyte limit became a one-megabyte log line. The URI is now capped at 256 bytes, several times a real hasJoined query, and printed with %q, so control, bidi and invalid bytes appear escaped. The handler assembly moved into nanoHandler so the logged handler can be tested on its own; cmdNano serves it unchanged. The new test sends a query carrying U+202E, a 0x9b byte and 4 KiB of padding, and expects a valid UTF-8 line with the override escaped and no more than twice the cap. Restoring the old unquoted line fails it.
-
Minseong Choi authored
The installer and the config template tell the operator to run systemctl restart felis-nano after editing the source list. nano had no signal handling, so SIGTERM killed it mid-request: a login waiting on an upstream had its connection reset, and Velocity disconnected that player with "authentication servers are down". felis api already drains on shutdown; nano did not. nano now listens itself, serves until SIGINT or SIGTERM, then shuts the server down gracefully with a 30-second limit. That outlasts the source scan of any realistic list, at five seconds per source, and stays well inside systemd's default 90-second stop timeout. The new test holds a request inside the handler, cancels the serve context, and checks that serveNano is still running 200 ms later, that the held request then gets its answer, and that serveNano returns 0. Replacing the graceful shutdown with Close fails it.
-
Minseong Choi authored
LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset, and a test pinned that value, but felis nano never reads it: it binds the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An operator moving nano off loopback by writing [server] listen in its config got connection refused from the proxy and no hint that the key did nothing. LoadNano no longer sets the default, and nano prints a line naming the ignored value and the address it actually binds whenever the key is set. It is a warning rather than a load error so a full felis.toml copied onto a nano host keeps starting. The assertion that pinned the unused default is removed along with it. The new test runs cmdNano against a config that sets [server] listen and one that does not, with an unbindable -listen so it returns after loading. The first must warn and the second must not; with the old default restored, the second prints a warning about 0.0.0.0:8080.
-
Minseong Choi authored
mojangProfileAPI defaults to https://api.mojang.com, and only the tests that call stubMojangNames or setProfileAPI swap it out. A new test that reaches a third-party login without doing so would query the real service: its result then depends on network access and on whether someone owns the name that day, and the shared premium cache can carry that answer into later tests. A TestMain now points the lookup at an address nothing listens on before any test runs, so a forgotten stub always takes the same fail-closed path. Tests that stub it restore this address, not the live one, when they finish.
-
Minseong Choi authored
The rewrite test computed its expected UUID from felisAuthNS itself, so a change to the namespace seed moved both sides together and still passed. Such a change gives every third-party player a new UUID on next login, orphaning their playerdata and account links and letting any squatter barred by the old UUID back in. The test now also compares felisAuthNS and the rewrite of littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e- 9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were computed independently with Python's uuid5/uuid3, not read back from the code. Prefixing the seed with https:// fails the test.
-
Minseong Choi authored
isPremiumName decides on every third-party login whether the player keeps their name, and none of its rules had a test that fails when the rule breaks: treating a 429 or 5xx from api.mojang.com as "free", swapping the free and taken TTLs, flipping the freshness comparison, answering "free" from an expired taken entry during an outage, or dropping the clear-at-4096 bound. Each of those leaves a squatter holding a name its owner has bought, or grows the cache without limit, with CI green. TestPremiumNameCache drives isPremiumName against a stub that answers with a fixed status and counts lookups, and seeds cache entries at chosen ages. Five mutants of handlers_hasjoined.go, one per rule above, each fail at least one subtest. It does not test an expired "free" entry during an outage; what that case should return is still open.
-
Minseong Choi authored
Three paths in handleHasJoined had no test that fails when they break: - A bar-list lookup error answers 500. Logging it and carrying on would admit a reclaimed squatter during a database outage. - An identity (Mojang) id that does not parse answers 204. Ignoring the parse error would emit the nil UUID for every such login, so they all share one player's data. - The ip parameter is relayed to each source. Dropping it turns off the sources' check that the session is used from the player's own address. One subtest each. Mutants that ignore the bar-list error, ignore the id parse error, or stop appending ip each fail their subtest.
-
Minseong Choi authored
TestLoadRejectsAuthSourceIdentityKey is the guard against a config line identity = true making a third-party source's UUIDs trusted as-is. Its fixture had no prefix, so Load failed on the prefix rule and the test passed on that error. With the unknown-key check in decodeConfig disabled, the test still passed. The fixture now carries a valid prefix, the error must mention unknown keys and identity, and LoadNano is checked alongside Load. With the unknown-key check disabled, both loaders now fail the test; the old version of the test passes against the same change.
-
Minseong Choi authored
felis nano serves the same hasJoined handler as felis api, but behind nanoStubRepo, which implements only the bar-list lookup and embeds a nil Repo for everything else. Only the full-api path was tested, against a complete fake store, so a second store call added to handleHasJoined would pass CI and panic on every nano login. The loopback default of -listen, the one thing keeping nano from being an open auth relay, was not pinned either. The default moves into a nanoDefaultListen constant, and two tests cover the path. One serves a login through api.HasJoinedHandler with nanoStubRepo and a fake identity source and expects the profile back. The other requires the default to parse as a loopback IP. Taking the bar-list method off the stub makes the first panic on the nil Repo; defaulting to 0.0.0.0:8081 or :8081 fails the second.
-
Minseong Choi authored
The comments around hasJoined still described an authlib client that is not in the path. Velocity reads -Dmojang.sessionserver and sends the request itself, and it turns a 204 into its online-mode-only kick, not authlib's "failed to verify username". The route comment in api.go also offered "a thin login hook" as an alternative that does not exist. Other comments had drifted from the code: - The [[auth_source]] doc said an empty list ships the multiplexer off. Mojang is always prepended, so an empty list means Mojang is the only source. - The premium-name cache said Mojang does not recycle names. A name frees up when its owner renames away. The day-long "taken" TTL still holds, because a stale "taken" costs a third-party player only a prefix. - The cache bound claimed entries come only from players who authenticated somewhere. Any third-party source that validates a login adds one, so a hostile source can force the map to clear. That costs repeat lookups, or a fail-closed prefix while Mojang is unreachable, never an identity. The rewrite rationale now states what it costs a backend operator. A chat-session key that a third-party source signed over its native UUID cannot verify against the canonical UUID, so chat from those players can only be accepted unsigned. In the tests, comments that repeated their subtest names are gone.
-
Minseong Choi authored
The hasJoined contract listed only 200 and 204 and named authlib as the caller. The handler now answers four more ways, and a proxy operator reading the contract could not tell a refused login from a down source. - 204 also covers a missing or oversized parameter (no source is asked), a third-party name that is not a legal Minecraft username, and an identity id that does not parse. - 400 for a request that declares a body. There is no response body, and the connection is closed. - 500 when the bar-list lookup fails, with the usual error body. - 503 when no source validated and at least one failed, since that source's player may be the one logging in. The three query parameters now carry the 64-byte cap. The profile name says a third-party player holding a registered Mojang name gets it back prefixed and cut to 16 characters. The description names Velocity, drops the "thin login hook" that does not exist, and says that a non-200, non-204 answer makes Velocity report the auth servers as down.
-
Minseong Choi authored
Four messages that reach an operator or an API client cited sections of a specification nobody outside the project can read: - the unimplemented archive store error from config load - the running-server cap refusal, from both the user wake and the internal wake - the missing memory ceiling guard, in the API and in felis apply The references are gone and the wording is otherwise unchanged. Each message still says what went wrong and, where there is one, what to do about it. The test for the archive store message checks for the tarLocal remediation, which is still there.
-
Minseong Choi authored
A re-run copies the operator's [[auth_source]] tables from the existing felis toml into the new one. The awk program that finds them matched only the literal header [[auth_source]], so a table written as [[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid TOML, was taken for some other section and dropped from the config. Each section header now decides afresh whether it opens an auth_source table, through one regex that allows inner whitespace and a single- or double-quoted key. The single quote is spelled \047, which gawk and mawk both honour inside a bracket expression. The harness carries each spelling and checks that the table still stops at the next section.
-
Minseong Choi authored
Re-running the installer is how a nano host updates. That re-run reset FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost its endpoint and every login through it failed. It also offered the full control plane as the default, which on a nano host means k3s and Postgres nobody asked for. The listen address is now settled by resolve_nano_listen, the first step of main, so the later checks see the result. The operator's value wins, then the -listen argument of the installed felis-nano unit, then loopback. The install mode defaults to nano, at the prompt and without a terminal, when the felis-nano unit exists and the full install's bootstrap.done marker does not. Only the full install writes that marker. The harness reads back the unit it wrote earlier, and checks the mode default on a nano-only host, a host with the full install, and a fresh host.
-
Minseong Choi authored
prompt_install_mode also runs inside felis setup. Setup then goes on to the Owner and edge setup, which need the control plane, so choosing nano there always ended in a setup error. Under felis setup the mode is now full before any prompt or default is considered, and an explicit FELIS_INSTALL_MODE=nano stops with a message pointing at deploy/bootstrap.sh. That leaves the felis setup branch of acquire_nano_binary unreachable, so it goes. install_embedded_binary stays, since the full install still uses it.
-
Minseong Choi authored
FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in the firewall while nano bound nothing, a bare 127.0.0.1 printed http://127.0.0.1:127.0.0.1/... in the summary, and the unit crash-looped either way. validate_settings now requires a ':' and a decimal port of 1-65535 after the last one. It runs after resolve_nano_listen, so an address read back from an existing unit is checked too, and the default always passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
-
Minseong Choi authored
summary_nano printed the node's primary IP for every non-loopback bind and 127.0.0.1 for every loopback one. A bind to a second private address, or to [::1], handed the operator a hasJoined URL that nothing listens on. The host is now the part of FELIS_NANO_LISTEN before the last ':'. The node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]), which names no address a proxy could dial. The loopback and public-bind notes are unchanged.
-
Minseong Choi authored
install_go_toolchain downloaded the tarball to a fixed /tmp name and unpacked it into /usr/local as root, with no digest check. Another local user could plant that file first, and nothing would notice a tampered download. The tarball is now staged in a mktemp -d directory that the exit cleanup removes, and its sha256 must match before the old toolchain is touched, so a refusal leaves the host as it was. The default 1.26.4 carries pinned amd64 and arm64 digests next to its version; they are the ones https://go.dev/dl/?mode=json&include=all publishes. Any other FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in the header, because no pin can cover a version chosen at run time. Where and which version gets installed is unchanged.
-
Minseong Choi authored
The header said to export tunables before running, but its own `curl ... | sudo bash` entrypoint resets the environment, so an exported FELIS_INSTALL_MODE or FELIS_NANO_LISTEN never reached the installer. The header now shows the two forms that do arrive: the variable named on the sudo line, or export followed by sudo -E. The nano summary's hint for a proxy on another machine now prints a sudo line that can be pasted as is, instead of "re-run with FELIS_NANO_LISTEN=...". Comment and log text only.
-
Minseong Choi authored
prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which never fires on Linux: /dev/tty is mode 0666 whether or not the process has a controlling terminal, and only opening it fails. Without a terminal the menu was printed, the read failed with "No such device or address", and the default was taken by accident rather than by the documented path. The guard now opens /dev/tty in a subshell and takes the "no terminal for a prompt" path when that fails.
-
Minseong Choi authored
A source build against a private repository with no FELIS_GITHUB_TOKEN, or a wrong one, made git ask for a username on /dev/tty, and a piped install sat there waiting. git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git fails at once with "terminal prompts disabled". Both fetch_source failures name FELIS_GITHUB_TOKEN in their message: the fresh clone, and the fetch into an existing checkout, which had no message of its own before.
-
Minseong Choi authored
Both config templates the installer writes, the nano felis.toml and the comment above [[auth_source]] in the generated felis tomls, now state two things an operator editing the list needs to know. A tag is hashed verbatim into every player UUID of its source, with no case folding, so renaming it gives all of those players new UUIDs and orphans their data, links and bans. The list is scanned in order and the first source that validates wins, so order is trust, and a compromised root has to be removed, not moved down. Comment text only.
-
Minseong Choi authored
write_nano_config creates a missing /etc/felis as 0755, but it left an existing one alone. On a nano-only host an older installer made that directory with a bare mkdir -p, so under a root umask of 027 it is 0750. The DynamicUser unit cannot search it, so felis-nano cannot read its config, and a re-run stops at the service check instead of repairing the directory. An existing directory is now set to 0755 unless it holds the full install's secrets.env or bootstrap.done. The full install locks the directory to 0700 and writes secrets.env right after, so its directory keeps that mode, and install_nano_service still reports the lockout rather than this widening it. The mode cases run only where chmod works; on a filesystem that ignores it the harness skips them.
-
Minseong Choi authored
nano_listen_is_loopback decides whether configure_nano_firewall opens the port, and hasJoined takes no token. A default that does not classify as loopback would make every fresh nano host a public auth relay. The harness now runs the classifier on four loopback binds and three routable ones, and feeds it the default resolve_nano_listen applies on a first install, with no operator value and no existing unit. Setting that default to 0.0.0.0:8081 or :8081, or counting 0.0.0.0 as loopback, now fails the harness. Test only.
-
Minseong Choi authored
Two installer comments still said authlib makes the hasJoined request and sends no token. Velocity reads -Dmojang.sessionserver and sends the request itself. Comment text only.
-
Minseong Choi authored
validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081 passed. Go refuses that form ("too many colons in address") and needs [::1]:8081, so the unit crash-looped on every start. A host part that contains a colon must now be in brackets. With that, the bare ::1 pattern in nano_listen_is_loopback can no longer match an address that gets this far, so it goes. [::1] stays. The harness adds ::1:8081 to the refused addresses, and [::]:8081 and :8081, both of which Go binds, to the accepted ones. -
Minseong Choi authored
When the premium-name lookup failed, isPremiumName fell back to any cached answer, however old. An expired "free" is exactly the answer that may have stopped being true: someone can buy the name after it was last seen free. For as long as api.mojang.com kept failing (429, 5xx, a timeout), a third-party player holding that name kept it on every reconnect, and the Velocity registry, keyed on the name, turned its new owner away as already connected. A hostile source could drive the host into Mojang's rate limit on purpose to hold names that way. A failed lookup now always counts as taken, so the player is renamed with the source's prefix. An expired "taken" already gave that answer, so only the stale "free" case changes. The cost is cosmetic: during an outage an ordinary third-party player may get a prefix they do not need, and their data follows the UUID, not the name. A new test gives the cache a free entry past its TTL and has Mojang answer 429. It fails on the old fallback. The two comments that described the fallback now describe the fail-closed rule.
-
Minseong Choi authored
An auth_source url could be http:// to any host. Anyone on the path to a public root, or anyone who can spoof its DNS name, can then answer hasJoined with a 200 and log in as any player of that source, including a third-party account linked to staff. The player's IP also travels in cleartext. Mojang logins are unaffected, since that source is built in over https. Config load now refuses http:// unless the host is localhost or a loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12, 192.168/16, fc00::/7), so a root on the same host or the LAN still works without TLS. The decision is made on the literal host because nothing is resolved at load time, so a LAN root named by hostname needs its IP address or https. The error says what to change. The new test covers public names and addresses, link-local, 0.0.0.0 and the first address past 172.16/12 (all refused over http, all accepted over https), and the loopback and private forms that stay allowed. It fails on the old check.
-
Minseong Choi authored
For a non-loopback bind, configure_nano_firewall opened the nano port in firewalld to every source, while the summary told the operator to restrict it to the proxy. hasJoined takes no token, so on a public host that port is an auth relay anyone can point a proxy at, spending this host's Mojang egress until Mojang rate-limits it and the operator's own players stop getting in. A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets one rich rule that admits the port from that source only, ipv4 or ipv6 by the address given. Without it, no port is opened and the summary prints the rule to add. A re-run closes the port an earlier installer opened to every source. A rule for a previous FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand. Hosts without firewalld are handled as before. The value goes into the rule text, so it is checked up front for an address with one prefix length and nothing else. firewalld's own parser accepts both rule forms and refuses an ipv6 address under the ipv4 family. The harness covers the rule for each family, the closed-by-default case, the re-run cleanup, the loopback case and the CIDR check.
-
Minseong Choi authored
The source build of the nano binary installed Go at /usr/local/go and replaced whatever version was already there. On a host that also builds other things, the operator's own toolchain was removed and swapped for Felis's pinned version without a word. GOROOT_DIR is now /opt/felis/go, next to the source, the Velocity install and the JRE Felis already keeps under /opt/felis, and install_go_toolchain creates the parent before unpacking. A host where an earlier run put Go at /usr/local/go downloads it once more on the next re-run and keeps the old tree untouched; removing it is the operator's call. The harness now requires the toolchain directory to be under /opt/felis.
-
Minseong Choi authored
docs/changes held 26 per-feature change notes and their index, written while each feature was built. They were working records, not documentation: they cite internal milestone numbers and plan steps, several describe designs that changed before they shipped (the nano note's config schema and a proxy plugin that was never built), and nothing in the code, the build or the other docs refers to them. New notes stopped being added a while ago; the commit messages carry that record now. The directory leaves the tree in this commit. Its contents stay reachable in history, and the files were kept outside the repository before removal. No code, build or test changes.
-
Lemon-miaow authored
Nine files had drifted from gofmt and nothing checked; nine staticcheck findings were live (three dead symbols, capitalization, a redundant Sprintf, two literal-to-conversion sites, a nil test context). Fix all of them and make CI fail on unformatted Go so this cannot re-drift.
-
Lemon-miaow authored
govulncheck flagged pgx v5.7.1 (GO-2026-5004, SQL-injection class) as reachable from pgrepo.go, plus the old x/net and x/text. Bump all three to the fixed versions; go vet/test stay green.
-
Lemon-miaow authored
An E2E audit on a live install found that a FAILED restore held its deterministic Job name for the rest of the 10-minute TTL, so the next restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was treated as success unconditionally). K8sJobs now inspects the colliding Job: in-flight still coalesces, finished (succeeded or failed) is deleted and replaced. The minecraft-namespace Role gains jobs:get/delete for exactly that replacement. The same audit found the backup Job mounts the felis-config Secret but the installer only provisions it in the control namespace, so every backup Job stranded on FailedMount. felis setup now replicates it into the minecraft namespace beside the service-token and forwarding secrets.
-
Lemon-miaow authored
Live verification of the previous commit showed the immediate retry STILL stranded: deleting a finished Job leaves it terminating (job-tracking finalizer), so the re-Create collided with the dying object and was mapped to ErrAlreadyExists a second time. Poll until the name actually frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck finalizer ever outlives the budget. Fake-client tests pin both the replace-finished and coalesce-in-flight branches.
-
Lemon-miaow authored
The gofmt/staticcheck commit staged handlers_user.go (singular) for the formatting fix but missed this sibling for its two struct-literal-to- conversion cleanups.
-
Lemon-miaow authored
fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500 The PG implementation was a single UPDATE ... WHERE code_hash that returned ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on the pre-session email-login door (and the op-login finish / migration confirm doors) fell through to writeError's unmapped-error 500, and attempts were never charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the intended contract; only the PG side had drifted. Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR UPDATE the newest live row, expiry + attempt cap before the hash compare, mismatch charges one attempt and returns ErrOTPInvalid without consuming, match consumes and commits. Verified live on the VM: 5 wrong guesses return 400 and stop at attempts=5 (correct code then also refused, unconsumed); fresh code redeems; replay returns 400. -
Lemon-miaow authored
The interface doc promised 'each joined to its staff username', the fake and the pending handler both project username and created_at, but the PG query selected neither — live internal /op-login/pending returned username:"" and created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based tests can't see PG-only regressions.
-
Lemon-miaow authored
fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges 6-way concurrent redeem of one code 500'd on users_username_key (each request generated a fresh user id but the same uuid-derived username), plus the rarer two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already locks its code row and handles the conflict. - SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted. - INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username: cross-code racers converge on the winner's row (role checked, staff still refused) instead of a unique-violation 500. - account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race. Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user; db clean; zero unmapped errors.
-
Lemon-miaow authored
A Pod can only mount PVCs from its own namespace; the CronJob referenced the minecraft-namespace backup PVC while being rendered under ControlNamespace, so it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim felis-backups not found'. The reaper Role/RoleBinding were already minecraft-scoped (the objects it touches live there), so the CronJob was the odd one out. The minecraft felis-config replica (felis setup, backup Job fix) supplies its config mount.
-
Lemon-miaow authored
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from another namespace either (live drill: 'error looking up service account minecraft/felis-reaper: serviceaccount not found'). Move the SA and its RoleBinding subject to the Minecraft namespace alongside the CronJob.
-
Lemon-miaow authored