Commits on Source 61

  • Minseong Choi's avatar
    Merge pull request #19 from MliroLirrorsIngenuity/chore/issue-sweep · 587f1831
    Minseong Choi authored
    chore: work the tracker items that need no cluster
    587f1831
  • Minseong Choi's avatar
  • Minseong Choi's avatar
  • Minseong Choi's avatar
    test: use placeholder domains in setup and system-server tests · 800a9042
    Minseong Choi authored
    Three tests carried the maintainer's production root domain, a personal
    mailbox and the public IP of a live demo host as fixture values. None of
    them needs the value to be real: the re-domain test only needs two
    different roots, and the setup flow only needs a well-formed address.
    
    Swap them for the placeholders the rest of the suite already uses
    (mc.example.net, [email protected]), and move the "before" root in the
    re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737
    documentation range, so the stale install the test models still has an
    IP-derived hostname, which is the case the refresh exists for.
    800a9042
  • Minseong Choi's avatar
    fix(api): relay Mojang logins when no auth source is configured · 8fe255e3
    Minseong Choi authored
    felis api wired the hasJoined multiplexer only when felis.toml had at
    least one [[auth_source]]. With none, the source list stayed nil and
    every hasJoined answer was a 204. That was harmless while nothing
    pointed at the route, but the installer now starts Velocity with
    -Dmojang.sessionserver aimed at felis-api unconditionally, and the
    generated felis.toml tells the operator to delete the LittleSkin block
    for a Mojang-only server. Doing exactly that turned every login away,
    premium accounts included, and felis-api logged nothing about it.
    
    Always build the list through authSourcesFromConfig, which prepends
    Mojang in code, so an empty config is a Mojang-only relay. felis nano
    already behaves this way with the same file.
    
    The new test pins authSourcesFromConfig itself: Mojang first, the only
    Identity source, and still present when nothing is configured. Marking
    a configured source Identity makes it fail. The call site in cmdAPI is
    now a single unconditional assignment and has no unit test of its own.
    8fe255e3
  • Minseong Choi's avatar
    fix(bootstrap): keep the operator's auth sources across re-runs · 07bafebf
    Minseong Choi authored
    write_felis_toml regenerates felis.host.toml and felis.pod.toml with a
    wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin
    block in that heredoc. Re-running the installer, which is also what
    `felis setup` does, threw away any edit to the list: a root the operator
    added stopped admitting logins, and a root they removed came back. The
    generated comment invited exactly that edit.
    
    Carry the tables forward the way [smtp] already is: read every
    [[auth_source]] table from the existing felis.host.toml (falling back to
    felis.pod.toml) and emit the LittleSkin default only when there is no
    earlier file at all. An earlier file with no tables stays empty, because
    that is a Mojang-only server rather than a missing value; felis-api now
    treats an empty list that way.
    
    The file header and the comment above the list now say what survives a
    re-run, and point at felis.host.toml, which is what the next run reads.
    
    bootstrap_test.sh extracts the new function from bootstrap.sh and checks
    the fresh-install default, an operator's own table carried without the
    default or the following section, an empty list staying empty, and the
    indented form the setup TUI writes. It passes under dash with gawk and
    with mawk; forcing the function to always return the default fails five
    of the new cases.
    07bafebf
  • Minseong Choi's avatar
    fix(nano): stop following redirects from upstream Yggdrasil roots · 28d36389
    Minseong Choi authored
    authHTTPClient kept net/http's default redirect policy, so a configured
    third-party root that answered hasJoined with a 3xx made this host fetch
    whatever URL it named, up to ten hops. That is a blind SSRF into
    anything the host can reach, and it includes the multiplexer's own
    listener: a root that redirects back to /session/minecraft/hasJoined
    re-enters the handler, which queries Mojang and every source again and
    gets redirected again, until the outer 5s client timeout fires. With a
    50ms Mojang stub, one login produced 86 nested handler calls and 86
    Mojang requests from this host's egress IP. The loopback default does
    not help, because the redirect target is resolved from this host.
    
    Return the 3xx as the response instead. resolveHasJoined already skips
    any non-200 answer and closes its body, so a redirecting source is now
    treated like one that is down, and the next source gets its turn. The
    same probe now makes one handler call and one Mojang request.
    Neither Mojang's nor LittleSkin's hasJoined redirects.
    
    The new subtest puts a redirecting root ahead of an honest one and
    checks that the redirect target is never contacted and the honest
    source's player is returned. The pre-fix handler fails it.
    28d36389
  • Minseong Choi's avatar
    fix(nano): always relay properties as an array · 1976fca8
    Minseong Choi authored
    sessionProfile tagged properties with omitempty, so an upstream answer
    of "properties": [] (or null, or no key at all) reached Velocity with
    no properties key. A Yggdrasil root may legitimately answer that way for
    a player without a skin. Velocity 3.5.1's GameProfile deserializer
    passes the missing key on as null and ImmutableList.copyOf throws, so
    that player hangs at login with nothing logged, even though the same
    answer sent straight to Velocity is accepted. Mojang always sends
    textures, which is why the premium path and the hardware runs never hit
    it.
    
    Drop omitempty and replace a nil slice with an empty one before the
    response is written. Removing omitempty alone is not enough: a nil
    slice marshals as null, which Velocity rejects the same way.
    
    The new subtest feeds the relay [], null and a missing key and expects
    "properties":[] every time. The previous handler fails all three.
    1976fca8
  • Minseong Choi's avatar
    fix(config): reject auth-source tags that contain a colon · 4e98ae6e
    Minseong Choi authored
    A third-party player's canonical UUID is UUIDv3 over tag+":"+nativeID,
    and the native id is whatever the source answers. Tags were only
    checked for being non-empty and unique, so both "guild" and "guild:eu"
    could be configured. The "guild" root could then answer hasJoined with
    id "eu:X" and receive exactly the UUID of "guild:eu"'s player X, along
    with their playerdata, permissions and account links. Real native ids
    are 32 hex digits, so only the shorter tag's source can do this, and
    only when the operator has configured such a pair; when they have, it
    is a full impersonation.
    
    Reject a ':' in a tag at load. With colon-free tags the join is
    unambiguous: two different (tag, id) pairs can no longer produce the
    same input, since equal inputs force equal tags and duplicate tags are
    already refused. The tag is deliberately not narrowed any further.
    It is a permanent UUID namespace, and forcing an operator to rename a
    tag that has no colon would move every one of its players to a new
    UUID. The hash input and the native id are left exactly as they were,
    so no existing player's UUID changes.
    
    Load and LoadNano share validateAuthSources; the new test runs both
    against the guild / guild:eu pair and fails on the previous config.go.
    4e98ae6e
  • Minseong Choi's avatar
    chore: drop tool-name markers from source comments · 2180e77c
    Minseong Choi authored
    Seventeen comments opened with a tag naming the tool that wrote them.
    The tag goes and each comment keeps its reasoning, now starting as a
    plain sentence. None of the reasoning changes.
    
    The AGENTS.md note in .gitignore drops the story of how the file got
    into the tree and keeps the one fact a reader needs: its advice to run
    go fmt is destructive on this CRLF working tree.
    
    Comments only; no code, build or test changes.
    2180e77c
  • Minseong Choi's avatar
    fix(nano): cap upstream response headers at 16 KiB · 1dd62a9b
    Minseong Choi authored
    The hasJoined and name-lookup clients limited the body to 64 KiB but
    left headers at the transport default of 1 MiB. A configured root could
    answer with a megabyte of headers and stall the body, holding a few MiB
    of heap per in-flight login for the full five seconds; enough parallel
    logins take down the host, and every source's logins with it.
    
    Both clients now share a transport with MaxResponseHeaderBytes set to
    16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends
    338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over
    the cap fails the request and the resolver moves on to the next one.
    
    The new subtest puts a source with 64 KiB of headers and a valid profile
    ahead of an honest one and expects the honest player. Without the cap
    the padded source wins.
    1dd62a9b
  • Minseong Choi's avatar
    fix(bootstrap): create the nano config dir world-searchable · 26f685be
    Minseong Choi authored
    felis-nano runs as a systemd DynamicUser, so it can read
    /etc/felis/felis.toml only if others may search /etc/felis.
    write_nano_config made the directory with a bare mkdir -p, which takes
    its mode from root's umask. On a host hardened to umask 027 that is
    0750: nano exits on "permission denied", the unit restarts every five
    seconds, and no login gets through.
    
    A missing directory is now created 0755 explicitly. An existing one
    keeps its mode, because the full install sets it to 0700 to protect its
    secrets and widening that from the nano path would expose them. A nano
    unit locked out that way is left for the install to report.
    
    The harness runs the extracted function under umask 027 and checks both
    cases. Reverting to the bare mkdir fails the first; an unconditional
    chmod 0755 fails the second. The mode checks skip on filesystems that
    ignore chmod, such as Git Bash on NTFS.
    26f685be
  • Minseong Choi's avatar
    fix(bootstrap): fail the nano install when the unit does not stay up · a0f54df2
    Minseong Choi authored
    install_nano_service printed "enabled and started" straight after
    systemctl restart, which returns as soon as the process is forked. An
    upgrade that keeps an old felis.toml the new binary rejects (an
    [[auth_source]] without a prefix, say) left the unit crash-looping in
    auto-restart while the installer reported success, and every login
    through the proxy failed.
    
    The install now waits two seconds and asks systemctl is-active. A unit
    that exited is in "activating (auto-restart)", which is-active does not
    count as active; on real systemd a unit whose process exits 1 under
    Restart=on-failure reads activating/auto-restart and is-active returns
    non-zero, while a running one reads active/running and returns 0. On
    failure the install prints the unit's last 20 journal lines and stops.
    This also surfaces a nano unit locked out of an existing 0700 /etc/felis.
    
    The harness runs the extracted function with systemctl stubbed both
    ways. Without the check, the dead-unit cases fail.
    a0f54df2
  • Minseong Choi's avatar
    fix(nano): report failing sources instead of treating them as a no · ff81295a
    Minseong Choi authored
    A source that timed out, answered 5xx or 429, redirected, or sent a 200
    without a usable profile was skipped exactly like one that answered 204.
    With nobody else validating, the login got a 204 and Velocity told the
    player their account is offline-mode. Nothing was logged, so a dead or
    mistyped source URL, or an http:// root that now redirects to https since
    redirects stopped being followed, failed every one of its players with
    no trace.
    
    Each such failure now logs the source tag and the cause; for a 3xx it
    names the Location to configure instead. When no source validates and at
    least one failed, the answer is 503, which Velocity reports as the auth
    servers being down and logs with the status. A source answering 204 is
    still a plain no, and a validating source still wins regardless of
    failures before it.
    
    The new subtest puts a 503 source, a redirecting source and an
    unreachable one each behind a Mojang that answers 204, and expects 503.
    Against the previous handler every case returns 204.
    ff81295a
  • Minseong Choi's avatar
    fix(nano): refuse hasJoined requests that declare a body · a4779186
    Minseong Choi authored
    A GET to hasJoined with a Content-Length and no body held its
    connection indefinitely. The handler returned, but net/http tries to
    drain an unread body before it writes the answer, and nothing bounds
    that wait: ReadHeaderTimeout ends with the headers. One such request
    per socket pins a goroutine and a descriptor on nano or on felis-api's
    internal face.
    
    Velocity never sends a body, so any request that declares one, including
    a chunked one, now gets a 400 with Connection: close, which skips the
    drain and releases the connection once the answer is written.
    
    The new subtest writes that request over a raw socket and waits three
    seconds for an answer. Before the change it times out with no response
    at all; now it reads a 400 marked close.
    a4779186
  • Minseong Choi's avatar
    fix(nano): drop oversized hasJoined parameters before asking sources · 3338d6f0
    Minseong Choi authored
    username, serverId and ip were forwarded to every configured source at
    whatever length the caller sent, up to the megabyte net/http allows in a
    request line. Velocity never sends more than a 16-character name, a
    41-character signed SHA-1 serverId and a textual IP address, so only a
    direct caller reaches those sizes, and each such request cost one
    oversized upstream call per source.
    
    Any of the three over 64 bytes is now answered 204 before a source is
    asked, the same as a missing username or serverId. 64 bytes still
    leaves room for a 16-character name in multi-byte UTF-8.
    
    The subtest behind this points a source that validates anything at the
    handler and sends missing and oversized fields, expecting 204 and zero
    upstream requests, then a well-formed login that gets 200. It replaces
    the old missing-username case, whose only source was unreachable, so
    the test passed even with the guard removed. Dropping the length check
    now fails it on the long username; dropping the whole guard fails it on
    the first missing field.
    3338d6f0
  • Minseong Choi's avatar
    fix(config): reject auth-source urls the resolver cannot query · 2c74080b
    Minseong Choi authored
    The url check only looked for an http:// or https:// prefix. Several
    shapes passed it and then left the source dead at login time: no host
    ("https://"), a bad port, surrounding whitespace (sent as %20 and
    answered 404), and any query or fragment. The resolver appends
    "?username=…&serverId=…" to the url as a string, so an existing query
    swallows those parameters and a fragment hides them from the request
    entirely. Each loaded green, and every login from that source failed.
    
    The url is now parsed and must be http or https with a host, no query,
    no fragment and no surrounding whitespace. Load and LoadNano share the
    check. The shipped LittleSkin default and plain http:// endpoints, such
    as a same-host root on loopback, still load.
    
    The new test feeds each rejected shape to LoadNano. Against the previous
    prefix check, six of the seven load; only ftp:// was refused.
    2c74080b
  • Minseong Choi's avatar
    fix(config): refuse mojang as an auth-source tag · 72a27504
    Minseong Choi authored
    Mojang is prepended in code as the first, identity source, and the
    config templates say not to list it. Nothing enforced that. A listed
    tag = "mojang" loaded, and nano's startup list printed it as if Mojang
    had been pointed at that url, while the real Mojang was still asked
    first. The listed entry was a separate third-party source: asked again
    on every login that got past Mojang, adding up to five seconds when its
    url was Mojang's own and it answered 204 each time.
    
    Any case of "mojang" is now rejected at load with a message saying
    Mojang is built in and must not be listed. The duplicate-tag check could
    not catch this because the built-in source never passes through it.
    
    The new test loads "mojang" and "Mojang" through LoadNano; both loaded
    before this change.
    72a27504
  • Minseong Choi's avatar
    fix(config): refuse auth-source tags padded with whitespace · 1905cac9
    Minseong Choi authored
    A third-party player's UUID is hashed from the source tag byte for byte,
    so the tag is a permanent namespace: change it and every player of that
    source comes back as someone new, with their playerdata, permissions,
    account links and reclaim bans left behind. Nothing said so, and a tag
    with a stray leading or trailing space, which nobody can see in the
    file, loaded as a brand new namespace.
    
    Such a tag is now rejected at load, and the AuthSourceConfig doc states
    that the tag is permanent, case included. The charset stays otherwise
    open: tightening it would force existing installs to rename, which is
    the very thing that rekeys their players.
    
    The new test loads a tag with a trailing space, a leading space and a
    trailing tab through LoadNano; all three loaded before this change.
    1905cac9
  • Minseong Choi's avatar
    fix(nano): quote and cap the request log line · fa3eda52
    Minseong Choi authored
    felis nano logged every request with the raw RequestURI and %s. That
    text is the caller's: a right-to-left override reordered the line as
    displayed, an invalid UTF-8 byte made journald store the entry as a
    binary blob that journalctl -f shows as "[N blob data]", and a query
    near net/http's one-megabyte limit became a one-megabyte log line.
    
    The URI is now capped at 256 bytes, several times a real hasJoined
    query, and printed with %q, so control, bidi and invalid bytes appear
    escaped. The handler assembly moved into nanoHandler so the logged
    handler can be tested on its own; cmdNano serves it unchanged.
    
    The new test sends a query carrying U+202E, a 0x9b byte and 4 KiB of
    padding, and expects a valid UTF-8 line with the override escaped and
    no more than twice the cap. Restoring the old unquoted line fails it.
    fa3eda52
  • Minseong Choi's avatar
    fix(nano): drain in-flight logins on shutdown · 1d6c7300
    Minseong Choi authored
    The installer and the config template tell the operator to run
    systemctl restart felis-nano after editing the source list. nano had no
    signal handling, so SIGTERM killed it mid-request: a login waiting on an
    upstream had its connection reset, and Velocity disconnected that
    player with "authentication servers are down". felis api already drains
    on shutdown; nano did not.
    
    nano now listens itself, serves until SIGINT or SIGTERM, then shuts the
    server down gracefully with a 30-second limit. That outlasts the source
    scan of any realistic list, at five seconds per source, and stays well
    inside systemd's default 90-second stop timeout.
    
    The new test holds a request inside the handler, cancels the serve
    context, and checks that serveNano is still running 200 ms later, that
    the held request then gets its answer, and that serveNano returns 0.
    Replacing the graceful shutdown with Close fails it.
    1d6c7300
  • Minseong Choi's avatar
    fix(nano): say that [server] listen is ignored instead of defaulting it · 8e9c8ca4
    Minseong Choi authored
    LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset,
    and a test pinned that value, but felis nano never reads it: it binds
    the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An
    operator moving nano off loopback by writing [server] listen in its
    config got connection refused from the proxy and no hint that the key
    did nothing.
    
    LoadNano no longer sets the default, and nano prints a line naming the
    ignored value and the address it actually binds whenever the key is
    set. It is a warning rather than a load error so a full felis.toml
    copied onto a nano host keeps starting. The assertion that pinned the
    unused default is removed along with it.
    
    The new test runs cmdNano against a config that sets [server] listen
    and one that does not, with an unbindable -listen so it returns after
    loading. The first must warn and the second must not; with the old
    default restored, the second prints a warning about 0.0.0.0:8080.
    8e9c8ca4
  • Minseong Choi's avatar
    test(api): keep the package's tests off the live mojang profile api · a7fe525b
    Minseong Choi authored
    mojangProfileAPI defaults to https://api.mojang.com, and only the tests
    that call stubMojangNames or setProfileAPI swap it out. A new test that
    reaches a third-party login without doing so would query the real
    service: its result then depends on network access and on whether
    someone owns the name that day, and the shared premium cache can carry
    that answer into later tests.
    
    A TestMain now points the lookup at an address nothing listens on
    before any test runs, so a forgotten stub always takes the same
    fail-closed path. Tests that stub it restore this address, not the live
    one, when they finish.
    a7fe525b
  • Minseong Choi's avatar
    test(nano): pin the auth namespace and one rewritten uuid as literals · 3f7274d2
    Minseong Choi authored
    The rewrite test computed its expected UUID from felisAuthNS itself, so
    a change to the namespace seed moved both sides together and still
    passed. Such a change gives every third-party player a new UUID on next
    login, orphaning their playerdata and account links and letting any
    squatter barred by the old UUID back in.
    
    The test now also compares felisAuthNS and the rewrite of
    littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e-
    9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were
    computed independently with Python's uuid5/uuid3, not read back from
    the code. Prefixing the seed with https:// fails the test.
    3f7274d2
  • Minseong Choi's avatar
    test(nano): cover the premium-name cache rules · 942e9a5f
    Minseong Choi authored
    isPremiumName decides on every third-party login whether the player
    keeps their name, and none of its rules had a test that fails when the
    rule breaks: treating a 429 or 5xx from api.mojang.com as "free",
    swapping the free and taken TTLs, flipping the freshness comparison,
    answering "free" from an expired taken entry during an outage, or
    dropping the clear-at-4096 bound. Each of those leaves a squatter
    holding a name its owner has bought, or grows the cache without limit,
    with CI green.
    
    TestPremiumNameCache drives isPremiumName against a stub that answers
    with a fixed status and counts lookups, and seeds cache entries at chosen
    ages. Five mutants of handlers_hasjoined.go, one per rule above, each
    fail at least one subtest. It does not test an expired "free" entry
    during an outage; what that case should return is still open.
    942e9a5f
  • Minseong Choi's avatar
    test(nano): cover the bar-list error, bad identity id and ip relay · 30b4e1df
    Minseong Choi authored
    Three paths in handleHasJoined had no test that fails when they break:
    
    - A bar-list lookup error answers 500. Logging it and carrying on would
      admit a reclaimed squatter during a database outage.
    - An identity (Mojang) id that does not parse answers 204. Ignoring the
      parse error would emit the nil UUID for every such login, so they all
      share one player's data.
    - The ip parameter is relayed to each source. Dropping it turns off the
      sources' check that the session is used from the player's own address.
    
    One subtest each. Mutants that ignore the bar-list error, ignore the id
    parse error, or stop appending ip each fail their subtest.
    30b4e1df
  • Minseong Choi's avatar
    test(config): make the identity-key test fail when the key is accepted · e0ad78af
    Minseong Choi authored
    TestLoadRejectsAuthSourceIdentityKey is the guard against a config line
    identity = true making a third-party source's UUIDs trusted as-is. Its
    fixture had no prefix, so Load failed on the prefix rule and the test
    passed on that error. With the unknown-key check in decodeConfig
    disabled, the test still passed.
    
    The fixture now carries a valid prefix, the error must mention unknown
    keys and identity, and LoadNano is checked alongside Load. With the
    unknown-key check disabled, both loaders now fail the test; the old
    version of the test passes against the same change.
    e0ad78af
  • Minseong Choi's avatar
    test(nano): cover the nano delivery path and its loopback default · 59ec23d4
    Minseong Choi authored
    felis nano serves the same hasJoined handler as felis api, but behind
    nanoStubRepo, which implements only the bar-list lookup and embeds a nil
    Repo for everything else. Only the full-api path was tested, against a
    complete fake store, so a second store call added to handleHasJoined
    would pass CI and panic on every nano login. The loopback default of
    -listen, the one thing keeping nano from being an open auth relay, was
    not pinned either.
    
    The default moves into a nanoDefaultListen constant, and two tests
    cover the path. One serves a login through api.HasJoinedHandler with
    nanoStubRepo and a fake identity source and expects the profile back.
    The other requires the default to parse as a loopback IP. Taking the
    bar-list method off the stub makes the first panic on the nil Repo;
    defaulting to 0.0.0.0:8081 or :8081 fails the second.
    59ec23d4
  • Minseong Choi's avatar
    docs(nano): describe the hasjoined path as it works · 1ebd73a3
    Minseong Choi authored
    The comments around hasJoined still described an authlib client that
    is not in the path. Velocity reads -Dmojang.sessionserver and sends
    the request itself, and it turns a 204 into its online-mode-only kick,
    not authlib's "failed to verify username". The route comment in api.go
    also offered "a thin login hook" as an alternative that does not
    exist.
    
    Other comments had drifted from the code:
    
    - The [[auth_source]] doc said an empty list ships the multiplexer
      off. Mojang is always prepended, so an empty list means Mojang is
      the only source.
    - The premium-name cache said Mojang does not recycle names. A name
      frees up when its owner renames away. The day-long "taken" TTL still
      holds, because a stale "taken" costs a third-party player only a
      prefix.
    - The cache bound claimed entries come only from players who
      authenticated somewhere. Any third-party source that validates a
      login adds one, so a hostile source can force the map to clear. That
      costs repeat lookups, or a fail-closed prefix while Mojang is
      unreachable, never an identity.
    
    The rewrite rationale now states what it costs a backend operator. A
    chat-session key that a third-party source signed over its native UUID
    cannot verify against the canonical UUID, so chat from those players
    can only be accepted unsigned.
    
    In the tests, comments that repeated their subtest names are gone.
    1ebd73a3
  • Minseong Choi's avatar
    docs(openapi): list every answer hasjoined gives · 9ee8c48f
    Minseong Choi authored
    The hasJoined contract listed only 200 and 204 and named authlib as
    the caller. The handler now answers four more ways, and a proxy
    operator reading the contract could not tell a refused login from a
    down source.
    
    - 204 also covers a missing or oversized parameter (no source is
      asked), a third-party name that is not a legal Minecraft username,
      and an identity id that does not parse.
    - 400 for a request that declares a body. There is no response body,
      and the connection is closed.
    - 500 when the bar-list lookup fails, with the usual error body.
    - 503 when no source validated and at least one failed, since that
      source's player may be the one logging in.
    
    The three query parameters now carry the 64-byte cap. The profile name
    says a third-party player holding a registered Mojang name gets it
    back prefixed and cut to 16 characters. The description names
    Velocity, drops the "thin login hook" that does not exist, and says
    that a non-200, non-204 answer makes Velocity report the auth servers
    as down.
    9ee8c48f
  • Minseong Choi's avatar
    fix: keep internal section numbers out of runtime messages · c2a5645c
    Minseong Choi authored
    Four messages that reach an operator or an API client cited sections
    of a specification nobody outside the project can read:
    
    - the unimplemented archive store error from config load
    - the running-server cap refusal, from both the user wake and the
      internal wake
    - the missing memory ceiling guard, in the API and in felis apply
    
    The references are gone and the wording is otherwise unchanged. Each
    message still says what went wrong and, where there is one, what to
    do about it. The test for the archive store message checks for the
    tarLocal remediation, which is still there.
    c2a5645c
  • Minseong Choi's avatar
    fix(bootstrap): carry auth_source tables with spaced or quoted headers · 17b43964
    Minseong Choi authored
    A re-run copies the operator's [[auth_source]] tables from the existing
    felis toml into the new one. The awk program that finds them matched
    only the literal header [[auth_source]], so a table written as
    [[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid
    TOML, was taken for some other section and dropped from the config.
    
    Each section header now decides afresh whether it opens an auth_source
    table, through one regex that allows inner whitespace and a single- or
    double-quoted key. The single quote is spelled \047, which gawk and
    mawk both honour inside a bracket expression. The harness carries each
    spelling and checks that the table still stops at the next section.
    17b43964
  • Minseong Choi's avatar
    fix(bootstrap): keep a nano host's listen address and mode on re-run · 515c4a64
    Minseong Choi authored
    Re-running the installer is how a nano host updates. That re-run reset
    FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost
    its endpoint and every login through it failed. It also offered the
    full control plane as the default, which on a nano host means k3s and
    Postgres nobody asked for.
    
    The listen address is now settled by resolve_nano_listen, the first
    step of main, so the later checks see the result. The operator's value
    wins, then the -listen argument of the installed felis-nano unit, then
    loopback. The install mode defaults to nano, at the prompt and without
    a terminal, when the felis-nano unit exists and the full install's
    bootstrap.done marker does not. Only the full install writes that
    marker.
    
    The harness reads back the unit it wrote earlier, and checks the mode
    default on a nano-only host, a host with the full install, and a fresh
    host.
    515c4a64
  • Minseong Choi's avatar
    fix(bootstrap): install only the full control plane under felis setup · 3918a4b1
    Minseong Choi authored
    prompt_install_mode also runs inside felis setup. Setup then goes on
    to the Owner and edge setup, which need the control plane, so choosing
    nano there always ended in a setup error.
    
    Under felis setup the mode is now full before any prompt or default is
    considered, and an explicit FELIS_INSTALL_MODE=nano stops with a
    message pointing at deploy/bootstrap.sh. That leaves the felis setup
    branch of acquire_nano_binary unreachable, so it goes.
    install_embedded_binary stays, since the full install still uses it.
    3918a4b1
  • Minseong Choi's avatar
    fix(bootstrap): refuse a nano listen address without a usable port · 404d1172
    Minseong Choi authored
    FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in
    the firewall while nano bound nothing, a bare 127.0.0.1 printed
    http://127.0.0.1:127.0.0.1/... in the summary, and the unit
    crash-looped either way.
    
    validate_settings now requires a ':' and a decimal port of 1-65535
    after the last one. It runs after resolve_nano_listen, so an address
    read back from an existing unit is checked too, and the default always
    passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
    404d1172
  • Minseong Choi's avatar
    fix(bootstrap): print the address nano binds in the install summary · 3b0fc7a3
    Minseong Choi authored
    summary_nano printed the node's primary IP for every non-loopback bind
    and 127.0.0.1 for every loopback one. A bind to a second private
    address, or to [::1], handed the operator a hasJoined URL that nothing
    listens on.
    
    The host is now the part of FELIS_NANO_LISTEN before the last ':'. The
    node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]),
    which names no address a proxy could dial. The loopback and
    public-bind notes are unchanged.
    3b0fc7a3
  • Minseong Choi's avatar
    fix(bootstrap): verify the go toolchain tarball against a pinned digest · 02c079c8
    Minseong Choi authored
    install_go_toolchain downloaded the tarball to a fixed /tmp name and
    unpacked it into /usr/local as root, with no digest check. Another
    local user could plant that file first, and nothing would notice a
    tampered download.
    
    The tarball is now staged in a mktemp -d directory that the exit
    cleanup removes, and its sha256 must match before the old toolchain is
    touched, so a refusal leaves the host as it was. The default 1.26.4
    carries pinned amd64 and arm64 digests next to its version; they are
    the ones https://go.dev/dl/?mode=json&include=all publishes. Any other
    FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in
    the header, because no pin can cover a version chosen at run time.
    Where and which version gets installed is unchanged.
    02c079c8
  • Minseong Choi's avatar
    docs(bootstrap): pass tunables on the sudo line, not by export · 0758b9c5
    Minseong Choi authored
    The header said to export tunables before running, but its own
    `curl ... | sudo bash` entrypoint resets the environment, so an
    exported FELIS_INSTALL_MODE or FELIS_NANO_LISTEN never reached the
    installer. The header now shows the two forms that do arrive: the
    variable named on the sudo line, or export followed by sudo -E.
    
    The nano summary's hint for a proxy on another machine now prints a
    sudo line that can be pasted as is, instead of "re-run with
    FELIS_NANO_LISTEN=...". Comment and log text only.
    0758b9c5
  • Minseong Choi's avatar
    fix(bootstrap): detect a missing terminal by opening /dev/tty · 34f73ba1
    Minseong Choi authored
    prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which
    never fires on Linux: /dev/tty is mode 0666 whether or not the process
    has a controlling terminal, and only opening it fails. Without a
    terminal the menu was printed, the read failed with "No such device or
    address", and the default was taken by accident rather than by the
    documented path.
    
    The guard now opens /dev/tty in a subshell and takes the "no terminal
    for a prompt" path when that fails.
    34f73ba1
  • Minseong Choi's avatar
    fix(bootstrap): fail a tokenless private clone instead of prompting · 6794e66c
    Minseong Choi authored
    A source build against a private repository with no FELIS_GITHUB_TOKEN,
    or a wrong one, made git ask for a username on /dev/tty, and a piped
    install sat there waiting.
    
    git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git
    fails at once with "terminal prompts disabled". Both fetch_source
    failures name FELIS_GITHUB_TOKEN in their message: the fresh clone,
    and the fetch into an existing checkout, which had no message of its
    own before.
    6794e66c
  • Minseong Choi's avatar
    docs(bootstrap): say auth_source tags are permanent and order is trust · 2458ee17
    Minseong Choi authored
    Both config templates the installer writes, the nano felis.toml and
    the comment above [[auth_source]] in the generated felis tomls, now
    state two things an operator editing the list needs to know.
    
    A tag is hashed verbatim into every player UUID of its source, with no
    case folding, so renaming it gives all of those players new UUIDs and
    orphans their data, links and bans. The list is scanned in order and
    the first source that validates wins, so order is trust, and a
    compromised root has to be removed, not moved down. Comment text only.
    2458ee17
  • Minseong Choi's avatar
    fix(bootstrap): open up a nano-only config dir an older run left 0750 · cf65ffda
    Minseong Choi authored
    write_nano_config creates a missing /etc/felis as 0755, but it left
    an existing one alone. On a nano-only host an older installer made
    that directory with a bare mkdir -p, so under a root umask of 027 it
    is 0750. The DynamicUser unit cannot search it, so felis-nano cannot
    read its config, and a re-run stops at the service check instead of
    repairing the directory.
    
    An existing directory is now set to 0755 unless it holds the full
    install's secrets.env or bootstrap.done. The full install locks the
    directory to 0700 and writes secrets.env right after, so its directory
    keeps that mode, and install_nano_service still reports the lockout
    rather than this widening it. The mode cases run only where chmod
    works; on a filesystem that ignores it the harness skips them.
    cf65ffda
  • Minseong Choi's avatar
    test(bootstrap): pin the nano listen default to loopback · b58c2031
    Minseong Choi authored
    nano_listen_is_loopback decides whether configure_nano_firewall opens
    the port, and hasJoined takes no token. A default that does not
    classify as loopback would make every fresh nano host a public auth
    relay.
    
    The harness now runs the classifier on four loopback binds and three
    routable ones, and feeds it the default resolve_nano_listen applies on
    a first install, with no operator value and no existing unit. Setting
    that default to 0.0.0.0:8081 or :8081, or counting 0.0.0.0 as
    loopback, now fails the harness. Test only.
    b58c2031
  • Minseong Choi's avatar
    docs(bootstrap): credit velocity, not authlib, with the hasjoined call · 928a1fdf
    Minseong Choi authored
    Two installer comments still said authlib makes the hasJoined request
    and sends no token. Velocity reads -Dmojang.sessionserver and sends
    the request itself. Comment text only.
    928a1fdf
  • Minseong Choi's avatar
    fix(bootstrap): refuse an unbracketed ipv6 nano listen address · fa7b54f5
    Minseong Choi authored
    validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081
    passed. Go refuses that form ("too many colons in address") and needs
    [::1]:8081, so the unit crash-looped on every start. A host part that
    contains a colon must now be in brackets.
    
    With that, the bare ::1 pattern in nano_listen_is_loopback can no
    longer match an address that gets this far, so it goes. [::1] stays.
    The harness adds ::1:8081 to the refused addresses, and [::]:8081 and
    :8081, both of which Go binds, to the accepted ones.
    fa7b54f5
  • Minseong Choi's avatar
    fix(nano): stop trusting an expired free name while mojang is failing · e9f74f3f
    Minseong Choi authored
    When the premium-name lookup failed, isPremiumName fell back to any
    cached answer, however old. An expired "free" is exactly the answer
    that may have stopped being true: someone can buy the name after it
    was last seen free. For as long as api.mojang.com kept failing (429,
    5xx, a timeout), a third-party player holding that name kept it on
    every reconnect, and the Velocity registry, keyed on the name, turned
    its new owner away as already connected. A hostile source could drive
    the host into Mojang's rate limit on purpose to hold names that way.
    
    A failed lookup now always counts as taken, so the player is renamed
    with the source's prefix. An expired "taken" already gave that answer,
    so only the stale "free" case changes. The cost is cosmetic: during an
    outage an ordinary third-party player may get a prefix they do not
    need, and their data follows the UUID, not the name.
    
    A new test gives the cache a free entry past its TTL and has Mojang
    answer 429. It fails on the old fallback. The two comments that
    described the fallback now describe the fail-closed rule.
    e9f74f3f
  • Minseong Choi's avatar
    fix(config): refuse plaintext auth-source urls to public hosts · 99c31c1d
    Minseong Choi authored
    An auth_source url could be http:// to any host. Anyone on the path
    to a public root, or anyone who can spoof its DNS name, can then
    answer hasJoined with a 200 and log in as any player of that source,
    including a third-party account linked to staff. The player's IP also
    travels in cleartext. Mojang logins are unaffected, since that source
    is built in over https.
    
    Config load now refuses http:// unless the host is localhost or a
    loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12,
    192.168/16, fc00::/7), so a root on the same host or the LAN still
    works without TLS. The decision is made on the literal host because
    nothing is resolved at load time, so a LAN root named by hostname
    needs its IP address or https. The error says what to change.
    
    The new test covers public names and addresses, link-local, 0.0.0.0
    and the first address past 172.16/12 (all refused over http, all
    accepted over https), and the loopback and private forms that stay
    allowed. It fails on the old check.
    99c31c1d
  • Minseong Choi's avatar
    fix(bootstrap): open the nano port to the proxy alone · 0faec2b0
    Minseong Choi authored
    For a non-loopback bind, configure_nano_firewall opened the nano port
    in firewalld to every source, while the summary told the operator to
    restrict it to the proxy. hasJoined takes no token, so on a public
    host that port is an auth relay anyone can point a proxy at, spending
    this host's Mojang egress until Mojang rate-limits it and the
    operator's own players stop getting in.
    
    A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets
    one rich rule that admits the port from that source only, ipv4 or
    ipv6 by the address given. Without it, no port is opened and the
    summary prints the rule to add. A re-run closes the port an earlier
    installer opened to every source. A rule for a previous
    FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand.
    Hosts without firewalld are handled as before.
    
    The value goes into the rule text, so it is checked up front for an
    address with one prefix length and nothing else. firewalld's own
    parser accepts both rule forms and refuses an ipv6 address under the
    ipv4 family. The harness covers the rule for each family, the
    closed-by-default case, the re-run cleanup, the loopback case and the
    CIDR check.
    0faec2b0
  • Minseong Choi's avatar
    fix(bootstrap): keep the nano build toolchain under /opt/felis · 7b5b28c5
    Minseong Choi authored
    The source build of the nano binary installed Go at /usr/local/go and
    replaced whatever version was already there. On a host that also
    builds other things, the operator's own toolchain was removed and
    swapped for Felis's pinned version without a word.
    
    GOROOT_DIR is now /opt/felis/go, next to the source, the Velocity
    install and the JRE Felis already keeps under /opt/felis, and
    install_go_toolchain creates the parent before unpacking. A host where
    an earlier run put Go at /usr/local/go downloads it once more on the
    next re-run and keeps the old tree untouched; removing it is the
    operator's call. The harness now requires the toolchain directory to
    be under /opt/felis.
    7b5b28c5
  • Minseong Choi's avatar
    chore: stop tracking the docs/changes ledger · a56c3265
    Minseong Choi authored
    docs/changes held 26 per-feature change notes and their index,
    written while each feature was built. They were working records, not
    documentation: they cite internal milestone numbers and plan steps,
    several describe designs that changed before they shipped (the nano
    note's config schema and a proxy plugin that was never built), and
    nothing in the code, the build or the other docs refers to them. New
    notes stopped being added a while ago; the commit messages carry that
    record now.
    
    The directory leaves the tree in this commit. Its contents stay
    reachable in history, and the files were kept outside the repository
    before removal. No code, build or test changes.
    a56c3265
  • Lemon-miaow's avatar
    chore: gofmt the tree, clear staticcheck, add a CI gofmt gate · a05edc93
    Lemon-miaow authored
    Nine files had drifted from gofmt and nothing checked; nine staticcheck
    findings were live (three dead symbols, capitalization, a redundant
    Sprintf, two literal-to-conversion sites, a nil test context). Fix all
    of them and make CI fail on unformatted Go so this cannot re-drift.
    a05edc93
  • Lemon-miaow's avatar
    chore(deps): pgx v5.9.2, x/net v0.55.0, x/text v0.39.0 · fd0794d0
    Lemon-miaow authored
    govulncheck flagged pgx v5.7.1 (GO-2026-5004, SQL-injection class) as
    reachable from pgrepo.go, plus the old x/net and x/text. Bump all three
    to the fixed versions; go vet/test stay green.
    fd0794d0
  • Lemon-miaow's avatar
    fix(restore): replace a finished Job so retries enqueue; replicate felis-config · 90ccbfed
    Lemon-miaow authored
    An E2E audit on a live install found that a FAILED restore held its
    deterministic Job name for the rest of the 10-minute TTL, so the next
    restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
    treated as success unconditionally). K8sJobs now inspects the colliding
    Job: in-flight still coalesces, finished (succeeded or failed) is
    deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
    for exactly that replacement.
    
    The same audit found the backup Job mounts the felis-config Secret but
    the installer only provisions it in the control namespace, so every
    backup Job stranded on FailedMount. felis setup now replicates it into
    the minecraft namespace beside the service-token and forwarding
    secrets.
    90ccbfed
  • Lemon-miaow's avatar
    fix(restore): wait for the tracking finalizer before recreating · e690b058
    Lemon-miaow authored
    Live verification of the previous commit showed the immediate retry STILL
    stranded: deleting a finished Job leaves it terminating (job-tracking
    finalizer), so the re-Create collided with the dying object and was
    mapped to ErrAlreadyExists a second time. Poll until the name actually
    frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck
    finalizer ever outlives the budget. Fake-client tests pin both the
    replace-finished and coalesce-in-flight branches.
    e690b058
  • Lemon-miaow's avatar
    chore: apply the missed S1016 conversions in handlers_users · 9309ff5a
    Lemon-miaow authored
    The gofmt/staticcheck commit staged handlers_user.go (singular) for the
    formatting fix but missed this sibling for its two struct-literal-to-
    conversion cleanups.
    9309ff5a
  • Lemon-miaow's avatar
    fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are... · 52549f7b
    Lemon-miaow authored
    fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
    
    The PG implementation was a single UPDATE ... WHERE code_hash that returned
    ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
    the pre-session email-login door (and the op-login finish / migration confirm
    doors) fell through to writeError's unmapped-error 500, and attempts were never
    charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
    Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
    intended contract; only the PG side had drifted.
    
    Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
    UPDATE the newest live row, expiry + attempt cap before the hash compare,
    mismatch charges one attempt and returns ErrOTPInvalid without consuming,
    match consumes and commits. Verified live on the VM: 5 wrong guesses return
    400 and stop at attempts=5 (correct code then also refused, unconsumed);
    fresh code redeems; replay returns 400.
    52549f7b
  • Lemon-miaow's avatar
    fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake) · dcc3b740
    Lemon-miaow authored
    The interface doc promised 'each joined to its staff username', the fake and
    the pending handler both project username and created_at, but the PG query
    selected neither — live internal /op-login/pending returned username:"" and
    created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
    tests can't see PG-only regressions.
    dcc3b740
  • Lemon-miaow's avatar
    fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean... · 0414913b
    Lemon-miaow authored
    fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
    
    6-way concurrent redeem of one code 500'd on users_username_key (each request
    generated a fresh user id but the same uuid-derived username), plus the rarer
    two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
    locks its code row and handles the conflict.
    
    - SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
      as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
    - INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
      cross-code racers converge on the winner's row (role checked, staff still
      refused) instead of a unique-violation 500.
    - account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.
    
    Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
    db clean; zero unmapped errors.
    0414913b
  • Lemon-miaow's avatar
    fix(manifests): render the retention reaper CronJob into the Minecraft namespace · e4f2cff5
    Lemon-miaow authored
    A Pod can only mount PVCs from its own namespace; the CronJob referenced the
    minecraft-namespace backup PVC while being rendered under ControlNamespace, so
    it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim
    felis-backups not found'. The reaper Role/RoleBinding were already
    minecraft-scoped (the objects it touches live there), so the CronJob was the
    odd one out. The minecraft felis-config replica (felis setup, backup Job fix)
    supplies its config mount.
    e4f2cff5
  • Lemon-miaow's avatar
    fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace · c839454a
    Lemon-miaow authored
    Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
    another namespace either (live drill: 'error looking up service account
    minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
    RoleBinding subject to the Minecraft namespace alongside the CronJob.
    c839454a
  • Lemon-miaow's avatar
Loading
Loading