Commit Graph
319 Commits
Author SHA1 Message Date
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow a2df2f242b fix(operator): re-probe RCON every 2s while Starting
The readiness gate is status-driven; at a 5s re-probe cadence the observed
'container Ready but API still 409 not_running' window was 6~10s. Halving the
cadence halves the worst case; probes still only run while unreachable.
2026-09-22 20:31:53 +08:00
Lemon-miaow 2a8f897e61 fix(api): a session-store outage answers 503, not 401
Resolving a session cookie failed identically whether the credential was
missing or Postgres was unreachable: local_auth_enabled read errors fell into
the fail-closed 'disabled' branch and SessionUser errors into 'invalid
session', both surfacing as 401 'authentication required' — a lie that reads
as 'log in again' during an outage. Split the enabled-read into
(enabled, error), tag non-ErrNotFound store failures with errAuthBackend, and
map that to a new 503 auth_unavailable in requireExternal. Fail-closed is
unchanged: missing setting / bad value / missing session stay 401.
2026-09-22 20:30:43 +08:00
Lemon-miaow abce381faa fix(tui): wrap the one-time setup URL so narrow terminals can't truncate it
The setup URL carries a 43-char token and overruns 80 columns; the TUI
renderer clipped it. Break it at the query '=' boundary (token on its own
line) with a shared wrapDisplayURL helper used by both the Owner wizard and
the mc-bind wizard; unit test pins the no-loss concatenation.
2026-09-22 20:28:01 +08:00
Lemon-miaow a415246adc fix(operator): give controller-runtime a logger instead of a goroutine stack
Without SetLogger, the first reconcile prints
'[controller-runtime] log.SetLogger(...) was never called; logs will not be
displayed' followed by a full stack trace (live-observed in felis-operator).
Route it through logr.FromSlogHandler(slog.Default()) so its messages are
ordinary stderr lines; go-logr/logr promoted to a direct dependency.
2026-09-22 20:25:39 +08:00
Lemon-miaow 1efa8a4b08 docs(audit): batch 2026-09-22 evening — auth E2E, concurrency, reaper drill, fixes #16-#19 2026-09-22 20:19:35 +08:00
Lemon-miaow c839454a1f fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
another namespace either (live drill: 'error looking up service account
minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
RoleBinding subject to the Minecraft namespace alongside the CronJob.
2026-09-22 20:18:36 +08:00
Lemon-miaow e4f2cff532 fix(manifests): render the retention reaper CronJob into the Minecraft namespace
A Pod can only mount PVCs from its own namespace; the CronJob referenced the
minecraft-namespace backup PVC while being rendered under ControlNamespace, so
it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim
felis-backups not found'. The reaper Role/RoleBinding were already
minecraft-scoped (the objects it touches live there), so the CronJob was the
odd one out. The minecraft felis-config replica (felis setup, backup Job fix)
supplies its config mount.
2026-09-22 20:14:19 +08:00
Lemon-miaow 0414913bc7 fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
6-way concurrent redeem of one code 500'd on users_username_key (each request
generated a fresh user id but the same uuid-derived username), plus the rarer
two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
locks its code row and handles the conflict.

- SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
  as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
- INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
  cross-code racers converge on the winner's row (role checked, staff still
  refused) instead of a unique-violation 500.
- account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.

Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
db clean; zero unmapped errors.
2026-09-22 20:08:38 +08:00
Lemon-miaow dcc3b7403e fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake)
The interface doc promised 'each joined to its staff username', the fake and
the pending handler both project username and created_at, but the PG query
selected neither — live internal /op-login/pending returned username:"" and
created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
tests can't see PG-only regressions.
2026-09-22 20:04:32 +08:00
Lemon-miaow 52549f7b3a fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
The PG implementation was a single UPDATE ... WHERE code_hash that returned
ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
the pre-session email-login door (and the op-login finish / migration confirm
doors) fell through to writeError's unmapped-error 500, and attempts were never
charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
intended contract; only the PG side had drifted.

Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
UPDATE the newest live row, expiry + attempt cap before the hash compare,
mismatch charges one attempt and returns ErrOTPInvalid without consuming,
match consumes and commits. Verified live on the VM: 5 wrong guesses return
400 and stop at attempts=5 (correct code then also refused, unconsumed);
fresh code redeems; replay returns 400.
2026-09-22 19:51:04 +08:00
Lemon-miaow 9309ff5a7f chore: apply the missed S1016 conversions in handlers_users
The gofmt/staticcheck commit staged handlers_user.go (singular) for the
formatting fix but missed this sibling for its two struct-literal-to-
conversion cleanups.
2026-09-22 17:57:03 +08:00
Lemon-miaow e690b058db fix(restore): wait for the tracking finalizer before recreating
Live verification of the previous commit showed the immediate retry STILL
stranded: deleting a finished Job leaves it terminating (job-tracking
finalizer), so the re-Create collided with the dying object and was
mapped to ErrAlreadyExists a second time. Poll until the name actually
frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck
finalizer ever outlives the budget. Fake-client tests pin both the
replace-finished and coalesce-in-flight branches.
2026-09-22 17:54:01 +08:00
Lemon-miaow 90ccbfede4 fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
2026-09-22 17:50:04 +08:00
Lemon-miaow fd0794d04d chore(deps): pgx v5.9.2, x/net v0.55.0, x/text v0.39.0
govulncheck flagged pgx v5.7.1 (GO-2026-5004, SQL-injection class) as
reachable from pgrepo.go, plus the old x/net and x/text. Bump all three
to the fixed versions; go vet/test stay green.
2026-09-22 17:50:04 +08:00
Lemon-miaow a05edc934c chore: gofmt the tree, clear staticcheck, add a CI gofmt gate
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
2026-09-22 17:49:54 +08:00
flyemoji a56c326518 chore: stop tracking the docs/changes ledger
docs/changes held 26 per-feature change notes and their index,
written while each feature was built. They were working records, not
documentation: they cite internal milestone numbers and plan steps,
several describe designs that changed before they shipped (the nano
note's config schema and a proxy plugin that was never built), and
nothing in the code, the build or the other docs refers to them. New
notes stopped being added a while ago; the commit messages carry that
record now.

The directory leaves the tree in this commit. Its contents stay
reachable in history, and the files were kept outside the repository
before removal. No code, build or test changes.
2026-09-22 15:02:47 +09:00
flyemoji 7b5b28c587 fix(bootstrap): keep the nano build toolchain under /opt/felis
The source build of the nano binary installed Go at /usr/local/go and
replaced whatever version was already there. On a host that also
builds other things, the operator's own toolchain was removed and
swapped for Felis's pinned version without a word.

GOROOT_DIR is now /opt/felis/go, next to the source, the Velocity
install and the JRE Felis already keeps under /opt/felis, and
install_go_toolchain creates the parent before unpacking. A host where
an earlier run put Go at /usr/local/go downloads it once more on the
next re-run and keeps the old tree untouched; removing it is the
operator's call. The harness now requires the toolchain directory to
be under /opt/felis.
2026-09-22 14:59:39 +09:00
flyemoji 0faec2b02a fix(bootstrap): open the nano port to the proxy alone
For a non-loopback bind, configure_nano_firewall opened the nano port
in firewalld to every source, while the summary told the operator to
restrict it to the proxy. hasJoined takes no token, so on a public
host that port is an auth relay anyone can point a proxy at, spending
this host's Mojang egress until Mojang rate-limits it and the
operator's own players stop getting in.

A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets
one rich rule that admits the port from that source only, ipv4 or
ipv6 by the address given. Without it, no port is opened and the
summary prints the rule to add. A re-run closes the port an earlier
installer opened to every source. A rule for a previous
FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand.
Hosts without firewalld are handled as before.

The value goes into the rule text, so it is checked up front for an
address with one prefix length and nothing else. firewalld's own
parser accepts both rule forms and refuses an ipv6 address under the
ipv4 family. The harness covers the rule for each family, the
closed-by-default case, the re-run cleanup, the loopback case and the
CIDR check.
2026-09-22 14:51:49 +09:00
flyemoji 99c31c1d4e fix(config): refuse plaintext auth-source urls to public hosts
An auth_source url could be http:// to any host. Anyone on the path
to a public root, or anyone who can spoof its DNS name, can then
answer hasJoined with a 200 and log in as any player of that source,
including a third-party account linked to staff. The player's IP also
travels in cleartext. Mojang logins are unaffected, since that source
is built in over https.

Config load now refuses http:// unless the host is localhost or a
loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12,
192.168/16, fc00::/7), so a root on the same host or the LAN still
works without TLS. The decision is made on the literal host because
nothing is resolved at load time, so a LAN root named by hostname
needs its IP address or https. The error says what to change.

The new test covers public names and addresses, link-local, 0.0.0.0
and the first address past 172.16/12 (all refused over http, all
accepted over https), and the loopback and private forms that stay
allowed. It fails on the old check.
2026-09-22 14:24:41 +09:00
flyemoji e9f74f3f0f fix(nano): stop trusting an expired free name while mojang is failing
When the premium-name lookup failed, isPremiumName fell back to any
cached answer, however old. An expired "free" is exactly the answer
that may have stopped being true: someone can buy the name after it
was last seen free. For as long as api.mojang.com kept failing (429,
5xx, a timeout), a third-party player holding that name kept it on
every reconnect, and the Velocity registry, keyed on the name, turned
its new owner away as already connected. A hostile source could drive
the host into Mojang's rate limit on purpose to hold names that way.

A failed lookup now always counts as taken, so the player is renamed
with the source's prefix. An expired "taken" already gave that answer,
so only the stale "free" case changes. The cost is cosmetic: during an
outage an ordinary third-party player may get a prefix they do not
need, and their data follows the UUID, not the name.

A new test gives the cache a free entry past its TTL and has Mojang
answer 429. It fails on the old fallback. The two comments that
described the fallback now describe the fail-closed rule.
2026-09-22 14:17:53 +09:00
flyemoji fa7b54f5ab fix(bootstrap): refuse an unbracketed ipv6 nano listen address
validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081
passed. Go refuses that form ("too many colons in address") and needs
[::1]:8081, so the unit crash-looped on every start. A host part that
contains a colon must now be in brackets.

With that, the bare ::1 pattern in nano_listen_is_loopback can no
longer match an address that gets this far, so it goes. [::1] stays.
The harness adds ::1:8081 to the refused addresses, and [::]:8081 and
:8081, both of which Go binds, to the accepted ones.
2026-09-22 14:08:11 +09:00
flyemoji 928a1fdfff docs(bootstrap): credit velocity, not authlib, with the hasjoined call
Two installer comments still said authlib makes the hasJoined request
and sends no token. Velocity reads -Dmojang.sessionserver and sends
the request itself. Comment text only.
2026-09-22 13:55:12 +09:00
flyemoji b58c20311c test(bootstrap): pin the nano listen default to loopback
nano_listen_is_loopback decides whether configure_nano_firewall opens
the port, and hasJoined takes no token. A default that does not
classify as loopback would make every fresh nano host a public auth
relay.

The harness now runs the classifier on four loopback binds and three
routable ones, and feeds it the default resolve_nano_listen applies on
a first install, with no operator value and no existing unit. Setting
that default to 0.0.0.0:8081 or :8081, or counting 0.0.0.0 as
loopback, now fails the harness. Test only.
2026-09-22 13:54:31 +09:00
flyemoji cf65ffdae5 fix(bootstrap): open up a nano-only config dir an older run left 0750
write_nano_config creates a missing /etc/felis as 0755, but it left
an existing one alone. On a nano-only host an older installer made
that directory with a bare mkdir -p, so under a root umask of 027 it
is 0750. The DynamicUser unit cannot search it, so felis-nano cannot
read its config, and a re-run stops at the service check instead of
repairing the directory.

An existing directory is now set to 0755 unless it holds the full
install's secrets.env or bootstrap.done. The full install locks the
directory to 0700 and writes secrets.env right after, so its directory
keeps that mode, and install_nano_service still reports the lockout
rather than this widening it. The mode cases run only where chmod
works; on a filesystem that ignores it the harness skips them.
2026-09-22 13:54:18 +09:00
flyemoji 2458ee1722 docs(bootstrap): say auth_source tags are permanent and order is trust
Both config templates the installer writes, the nano felis.toml and
the comment above [[auth_source]] in the generated felis tomls, now
state two things an operator editing the list needs to know.

A tag is hashed verbatim into every player UUID of its source, with no
case folding, so renaming it gives all of those players new UUIDs and
orphans their data, links and bans. The list is scanned in order and
the first source that validates wins, so order is trust, and a
compromised root has to be removed, not moved down. Comment text only.
2026-09-22 13:53:15 +09:00
flyemoji 6794e66c4d fix(bootstrap): fail a tokenless private clone instead of prompting
A source build against a private repository with no FELIS_GITHUB_TOKEN,
or a wrong one, made git ask for a username on /dev/tty, and a piped
install sat there waiting.

git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git
fails at once with "terminal prompts disabled". Both fetch_source
failures name FELIS_GITHUB_TOKEN in their message: the fresh clone,
and the fetch into an existing checkout, which had no message of its
own before.
2026-09-22 13:53:03 +09:00
flyemoji 34f73ba19f fix(bootstrap): detect a missing terminal by opening /dev/tty
prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which
never fires on Linux: /dev/tty is mode 0666 whether or not the process
has a controlling terminal, and only opening it fails. Without a
terminal the menu was printed, the read failed with "No such device or
address", and the default was taken by accident rather than by the
documented path.

The guard now opens /dev/tty in a subshell and takes the "no terminal
for a prompt" path when that fails.
2026-09-22 13:52:51 +09:00
flyemoji 0758b9c5d7 docs(bootstrap): pass tunables on the sudo line, not by export
The header said to export tunables before running, but its own
`curl ... | sudo bash` entrypoint resets the environment, so an
exported FELIS_INSTALL_MODE or FELIS_NANO_LISTEN never reached the
installer. The header now shows the two forms that do arrive: the
variable named on the sudo line, or export followed by sudo -E.

The nano summary's hint for a proxy on another machine now prints a
sudo line that can be pasted as is, instead of "re-run with
FELIS_NANO_LISTEN=...". Comment and log text only.
2026-09-22 13:52:40 +09:00
flyemoji 02c079c893 fix(bootstrap): verify the go toolchain tarball against a pinned digest
install_go_toolchain downloaded the tarball to a fixed /tmp name and
unpacked it into /usr/local as root, with no digest check. Another
local user could plant that file first, and nothing would notice a
tampered download.

The tarball is now staged in a mktemp -d directory that the exit
cleanup removes, and its sha256 must match before the old toolchain is
touched, so a refusal leaves the host as it was. The default 1.26.4
carries pinned amd64 and arm64 digests next to its version; they are
the ones https://go.dev/dl/?mode=json&include=all publishes. Any other
FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in
the header, because no pin can cover a version chosen at run time.
Where and which version gets installed is unchanged.
2026-09-22 13:52:29 +09:00
flyemoji 3b0fc7a3e0 fix(bootstrap): print the address nano binds in the install summary
summary_nano printed the node's primary IP for every non-loopback bind
and 127.0.0.1 for every loopback one. A bind to a second private
address, or to [::1], handed the operator a hasJoined URL that nothing
listens on.

The host is now the part of FELIS_NANO_LISTEN before the last ':'. The
node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]),
which names no address a proxy could dial. The loopback and
public-bind notes are unchanged.
2026-09-22 13:52:18 +09:00
flyemoji 404d1172a6 fix(bootstrap): refuse a nano listen address without a usable port
FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in
the firewall while nano bound nothing, a bare 127.0.0.1 printed
http://127.0.0.1:127.0.0.1/... in the summary, and the unit
crash-looped either way.

validate_settings now requires a ':' and a decimal port of 1-65535
after the last one. It runs after resolve_nano_listen, so an address
read back from an existing unit is checked too, and the default always
passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
2026-09-22 13:52:08 +09:00
flyemoji 3918a4b11a fix(bootstrap): install only the full control plane under felis setup
prompt_install_mode also runs inside felis setup. Setup then goes on
to the Owner and edge setup, which need the control plane, so choosing
nano there always ended in a setup error.

Under felis setup the mode is now full before any prompt or default is
considered, and an explicit FELIS_INSTALL_MODE=nano stops with a
message pointing at deploy/bootstrap.sh. That leaves the felis setup
branch of acquire_nano_binary unreachable, so it goes.
install_embedded_binary stays, since the full install still uses it.
2026-09-22 13:52:00 +09:00
flyemoji 515c4a6496 fix(bootstrap): keep a nano host's listen address and mode on re-run
Re-running the installer is how a nano host updates. That re-run reset
FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost
its endpoint and every login through it failed. It also offered the
full control plane as the default, which on a nano host means k3s and
Postgres nobody asked for.

The listen address is now settled by resolve_nano_listen, the first
step of main, so the later checks see the result. The operator's value
wins, then the -listen argument of the installed felis-nano unit, then
loopback. The install mode defaults to nano, at the prompt and without
a terminal, when the felis-nano unit exists and the full install's
bootstrap.done marker does not. Only the full install writes that
marker.

The harness reads back the unit it wrote earlier, and checks the mode
default on a nano-only host, a host with the full install, and a fresh
host.
2026-09-22 13:51:52 +09:00
flyemoji 17b4396460 fix(bootstrap): carry auth_source tables with spaced or quoted headers
A re-run copies the operator's [[auth_source]] tables from the existing
felis toml into the new one. The awk program that finds them matched
only the literal header [[auth_source]], so a table written as
[[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid
TOML, was taken for some other section and dropped from the config.

Each section header now decides afresh whether it opens an auth_source
table, through one regex that allows inner whitespace and a single- or
double-quoted key. The single quote is spelled \047, which gawk and
mawk both honour inside a bracket expression. The harness carries each
spelling and checks that the table still stops at the next section.
2026-09-22 13:51:45 +09:00
flyemoji c2a5645c55 fix: keep internal section numbers out of runtime messages
Four messages that reach an operator or an API client cited sections
of a specification nobody outside the project can read:

- the unimplemented archive store error from config load
- the running-server cap refusal, from both the user wake and the
  internal wake
- the missing memory ceiling guard, in the API and in felis apply

The references are gone and the wording is otherwise unchanged. Each
message still says what went wrong and, where there is one, what to
do about it. The test for the archive store message checks for the
tarLocal remediation, which is still there.
2026-09-22 13:44:38 +09:00
flyemoji 9ee8c48fff docs(openapi): list every answer hasjoined gives
The hasJoined contract listed only 200 and 204 and named authlib as
the caller. The handler now answers four more ways, and a proxy
operator reading the contract could not tell a refused login from a
down source.

- 204 also covers a missing or oversized parameter (no source is
  asked), a third-party name that is not a legal Minecraft username,
  and an identity id that does not parse.
- 400 for a request that declares a body. There is no response body,
  and the connection is closed.
- 500 when the bar-list lookup fails, with the usual error body.
- 503 when no source validated and at least one failed, since that
  source's player may be the one logging in.

The three query parameters now carry the 64-byte cap. The profile name
says a third-party player holding a registered Mojang name gets it
back prefixed and cut to 16 characters. The description names
Velocity, drops the "thin login hook" that does not exist, and says
that a non-200, non-204 answer makes Velocity report the auth servers
as down.
2026-09-22 13:43:08 +09:00
flyemoji 1ebd73a309 docs(nano): describe the hasjoined path as it works
The comments around hasJoined still described an authlib client that
is not in the path. Velocity reads -Dmojang.sessionserver and sends
the request itself, and it turns a 204 into its online-mode-only kick,
not authlib's "failed to verify username". The route comment in api.go
also offered "a thin login hook" as an alternative that does not
exist.

Other comments had drifted from the code:

- The [[auth_source]] doc said an empty list ships the multiplexer
  off. Mojang is always prepended, so an empty list means Mojang is
  the only source.
- The premium-name cache said Mojang does not recycle names. A name
  frees up when its owner renames away. The day-long "taken" TTL still
  holds, because a stale "taken" costs a third-party player only a
  prefix.
- The cache bound claimed entries come only from players who
  authenticated somewhere. Any third-party source that validates a
  login adds one, so a hostile source can force the map to clear. That
  costs repeat lookups, or a fail-closed prefix while Mojang is
  unreachable, never an identity.

The rewrite rationale now states what it costs a backend operator. A
chat-session key that a third-party source signed over its native UUID
cannot verify against the canonical UUID, so chat from those players
can only be accepted unsigned.

In the tests, comments that repeated their subtest names are gone.
2026-09-22 13:42:03 +09:00
flyemoji 59ec23d4a8 test(nano): cover the nano delivery path and its loopback default
felis nano serves the same hasJoined handler as felis api, but behind
nanoStubRepo, which implements only the bar-list lookup and embeds a nil
Repo for everything else. Only the full-api path was tested, against a
complete fake store, so a second store call added to handleHasJoined
would pass CI and panic on every nano login. The loopback default of
-listen, the one thing keeping nano from being an open auth relay, was
not pinned either.

The default moves into a nanoDefaultListen constant, and two tests
cover the path. One serves a login through api.HasJoinedHandler with
nanoStubRepo and a fake identity source and expects the profile back.
The other requires the default to parse as a loopback IP. Taking the
bar-list method off the stub makes the first panic on the nil Repo;
defaulting to 0.0.0.0:8081 or :8081 fails the second.
2026-09-22 13:37:51 +09:00
flyemoji e0ad78af98 test(config): make the identity-key test fail when the key is accepted
TestLoadRejectsAuthSourceIdentityKey is the guard against a config line
identity = true making a third-party source's UUIDs trusted as-is. Its
fixture had no prefix, so Load failed on the prefix rule and the test
passed on that error. With the unknown-key check in decodeConfig
disabled, the test still passed.

The fixture now carries a valid prefix, the error must mention unknown
keys and identity, and LoadNano is checked alongside Load. With the
unknown-key check disabled, both loaders now fail the test; the old
version of the test passes against the same change.
2026-09-22 13:36:29 +09:00
flyemoji 30b4e1dfb2 test(nano): cover the bar-list error, bad identity id and ip relay
Three paths in handleHasJoined had no test that fails when they break:

- A bar-list lookup error answers 500. Logging it and carrying on would
  admit a reclaimed squatter during a database outage.
- An identity (Mojang) id that does not parse answers 204. Ignoring the
  parse error would emit the nil UUID for every such login, so they all
  share one player's data.
- The ip parameter is relayed to each source. Dropping it turns off the
  sources' check that the session is used from the player's own address.

One subtest each. Mutants that ignore the bar-list error, ignore the id
parse error, or stop appending ip each fail their subtest.
2026-09-22 13:35:50 +09:00
flyemoji 942e9a5ff8 test(nano): cover the premium-name cache rules
isPremiumName decides on every third-party login whether the player
keeps their name, and none of its rules had a test that fails when the
rule breaks: treating a 429 or 5xx from api.mojang.com as "free",
swapping the free and taken TTLs, flipping the freshness comparison,
answering "free" from an expired taken entry during an outage, or
dropping the clear-at-4096 bound. Each of those leaves a squatter
holding a name its owner has bought, or grows the cache without limit,
with CI green.

TestPremiumNameCache drives isPremiumName against a stub that answers
with a fixed status and counts lookups, and seeds cache entries at chosen
ages. Five mutants of handlers_hasjoined.go, one per rule above, each
fail at least one subtest. It does not test an expired "free" entry
during an outage; what that case should return is still open.
2026-09-22 13:34:43 +09:00
flyemoji 3f7274d29f test(nano): pin the auth namespace and one rewritten uuid as literals
The rewrite test computed its expected UUID from felisAuthNS itself, so
a change to the namespace seed moved both sides together and still
passed. Such a change gives every third-party player a new UUID on next
login, orphaning their playerdata and account links and letting any
squatter barred by the old UUID back in.

The test now also compares felisAuthNS and the rewrite of
littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e-
9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were
computed independently with Python's uuid5/uuid3, not read back from
the code. Prefixing the seed with https:// fails the test.
2026-09-22 13:32:45 +09:00
flyemoji a7fe525bfc test(api): keep the package's tests off the live mojang profile api
mojangProfileAPI defaults to https://api.mojang.com, and only the tests
that call stubMojangNames or setProfileAPI swap it out. A new test that
reaches a third-party login without doing so would query the real
service: its result then depends on network access and on whether
someone owns the name that day, and the shared premium cache can carry
that answer into later tests.

A TestMain now points the lookup at an address nothing listens on
before any test runs, so a forgotten stub always takes the same
fail-closed path. Tests that stub it restore this address, not the live
one, when they finish.
2026-09-22 13:31:53 +09:00
flyemoji 8e9c8ca4e6 fix(nano): say that [server] listen is ignored instead of defaulting it
LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset,
and a test pinned that value, but felis nano never reads it: it binds
the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An
operator moving nano off loopback by writing [server] listen in its
config got connection refused from the proxy and no hint that the key
did nothing.

LoadNano no longer sets the default, and nano prints a line naming the
ignored value and the address it actually binds whenever the key is
set. It is a warning rather than a load error so a full felis.toml
copied onto a nano host keeps starting. The assertion that pinned the
unused default is removed along with it.

The new test runs cmdNano against a config that sets [server] listen
and one that does not, with an unbindable -listen so it returns after
loading. The first must warn and the second must not; with the old
default restored, the second prints a warning about 0.0.0.0:8080.
2026-09-22 13:30:23 +09:00
flyemoji 1d6c73007e fix(nano): drain in-flight logins on shutdown
The installer and the config template tell the operator to run
systemctl restart felis-nano after editing the source list. nano had no
signal handling, so SIGTERM killed it mid-request: a login waiting on an
upstream had its connection reset, and Velocity disconnected that
player with "authentication servers are down". felis api already drains
on shutdown; nano did not.

nano now listens itself, serves until SIGINT or SIGTERM, then shuts the
server down gracefully with a 30-second limit. That outlasts the source
scan of any realistic list, at five seconds per source, and stays well
inside systemd's default 90-second stop timeout.

The new test holds a request inside the handler, cancels the serve
context, and checks that serveNano is still running 200 ms later, that
the held request then gets its answer, and that serveNano returns 0.
Replacing the graceful shutdown with Close fails it.
2026-09-22 13:28:59 +09:00
flyemoji fa3eda5228 fix(nano): quote and cap the request log line
felis nano logged every request with the raw RequestURI and %s. That
text is the caller's: a right-to-left override reordered the line as
displayed, an invalid UTF-8 byte made journald store the entry as a
binary blob that journalctl -f shows as "[N blob data]", and a query
near net/http's one-megabyte limit became a one-megabyte log line.

The URI is now capped at 256 bytes, several times a real hasJoined
query, and printed with %q, so control, bidi and invalid bytes appear
escaped. The handler assembly moved into nanoHandler so the logged
handler can be tested on its own; cmdNano serves it unchanged.

The new test sends a query carrying U+202E, a 0x9b byte and 4 KiB of
padding, and expects a valid UTF-8 line with the override escaped and
no more than twice the cap. Restoring the old unquoted line fails it.
2026-09-22 13:27:28 +09:00
flyemoji 1905cac950 fix(config): refuse auth-source tags padded with whitespace
A third-party player's UUID is hashed from the source tag byte for byte,
so the tag is a permanent namespace: change it and every player of that
source comes back as someone new, with their playerdata, permissions,
account links and reclaim bans left behind. Nothing said so, and a tag
with a stray leading or trailing space, which nobody can see in the
file, loaded as a brand new namespace.

Such a tag is now rejected at load, and the AuthSourceConfig doc states
that the tag is permanent, case included. The charset stays otherwise
open: tightening it would force existing installs to rename, which is
the very thing that rekeys their players.

The new test loads a tag with a trailing space, a leading space and a
trailing tab through LoadNano; all three loaded before this change.
2026-09-22 13:24:23 +09:00
flyemoji 72a2750461 fix(config): refuse mojang as an auth-source tag
Mojang is prepended in code as the first, identity source, and the
config templates say not to list it. Nothing enforced that. A listed
tag = "mojang" loaded, and nano's startup list printed it as if Mojang
had been pointed at that url, while the real Mojang was still asked
first. The listed entry was a separate third-party source: asked again
on every login that got past Mojang, adding up to five seconds when its
url was Mojang's own and it answered 204 each time.

Any case of "mojang" is now rejected at load with a message saying
Mojang is built in and must not be listed. The duplicate-tag check could
not catch this because the built-in source never passes through it.

The new test loads "mojang" and "Mojang" through LoadNano; both loaded
before this change.
2026-09-22 13:23:37 +09:00
flyemoji 2c74080b78 fix(config): reject auth-source urls the resolver cannot query
The url check only looked for an http:// or https:// prefix. Several
shapes passed it and then left the source dead at login time: no host
("https://"), a bad port, surrounding whitespace (sent as %20 and
answered 404), and any query or fragment. The resolver appends
"?username=…&serverId=…" to the url as a string, so an existing query
swallows those parameters and a fragment hides them from the request
entirely. Each loaded green, and every login from that source failed.

The url is now parsed and must be http or https with a host, no query,
no fragment and no surrounding whitespace. Load and LoadNano share the
check. The shipped LittleSkin default and plain http:// endpoints, such
as a same-host root on loopback, still load.

The new test feeds each rejected shape to LoadNano. Against the previous
prefix check, six of the seven load; only ftp:// was refused.
2026-09-22 13:23:02 +09:00