Author SHA1 Message Date
Lemon-miaow 1efa8a4b08 docs(audit): batch 2026-09-22 evening — auth E2E, concurrency, reaper drill, fixes #16-#19 2026-09-22 20:19:35 +08:00
Lemon-miaow c839454a1f fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
another namespace either (live drill: 'error looking up service account
minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
RoleBinding subject to the Minecraft namespace alongside the CronJob.
2026-09-22 20:18:36 +08:00
Lemon-miaow e4f2cff532 fix(manifests): render the retention reaper CronJob into the Minecraft namespace
A Pod can only mount PVCs from its own namespace; the CronJob referenced the
minecraft-namespace backup PVC while being rendered under ControlNamespace, so
it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim
felis-backups not found'. The reaper Role/RoleBinding were already
minecraft-scoped (the objects it touches live there), so the CronJob was the
odd one out. The minecraft felis-config replica (felis setup, backup Job fix)
supplies its config mount.
2026-09-22 20:14:19 +08:00
Lemon-miaow 0414913bc7 fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
6-way concurrent redeem of one code 500'd on users_username_key (each request
generated a fresh user id but the same uuid-derived username), plus the rarer
two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
locks its code row and handles the conflict.

- SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
  as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
- INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
  cross-code racers converge on the winner's row (role checked, staff still
  refused) instead of a unique-violation 500.
- account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.

Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
db clean; zero unmapped errors.
2026-09-22 20:08:38 +08:00
Lemon-miaow dcc3b7403e fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake)
The interface doc promised 'each joined to its staff username', the fake and
the pending handler both project username and created_at, but the PG query
selected neither — live internal /op-login/pending returned username:"" and
created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
tests can't see PG-only regressions.
2026-09-22 20:04:32 +08:00
Lemon-miaow 52549f7b3a fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
The PG implementation was a single UPDATE ... WHERE code_hash that returned
ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
the pre-session email-login door (and the op-login finish / migration confirm
doors) fell through to writeError's unmapped-error 500, and attempts were never
charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
intended contract; only the PG side had drifted.

Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
UPDATE the newest live row, expiry + attempt cap before the hash compare,
mismatch charges one attempt and returns ErrOTPInvalid without consuming,
match consumes and commits. Verified live on the VM: 5 wrong guesses return
400 and stop at attempts=5 (correct code then also refused, unconsumed);
fresh code redeems; replay returns 400.
2026-09-22 19:51:04 +08:00
Lemon-miaow 9309ff5a7f chore: apply the missed S1016 conversions in handlers_users
The gofmt/staticcheck commit staged handlers_user.go (singular) for the
formatting fix but missed this sibling for its two struct-literal-to-
conversion cleanups.
2026-09-22 17:57:03 +08:00
Lemon-miaow e690b058db fix(restore): wait for the tracking finalizer before recreating
Live verification of the previous commit showed the immediate retry STILL
stranded: deleting a finished Job leaves it terminating (job-tracking
finalizer), so the re-Create collided with the dying object and was
mapped to ErrAlreadyExists a second time. Poll until the name actually
frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck
finalizer ever outlives the budget. Fake-client tests pin both the
replace-finished and coalesce-in-flight branches.
2026-09-22 17:54:01 +08:00
Lemon-miaow 90ccbfede4 fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
2026-09-22 17:50:04 +08:00
Lemon-miaow fd0794d04d chore(deps): pgx v5.9.2, x/net v0.55.0, x/text v0.39.0
govulncheck flagged pgx v5.7.1 (GO-2026-5004, SQL-injection class) as
reachable from pgrepo.go, plus the old x/net and x/text. Bump all three
to the fixed versions; go vet/test stay green.
2026-09-22 17:50:04 +08:00
Lemon-miaow a05edc934c chore: gofmt the tree, clear staticcheck, add a CI gofmt gate
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
2026-09-22 17:49:54 +08:00
flyemoji a56c326518 chore: stop tracking the docs/changes ledger
docs/changes held 26 per-feature change notes and their index,
written while each feature was built. They were working records, not
documentation: they cite internal milestone numbers and plan steps,
several describe designs that changed before they shipped (the nano
note's config schema and a proxy plugin that was never built), and
nothing in the code, the build or the other docs refers to them. New
notes stopped being added a while ago; the commit messages carry that
record now.

The directory leaves the tree in this commit. Its contents stay
reachable in history, and the files were kept outside the repository
before removal. No code, build or test changes.
2026-09-22 15:02:47 +09:00
flyemoji 7b5b28c587 fix(bootstrap): keep the nano build toolchain under /opt/felis
The source build of the nano binary installed Go at /usr/local/go and
replaced whatever version was already there. On a host that also
builds other things, the operator's own toolchain was removed and
swapped for Felis's pinned version without a word.

GOROOT_DIR is now /opt/felis/go, next to the source, the Velocity
install and the JRE Felis already keeps under /opt/felis, and
install_go_toolchain creates the parent before unpacking. A host where
an earlier run put Go at /usr/local/go downloads it once more on the
next re-run and keeps the old tree untouched; removing it is the
operator's call. The harness now requires the toolchain directory to
be under /opt/felis.
2026-09-22 14:59:39 +09:00
flyemoji 0faec2b02a fix(bootstrap): open the nano port to the proxy alone
For a non-loopback bind, configure_nano_firewall opened the nano port
in firewalld to every source, while the summary told the operator to
restrict it to the proxy. hasJoined takes no token, so on a public
host that port is an auth relay anyone can point a proxy at, spending
this host's Mojang egress until Mojang rate-limits it and the
operator's own players stop getting in.

A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets
one rich rule that admits the port from that source only, ipv4 or
ipv6 by the address given. Without it, no port is opened and the
summary prints the rule to add. A re-run closes the port an earlier
installer opened to every source. A rule for a previous
FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand.
Hosts without firewalld are handled as before.

The value goes into the rule text, so it is checked up front for an
address with one prefix length and nothing else. firewalld's own
parser accepts both rule forms and refuses an ipv6 address under the
ipv4 family. The harness covers the rule for each family, the
closed-by-default case, the re-run cleanup, the loopback case and the
CIDR check.
2026-09-22 14:51:49 +09:00
flyemoji 99c31c1d4e fix(config): refuse plaintext auth-source urls to public hosts
An auth_source url could be http:// to any host. Anyone on the path
to a public root, or anyone who can spoof its DNS name, can then
answer hasJoined with a 200 and log in as any player of that source,
including a third-party account linked to staff. The player's IP also
travels in cleartext. Mojang logins are unaffected, since that source
is built in over https.

Config load now refuses http:// unless the host is localhost or a
loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12,
192.168/16, fc00::/7), so a root on the same host or the LAN still
works without TLS. The decision is made on the literal host because
nothing is resolved at load time, so a LAN root named by hostname
needs its IP address or https. The error says what to change.

The new test covers public names and addresses, link-local, 0.0.0.0
and the first address past 172.16/12 (all refused over http, all
accepted over https), and the loopback and private forms that stay
allowed. It fails on the old check.
2026-09-22 14:24:41 +09:00
flyemoji e9f74f3f0f fix(nano): stop trusting an expired free name while mojang is failing
When the premium-name lookup failed, isPremiumName fell back to any
cached answer, however old. An expired "free" is exactly the answer
that may have stopped being true: someone can buy the name after it
was last seen free. For as long as api.mojang.com kept failing (429,
5xx, a timeout), a third-party player holding that name kept it on
every reconnect, and the Velocity registry, keyed on the name, turned
its new owner away as already connected. A hostile source could drive
the host into Mojang's rate limit on purpose to hold names that way.

A failed lookup now always counts as taken, so the player is renamed
with the source's prefix. An expired "taken" already gave that answer,
so only the stale "free" case changes. The cost is cosmetic: during an
outage an ordinary third-party player may get a prefix they do not
need, and their data follows the UUID, not the name.

A new test gives the cache a free entry past its TTL and has Mojang
answer 429. It fails on the old fallback. The two comments that
described the fallback now describe the fail-closed rule.
2026-09-22 14:17:53 +09:00
flyemoji fa7b54f5ab fix(bootstrap): refuse an unbracketed ipv6 nano listen address
validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081
passed. Go refuses that form ("too many colons in address") and needs
[::1]:8081, so the unit crash-looped on every start. A host part that
contains a colon must now be in brackets.

With that, the bare ::1 pattern in nano_listen_is_loopback can no
longer match an address that gets this far, so it goes. [::1] stays.
The harness adds ::1:8081 to the refused addresses, and [::]:8081 and
:8081, both of which Go binds, to the accepted ones.
2026-09-22 14:08:11 +09:00
flyemoji 928a1fdfff docs(bootstrap): credit velocity, not authlib, with the hasjoined call
Two installer comments still said authlib makes the hasJoined request
and sends no token. Velocity reads -Dmojang.sessionserver and sends
the request itself. Comment text only.
2026-09-22 13:55:12 +09:00
flyemoji b58c20311c test(bootstrap): pin the nano listen default to loopback
nano_listen_is_loopback decides whether configure_nano_firewall opens
the port, and hasJoined takes no token. A default that does not
classify as loopback would make every fresh nano host a public auth
relay.

The harness now runs the classifier on four loopback binds and three
routable ones, and feeds it the default resolve_nano_listen applies on
a first install, with no operator value and no existing unit. Setting
that default to 0.0.0.0:8081 or :8081, or counting 0.0.0.0 as
loopback, now fails the harness. Test only.
2026-09-22 13:54:31 +09:00
flyemoji cf65ffdae5 fix(bootstrap): open up a nano-only config dir an older run left 0750
write_nano_config creates a missing /etc/felis as 0755, but it left
an existing one alone. On a nano-only host an older installer made
that directory with a bare mkdir -p, so under a root umask of 027 it
is 0750. The DynamicUser unit cannot search it, so felis-nano cannot
read its config, and a re-run stops at the service check instead of
repairing the directory.

An existing directory is now set to 0755 unless it holds the full
install's secrets.env or bootstrap.done. The full install locks the
directory to 0700 and writes secrets.env right after, so its directory
keeps that mode, and install_nano_service still reports the lockout
rather than this widening it. The mode cases run only where chmod
works; on a filesystem that ignores it the harness skips them.
2026-09-22 13:54:18 +09:00
flyemoji 2458ee1722 docs(bootstrap): say auth_source tags are permanent and order is trust
Both config templates the installer writes, the nano felis.toml and
the comment above [[auth_source]] in the generated felis tomls, now
state two things an operator editing the list needs to know.

A tag is hashed verbatim into every player UUID of its source, with no
case folding, so renaming it gives all of those players new UUIDs and
orphans their data, links and bans. The list is scanned in order and
the first source that validates wins, so order is trust, and a
compromised root has to be removed, not moved down. Comment text only.
2026-09-22 13:53:15 +09:00
flyemoji 6794e66c4d fix(bootstrap): fail a tokenless private clone instead of prompting
A source build against a private repository with no FELIS_GITHUB_TOKEN,
or a wrong one, made git ask for a username on /dev/tty, and a piped
install sat there waiting.

git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git
fails at once with "terminal prompts disabled". Both fetch_source
failures name FELIS_GITHUB_TOKEN in their message: the fresh clone,
and the fetch into an existing checkout, which had no message of its
own before.
2026-09-22 13:53:03 +09:00
flyemoji 34f73ba19f fix(bootstrap): detect a missing terminal by opening /dev/tty
prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which
never fires on Linux: /dev/tty is mode 0666 whether or not the process
has a controlling terminal, and only opening it fails. Without a
terminal the menu was printed, the read failed with "No such device or
address", and the default was taken by accident rather than by the
documented path.

The guard now opens /dev/tty in a subshell and takes the "no terminal
for a prompt" path when that fails.
2026-09-22 13:52:51 +09:00
flyemoji 0758b9c5d7 docs(bootstrap): pass tunables on the sudo line, not by export
The header said to export tunables before running, but its own
`curl ... | sudo bash` entrypoint resets the environment, so an
exported FELIS_INSTALL_MODE or FELIS_NANO_LISTEN never reached the
installer. The header now shows the two forms that do arrive: the
variable named on the sudo line, or export followed by sudo -E.

The nano summary's hint for a proxy on another machine now prints a
sudo line that can be pasted as is, instead of "re-run with
FELIS_NANO_LISTEN=...". Comment and log text only.
2026-09-22 13:52:40 +09:00
flyemoji 02c079c893 fix(bootstrap): verify the go toolchain tarball against a pinned digest
install_go_toolchain downloaded the tarball to a fixed /tmp name and
unpacked it into /usr/local as root, with no digest check. Another
local user could plant that file first, and nothing would notice a
tampered download.

The tarball is now staged in a mktemp -d directory that the exit
cleanup removes, and its sha256 must match before the old toolchain is
touched, so a refusal leaves the host as it was. The default 1.26.4
carries pinned amd64 and arm64 digests next to its version; they are
the ones https://go.dev/dl/?mode=json&include=all publishes. Any other
FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in
the header, because no pin can cover a version chosen at run time.
Where and which version gets installed is unchanged.
2026-09-22 13:52:29 +09:00
flyemoji 3b0fc7a3e0 fix(bootstrap): print the address nano binds in the install summary
summary_nano printed the node's primary IP for every non-loopback bind
and 127.0.0.1 for every loopback one. A bind to a second private
address, or to [::1], handed the operator a hasJoined URL that nothing
listens on.

The host is now the part of FELIS_NANO_LISTEN before the last ':'. The
node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]),
which names no address a proxy could dial. The loopback and
public-bind notes are unchanged.
2026-09-22 13:52:18 +09:00
flyemoji 404d1172a6 fix(bootstrap): refuse a nano listen address without a usable port
FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in
the firewall while nano bound nothing, a bare 127.0.0.1 printed
http://127.0.0.1:127.0.0.1/... in the summary, and the unit
crash-looped either way.

validate_settings now requires a ':' and a decimal port of 1-65535
after the last one. It runs after resolve_nano_listen, so an address
read back from an existing unit is checked too, and the default always
passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
2026-09-22 13:52:08 +09:00
flyemoji 3918a4b11a fix(bootstrap): install only the full control plane under felis setup
prompt_install_mode also runs inside felis setup. Setup then goes on
to the Owner and edge setup, which need the control plane, so choosing
nano there always ended in a setup error.

Under felis setup the mode is now full before any prompt or default is
considered, and an explicit FELIS_INSTALL_MODE=nano stops with a
message pointing at deploy/bootstrap.sh. That leaves the felis setup
branch of acquire_nano_binary unreachable, so it goes.
install_embedded_binary stays, since the full install still uses it.
2026-09-22 13:52:00 +09:00
flyemoji 515c4a6496 fix(bootstrap): keep a nano host's listen address and mode on re-run
Re-running the installer is how a nano host updates. That re-run reset
FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost
its endpoint and every login through it failed. It also offered the
full control plane as the default, which on a nano host means k3s and
Postgres nobody asked for.

The listen address is now settled by resolve_nano_listen, the first
step of main, so the later checks see the result. The operator's value
wins, then the -listen argument of the installed felis-nano unit, then
loopback. The install mode defaults to nano, at the prompt and without
a terminal, when the felis-nano unit exists and the full install's
bootstrap.done marker does not. Only the full install writes that
marker.

The harness reads back the unit it wrote earlier, and checks the mode
default on a nano-only host, a host with the full install, and a fresh
host.
2026-09-22 13:51:52 +09:00
flyemoji 17b4396460 fix(bootstrap): carry auth_source tables with spaced or quoted headers
A re-run copies the operator's [[auth_source]] tables from the existing
felis toml into the new one. The awk program that finds them matched
only the literal header [[auth_source]], so a table written as
[[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid
TOML, was taken for some other section and dropped from the config.

Each section header now decides afresh whether it opens an auth_source
table, through one regex that allows inner whitespace and a single- or
double-quoted key. The single quote is spelled \047, which gawk and
mawk both honour inside a bracket expression. The harness carries each
spelling and checks that the table still stops at the next section.
2026-09-22 13:51:45 +09:00
flyemoji c2a5645c55 fix: keep internal section numbers out of runtime messages
Four messages that reach an operator or an API client cited sections
of a specification nobody outside the project can read:

- the unimplemented archive store error from config load
- the running-server cap refusal, from both the user wake and the
  internal wake
- the missing memory ceiling guard, in the API and in felis apply

The references are gone and the wording is otherwise unchanged. Each
message still says what went wrong and, where there is one, what to
do about it. The test for the archive store message checks for the
tarLocal remediation, which is still there.
2026-09-22 13:44:38 +09:00
flyemoji 9ee8c48fff docs(openapi): list every answer hasjoined gives
The hasJoined contract listed only 200 and 204 and named authlib as
the caller. The handler now answers four more ways, and a proxy
operator reading the contract could not tell a refused login from a
down source.

- 204 also covers a missing or oversized parameter (no source is
  asked), a third-party name that is not a legal Minecraft username,
  and an identity id that does not parse.
- 400 for a request that declares a body. There is no response body,
  and the connection is closed.
- 500 when the bar-list lookup fails, with the usual error body.
- 503 when no source validated and at least one failed, since that
  source's player may be the one logging in.

The three query parameters now carry the 64-byte cap. The profile name
says a third-party player holding a registered Mojang name gets it
back prefixed and cut to 16 characters. The description names
Velocity, drops the "thin login hook" that does not exist, and says
that a non-200, non-204 answer makes Velocity report the auth servers
as down.
2026-09-22 13:43:08 +09:00
flyemoji 1ebd73a309 docs(nano): describe the hasjoined path as it works
The comments around hasJoined still described an authlib client that
is not in the path. Velocity reads -Dmojang.sessionserver and sends
the request itself, and it turns a 204 into its online-mode-only kick,
not authlib's "failed to verify username". The route comment in api.go
also offered "a thin login hook" as an alternative that does not
exist.

Other comments had drifted from the code:

- The [[auth_source]] doc said an empty list ships the multiplexer
  off. Mojang is always prepended, so an empty list means Mojang is
  the only source.
- The premium-name cache said Mojang does not recycle names. A name
  frees up when its owner renames away. The day-long "taken" TTL still
  holds, because a stale "taken" costs a third-party player only a
  prefix.
- The cache bound claimed entries come only from players who
  authenticated somewhere. Any third-party source that validates a
  login adds one, so a hostile source can force the map to clear. That
  costs repeat lookups, or a fail-closed prefix while Mojang is
  unreachable, never an identity.

The rewrite rationale now states what it costs a backend operator. A
chat-session key that a third-party source signed over its native UUID
cannot verify against the canonical UUID, so chat from those players
can only be accepted unsigned.

In the tests, comments that repeated their subtest names are gone.
2026-09-22 13:42:03 +09:00
flyemoji 59ec23d4a8 test(nano): cover the nano delivery path and its loopback default
felis nano serves the same hasJoined handler as felis api, but behind
nanoStubRepo, which implements only the bar-list lookup and embeds a nil
Repo for everything else. Only the full-api path was tested, against a
complete fake store, so a second store call added to handleHasJoined
would pass CI and panic on every nano login. The loopback default of
-listen, the one thing keeping nano from being an open auth relay, was
not pinned either.

The default moves into a nanoDefaultListen constant, and two tests
cover the path. One serves a login through api.HasJoinedHandler with
nanoStubRepo and a fake identity source and expects the profile back.
The other requires the default to parse as a loopback IP. Taking the
bar-list method off the stub makes the first panic on the nil Repo;
defaulting to 0.0.0.0:8081 or :8081 fails the second.
2026-09-22 13:37:51 +09:00
flyemoji e0ad78af98 test(config): make the identity-key test fail when the key is accepted
TestLoadRejectsAuthSourceIdentityKey is the guard against a config line
identity = true making a third-party source's UUIDs trusted as-is. Its
fixture had no prefix, so Load failed on the prefix rule and the test
passed on that error. With the unknown-key check in decodeConfig
disabled, the test still passed.

The fixture now carries a valid prefix, the error must mention unknown
keys and identity, and LoadNano is checked alongside Load. With the
unknown-key check disabled, both loaders now fail the test; the old
version of the test passes against the same change.
2026-09-22 13:36:29 +09:00
flyemoji 30b4e1dfb2 test(nano): cover the bar-list error, bad identity id and ip relay
Three paths in handleHasJoined had no test that fails when they break:

- A bar-list lookup error answers 500. Logging it and carrying on would
  admit a reclaimed squatter during a database outage.
- An identity (Mojang) id that does not parse answers 204. Ignoring the
  parse error would emit the nil UUID for every such login, so they all
  share one player's data.
- The ip parameter is relayed to each source. Dropping it turns off the
  sources' check that the session is used from the player's own address.

One subtest each. Mutants that ignore the bar-list error, ignore the id
parse error, or stop appending ip each fail their subtest.
2026-09-22 13:35:50 +09:00
flyemoji 942e9a5ff8 test(nano): cover the premium-name cache rules
isPremiumName decides on every third-party login whether the player
keeps their name, and none of its rules had a test that fails when the
rule breaks: treating a 429 or 5xx from api.mojang.com as "free",
swapping the free and taken TTLs, flipping the freshness comparison,
answering "free" from an expired taken entry during an outage, or
dropping the clear-at-4096 bound. Each of those leaves a squatter
holding a name its owner has bought, or grows the cache without limit,
with CI green.

TestPremiumNameCache drives isPremiumName against a stub that answers
with a fixed status and counts lookups, and seeds cache entries at chosen
ages. Five mutants of handlers_hasjoined.go, one per rule above, each
fail at least one subtest. It does not test an expired "free" entry
during an outage; what that case should return is still open.
2026-09-22 13:34:43 +09:00
flyemoji 3f7274d29f test(nano): pin the auth namespace and one rewritten uuid as literals
The rewrite test computed its expected UUID from felisAuthNS itself, so
a change to the namespace seed moved both sides together and still
passed. Such a change gives every third-party player a new UUID on next
login, orphaning their playerdata and account links and letting any
squatter barred by the old UUID back in.

The test now also compares felisAuthNS and the rewrite of
littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e-
9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were
computed independently with Python's uuid5/uuid3, not read back from
the code. Prefixing the seed with https:// fails the test.
2026-09-22 13:32:45 +09:00
flyemoji a7fe525bfc test(api): keep the package's tests off the live mojang profile api
mojangProfileAPI defaults to https://api.mojang.com, and only the tests
that call stubMojangNames or setProfileAPI swap it out. A new test that
reaches a third-party login without doing so would query the real
service: its result then depends on network access and on whether
someone owns the name that day, and the shared premium cache can carry
that answer into later tests.

A TestMain now points the lookup at an address nothing listens on
before any test runs, so a forgotten stub always takes the same
fail-closed path. Tests that stub it restore this address, not the live
one, when they finish.
2026-09-22 13:31:53 +09:00
flyemoji 8e9c8ca4e6 fix(nano): say that [server] listen is ignored instead of defaulting it
LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset,
and a test pinned that value, but felis nano never reads it: it binds
the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An
operator moving nano off loopback by writing [server] listen in its
config got connection refused from the proxy and no hint that the key
did nothing.

LoadNano no longer sets the default, and nano prints a line naming the
ignored value and the address it actually binds whenever the key is
set. It is a warning rather than a load error so a full felis.toml
copied onto a nano host keeps starting. The assertion that pinned the
unused default is removed along with it.

The new test runs cmdNano against a config that sets [server] listen
and one that does not, with an unbindable -listen so it returns after
loading. The first must warn and the second must not; with the old
default restored, the second prints a warning about 0.0.0.0:8080.
2026-09-22 13:30:23 +09:00
flyemoji 1d6c73007e fix(nano): drain in-flight logins on shutdown
The installer and the config template tell the operator to run
systemctl restart felis-nano after editing the source list. nano had no
signal handling, so SIGTERM killed it mid-request: a login waiting on an
upstream had its connection reset, and Velocity disconnected that
player with "authentication servers are down". felis api already drains
on shutdown; nano did not.

nano now listens itself, serves until SIGINT or SIGTERM, then shuts the
server down gracefully with a 30-second limit. That outlasts the source
scan of any realistic list, at five seconds per source, and stays well
inside systemd's default 90-second stop timeout.

The new test holds a request inside the handler, cancels the serve
context, and checks that serveNano is still running 200 ms later, that
the held request then gets its answer, and that serveNano returns 0.
Replacing the graceful shutdown with Close fails it.
2026-09-22 13:28:59 +09:00
flyemoji fa3eda5228 fix(nano): quote and cap the request log line
felis nano logged every request with the raw RequestURI and %s. That
text is the caller's: a right-to-left override reordered the line as
displayed, an invalid UTF-8 byte made journald store the entry as a
binary blob that journalctl -f shows as "[N blob data]", and a query
near net/http's one-megabyte limit became a one-megabyte log line.

The URI is now capped at 256 bytes, several times a real hasJoined
query, and printed with %q, so control, bidi and invalid bytes appear
escaped. The handler assembly moved into nanoHandler so the logged
handler can be tested on its own; cmdNano serves it unchanged.

The new test sends a query carrying U+202E, a 0x9b byte and 4 KiB of
padding, and expects a valid UTF-8 line with the override escaped and
no more than twice the cap. Restoring the old unquoted line fails it.
2026-09-22 13:27:28 +09:00
flyemoji 1905cac950 fix(config): refuse auth-source tags padded with whitespace
A third-party player's UUID is hashed from the source tag byte for byte,
so the tag is a permanent namespace: change it and every player of that
source comes back as someone new, with their playerdata, permissions,
account links and reclaim bans left behind. Nothing said so, and a tag
with a stray leading or trailing space, which nobody can see in the
file, loaded as a brand new namespace.

Such a tag is now rejected at load, and the AuthSourceConfig doc states
that the tag is permanent, case included. The charset stays otherwise
open: tightening it would force existing installs to rename, which is
the very thing that rekeys their players.

The new test loads a tag with a trailing space, a leading space and a
trailing tab through LoadNano; all three loaded before this change.
2026-09-22 13:24:23 +09:00
flyemoji 72a2750461 fix(config): refuse mojang as an auth-source tag
Mojang is prepended in code as the first, identity source, and the
config templates say not to list it. Nothing enforced that. A listed
tag = "mojang" loaded, and nano's startup list printed it as if Mojang
had been pointed at that url, while the real Mojang was still asked
first. The listed entry was a separate third-party source: asked again
on every login that got past Mojang, adding up to five seconds when its
url was Mojang's own and it answered 204 each time.

Any case of "mojang" is now rejected at load with a message saying
Mojang is built in and must not be listed. The duplicate-tag check could
not catch this because the built-in source never passes through it.

The new test loads "mojang" and "Mojang" through LoadNano; both loaded
before this change.
2026-09-22 13:23:37 +09:00
flyemoji 2c74080b78 fix(config): reject auth-source urls the resolver cannot query
The url check only looked for an http:// or https:// prefix. Several
shapes passed it and then left the source dead at login time: no host
("https://"), a bad port, surrounding whitespace (sent as %20 and
answered 404), and any query or fragment. The resolver appends
"?username=…&serverId=…" to the url as a string, so an existing query
swallows those parameters and a fragment hides them from the request
entirely. Each loaded green, and every login from that source failed.

The url is now parsed and must be http or https with a host, no query,
no fragment and no surrounding whitespace. Load and LoadNano share the
check. The shipped LittleSkin default and plain http:// endpoints, such
as a same-host root on loopback, still load.

The new test feeds each rejected shape to LoadNano. Against the previous
prefix check, six of the seven load; only ftp:// was refused.
2026-09-22 13:23:02 +09:00
flyemoji 3338d6f0fe fix(nano): drop oversized hasJoined parameters before asking sources
username, serverId and ip were forwarded to every configured source at
whatever length the caller sent, up to the megabyte net/http allows in a
request line. Velocity never sends more than a 16-character name, a
41-character signed SHA-1 serverId and a textual IP address, so only a
direct caller reaches those sizes, and each such request cost one
oversized upstream call per source.

Any of the three over 64 bytes is now answered 204 before a source is
asked, the same as a missing username or serverId. 64 bytes still
leaves room for a 16-character name in multi-byte UTF-8.

The subtest behind this points a source that validates anything at the
handler and sends missing and oversized fields, expecting 204 and zero
upstream requests, then a well-formed login that gets 200. It replaces
the old missing-username case, whose only source was unreachable, so
the test passed even with the guard removed. Dropping the length check
now fails it on the long username; dropping the whole guard fails it on
the first missing field.
2026-09-22 13:21:43 +09:00
flyemoji a4779186a4 fix(nano): refuse hasJoined requests that declare a body
A GET to hasJoined with a Content-Length and no body held its
connection indefinitely. The handler returned, but net/http tries to
drain an unread body before it writes the answer, and nothing bounds
that wait: ReadHeaderTimeout ends with the headers. One such request
per socket pins a goroutine and a descriptor on nano or on felis-api's
internal face.

Velocity never sends a body, so any request that declares one, including
a chunked one, now gets a 400 with Connection: close, which skips the
drain and releases the connection once the answer is written.

The new subtest writes that request over a raw socket and waits three
seconds for an answer. Before the change it times out with no response
at all; now it reads a 400 marked close.
2026-09-22 13:19:55 +09:00
flyemoji ff81295aa9 fix(nano): report failing sources instead of treating them as a no
A source that timed out, answered 5xx or 429, redirected, or sent a 200
without a usable profile was skipped exactly like one that answered 204.
With nobody else validating, the login got a 204 and Velocity told the
player their account is offline-mode. Nothing was logged, so a dead or
mistyped source URL, or an http:// root that now redirects to https since
redirects stopped being followed, failed every one of its players with
no trace.

Each such failure now logs the source tag and the cause; for a 3xx it
names the Location to configure instead. When no source validates and at
least one failed, the answer is 503, which Velocity reports as the auth
servers being down and logs with the status. A source answering 204 is
still a plain no, and a validating source still wins regardless of
failures before it.

The new subtest puts a 503 source, a redirecting source and an
unreachable one each behind a Mojang that answers 204, and expects 503.
Against the previous handler every case returns 204.
2026-09-22 13:18:33 +09:00
flyemoji a0f54df2a6 fix(bootstrap): fail the nano install when the unit does not stay up
install_nano_service printed "enabled and started" straight after
systemctl restart, which returns as soon as the process is forked. An
upgrade that keeps an old felis.toml the new binary rejects (an
[[auth_source]] without a prefix, say) left the unit crash-looping in
auto-restart while the installer reported success, and every login
through the proxy failed.

The install now waits two seconds and asks systemctl is-active. A unit
that exited is in "activating (auto-restart)", which is-active does not
count as active; on real systemd a unit whose process exits 1 under
Restart=on-failure reads activating/auto-restart and is-active returns
non-zero, while a running one reads active/running and returns 0. On
failure the install prints the unit's last 20 journal lines and stops.
This also surfaces a nano unit locked out of an existing 0700 /etc/felis.

The harness runs the extracted function with systemctl stubbed both
ways. Without the check, the dead-unit cases fail.
2026-09-22 13:05:42 +09:00
flyemoji 26f685be0e fix(bootstrap): create the nano config dir world-searchable
felis-nano runs as a systemd DynamicUser, so it can read
/etc/felis/felis.toml only if others may search /etc/felis.
write_nano_config made the directory with a bare mkdir -p, which takes
its mode from root's umask. On a host hardened to umask 027 that is
0750: nano exits on "permission denied", the unit restarts every five
seconds, and no login gets through.

A missing directory is now created 0755 explicitly. An existing one
keeps its mode, because the full install sets it to 0700 to protect its
secrets and widening that from the nano path would expose them. A nano
unit locked out that way is left for the install to report.

The harness runs the extracted function under umask 027 and checks both
cases. Reverting to the bare mkdir fails the first; an unconditional
chmod 0755 fails the second. The mode checks skip on filesystems that
ignore chmod, such as Git Bash on NTFS.
2026-09-22 13:04:34 +09:00
flyemoji 1dd62a9bdc fix(nano): cap upstream response headers at 16 KiB
The hasJoined and name-lookup clients limited the body to 64 KiB but
left headers at the transport default of 1 MiB. A configured root could
answer with a megabyte of headers and stall the body, holding a few MiB
of heap per in-flight login for the full five seconds; enough parallel
logins take down the host, and every source's logins with it.

Both clients now share a transport with MaxResponseHeaderBytes set to
16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends
338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over
the cap fails the request and the resolver moves on to the next one.

The new subtest puts a source with 64 KiB of headers and a valid profile
ahead of an honest one and expects the honest player. Without the cap
the padded source wins.
2026-09-22 13:02:48 +09:00
flyemoji 2180e77cf5 chore: drop tool-name markers from source comments
Seventeen comments opened with a tag naming the tool that wrote them.
The tag goes and each comment keeps its reasoning, now starting as a
plain sentence. None of the reasoning changes.

The AGENTS.md note in .gitignore drops the story of how the file got
into the tree and keeps the one fact a reader needs: its advice to run
go fmt is destructive on this CRLF working tree.

Comments only; no code, build or test changes.
2026-09-22 12:57:19 +09:00
flyemoji 4e98ae6e56 fix(config): reject auth-source tags that contain a colon
A third-party player's canonical UUID is UUIDv3 over tag+":"+nativeID,
and the native id is whatever the source answers. Tags were only
checked for being non-empty and unique, so both "guild" and "guild:eu"
could be configured. The "guild" root could then answer hasJoined with
id "eu:X" and receive exactly the UUID of "guild:eu"'s player X, along
with their playerdata, permissions and account links. Real native ids
are 32 hex digits, so only the shorter tag's source can do this, and
only when the operator has configured such a pair; when they have, it
is a full impersonation.

Reject a ':' in a tag at load. With colon-free tags the join is
unambiguous: two different (tag, id) pairs can no longer produce the
same input, since equal inputs force equal tags and duplicate tags are
already refused. The tag is deliberately not narrowed any further.
It is a permanent UUID namespace, and forcing an operator to rename a
tag that has no colon would move every one of its players to a new
UUID. The hash input and the native id are left exactly as they were,
so no existing player's UUID changes.

Load and LoadNano share validateAuthSources; the new test runs both
against the guild / guild:eu pair and fails on the previous config.go.
2026-09-22 12:53:35 +09:00
flyemoji 1976fca809 fix(nano): always relay properties as an array
sessionProfile tagged properties with omitempty, so an upstream answer
of "properties": [] (or null, or no key at all) reached Velocity with
no properties key. A Yggdrasil root may legitimately answer that way for
a player without a skin. Velocity 3.5.1's GameProfile deserializer
passes the missing key on as null and ImmutableList.copyOf throws, so
that player hangs at login with nothing logged, even though the same
answer sent straight to Velocity is accepted. Mojang always sends
textures, which is why the premium path and the hardware runs never hit
it.

Drop omitempty and replace a nil slice with an empty one before the
response is written. Removing omitempty alone is not enough: a nil
slice marshals as null, which Velocity rejects the same way.

The new subtest feeds the relay [], null and a missing key and expects
"properties":[] every time. The previous handler fails all three.
2026-09-22 12:52:04 +09:00
flyemoji 28d3638952 fix(nano): stop following redirects from upstream Yggdrasil roots
authHTTPClient kept net/http's default redirect policy, so a configured
third-party root that answered hasJoined with a 3xx made this host fetch
whatever URL it named, up to ten hops. That is a blind SSRF into
anything the host can reach, and it includes the multiplexer's own
listener: a root that redirects back to /session/minecraft/hasJoined
re-enters the handler, which queries Mojang and every source again and
gets redirected again, until the outer 5s client timeout fires. With a
50ms Mojang stub, one login produced 86 nested handler calls and 86
Mojang requests from this host's egress IP. The loopback default does
not help, because the redirect target is resolved from this host.

Return the 3xx as the response instead. resolveHasJoined already skips
any non-200 answer and closes its body, so a redirecting source is now
treated like one that is down, and the next source gets its turn. The
same probe now makes one handler call and one Mojang request.
Neither Mojang's nor LittleSkin's hasJoined redirects.

The new subtest puts a redirecting root ahead of an honest one and
checks that the redirect target is never contacted and the honest
source's player is returned. The pre-fix handler fails it.
2026-09-22 12:50:56 +09:00
flyemoji 07bafebf0d fix(bootstrap): keep the operator's auth sources across re-runs
write_felis_toml regenerates felis.host.toml and felis.pod.toml with a
wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin
block in that heredoc. Re-running the installer, which is also what
`felis setup` does, threw away any edit to the list: a root the operator
added stopped admitting logins, and a root they removed came back. The
generated comment invited exactly that edit.

Carry the tables forward the way [smtp] already is: read every
[[auth_source]] table from the existing felis.host.toml (falling back to
felis.pod.toml) and emit the LittleSkin default only when there is no
earlier file at all. An earlier file with no tables stays empty, because
that is a Mojang-only server rather than a missing value; felis-api now
treats an empty list that way.

The file header and the comment above the list now say what survives a
re-run, and point at felis.host.toml, which is what the next run reads.

bootstrap_test.sh extracts the new function from bootstrap.sh and checks
the fresh-install default, an operator's own table carried without the
default or the following section, an empty list staying empty, and the
indented form the setup TUI writes. It passes under dash with gawk and
with mawk; forcing the function to always return the default fails five
of the new cases.
2026-09-22 12:49:25 +09:00
flyemoji 8fe255e38f fix(api): relay Mojang logins when no auth source is configured
felis api wired the hasJoined multiplexer only when felis.toml had at
least one [[auth_source]]. With none, the source list stayed nil and
every hasJoined answer was a 204. That was harmless while nothing
pointed at the route, but the installer now starts Velocity with
-Dmojang.sessionserver aimed at felis-api unconditionally, and the
generated felis.toml tells the operator to delete the LittleSkin block
for a Mojang-only server. Doing exactly that turned every login away,
premium accounts included, and felis-api logged nothing about it.

Always build the list through authSourcesFromConfig, which prepends
Mojang in code, so an empty config is a Mojang-only relay. felis nano
already behaves this way with the same file.

The new test pins authSourcesFromConfig itself: Mojang first, the only
Identity source, and still present when nothing is configured. Marking
a configured source Identity makes it fail. The call site in cmdAPI is
now a single unconditional assignment and has no unit test of its own.
2026-09-22 12:46:00 +09:00
flyemoji 800a9042a1 test: use placeholder domains in setup and system-server tests
Three tests carried the maintainer's production root domain, a personal
mailbox and the public IP of a live demo host as fixture values. None of
them needs the value to be real: the re-domain test only needs two
different roots, and the setup flow only needs a well-formed address.

Swap them for the placeholders the rest of the suite already uses
(mc.example.net, [email protected]), and move the "before" root in the
re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737
documentation range, so the stale install the test models still has an
IP-derived hostname, which is the case the refresh exists for.
2026-09-22 12:43:54 +09:00
flyemoji d9246ddae6 feat(bootstrap): verify Paper and Velocity jars against Fill's digest 2026-08-04 17:36:47 +09:00
flyemoji 3f2b28d0ec fix(bootstrap): ship deploy/paper in the embedded game-stack tar 2026-08-04 16:34:09 +09:00
flyemoji 587f183191 Merge pull request #19 from MliroLirrorsIngenuity/chore/issue-sweep
chore: work the tracker items that need no cluster
2026-07-29 01:52:15 +09:00
87 changed files with 2313 additions and 2248 deletions

No files matched your search

+6
View File
@@ -36,6 +36,12 @@ jobs:
with:
go-version-file: go.mod
- name: gofmt
run: |
unformatted=$(gofmt -l .)
if [ -n "$unformatted" ]; then
echo "gofmt needed on:"; echo "$unformatted"; exit 1
fi
- run: go vet ./...
- run: go test ./...
+2 -3
View File
@@ -41,9 +41,8 @@ plugins/*/bin/
# ---- Local agent / loop state ----
.claude/
# Autohand-generated agent guide — kept on disk for local tooling, never tracked.
# It rode in via 0c1cc59, claims precedence over CLAUDE.md, and tells agents to
# run `go fmt` (destructive on this CRLF working tree).
# Generated tooling guide, kept on disk for local use and never tracked. Its advice
# to run `go fmt` is destructive on this CRLF working tree.
AGENTS.md
# ---- Internal planning & design docs (excluded from the public remote per
+61
View File
@@ -0,0 +1,61 @@
# Felis 生产就绪审计 — 2026-09-22(真机 E2E + 混沌)
分支:`audit-fixes-20260922`(已推送)。环境:CentOS Stream 9 / aarch64 / k3s v1.36.4,
IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:fede:69ec`),
面板经 `socat TCP6:443 → 127.0.0.1:30443` 中继(手动启动,重启 VM 后需重开)。
## 已修复并验证(分支内)
| # | 缺陷 | 证据 | 修复 |
|---|------|------|------|
| 1 | **失败恢复后重试被静默吞掉**:restore Job 固定名 + `ErrAlreadyExists` 一律当"幂等成功";失败 Job 占名 10 分钟(TTL),期间重试返回 202 `restoring` 但什么都不跑 | 真机:坏 ref 制造失败 → 立刻合法重试 → Job 原地不动、无新 Pod | `k8sjobs.go`:撞名时检查已完成(成功/失败)→ 删除+**等 finalizer 释放**(有界 10s)+ 重建;进行中仍幂等吸收。单测 3 个。**真机复验:重试 4s 完成恢复** |
| 2 | **备份 Job 全部 FailedMount**:Job 在 minecraft 命名空间挂 `felis-config`,而 bootstrap 只在 felis 命名空间创建该 Secret | 事件:`MountVolume.SetUp failed: secret "felis-config" not found` | `felis setup` 用既有 `ensureSecretReplica` 把 felis-config(`felis.toml`) 复制到 minecraft;VM 上手工复制后备份成功(167MB 归档) |
| 3 | RBAC 缺 `jobs:get/delete`(修复 #1 需要) | Role 检查 | `APIMinecraftRole` jobs → create/get/delete,测试同步更新 |
| 4 | gofmt 9 文件漂移 + CI 无 gofmt 门禁;staticcheck 9 处 | 基线扫描 | 全部修复;CI 加 gofmt job |
| 5 | 依赖漏洞:pgx v5.7.1(GO-2026-5004,可达 pgrepo.go)、x/net、x/text | govulncheck | 升 pgx v5.9.2 / x/net v0.55.0 / x/text v0.39.0 |
| 16 | **`ConsumeLoginEmailOTP` PG 实现与接口契约漂移**:契约/`fakeRepo` 说「同 VerifyEmailOTP 生命周期(扣尝试/锁定/ErrOTPInvalid/ErrOTPLocked)」,PG 却是单条 UPDATE+`ErrNotFound` → 错码/重放/过期在 login 门、op-login finish、migration confirm 三个入口全部 500;且尝试次数永不累计、`otpMaxAttempts` 锁定失效 | 真机:login 门重放正确码 → **HTTP 500**;修复前错码不扣次。修复后复测:5 次错码 400 且 attempts=5(正确码因锁定也 400、码未消费)、新码可用、重放 400 | `pgrepo.go` 改置为 VerifyEmailOTP 同构事务(FOR UPDATE、先锁后比、mismatch 扣次、match 消费),无 users 写副作用。提交 `52549f7` |
| 17 | **`ListPendingOpLogins` PG 少列**:接口注释承诺「joined to its staff username」,handler 输出 `username`/`created_at`,fake 正确填充;PG SQL 未 JOIN 也未取 `created_at` → 真机 pending 列表 username 为空、created_at 为 `0001-01-01` | 真机 internal `/op-login/pending` 响应 | `pgrepo.go` SQL 改为 JOIN users + 取 created_at。提交见分支 |
| 18 | **绑定码并发兑换 500**:`RedeemPlayerBindCode` 无 `FOR UPDATE`(同文件 `VerifyLinkCode` 有),且裸 INSERT。6 路并发同码兑换 → 3×HTTP 500(`users_username_key` 唯一冲突)+ 1×400 + 2×200;跨码并发同 UUID 同样会撞 | 真机四组并发测试 + API 日志 `unmapped error ... duplicate key` | 同码:码行 `FOR UPDATE`(输家干净地 400 invalid_code);跨码:`INSERT users ... ON CONFLICT (username) DO NOTHING`+重读、`account_links ON CONFLICT (mc_uuid) DO NOTHING`(两路汇聚同一 user)。复测:同码×6=1×200+5×400、双码×2=2×200 同 ID、DB 干净、日志零 unmapped |
| 19 | **reaper CronJob 渲染位置错误,永远无法调度**:`felis manifests` 把 CronJob 渲染到控制 ns,却引用 minecraft ns 的备份 PVC(Pod 不能跨 ns 挂 PVC:真机 `FailedScheduling: persistentvolumeclaim "felis-backups" not found`);修正 ns 后又发现 ServerAccount 也不能跨 ns 使用(`serviceaccount "felis-reaper" not found`)。而 reaper 的 Role/RoleBinding 本就在 minecraft ns | 真机三层取证(PVC/SA/调度)| CronJob 与 SA、RoleBinding subject 全部移到 `MinecraftNamespace`(提交 `e4f2cff`+`c839454`)。**修复后完整演练通过**(见下) |
## 待决策台账(未修)
| # | 主题 | 说明 |
|---|------|------|
| 6 | 默认安装无备份能力 | backup/restore 端点默认 503(需 `FELIS_BACKUP_PVC`+PVC),reaper CronJob 需 `--backup-pvc/--archive-local-path/--worlds-host-path` 三旗标渲染,bootstrap 一个都不传;文档无说明;README 与"自动备份"口径不符。另:归档 3 个月过期依赖 reaper 清理,未启用则磁盘只增不减 |
| 7 | 异步失败不可感知 | backup/restore 失败后无状态出口:restore 行不变、backup 无行;只有集群侧 Job/日志可查。建议状态字段或 `?failed` 查询 |
| 8 | 磁盘打满灾难链 | DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后约 8 分钟才恢复调度。恢复靠 `docker save felis:* | k3s ctr images import -`(docker 守护进程存储是唯一副本,需固化回源路径)。建议:PriorityClass、镜像入内置 registry、磁盘告警 |
| 9 | 升级策略 Recreate | 单副本 + Recreate:任何控制面升级=停机;坏升级(实测错 tag)服务中断约 95s 且需人工 `rollout undo`(无自动回滚)。建议 runbook/文档化 |
| 10 | 备份语义 | 归档包含整个 /data(jar、libraries、cache),167MB;是否符合"world backup"定位待评估 |
| 11 | PG 断连表现 | 会话查询失败报 401 而非 503(fail-closed 但误导;用户以为没登录) |
| 12 | ready 门滞后 | 容器 Ready 后 6~10s 内 API 仍 409 not_running |
| 13 | setup token 截断 | 43 字符 token + 长域名,80 列终端下 TUI 截断显示(复现:tmux 80 列) |
| 14 | 日志噪音 | controller-runtime 未 SetLogger,首用打印整段堆栈;TLS handshake EOF 噪音(kubelet 探针) |
| 15 | reaper 启用未演练 | **已演练完毕(2026-09-22)**:绑定挂载演练 world → 一次 pass 归档+删 PVC+保留 servers 行/CRD;过期归档驱逐(行转 deleted、文件删除、合法归档未动)。启用仍受 #6 三旗标约束;另注意 `--worlds-host-path` 的 `<path>/<pvc>` 布局在 stock local-path 下不成立(编排责任),且 VM 上演练用的 CronJob 已 **suspend** 防误删(真实世界目录不在 /srv/worlds-root)|
## 已验证事实(正向清单)
- 安装→hook 发码→Owner 绑定→passkey(虚拟认证器)→面板管理员全链路 ✅
- 建服(POST /servers)→ 唤醒(operator 拉 StatefulSet pod)→ RCON `list` → SSE 控制台 → 停止 ✅
- 文件编辑:列目录/读/写(wire 为 base64)/256KiB 413/路径穿越 5 变体全拦截/运行中 409 ✅
- **备份→篡改→恢复数据演练**:v1→备份→v2→恢复→读回 v1 ✅(G2 数据可恢复)
- **毒档案 fail-closed**:穿越/绝对路径/符号链接条目 → `archive entry escapes target` 退出码 1,零写入 ✅
- 混沌:PG 掉线(healthz 仍 200、恢复后连接池自愈)、API pod 击杀(~2s 中断)、整机重启(32s 回归、会话/CRD/停止态全保留)✅
- `felis update` 报告(k3s/velocity 有更新、私有仓库 404 优雅处理)✅
- 账户全套真机 E2E:绑定码新玩家注册(幂等/并发见 #18)、邮箱验证(onboarding 门)、邮箱 OTP 登录(错码扣次/5 次锁定后正确码也 400、重放 400、staff 账号 403 拒绝)、Passkey 注册+discoverable 登录+邮箱优先登录(虚拟认证器;Chrome 要求 `Page.bringToFront` 才能过 focus 检查)、登出吊销会话(旧 cookie 401)
- op-login 全状态机(start→status→approve→finish;早 finish 不烧码、错码扣次且请求保留、重放/重复批准/非管理员批准/未知 handle 全部按契约返回)✅
- 并发:OTP 风暴 8 路 = 1×202 + 7×429 且仅铸 1 码;绑定码并发(同码/双码)见 #18 ✅
- **reaper 全链路真机演练**(修复 #19 后):绑定挂载假世界 → `felis reaper` 一 pass:归档 tar 落盘(内容含标记文件)、`world_backups` 行 `inactive_15d`+90d 过期、**PVC 删除且宿主目录回收**、servers 行保留且 activity 时钟重置(红线②)、CRD 保留;第二 pass:伪造过期归档被驱逐(`expired=1`,文件删、行转 `deleted`),合法归档未动;`evaluated=2 reaped=1` 只动到期的世界 ✅
## 系统性观察
- **PGRepo 与接口契约/fake 漂移**(#16/#17/#18):`fakeRepo` 与接口注释是对的、PG 实现在细节上落后,单测全绿也发现不了。建议后续引入 PG 级契约测试(testcontainers 或针对关键写路径的集成测试),重点覆盖「接口注释承诺了字段/错误码/生命周期」的方法。
- 部署知识:交叉编译必须先把 `panel/dist` 拷进 `internal/panel/static` 再 `go build`,否则镜像内 SPA 缺失(页面显示 “assets were not built”)。本批镜像为 `felis:auditfix7`。
## 复现入口速查
- 面板会话 cookie:`/tmp/felis-cookies.json`;API 助手:`/tmp/fcurl.sh`
- 测试服:`test-one`(minecraft ns,stopped);合法备份 `bk-47ee2e7e96e5a4ca9d0e51b805518bac`
- CDP 调试口:Mac `127.0.0.1:9333`(独立 Chrome,profile `/tmp/felis-chrome2`);WebAuthn 虚拟认证器需在**同一 CDP 会话**内完成仪式,且先 `Page.bringToFront`(否则 NotAllowedError: page does not have focus)
- 分支已部署到 VM:`felis-api` 镜像 = `felis:auditfix7`(含全部修复)
- VM 内部面:`k3s kubectl -n felis port-forward svc/felis-api-internal 18081:8081`(Pod 重建后转发会悬死,需重启);reaper CronJob(minecraft ns)已应用但 `suspend=true`
+6 -5
View File
@@ -13,8 +13,8 @@ var bootstrapScript string
//go:embed deploy/crd/*.yaml
var bootstrapAssets embed.FS
// gameStackAssets carries everything deploy/bootstrap.sh needs to build the two
// always-on game images (login limbo + lobby) and the Velocity plugin, for the TUI
// gameStackAssets carries everything deploy/bootstrap.sh needs to build the three
// game images (login limbo, lobby, plain Paper) and the Velocity plugin, for the TUI
// install path — which pipes the embedded bootstrap.sh into bash and therefore has
// NO source checkout on disk to build from.
//
@@ -26,6 +26,7 @@ var bootstrapAssets embed.FS
//
//go:embed deploy/limbo/Dockerfile deploy/limbo/entrypoint.sh
//go:embed deploy/lobby/Dockerfile deploy/lobby/entrypoint.sh
//go:embed deploy/paper/Dockerfile deploy/paper/entrypoint.sh
//go:embed plugins/limbo/build.gradle plugins/limbo/settings.gradle plugins/limbo/src
//go:embed plugins/paper/build.gradle plugins/paper/settings.gradle plugins/paper/src
//go:embed plugins/velocity/build.gradle plugins/velocity/settings.gradle plugins/velocity/src
@@ -46,9 +47,9 @@ func GameStackTar(w io.Writer) error {
if err != nil {
return err
}
// Mode 0644 for everything: entrypoint.sh is invoked as `sh <file>` by both
// Dockerfiles precisely because the +x bit does not survive a Windows checkout,
// so nothing here needs to be executable.
// Mode 0644 for everything: entrypoint.sh is invoked as `sh <file>` by all
// three Dockerfiles precisely because the +x bit does not survive a Windows
// checkout, so nothing here needs to be executable.
if err := tw.WriteHeader(&tar.Header{
Name: path,
Mode: 0o644,
+101
View File
@@ -1,6 +1,8 @@
package felis
import (
"io/fs"
"regexp"
"strings"
"testing"
)
@@ -64,6 +66,37 @@ func TestLobbyLuckPermsWiringIsConsistent(t *testing.T) {
}
}
// The Paper jar digest rides the same cross-file contract as LuckPerms above: bootstrap.sh
// resolves "url sha256" out of Fill's content-addressed download URL and passes the digest
// as a build-arg the Dockerfile must require and verify. docker only WARNS about an unknown
// --build-arg, so a renamed arg would surface as a required-arg failure on a real host
// mid-install — this test is the only compile step the pairing gets.
//
// Both images pull the same jar from the same URL, so both have to check it: a gate on one
// of them leaves the other booting on whatever bytes happened to arrive.
func TestPaperJarDigestWiringIsConsistent(t *testing.T) {
const arg = "PAPER_JAR_SHA256"
if n := strings.Count(BootstrapScript(), "--build-arg "+arg+"="); n < 2 {
t.Errorf("bootstrap.sh passes --build-arg %s %d time(s); the lobby and the "+
"plain-Paper build each need it", arg, n)
}
for _, name := range []string{"deploy/lobby/Dockerfile", "deploy/paper/Dockerfile"} {
dockerfile := readGameStackFile(t, name)
if !strings.Contains(dockerfile, "ARG "+arg) {
t.Errorf("%s declares no ARG %s", name, arg)
}
if !strings.Contains(dockerfile, `if [ -z "${PAPER_JAR_SHA256:-}" ]`) {
t.Errorf("%s does not fail the build when %s is unset", name, arg)
}
// Requiring the arg is not the same as spending it, and which file gets hashed
// matters as much as the command: a `sha256sum -c` over some other download
// would satisfy a bare substring check while paper.jar still arrives unchecked.
if !strings.Contains(dockerfile, `echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c`) {
t.Errorf("%s never verifies /paper/paper.jar against %s", name, arg)
}
}
}
// A 1.8 client joining a protocol-47 backend dies on the first chunk unless ViaVersion's
// serverside block-connection tracking is off: under modern forwarding the Velocity injector
// reports 1.13 as the lowest supported protocol, ConnectionData.init() returns early on that,
@@ -104,6 +137,74 @@ func TestBootstrapPinsViaBlockConnectionsOff(t *testing.T) {
}
}
// The embed list and the images bootstrap.sh builds are two lists nobody reconciles.
// deploy/paper shipped an image build without ever being added to gameStackAssets, and
// nothing said so: a checkout on disk satisfies the build either way, and the tar is
// only the build context on the path that has no checkout — `curl | bash`, where the
// third `docker build -f` then names a file that was never unpacked. So derive the
// inputs from the script and from each Dockerfile's own COPY lines instead of restating
// them here; a fourth image inherits the check for free.
func TestGameStackTarCarriesEveryBuildInput(t *testing.T) {
// Matches the path only when GAME_STACK_DIR is followed by one, which skips the
// build-context arguments (`"$GAME_STACK_DIR"`, `"${GAME_STACK_DIR}:/src:z"`) and
// the glob for gradle's output, none of which are inputs this tar has to carry.
found := regexp.MustCompile(`\$\{GAME_STACK_DIR\}/(\S+?)"`).FindAllStringSubmatch(BootstrapScript(), -1)
var paths []string
seen := map[string]bool{}
for _, m := range found {
if !seen[m[1]] {
seen[m[1]] = true
paths = append(paths, m[1])
}
}
// Guards the regex itself: a rewrite of how bootstrap.sh spells the build context
// would otherwise turn this test into an unconditional pass. It has to come before
// the loop — a missing file in there is fatal, and a floor placed after it would
// never be reached to say that the regex, not the tar, is what went wrong.
if len(paths) < 3 {
t.Fatalf("only %d game-stack path(s) resolved out of bootstrap.sh; the limbo, "+
"lobby and paper Dockerfiles are all built from ${GAME_STACK_DIR}", len(paths))
}
for _, path := range paths {
requireEmbedded(t, path)
// A Dockerfile that arrives without the files it COPYs fails just as late and
// just as far from here; the deploy/paper gap was missing its entrypoint too.
for _, src := range copySources(t, path) {
requireEmbedded(t, src)
}
}
}
// copySources lists the build-context paths a Dockerfile COPYs in, skipping the
// --from=<stage> copies, whose sources are produced by an earlier stage rather than
// unpacked from the tar.
func copySources(t *testing.T, dockerfile string) []string {
t.Helper()
var out []string
// Continuations are joined first: a COPY split across lines would otherwise be two
// fragments, neither of them starting with COPY followed by a source, and its
// source would slip past unchecked.
body := strings.ReplaceAll(readGameStackFile(t, dockerfile), "\\\n", " ")
for line := range strings.SplitSeq(body, "\n") {
f := strings.Fields(line)
if len(f) < 2 || f[0] != "COPY" || strings.HasPrefix(f[1], "--") {
continue
}
out = append(out, strings.TrimSuffix(f[1], "/"))
}
return out
}
// fs.Stat rather than ReadFile: half of these are directories (`COPY plugins/shared/`),
// and embed.FS answers for those too.
func requireEmbedded(t *testing.T, path string) {
t.Helper()
if _, err := fs.Stat(gameStackAssets, path); err != nil {
t.Errorf("%s is a game-stack build input but is not in gameStackAssets; an "+
"install with no source checkout dies on it: %v", path, err)
}
}
func readGameStackFile(t *testing.T, name string) string {
t.Helper()
b, err := gameStackAssets.ReadFile(name)
+8 -9
View File
@@ -286,15 +286,14 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
}
fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)")
// Felis-nano: wire the multi-source hasJoined multiplexer only when third-party auth
// sources are configured. Mojang leads as the code-owned identity anchor (正版优先);
// config can only append namespace-rewritten third-party sources, never a trusted one,
// so a misconfig cannot reopen the impersonation hole. No sources = a.AuthSources stays
// nil = the endpoint 204s every login (ships off).
if len(cfg.AuthSources) > 0 {
a.AuthSources = authSourcesFromConfig(cfg.AuthSources)
fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources))
}
// Felis-nano: the multi-source hasJoined multiplexer. Mojang leads as the code-owned
// identity anchor (正版优先); config can only append namespace-rewritten third-party
// sources, never a trusted one, so a misconfig cannot reopen the impersonation hole.
// Wired unconditionally: the installer points Velocity at this route whether or not any
// [[auth_source]] is configured, so an empty list has to mean a Mojang-only relay, the
// same as under `felis nano`. A nil list would 204 every login, premium ones included.
a.AuthSources = authSourcesFromConfig(cfg.AuthSources)
fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources))
// Passkey (WebAuthn) verifier (spec §14, Phase 6). One relying party spans BOTH
// web faces: the RP id is the panel hostname (console.<root>), and because that is
+39
View File
@@ -3,8 +3,47 @@ package main
import (
"net/http"
"testing"
"felis.lolicon.best/internal/config"
)
// TestAuthSourcesFromConfig pins the one place the hasJoined identity anchor is decided:
// Mojang is prepended in code, first, and is the only source whose UUIDs are trusted as-is.
// The empty case matters on its own — both `felis api` and `felis nano` call this with a
// config that has no [[auth_source]] at all, and that has to be a Mojang-only relay rather
// than an empty list that rejects every login.
func TestAuthSourcesFromConfig(t *testing.T) {
for _, tc := range []struct {
name string
configured []config.AuthSourceConfig
}{
{"no configured sources", nil},
{"configured sources", []config.AuthSourceConfig{
{Tag: "littleskin", Prefix: "LS", URL: "https://littleskin.example/hasJoined"},
{Tag: "guild", Prefix: "GD", URL: "https://guild.example/hasJoined"},
}},
} {
t.Run(tc.name, func(t *testing.T) {
got := authSourcesFromConfig(tc.configured)
if len(got) != len(tc.configured)+1 {
t.Fatalf("got %d sources, want Mojang + %d configured", len(got), len(tc.configured))
}
if got[0].Tag != "mojang" || got[0].URL != mojangSessionServer || !got[0].Identity {
t.Errorf("first source = %+v, want the Mojang identity anchor", got[0])
}
for i, c := range tc.configured {
s := got[i+1]
if s.Identity {
t.Errorf("configured source %q is marked Identity; only Mojang may be", c.Tag)
}
if s.Tag != c.Tag || s.Prefix != c.Prefix || s.URL != c.URL {
t.Errorf("source %d = %+v, want %+v in config order", i+1, s, c)
}
}
})
}
}
// TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every
// felis-api listener: the shared factory must bound the header and idle phases
// (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because
+1 -1
View File
@@ -220,7 +220,7 @@ func buildMinecraftServerFromApplyRequest(req applyRequest, namespace string) (*
}
memLim, ok := limits[corev1.ResourceMemory]
if !ok || memLim.IsZero() {
return nil, fmt.Errorf("internal error: refusing to create a server without a memory ceiling (§22)")
return nil, fmt.Errorf("internal error: refusing to create a server without a memory ceiling")
}
// ---- storage ----
+1 -1
View File
@@ -116,7 +116,7 @@ func newBackupID() string {
var b [16]byte
if _, err := rand.Read(b[:]); err != nil {
// crypto/rand failure is fatal and unrecoverable; a time-based fallback would
// be a weaker ID for no benefit. ponytail: panic is the honest failure here.
// be a weaker ID for no benefit. A panic is the honest failure here.
panic("felis backup: crypto/rand: " + err.Error())
}
return "bk-" + hex.EncodeToString(b[:])
+1 -1
View File
@@ -51,7 +51,7 @@ func resolveInternalAPI(ctx context.Context, cl client.Client, controlNamespace
}
token = string(sec.Data[naming.ServiceTokenSecretKey])
if token == "" {
return "", "", fmt.Errorf("Secret %s has no %s key", naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey)
return "", "", fmt.Errorf("secret %s has no %s key", naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey)
}
return fmt.Sprintf("http://%s:%d", ip, platform.APIInternalPort), token, nil
+4 -6
View File
@@ -571,12 +571,10 @@ type breakGlassResult struct {
backupStatus string
// Cloudflare-specific edge detail (set only when connectMethod is Cloudflare)
edgeConfigured bool
edgeAud string
edgeRoutedHosts []string
edgeConfigPath string
edgePanelHostname string
edgeAdminHostname string
edgeConfigured bool
edgeAud string
edgeRoutedHosts []string
edgeConfigPath string
}
type consoleMode string
+1 -1
View File
@@ -26,7 +26,7 @@ const defaultForwardingDataDir = "/data"
// volume of unknown ownership; 0666/0777 then let a non-root Paper rewrite the
// same files on boot.
//
// ponytail: relies on the initContainer running as root to write into a volume of
// This relies on the initContainer running as root to write into a volume of
// unknown ownership; that is how the operator schedules it. If that ever changes,
// give the server pod an fsGroup so the shared volume is group-writable instead.
const (
+71 -17
View File
@@ -29,7 +29,12 @@ import (
"flag"
"fmt"
"io"
"net"
"net/http"
"os"
"os/signal"
"syscall"
"time"
"felis.lolicon.best/internal/api"
"felis.lolicon.best/internal/config"
@@ -37,21 +42,27 @@ import (
// nanoStubRepo satisfies api.Repo but implements only the one method handleHasJoined calls.
// The reclaim username blacklist is a felis-api/DB concern; a nano host has no Postgres, so
// nothing is barred here. ponytail: a real blacklist would need the very DB nano exists to
// nothing is barred here. A real blacklist would need the very DB nano exists to
// avoid — YAGNI until a nano host grows a reclaim store.
type nanoStubRepo struct{ api.Repo }
func (nanoStubRepo) IsUsernameBlacklisted(context.Context, string) (bool, error) { return false, nil }
// nanoDefaultListen is loopback because hasJoined carries no auth token (Velocity speaks the
// vanilla sessionserver protocol), so a public bind is an open auth relay: anyone can point
// their proxy at it and spend this host's egress IP on Mojang. A same-host Velocity reaches
// 127.0.0.1; serving an off-host proxy is an explicit -listen opt-in.
const nanoDefaultListen = "127.0.0.1:8081"
// nanoLogURIMax is room for a real hasJoined query (a 16-character name, a 41-character
// serverId, an address) several times over.
const nanoLogURIMax = 256
func cmdNano(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("nano", flag.ContinueOnError)
fs.SetOutput(stderr)
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (reads [[auth_source]])")
// Loopback default: hasJoined carries no auth token (authlib speaks the vanilla
// sessionserver protocol), so a public bind is an open auth relay — anyone can point
// their proxy at it and spend this host's egress IP on Mojang. A same-host Velocity
// reaches 127.0.0.1; serving an off-host proxy is an explicit -listen opt-in.
listen := fs.String("listen", "127.0.0.1:8081", "listen address for the hasJoined endpoint")
listen := fs.String("listen", nanoDefaultListen, "listen address for the hasJoined endpoint")
if err := fs.Parse(args); err != nil {
return 2
}
@@ -61,24 +72,67 @@ func cmdNano(args []string, stdout, stderr io.Writer) int {
fmt.Fprintln(stderr, "felis nano:", err)
return 1
}
// [server] listen belongs to felis api. Someone moving nano off loopback naturally reaches
// for it, and without this line would get connection refused with no hint why.
if cfg.Server.Listen != "" {
fmt.Fprintf(stderr, "felis nano: [server] listen = %q is ignored; nano binds -listen (%s), which the installer sets from FELIS_NANO_LISTEN\n", cfg.Server.Listen, *listen)
}
handler := api.HasJoinedHandler(authSourcesFromConfig(cfg.AuthSources), nanoStubRepo{})
fmt.Fprintf(stderr, "felis nano: hasJoined multiplexer on %s — Mojang + %d third-party source(s)\n", *listen, len(cfg.AuthSources))
for i, s := range cfg.AuthSources {
fmt.Fprintf(stderr, " [%d] %s -> %s\n", i+1, s.Tag, s.URL)
}
// Log each request so a live login attempt is visible while testing against a real
// Velocity — "is authlib even reaching me?" is the first question during verification.
logged := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
fmt.Fprintf(stderr, "felis nano: %s %s\n", r.Method, r.RequestURI)
handler.ServeHTTP(w, r)
})
srv := newAPIServer(*listen, logged)
if err := srv.ListenAndServe(); err != nil {
ln, err := net.Listen("tcp", *listen)
if err != nil {
fmt.Fprintln(stderr, "felis nano:", err)
return 1
}
return 0
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
return serveNano(ctx, newAPIServer(*listen, nanoHandler(cfg.AuthSources, stderr)), ln, stderr)
}
// nanoDrainTimeout outlasts the source scan of any realistic list (each source is given
// five seconds) and stays well inside systemd's default 90-second stop timeout.
const nanoDrainTimeout = 30 * time.Second
// serveNano serves until ctx ends, then drains. A restart, the documented way to pick up a
// config edit, sends SIGTERM; without the drain a login already waiting on an upstream has
// its connection reset, and Velocity tells that player the auth servers are down.
func serveNano(ctx context.Context, srv *http.Server, ln net.Listener, stderr io.Writer) int {
errc := make(chan error, 1)
go func() { errc <- srv.Serve(ln) }()
select {
case err := <-errc:
fmt.Fprintln(stderr, "felis nano:", err)
return 1
case <-ctx.Done():
shutdownCtx, cancel := context.WithTimeout(context.Background(), nanoDrainTimeout)
defer cancel()
if err := srv.Shutdown(shutdownCtx); err != nil {
fmt.Fprintln(stderr, "felis nano: shutdown:", err)
return 1
}
return 0
}
}
// nanoHandler is what felis nano serves: the shared hasJoined handler, Mojang first, behind
// a request log.
func nanoHandler(sources []config.AuthSourceConfig, stderr io.Writer) http.Handler {
handler := api.HasJoinedHandler(authSourcesFromConfig(sources), nanoStubRepo{})
// Log each request so a live login attempt is visible while testing against a real
// Velocity — "is Velocity even reaching me?" is the first question during verification.
// The URI is the caller's text: quoted so a control or bidi character cannot rewrite the
// line and invalid UTF-8 cannot turn the journal entry into a blob, and capped so one
// request cannot write a megabyte of log.
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
uri := r.RequestURI
if len(uri) > nanoLogURIMax {
uri = uri[:nanoLogURIMax] + "..."
}
fmt.Fprintf(stderr, "felis nano: %s %q\n", r.Method, uri)
handler.ServeHTTP(w, r)
})
}
+139
View File
@@ -0,0 +1,139 @@
package main
import (
"bytes"
"context"
"io"
"net"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
"time"
"unicode/utf8"
"felis.lolicon.best/internal/api"
)
// [server] listen in a nano config reads like the bind address but is not one; nano must
// say so. The -listen value cannot be bound, so cmdNano returns right after loading.
func TestNanoWarnsThatServerListenIsIgnored(t *testing.T) {
cfg := filepath.Join(t.TempDir(), "felis.toml")
if err := os.WriteFile(cfg, []byte("[server]\nlisten = \"0.0.0.0:9999\"\n"), 0o600); err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
if rc := cmdNano([]string{"-config", cfg, "-listen", "127.0.0.1:-1"}, io.Discard, &stderr); rc != 1 {
t.Fatalf("cmdNano = %d, want 1 from the unbindable -listen", rc)
}
if !strings.Contains(stderr.String(), `listen = "0.0.0.0:9999" is ignored`) {
t.Fatalf("stderr %q should say the configured listen is ignored", stderr.String())
}
// With no [server] table at all there is nothing to warn about.
if err := os.WriteFile(cfg, nil, 0o600); err != nil {
t.Fatal(err)
}
stderr.Reset()
_ = cmdNano([]string{"-config", cfg, "-listen", "127.0.0.1:-1"}, io.Discard, &stderr)
if strings.Contains(stderr.String(), "is ignored") {
t.Fatalf("stderr %q warns about a listen the operator never set", stderr.String())
}
}
// A stop signal that lands while a login is waiting on an upstream must let that login
// finish: the request is answered, and serveNano returns only afterwards.
func TestNanoDrainsInFlightLoginOnShutdown(t *testing.T) {
entered, release := make(chan struct{}), make(chan struct{})
srv := newAPIServer("", http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
close(entered)
<-release
w.WriteHeader(http.StatusNoContent)
}))
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
ctx, stop := context.WithCancel(context.Background())
done := make(chan int, 1)
go func() { done <- serveNano(ctx, srv, ln, io.Discard) }()
got := make(chan int, 1)
go func() {
resp, err := http.Get("http://" + ln.Addr().String() + "/session/minecraft/hasJoined")
if err != nil {
got <- -1
return
}
resp.Body.Close()
got <- resp.StatusCode
}()
<-entered
stop()
select {
case <-done:
t.Fatal("serveNano returned while a login was still in flight")
case <-time.After(200 * time.Millisecond):
}
close(release)
if code := <-got; code != http.StatusNoContent {
t.Fatalf("in-flight login got %d, want its answer (204)", code)
}
if rc := <-done; rc != 0 {
t.Fatalf("serveNano = %d after a clean drain, want 0", rc)
}
}
// The nano delivery path: the shared handler behind nano's stub store must admit a login its
// source validated. nanoStubRepo implements only the bar-list lookup, so a new store call in
// handleHasJoined would reach its nil embedded Repo and panic here, while the full-api tests,
// which use a complete fake store, stay green.
func TestNanoAdmitsAValidatedLogin(t *testing.T) {
const id = "069a79f444e94726a5befca90e38aaf5"
ygg := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = io.WriteString(w, `{"id":"`+id+`","name":"Notch"}`)
}))
defer ygg.Close()
h := api.HasJoinedHandler([]api.AuthSource{{Tag: "mojang", URL: ygg.URL, Identity: true}}, nanoStubRepo{})
w := httptest.NewRecorder()
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, "/session/minecraft/hasJoined?username=Notch&serverId=abc", nil))
if w.Code != http.StatusOK || !strings.Contains(w.Body.String(), id) {
t.Fatalf("code = %d body = %q, want the validated profile", w.Code, w.Body.String())
}
}
// An unauthenticated relay on a public address spends this host's Mojang rate limit for
// anyone who finds it, so the default bind has to stay loopback.
func TestNanoListensOnLoopbackByDefault(t *testing.T) {
host, _, err := net.SplitHostPort(nanoDefaultListen)
if ip := net.ParseIP(host); err != nil || ip == nil || !ip.IsLoopback() {
t.Fatalf("default -listen %q is not a loopback address", nanoDefaultListen)
}
}
// The request log prints text the caller chose. A bidi override must not reorder the line,
// an invalid byte must not make journald store the entry as a blob, and a huge query must
// not become a huge log line. serverId is left out so the handler answers without asking
// any source.
func TestNanoRequestLogIsQuotedAndCapped(t *testing.T) {
const rlo = rune(0x202e) // RIGHT-TO-LEFT OVERRIDE
var log bytes.Buffer
h := nanoHandler(nil, &log)
target := "/session/minecraft/hasJoined?username=" + string(rlo) + "evil" + string([]byte{0x9b}) + "31m" + strings.Repeat("a", 4096)
w := httptest.NewRecorder()
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, target, nil))
line := log.String()
if strings.ContainsRune(line, rlo) || !utf8.ValidString(line) {
t.Fatalf("raw caller bytes reached the log: %q", line)
}
if escaped := strings.Trim(strconv.QuoteRune(rlo), "'"); !strings.Contains(line, escaped) {
t.Fatalf("log line %q should show the override escaped as %s", line, escaped)
}
if len(line) > 2*nanoLogURIMax {
t.Fatalf("log line is %d bytes for a %d-byte URI; want it capped", len(line), len(target))
}
}
+10 -5
View File
@@ -225,16 +225,21 @@ func provisionSystemServers(ctx context.Context, cfg *config.Config, out io.Writ
// renamed it must replicate the Secret by hand.
controlNS := platform.DefaultControlNamespace
apiBaseURL := platform.InternalAPIBaseURL(controlNS)
// Both Secrets must land in the minecraft namespace before the pods that mount
// them are created: the service token (login authenticates to felis-api with it)
// and the Velocity forwarding secret (every backend verifies the proxy's signed
// handshake with it — without it the login gate would derive an OFFLINE UUID and
// the Owner would bind the wrong Minecraft identity).
// These Secrets must land in the minecraft namespace before the pods that
// mount them are created: the service token (login authenticates to felis-api
// with it), the Velocity forwarding secret (every backend verifies the proxy's
// signed handshake with it — without it the login gate would derive an OFFLINE
// UUID and the Owner would bind the wrong Minecraft identity), and felis-config
// (the on-demand BACKUP Job runs in the minecraft namespace and mounts it to
// self-record its world_backups row; without the replica the Job's volume
// mount fails and every backup request strands in the cluster).
secretOutcomes := []systemServerOutcome{
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token"),
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
naming.ForwardingSecretName, naming.ForwardingSecretKey, "forwarding-secret"),
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
"felis-config", "felis.toml", "config"),
}
outcomes := ensureSystemServers(ctx, cl, cfg.K8s.Namespace, cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage, apiBaseURL, cfg.Server.RootDomain, defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname))
outcomes = append(secretOutcomes, outcomes...)
+5 -5
View File
@@ -497,7 +497,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
// human would have added and a memory bump off the built-in default.
stale := func() *v1alpha1.MinecraftServer {
ms, err := loginSystemServer("felis-limbo:demo", "minecraft",
"http://old.internal:8081", "159.223.32.51.nip.io", "console.159.223.32.51.nip.io")
"http://old.internal:8081", "203.0.113.10.nip.io", "console.203.0.113.10.nip.io")
if err != nil {
t.Fatalf("build stale login server: %v", err)
}
@@ -509,7 +509,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
run := func(cl client.Client) []systemServerOutcome {
return ensureSystemServers(ctx, cl, "minecraft", "felis-limbo:demo", "felis-lobby:demo",
"http://felis-api-internal.felis.svc.cluster.local:8081",
"mc.flyemoji.network", "console.mc.flyemoji.network")
"mc.example.net", "console.mc.example.net")
}
envOf := func(t *testing.T, cl client.Client) map[string]string {
@@ -537,11 +537,11 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
}
}
env := envOf(t, cl)
if env[envPanelHostname] != "console.mc.flyemoji.network" {
if env[envPanelHostname] != "console.mc.example.net" {
t.Errorf("%s = %q — players are still being sent to the old console",
envPanelHostname, env[envPanelHostname])
}
if env[envRootDomain] != "mc.flyemoji.network" {
if env[envRootDomain] != "mc.example.net" {
t.Errorf("%s = %q, want the new root domain", envRootDomain, env[envRootDomain])
}
})
@@ -568,7 +568,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
t.Run("reports no refresh when config already matches", func(t *testing.T) {
fresh, err := loginSystemServer("felis-limbo:demo", "minecraft",
"http://felis-api-internal.felis.svc.cluster.local:8081",
"mc.flyemoji.network", "console.mc.flyemoji.network")
"mc.example.net", "console.mc.example.net")
if err != nil {
t.Fatalf("build fresh login server: %v", err)
}
+233 -58
View File
@@ -26,10 +26,18 @@
# The script is idempotent: re-running it converges rather than duplicating, and
# generated secrets are persisted to /etc/felis/secrets.env so reruns reuse them.
#
# Tunables (export before running to override the demo defaults):
# FELIS_INSTALL_MODE full|nano — skip the prompt (default: ask on a tty, else full)
# FELIS_NANO_LISTEN listen addr for `felis nano` (default: 127.0.0.1:8081 — loopback
# only; set a private-network IP to serve an off-host proxy)
# Tunables override the demo defaults. sudo resets the environment, so a variable exported
# before `curl ... | sudo bash` never arrives. Name it on the sudo line, or keep it with -E:
# curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo FELIS_INSTALL_MODE=nano FELIS_NANO_LISTEN=10.0.0.5:8081 bash
# export FELIS_INSTALL_MODE=nano; curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo -E bash
# FELIS_INSTALL_MODE full|nano — skip the prompt (default: ask on a tty, else full; nano
# instead on a host that runs felis-nano and no full install)
# FELIS_NANO_LISTEN listen addr for `felis nano` (default: the address an installed
# felis-nano already uses, else 127.0.0.1:8081 — loopback only; set a
# private-network IP to serve an off-host proxy)
# FELIS_NANO_PROXY_CIDR the proxy allowed to reach a non-loopback nano bind, as an address
# with a prefix length (for example 10.0.0.7/32). firewalld opens the
# port to that source only; unset, it opens nothing
# FELIS_LEGACY_FORWARDING_SERVERS comma-separated backends that receive their identity
# through the handshake address instead of modern forwarding
# (default: legacy18). Read once at Velocity start, so changing it
@@ -39,6 +47,8 @@
# FELIS_VELOCITY_FORK_JAR_SHA256 expected sha256 of that jar. REQUIRED whenever the jar
# above is set; the install refuses on a mismatch.
# FELIS_GO_VERSION Go toolchain used to build the nano binary (default: 1.26.4)
# FELIS_GO_SHA256 sha256 of that version's linux tarball for this host's architecture.
# REQUIRED for a non-default FELIS_GO_VERSION; the default's is pinned.
# FELIS_REPO_URL git URL to build from (raw script mode only)
# FELIS_VERSION_BOOTSTRAP release|dev — which version to install (default: release).
# release DOWNLOADS the prebuilt felis binary published for the newest
@@ -97,18 +107,28 @@ FELIS_IMAGE="${FELIS_IMAGE:-felis:demo}"
FELIS_EGRESS_MODE="${FELIS_EGRESS_MODE:-nodeport}"
FELIS_PANEL_NODEPORT="${FELIS_PANEL_NODEPORT:-30443}"
INSTALL_MODE="${FELIS_INSTALL_MODE:-}"
# Loopback by default: hasJoined is an unauthenticated endpoint by protocol (authlib
# Loopback by default: hasJoined is an unauthenticated endpoint by protocol (Velocity
# sends no token), so a public bind is a free auth relay — anyone can point their own
# proxy at it and spend YOUR egress IP on Mojang, until Mojang rate-limits you and your
# own players stop getting in. Same-host Velocity reaches 127.0.0.1 fine; a proxy on
# another machine must opt in explicitly with FELIS_NANO_LISTEN=<private-ip>:8081.
FELIS_NANO_LISTEN="${FELIS_NANO_LISTEN:-127.0.0.1:8081}"
# Left empty here: resolve_nano_listen applies that default only after an existing unit's
# address has had its say.
FELIS_NANO_LISTEN="${FELIS_NANO_LISTEN:-}"
FELIS_NANO_PROXY_CIDR="${FELIS_NANO_PROXY_CIDR:-}"
# Backends that take their forwarded identity through the handshake address instead of
# proxy-wide modern forwarding. See write_velocity_service for why a protocol-47 backend
# needs this. Overridable because adding a second 1.8 backend otherwise means editing this
# script; it is still a restart-time list, not one that follows the CRs.
FELIS_LEGACY_FORWARDING_SERVERS="${FELIS_LEGACY_FORWARDING_SERVERS:-legacy18}"
FELIS_GO_VERSION="${FELIS_GO_VERSION:-1.26.4}"
# The Go tarball is unpacked and run as root, so the default version is pinned by the sha256
# go.dev/dl publishes for each architecture install_go_toolchain handles. Move all three
# together; any other FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256.
GO_PINNED_VERSION="1.26.4"
GO_PINNED_SHA256_AMD64="1153d3d50e0ac764b447adfe05c2bcf08e889d42a02e0fe0259bd47f6733ad7f"
GO_PINNED_SHA256_ARM64="ef758ae7c6cf9267c9c0ef080b8965f453d89ab2d25d9eb22de4405925238768"
FELIS_GO_VERSION="${FELIS_GO_VERSION:-$GO_PINNED_VERSION}"
FELIS_GO_SHA256="${FELIS_GO_SHA256:-}"
PKG_LOCK_TIMEOUT="${PKG_LOCK_TIMEOUT:-${APT_LOCK_TIMEOUT:-900}}"
APT_LOCK_TIMEOUT="${APT_LOCK_TIMEOUT:-$PKG_LOCK_TIMEOUT}"
@@ -180,7 +200,9 @@ PANEL_TLS_CERT="${STATE_DIR}/panel-tls.crt"
PANEL_TLS_KEY="${STATE_DIR}/panel-tls.key"
SRC_DIR="/opt/felis/src"
HOST_BIN="/usr/local/bin/felis"
GOROOT_DIR="/usr/local/go"
# Felis's own build toolchain, not /usr/local/go: install_go_toolchain replaces whatever
# version sits here, and an operator's Go at the conventional path is not ours to swap.
GOROOT_DIR="/opt/felis/go"
NANO_SERVICE="/etc/systemd/system/felis-nano.service"
VELOCITY_DIR="/opt/felis/velocity"
VELOCITY_USER="felis-velocity"
@@ -433,10 +455,39 @@ validate_nodeport() {
fi
}
# A value without a usable port would reach the firewall and the summary as-is: 8081 opens
# port 8081 while nano binds nothing, 127.0.0.1 prints http://127.0.0.1:127.0.0.1/..., and
# the unit crash-loops either way.
validate_listen() {
local name="$1" value="$2" port="${2##*:}" host="${2%:*}"
case "$value" in *:*) ;; *) port="" ;; esac
case "$port" in
''|*[!0-9]*) die "${name} must be host:port (for example 127.0.0.1:8081), got: ${value}" ;;
esac
[ "$port" -ge 1 ] && [ "$port" -le 65535 ] || die "${name} port must be 1-65535, got: ${value}"
# Go takes a colon in the host only inside brackets; ::1:8081 would crash-loop the unit.
case "$host" in
*:*) case "$host" in "["*"]") ;; *) die "${name} needs an IPv6 host in brackets (for example [::1]:8081), got: ${value}" ;; esac ;;
esac
}
# The value lands inside a firewalld rich rule, so anything but address characters and one
# prefix length is refused here rather than handed to firewall-cmd.
validate_cidr() {
case "$2" in
"") return 0 ;;
*[!0-9A-Fa-f.:/]*|*/*/*|*/|/*) ;;
*/[0-9]*) return 0 ;;
esac
die "$1 must be an address with a prefix length (for example 10.0.0.7/32 or fd00::7/128), got: $2"
}
validate_settings() {
validate_timeout PKG_LOCK_TIMEOUT "$PKG_LOCK_TIMEOUT"
validate_timeout APT_LOCK_TIMEOUT "$APT_LOCK_TIMEOUT"
validate_nodeport FELIS_PANEL_NODEPORT "$FELIS_PANEL_NODEPORT"
validate_listen FELIS_NANO_LISTEN "$FELIS_NANO_LISTEN"
validate_cidr FELIS_NANO_PROXY_CIDR "$FELIS_NANO_PROXY_CIDR"
}
# ---------------------------------------------------------------------------
@@ -959,12 +1010,16 @@ download_release_binary() {
# approve that follows a successful auth, so it outlives the install after all. An empty
# value is git's documented list reset. Verified on Fedora: with store configured the single
# -c form persists the token to disk, the reset form writes nothing.
#
# GIT_TERMINAL_PROMPT=0 on both arms: GitHub answers a private repo with no or a bad token
# by asking for credentials, and git would put that prompt on /dev/tty, where a piped
# install sits waiting instead of failing with the FELIS_GITHUB_TOKEN hint.
git_auth() {
if [ -n "$FELIS_GITHUB_TOKEN" ]; then
git -c 'credential.helper=' \
GIT_TERMINAL_PROMPT=0 git -c 'credential.helper=' \
-c 'credential.helper=!f() { printf "username=x-access-token\npassword=%s\n" "$FELIS_GITHUB_TOKEN"; }; f' "$@"
else
git "$@"
GIT_TERMINAL_PROMPT=0 git "$@"
fi
}
@@ -1041,7 +1096,8 @@ fetch_source() {
resolve_install_ref
if [ -d "${SRC_DIR}/.git" ]; then
log "updating source in ${SRC_DIR}"
git_auth -C "$SRC_DIR" fetch --depth 1 origin "$FELIS_REF"
git_auth -C "$SRC_DIR" fetch --depth 1 origin "$FELIS_REF" \
|| die "could not fetch ${FELIS_REF} from ${FELIS_REPO_URL}; if the repository is private, set FELIS_GITHUB_TOKEN to a token with read access to it"
git -C "$SRC_DIR" checkout -f FETCH_HEAD
else
log "cloning ${FELIS_REPO_URL} (${FELIS_REF})"
@@ -1053,7 +1109,7 @@ fetch_source() {
git_auth clone --depth 1 --branch "$FELIS_REF" "$FELIS_REPO_URL" "$SRC_DIR" 2>/dev/null \
|| { git_auth clone "$FELIS_REPO_URL" "$SRC_DIR" \
&& git_auth -C "$SRC_DIR" checkout -f "$FELIS_REF"; } \
|| die "could not check out ${FELIS_REF} from ${FELIS_REPO_URL}"
|| die "could not check out ${FELIS_REF} from ${FELIS_REPO_URL}; if the repository is private, set FELIS_GITHUB_TOKEN to a token with read access to it"
fi
stamp_version
ok "source ready at ${SRC_DIR}"
@@ -1200,7 +1256,7 @@ game_stack_source() {
# MC_VERSION is read off Limbo's CI artifact name (Limbo-<limbo-ver>-<mc-ver>.jar), which
# is the only place the pairing is published.
resolve_game_jars() {
local ci="https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild" meta file base rest
local ci="https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild" meta file base rest paper
log "resolving the newest LOOHP/Limbo CI build"
# Fetch first, filter second: `curl | grep | head` dies of SIGPIPE under `set -o pipefail`
# the moment head closes the pipe early. Same shape everywhere below.
@@ -1221,8 +1277,10 @@ resolve_game_jars() {
# PaperMC Fill v3. The old api.papermc.io v2 has returned HTTP 410 since 2026-07-01 and
# is never coming back; Fill wants a descriptive User-Agent.
log "resolving the newest Paper ${MC_VERSION} build"
PAPER_JAR_URL="$(papermc_latest_jar paper "$MC_VERSION")" \
|| die "could not resolve a Paper build for Minecraft ${MC_VERSION} (the login gate pins this protocol; the build likely exists — Fill upstream is down or flapping)"
paper="$(papermc_latest_jar paper "$MC_VERSION")" \
|| die "could not resolve a Paper build for Minecraft ${MC_VERSION} (the login gate pins this protocol; the build likely exists — Fill upstream is down, flapping, or no longer content-addressed)"
PAPER_JAR_URL="${paper% *}"
PAPER_JAR_SHA256="${paper##* }"
# LuckPerms is not version-matched to MC_VERSION the way Paper is: it ships one
# current Bukkit build that supports the whole supported Minecraft range, so there is
# no per-version endpoint to ask.
@@ -1250,20 +1308,29 @@ luckperms_latest_jar() {
printf '%s\n' "$url"
}
# papermc_latest_jar prints the download URL of the newest build of <project> <version>.
# papermc_latest_jar prints "<url> <sha256>" for the newest build of <project> <version>.
# --retry rides out Fill's transient gateway errors (502/503/504 are in curl's retry
# set): a single blip must not abort the whole bootstrap claiming the build is missing.
# Plain --retry only, deliberately: --retry-connrefused needs curl 7.52+, which the yum
# (el7) path does not have, and it would only add ECONNREFUSED to an already-covered set.
# The digest is not fished out of the JSON separately: Fill's download URLs are
# content-addressed (/v1/objects/<sha256>/<name>.jar), so the path segment names the
# bytes the URL serves and both halves come from the same grep of the same response. A
# URL without that shape fails the resolve rather than waving the download through
# unchecked.
papermc_latest_jar() {
local project="$1" version="$2" json urls url
local project="$1" version="$2" json urls url sha
json="$(curl -fsSL --retry 5 --retry-delay 2 \
-A "felis-bootstrap (+https://github.com/MliroLirrorsIngenuity/Felis)" \
"https://fill.papermc.io/v3/projects/${project}/versions/${version}/builds/latest")" || return 1
urls="$(printf '%s' "$json" | grep -o 'https://fill-data\.papermc\.io/[^"]*\.jar' || true)"
url="${urls%%$'\n'*}"
[ -n "$url" ] || return 1
printf '%s\n' "$url"
sha="${url#*/objects/}"
sha="${sha%%/*}"
case "$sha" in *[!0-9a-f]*|"") return 1 ;; esac
[ "${#sha}" -eq 64 ] || return 1
printf '%s %s\n' "$url" "$sha"
}
build_game_stack() {
@@ -1281,6 +1348,7 @@ build_game_stack() {
log "building ${FELIS_LOBBY_IMAGE} (Paper ${MC_VERSION} + felis-paper /menu + LuckPerms)"
docker build -f "${GAME_STACK_DIR}/deploy/lobby/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
--build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
--build-arg LUCKPERMS_JAR_URL="$LUCKPERMS_JAR_URL" \
-t "$FELIS_LOBBY_IMAGE" "$GAME_STACK_DIR"
@@ -1289,6 +1357,7 @@ build_game_stack() {
log "building ${FELIS_PAPER_IMAGE} (plain Paper ${MC_VERSION}, forwarding via the operator initContainer)"
docker build -f "${GAME_STACK_DIR}/deploy/paper/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
--build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
-t "$FELIS_PAPER_IMAGE" "$GAME_STACK_DIR"
local img
@@ -1486,7 +1555,7 @@ install_jre() {
install_velocity() {
install_jre
local url tmp have want
local url tmp have want resolved
prepare_velocity_layout
if [ -n "$FELIS_VELOCITY_FORK_JAR" ]; then
[ -f "$FELIS_VELOCITY_FORK_JAR" ] \
@@ -1514,12 +1583,20 @@ install_velocity() {
atomic_install_file "$FELIS_VELOCITY_FORK_JAR" "${VELOCITY_DIR}/velocity.jar" 0644 root root
else
log "resolving the newest Velocity ${FELIS_VELOCITY_VERSION} build"
url="$(papermc_latest_jar velocity "$FELIS_VELOCITY_VERSION")" \
resolved="$(papermc_latest_jar velocity "$FELIS_VELOCITY_VERSION")" \
|| die "no Velocity build for ${FELIS_VELOCITY_VERSION} (override with FELIS_VELOCITY_VERSION)"
url="${resolved% *}"
want="${resolved##* }"
log "downloading Velocity ${FELIS_VELOCITY_VERSION}"
tmp="$(mktemp "${VELOCITY_DIR}/.velocity.jar.XXXXXX")"
remember_temp "$tmp"
curl -fsSL "$url" -o "$tmp" || die "failed to download Velocity: ${url}"
# The same gate the Via plugins and the fork jar pass: this jar is the proxy every
# player connects through, and Fill already promised its digest in the URL — a
# truncated or tampered download becomes a refusal here, not a proxy that won't boot.
have="$(sha256sum <"$tmp" | cut -d' ' -f1)"
[ "$have" = "$want" ] \
|| die "Velocity ${FELIS_VELOCITY_VERSION} checksum mismatch: got ${have}, expected ${want}"
atomic_install_file "$tmp" "${VELOCITY_DIR}/velocity.jar" 0644 root root
fi
@@ -1624,7 +1701,7 @@ EOF
install_velocity_service() {
local api_ip
# Point Velocity's authlib (mojang.sessionserver) at the felis-api hasJoined multiplexer so a
# Point Velocity (-Dmojang.sessionserver) at the felis-api hasJoined multiplexer so a
# full install federates Mojang + the configured [[auth_source]] set (LittleSkin by default)
# out of the box — not just the standalone `felis nano`. felis-api enforces the reclaim
# blacklist on this route; a loopback nano would bypass it.
@@ -1914,15 +1991,39 @@ persisted_smtp_block() {
printf '%s' "$SMTP_BLOCK"
}
# persisted_auth_source_blocks echoes the [[auth_source]] tables an earlier run left
# behind, or the LittleSkin default when there is no earlier felis.toml at all. The
# list is the operator's: it is the only way to add or drop a Yggdrasil root on a full
# install, and nothing in this script's inputs derives it. Without the carry-forward a
# re-run would put LittleSkin back after the operator removed it and silently drop any
# root they added. An earlier file with no tables stays that way — that is a Mojang-only
# server, not a missing value. Same first-readable-file rule as persisted_smtp_block.
persisted_auth_source_blocks() {
local f
for f in "${STATE_DIR}/felis.host.toml" "${STATE_DIR}/felis.pod.toml"; do
[ -r "$f" ] || continue
# Every [[auth_source]] table, up to (not including) the next other section header.
# TOML also accepts [[ auth_source ]] and a quoted key; a header this does not
# recognise would silently drop that table.
awk '/^[[:space:]]*\[/ { f = /^[[:space:]]*\[\[[[:space:]]*["\047]?auth_source["\047]?[[:space:]]*\]\]/ }
f { print }' "$f"
return 0
done
printf '%s\n' '[[auth_source]]' 'tag = "littleskin"' 'prefix = "LS"' \
'url = "https://littleskin.cn/api/yggdrasil/sessionserver/session/minecraft/hasJoined"'
}
write_felis_toml() {
local target="$1" db_host="$2" smtp_block
local target="$1" db_host="$2" smtp_block auth_source_blocks
smtp_block="$(persisted_smtp_block)"
if [ -n "$smtp_block" ]; then
log "carrying forward the configured [smtp] relay"
smtp_block="${smtp_block}"$'\n' # keep a blank line before the next section
fi
auth_source_blocks="$(persisted_auth_source_blocks)"
cat > "$target" <<EOF
# Generated by deploy/bootstrap.sh — do not edit by hand; rerun the installer.
# Generated by deploy/bootstrap.sh; rerun the installer to regenerate. Hand edits are
# overwritten, except [smtp] and [[auth_source]], which carry forward.
[server]
listen = "0.0.0.0:8080"
root_domain = "${FELIS_ROOT_DOMAIN}"
@@ -1955,12 +2056,14 @@ panel_hostname = "console.${FELIS_ROOT_DOMAIN}"
${smtp_block}
# Third-party Yggdrasil sources federated by the hasJoined multiplexer. Mojang is
# always the code-owned identity anchor (premium-first), prepended in Go; sources here
# append as namespace-rewritten guests. Shipping LittleSkin by default lets Mojang and
# LittleSkin both log in out of the box. Delete this block for a Mojang-only server.
[[auth_source]]
tag = "littleskin"
prefix = "LS"
url = "https://littleskin.cn/api/yggdrasil/sessionserver/session/minecraft/hasJoined"
# append as namespace-rewritten guests. A fresh install federates LittleSkin. Edit the
# list in ${STATE_DIR}/felis.host.toml and rerun the installer; re-runs keep it as it
# is, and with no [[auth_source]] at all the server is Mojang-only.
# Order is trust: the first source that answers 200 wins, so list the most trusted roots
# first, and remove a compromised root rather than just moving it down.
# A tag is permanent: it is hashed into every player UUID of its source, so changing it
# (even its case) gives all of them new UUIDs and orphans their data, links and bans.
${auth_source_blocks}
EOF
}
@@ -2147,18 +2250,32 @@ summary() {
# prompt (or FELIS_INSTALL_MODE=nano).
# ---------------------------------------------------------------------------
prompt_install_mode() {
# felis setup carries on to the Owner and edge setup, which needs the control plane, so
# a nano install under it could only end in a setup error.
if bootstrap_from_tui; then
[ "$INSTALL_MODE" != nano ] || die "felis setup installs the full control plane; for Felis-nano run deploy/bootstrap.sh with FELIS_INSTALL_MODE=nano"
INSTALL_MODE="full"
log "install mode: full (felis setup)"
return 0
fi
case "$INSTALL_MODE" in
full|nano) log "install mode: ${INSTALL_MODE} (from FELIS_INSTALL_MODE)"; return 0 ;;
"") ;;
*) die "FELIS_INSTALL_MODE must be 'full' or 'nano', got: ${INSTALL_MODE}" ;;
esac
# A felis-nano unit with no full install beside it makes this re-run a nano update;
# defaulting to full there would put k3s and Postgres on a host that asked for neither.
local def=full n=1 reply
if [ -e "$NANO_SERVICE" ] && [ ! -e "$BOOTSTRAP_DONE" ]; then def=nano n=2; fi
# No override: ask on the controlling terminal. Under `curl | sudo bash` stdin
# is the script, so we must read /dev/tty, not stdin. No tty (CI/cloud-init) →
# default to a full install.
if [ ! -r /dev/tty ]; then
INSTALL_MODE="full"
log "no terminal for a prompt; defaulting to a full Felis install (set FELIS_INSTALL_MODE=nano to override)"
# take the default. Open it to find out: /dev/tty is mode 0666 on every Linux host, so
# `-r` passes even when there is no controlling terminal and only the open fails.
if ! (: </dev/tty) 2>/dev/null; then
INSTALL_MODE="$def"
log "no terminal for a prompt; defaulting to a ${def} install (set FELIS_INSTALL_MODE=full or nano to override)"
return 0
fi
@@ -2166,12 +2283,12 @@ prompt_install_mode() {
printf 'What do you want to install on this host?\n'
printf ' [1] Felis — full control plane (k3s + Postgres + panel; orchestrates Minecraft servers)\n'
printf ' [2] Felis-nano — auth multiplexer only (federates Mojang + third-party Yggdrasil; no k3s/DB)\n'
local reply
while :; do
printf 'Choose [1/2] (default 1): '
printf 'Choose [1/2] (default %s): ' "$n"
IFS= read -r reply </dev/tty || reply=""
case "$reply" in
""|1|full|Felis|felis) INSTALL_MODE="full"; break ;;
"") INSTALL_MODE="$def"; break ;;
1|full|Felis|felis) INSTALL_MODE="full"; break ;;
2|nano|felis-nano|Felis-nano) INSTALL_MODE="nano"; break ;;
*) printf 'Please enter 1 or 2.\n' ;;
esac
@@ -2180,25 +2297,34 @@ prompt_install_mode() {
}
install_go_toolchain() {
local arch tarball url
local arch tarball url tmp want have
if [ -x "${GOROOT_DIR}/bin/go" ] && "${GOROOT_DIR}/bin/go" version | grep -q "go${FELIS_GO_VERSION} "; then
ok "go ${FELIS_GO_VERSION} already installed at ${GOROOT_DIR}"
return 0
fi
case "$(uname -m)" in
x86_64|amd64) arch="amd64" ;;
aarch64|arm64) arch="arm64" ;;
x86_64|amd64) arch="amd64"; want="$GO_PINNED_SHA256_AMD64" ;;
aarch64|arm64) arch="arm64"; want="$GO_PINNED_SHA256_ARM64" ;;
*) die "no Go toolchain build for architecture $(uname -m); set FELIS_GO_VERSION or pre-stage ${GOROOT_DIR}" ;;
esac
[ "$FELIS_GO_VERSION" = "$GO_PINNED_VERSION" ] || want="$FELIS_GO_SHA256"
[ -n "$want" ] || die "no pinned sha256 for Go ${FELIS_GO_VERSION}; set FELIS_GO_SHA256 to the linux-${arch} digest https://go.dev/dl/ lists for it, or pre-stage ${GOROOT_DIR}"
tarball="go${FELIS_GO_VERSION}.linux-${arch}.tar.gz"
url="https://go.dev/dl/${tarball}"
log "installing Go ${FELIS_GO_VERSION} (${arch}) to ${GOROOT_DIR}"
curl -fsSL "$url" -o "/tmp/${tarball}" || die "failed to download the Go toolchain: ${url}"
# A private directory, not a fixed /tmp name another local user could have planted first.
tmp="$(mktemp -d)"
remember_temp "$tmp"
curl -fsSL "$url" -o "${tmp}/${tarball}" || die "failed to download the Go toolchain: ${url}"
# Checked before the old toolchain is removed, so a refusal leaves the host as it was.
# Hash stdin, never the path — same reason as install_via_plugins.
have="$(sha256sum <"${tmp}/${tarball}" | cut -d' ' -f1)"
[ "$have" = "$want" ] || die "Go ${FELIS_GO_VERSION} (${arch}) checksum mismatch: got ${have}, expected ${want}"
rm -rf "$GOROOT_DIR"
tar -C "$(dirname "$GOROOT_DIR")" -xzf "/tmp/${tarball}" || die "failed to unpack ${tarball}"
rm -f "/tmp/${tarball}"
mkdir -p "$(dirname "$GOROOT_DIR")"
tar -C "$(dirname "$GOROOT_DIR")" -xzf "${tmp}/${tarball}" || die "failed to unpack ${tarball}"
ok "go toolchain at ${GOROOT_DIR}/bin/go"
}
@@ -2233,10 +2359,6 @@ build_nano_binary() {
}
acquire_nano_binary() {
if bootstrap_from_tui; then
install_embedded_binary
return 0
fi
# The release channel takes the same prebuilt binary the control plane does. This is the
# biggest win on this path: a host that only wants the auth multiplexer stops needing a Go
# toolchain and a checkout at all.
@@ -2258,7 +2380,16 @@ acquire_nano_binary() {
write_nano_config() {
local target="${STATE_DIR}/felis.toml"
mkdir -p "$STATE_DIR"
# The unit is a DynamicUser, so it can read felis.toml only if it can search this
# directory. The mode is explicit because a hardened root umask (027) would leave it 0750,
# which is what older installers did to nano-only hosts. Only the full install keeps its
# own 0700: that directory holds secrets, and install_nano_service reports the lockout
# rather than this widening it.
if [ ! -d "$STATE_DIR" ]; then
mkdir -p -m 0755 "$STATE_DIR"
elif [ ! -e "$SECRETS_ENV" ] && [ ! -e "$BOOTSTRAP_DONE" ]; then
chmod 0755 "$STATE_DIR"
fi
if [ -e "$target" ]; then
ok "config already present at ${target}; leaving it (edit it to add [[auth_source]] roots)"
return 0
@@ -2269,6 +2400,11 @@ write_nano_config() {
# Add each third-party Yggdrasil root below (priority = order). url is the FULL
# hasJoined endpoint. After editing: sudo systemctl restart felis-nano
#
# Order is trust: the first source that answers 200 wins, so list the most trusted roots
# first, and remove a compromised root rather than just moving it down.
# tag is permanent: it is hashed into every player UUID of its source, so changing it
# (even its case) gives all of them new UUIDs and orphans their data, links and bans.
#
# prefix is required, 1-4 letters/digits, unique per source. A player of this source
# whose name belongs to a Mojang account joins as PREFIX_name (LS_steve) instead —
# otherwise the proxy, which keys its player list on the NAME, refuses to have both
@@ -2284,9 +2420,20 @@ EOF
ok "wrote nano config template ${target} (edit it to add your Yggdrasil sources)"
}
# resolve_nano_listen settles FELIS_NANO_LISTEN: the operator's value, else the address the
# installed felis-nano unit listens on, else loopback. Re-running this script is how a nano
# host updates, and without the middle step that re-run moved an off-host proxy's endpoint
# back to 127.0.0.1, so every login through it failed.
resolve_nano_listen() {
if [ -z "$FELIS_NANO_LISTEN" ] && [ -r "$NANO_SERVICE" ]; then
FELIS_NANO_LISTEN="$(sed -n 's/^ExecStart=.* -listen \([^ ]*\).*$/\1/p' "$NANO_SERVICE")"
fi
FELIS_NANO_LISTEN="${FELIS_NANO_LISTEN:-127.0.0.1:8081}"
}
nano_listen_is_loopback() {
case "${FELIS_NANO_LISTEN%:*}" in
127.*|localhost|::1|"[::1]") return 0 ;;
127.*|localhost|"[::1]") return 0 ;;
*) return 1 ;;
esac
}
@@ -2300,9 +2447,21 @@ configure_nano_firewall() {
fi
command -v firewall-cmd >/dev/null 2>&1 || return 0
systemctl is-active --quiet firewalld || return 0
local port="${FELIS_NANO_LISTEN##*:}"
log "opening firewalld port ${port}/tcp for felis-nano"
firewall-cmd --permanent --add-port="${port}/tcp"
local port="${FELIS_NANO_LISTEN##*:}" family=ipv4
# hasJoined takes no token, so the port is opened to the proxy alone. Earlier installers
# opened it to every source, and a re-run must not leave that behind. A rule for a previous
# FELIS_NANO_PROXY_CIDR is not tracked; it stays until removed by hand.
if firewall-cmd --permanent --query-port="${port}/tcp" >/dev/null 2>&1; then
log "closing firewalld port ${port}/tcp, which an earlier install opened to every source"
firewall-cmd --permanent --remove-port="${port}/tcp"
fi
if [ -n "$FELIS_NANO_PROXY_CIDR" ]; then
case "$FELIS_NANO_PROXY_CIDR" in *:*) family=ipv6 ;; esac
log "opening firewalld port ${port}/tcp to ${FELIS_NANO_PROXY_CIDR} only"
firewall-cmd --permanent --add-rich-rule="rule family=\"${family}\" source address=\"${FELIS_NANO_PROXY_CIDR}\" port port=\"${port}\" protocol=\"tcp\" accept"
else
warn "no FELIS_NANO_PROXY_CIDR, so firewalld keeps ${port}/tcp closed; the summary shows how to admit your proxy"
fi
firewall-cmd --reload
}
@@ -2333,12 +2492,20 @@ EOF
# restart, not `enable --now`: on a re-run the service is already active and --now would
# leave the OLD binary running against the NEW unit. Converge means converge.
systemctl restart felis-nano
# restart returns as soon as the process is forked. A config the new binary rejects, or a
# file it cannot open, only shows once it has exited and the unit sits in auto-restart.
sleep 2
if ! systemctl is-active --quiet felis-nano; then
journalctl -u felis-nano -n 20 --no-pager || true
die "felis-nano did not stay up; its last log lines are above"
fi
ok "felis-nano.service enabled and started (listen ${FELIS_NANO_LISTEN})"
}
summary_nano() {
local port="${FELIS_NANO_LISTEN##*:}" host
if nano_listen_is_loopback; then host="127.0.0.1"; else host="${NODE_IP}"; fi
local port="${FELIS_NANO_LISTEN##*:}" host="${FELIS_NANO_LISTEN%:*}"
# A wildcard bind names no address a proxy could dial; the node's own is the useful one.
case "$host" in ""|0.0.0.0|"[::]") host="${NODE_IP}" ;; esac
echo
ok "Felis-nano deployed."
echo
@@ -2351,12 +2518,19 @@ summary_nano() {
log " the full https://sessionserver.mojang.com/session/minecraft/hasJoined)"
if nano_listen_is_loopback; then
log "Bound to loopback: reachable from Velocity on THIS host, and from nowhere else."
log "Proxy on another machine? Re-run with FELIS_NANO_LISTEN=<private-ip>:${port} and"
log "allow ${port}/tcp ONLY from that proxy — hasJoined takes no auth token, so an"
log "internet-facing one is a free auth relay burning your Mojang egress IP."
log "Proxy on another machine? Re-run with the address on the sudo line (sudo drops"
log "exported variables):"
log " curl -fsSL <raw-url>/deploy/bootstrap.sh | sudo FELIS_NANO_LISTEN=<private-ip>:${port} FELIS_NANO_PROXY_CIDR=<proxy-ip>/32 bash"
log "firewalld then admits ${port}/tcp ONLY from that proxy — hasJoined takes no auth token,"
log "so an internet-facing one is a free auth relay burning your Mojang egress IP."
elif [ -n "$FELIS_NANO_PROXY_CIDR" ]; then
log "Bound to ${FELIS_NANO_LISTEN}. firewalld, where it runs, admits ${port}/tcp only from"
log "${FELIS_NANO_PROXY_CIDR}; any other firewall in front of this host must do the same."
else
log "WARNING: bound to ${FELIS_NANO_LISTEN} — hasJoined takes no auth token, so restrict"
log "${port}/tcp to your proxy's source IP or anyone can relay their logins through you."
log "WARNING: bound to ${FELIS_NANO_LISTEN} with no FELIS_NANO_PROXY_CIDR. hasJoined takes no auth"
log "token, so admit ${port}/tcp from your proxy alone, or anyone can relay their logins"
log "through you. firewalld, where it runs, keeps the port closed until you add:"
log " firewall-cmd --permanent --add-rich-rule='rule family=\"ipv4\" source address=\"<proxy-ip>/32\" port port=\"${port}\" protocol=\"tcp\" accept' && firewall-cmd --reload"
fi
log "Then edit ${STATE_DIR}/felis.toml to add your [[auth_source]] roots and run:"
log " sudo systemctl restart felis-nano"
@@ -2375,6 +2549,7 @@ main_nano() {
}
main() {
resolve_nano_listen
validate_settings
detect_os
prompt_install_mode
+513
View File
@@ -61,6 +61,519 @@ expect "an uppercase digest is the same digest" "LOG: installing the Felis-Legac
out="$(run_gate "$(printf '%s' "$want" | sed 's/../& /g')")"
expect "a space-separated digest is the same digest" "LOG: installing the Felis-Legacy Velocity fork" "$out"
# --- papermc_latest_jar answers "url sha256" from one response --------------------------
# Fill's download URLs are content-addressed (/v1/objects/<sha256>/<name>.jar), and the
# resolver's contract is to hand both halves back from the same grep — or refuse a URL
# that carries no digest, rather than wave the download through unchecked. Run under
# bash, not sh: bootstrap.sh is bash and the function uses $'\n'.
fn="$(awk '/^papermc_latest_jar\(\)/,/^}/' "$BS")"
[ -n "$fn" ] || { echo "FAIL: no papermc_latest_jar in $BS"; exit 1; }
[ "$(printf '%s\n' "$fn" | wc -l)" -lt 30 ] \
|| { echo "FAIL: the extracted papermc_latest_jar is not just the function -- did its closing brace move?"; exit 1; }
rsha=0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef
run_resolver() { # canned-fill-response
CANNED="$1" bash -c '
curl() { printf "%s" "$CANNED"; }
'"$fn"'
if out="$(papermc_latest_jar velocity 3.5.1)"; then
printf "RESOLVED %s\n" "$out"
else
printf "REFUSED\n"
fi
'
}
out="$(run_resolver "{\"url\":\"https://fill-data.papermc.io/v1/objects/${rsha}/velocity-3.5.1-615.jar\"}")"
expect "the resolver pairs the url with its own digest" \
"RESOLVED https://fill-data.papermc.io/v1/objects/${rsha}/velocity-3.5.1-615.jar ${rsha}" "$out"
out="$(run_resolver '{"url":"https://fill-data.papermc.io/mirror/velocity-3.5.1-615.jar"}')"
expect "a URL that carries no digest is refused" "REFUSED" "$out"
# --- the resolved-Velocity digest gate --------------------------------------------------
# The download must hash to what the content-addressed URL promised, BEFORE
# atomic_install_file — the same refusal the Via plugins and the fork jar already get.
# The end pattern spells ${VELOCITY_DIR} with dots: escaped braces are literal in gawk
# and mawk but undefined in POSIX awk, and CI's awk is whatever ubuntu ships.
vblock="$(awk '/log "resolving the newest Velocity/,/atomic_install_file "\$tmp" "\$.VELOCITY_DIR.\/velocity\.jar"/' "$BS")"
[ -n "$vblock" ] || { echo "FAIL: no resolved-Velocity install block found in $BS"; exit 1; }
[ "$(printf '%s\n' "$vblock" | wc -l)" -lt 30 ] \
|| { echo "FAIL: the extracted block is not the velocity install -- did its last line move?"; exit 1; }
vdir="$(mktemp -d)"
trap 'rm -f "$jar"; rm -rf "$vdir"' EXIT
vwant="$(printf 'stand-in velocity build\n' | sha256sum | cut -d' ' -f1)"
run_velocity_install() { # digest-the-resolver-reports
WANT="$1" VELOCITY_DIR="$vdir" FELIS_VELOCITY_VERSION=3.5.1 bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { printf "LOG: %s\n" "$*"; }
remember_temp() { :; }
papermc_latest_jar() {
printf "%s %s\n" "https://fill-data.papermc.io/v1/objects/${WANT}/velocity-3.5.1-615.jar" "$WANT"
}
curl() { while [ "$#" -gt 1 ] && [ "$1" != "-o" ]; do shift; done; printf "stand-in velocity build\n" > "$2"; }
atomic_install_file() { printf "INSTALL: %s\n" "$2"; }
'"$vblock"
}
out="$(run_velocity_install deadbeef)"
expect "a download that does not hash to the promised digest is refused" \
"DIE: Velocity 3.5.1 checksum mismatch: got ${vwant}, expected deadbeef" "$out"
case "$out" in
*INSTALL:*) echo "FAIL a refused download must not reach atomic_install_file"; fails=$((fails + 1)) ;;
*) echo "PASS a refused download is not installed" ;;
esac
out="$(run_velocity_install "$vwant")"
expect "the matching download installs" "INSTALL: ${vdir}/velocity.jar" "$out"
# --- [[auth_source]] carry-forward -----------------------------------------------------
# write_felis_toml regenerates felis.toml wholesale on every run; this is what keeps the
# operator's Yggdrasil roots from being reset to the shipped default.
ablock="$(awk '/^persisted_auth_source_blocks\(\) \{/,/^}/' "$BS")"
[ -n "$ablock" ] || { echo "FAIL: no persisted_auth_source_blocks found in $BS"; exit 1; }
[ "$(printf '%s\n' "$ablock" | wc -l)" -lt 20 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
sdir="$(mktemp -d)"
trap 'rm -f "$jar"; rm -rf "$vdir" "$sdir"' EXIT
run_carry() {
STATE_DIR="$sdir" bash -c "$ablock"'
persisted_auth_source_blocks'
}
out="$(run_carry)"
expect "a first install gets the LittleSkin default" 'tag = "littleskin"' "$out"
printf '%s\n' '[server]' 'listen = "0.0.0.0:8080"' '' '[[auth_source]]' 'tag = "guild"' \
'prefix = "GD"' 'url = "https://guild.example/hasJoined"' '' '[smtp]' 'host = "mail.example"' \
> "$sdir/felis.host.toml"
out="$(run_carry)"
expect "an operator's root is carried forward" 'tag = "guild"' "$out"
case "$out" in
*littleskin*|*"[smtp]"*) echo "FAIL the carried list must be exactly the operator's tables:"; echo "$out"; fails=$((fails + 1)) ;;
*) echo "PASS the carried list stops at the next section and adds no default" ;;
esac
printf '%s\n' '[server]' 'listen = "0.0.0.0:8080"' > "$sdir/felis.host.toml"
out="$(run_carry)"
if [ -z "$out" ]; then
echo "PASS a config with no sources stays Mojang-only"
else
echo "FAIL a config with no sources must not get the default back:"; echo "$out"; fails=$((fails + 1))
fi
# The felis setup TUI re-encodes the whole file, which indents keys under each table.
rm -f "$sdir/felis.host.toml"
printf '%s\n' '[[auth_source]]' ' tag = "littleskin"' ' prefix = "LS"' ' url = "https://a.example"' \
'' '[[auth_source]]' ' tag = "guild"' ' prefix = "GD"' ' url = "https://b.example"' \
> "$sdir/felis.pod.toml"
out="$(run_carry)"
expect "both encoder-written tables are carried (first)" ' tag = "littleskin"' "$out"
expect "both encoder-written tables are carried (second)" ' tag = "guild"' "$out"
# TOML allows spaces inside the brackets and a quoted key. Each is still the operator's table.
for hdr in '[[ auth_source ]]' '[["auth_source"]]' "[['auth_source']]"; do
printf '%s\n' "$hdr" 'tag = "guild"' 'prefix = "GD"' 'url = "https://b.example"' '' \
'[smtp]' 'host = "mail.example"' > "$sdir/felis.host.toml"
out="$(run_carry)"
expect "a $hdr header is carried" "$hdr" "$out"
expect "a $hdr table keeps its keys" 'tag = "guild"' "$out"
case "$out" in
*"[smtp]"*) echo "FAIL a $hdr table must stop at the next section:"; echo "$out"; fails=$((fails + 1)) ;;
*) echo "PASS a $hdr table stops at the next section" ;;
esac
done
# --- write_nano_config leaves the unit able to read its config ---------------------------
# felis-nano runs as a DynamicUser, so the directory must be searchable by others under a
# hardened umask too, including one an older installer left at 0750 -- but the full
# install's, locked to 0700 for its secrets, must not be widened.
wblock="$(awk '/^write_nano_config\(\) \{/,/^}/' "$BS")"
[ -n "$wblock" ] || { echo "FAIL: no write_nano_config found in $BS"; exit 1; }
[ "$(printf '%s\n' "$wblock" | wc -l)" -lt 50 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_nano_config() { # state-dir
STATE_DIR="$1" SECRETS_ENV="$1/secrets.env" BOOTSTRAP_DONE="$1/bootstrap.done" bash -c 'umask 027
ok() { printf "OK: %s\n" "$*"; }
'"$wblock"'
write_nano_config'
}
mkdir "$sdir/probe" && chmod 0700 "$sdir/probe"
if [ "$(stat -c %a "$sdir/probe")" = 700 ]; then
run_nano_config "$sdir/nano" >/dev/null
expect "a fresh config dir is searchable under umask 027" 755 "$(stat -c %a "$sdir/nano")"
mkdir "$sdir/old" && chmod 0750 "$sdir/old"
run_nano_config "$sdir/old" >/dev/null
expect "a nano-only 0750 dir is opened up" 755 "$(stat -c %a "$sdir/old")"
: > "$sdir/probe/secrets.env"
run_nano_config "$sdir/probe" >/dev/null
expect "a dir holding the full install's secrets is not widened" 700 "$(stat -c %a "$sdir/probe")"
mkdir "$sdir/done" && chmod 0700 "$sdir/done" && : > "$sdir/done/bootstrap.done"
run_nano_config "$sdir/done" >/dev/null
expect "a dir marked as a full install is not widened" 700 "$(stat -c %a "$sdir/done")"
else
echo "SKIP directory modes: this filesystem ignores chmod"
fi
# --- install_nano_service reports a unit that dies at once ------------------------------
# A config the new binary rejects leaves the unit in auto-restart; the install must say so
# instead of printing "started" over a proxy whose every login now fails.
iblock="$(awk '/^install_nano_service\(\) \{/,/^}/' "$BS")"
[ -n "$iblock" ] || { echo "FAIL: no install_nano_service found in $BS"; exit 1; }
[ "$(printf '%s\n' "$iblock" | wc -l)" -lt 50 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_nano_service() { # exit status systemctl is-active reports
ACTIVE="$1" NANO_SERVICE="$sdir/felis-nano.service" HOST_BIN=/usr/local/bin/felis \
STATE_DIR=/etc/felis FELIS_NANO_LISTEN=127.0.0.1:25580 bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
ok() { printf "OK: %s\n" "$*"; }
sleep() { :; }
systemctl() { if [ "$1" = is-active ]; then return "$ACTIVE"; fi; }
journalctl() { printf "JOURNAL: config: needs prefix\n"; }
'"$iblock"'
install_nano_service'
}
out="$(run_nano_service 3)"
expect "a unit that dies at once fails the install" "DIE: felis-nano did not stay up" "$out"
expect "the failure shows the unit's own log" "JOURNAL: config: needs prefix" "$out"
case "$out" in
*"OK: felis-nano.service"*) echo "FAIL a dead unit must not be reported as started"; fails=$((fails + 1)) ;;
*) echo "PASS a dead unit is not reported as started" ;;
esac
out="$(run_nano_service 0)"
expect "a unit that stays up is reported as started" "OK: felis-nano.service enabled and started" "$out"
# --- a re-run on a nano host keeps what that host is --------------------------------------
# Re-running the installer is how a nano host updates. It must not move the endpoint an
# off-host proxy points at, nor default a nano-only host to the full control plane. The unit
# read back here is the one install_nano_service wrote above.
rblock="$(awk '/^resolve_nano_listen\(\) \{/,/^}/' "$BS")"
[ -n "$rblock" ] || { echo "FAIL: no resolve_nano_listen found in $BS"; exit 1; }
[ "$(printf '%s\n' "$rblock" | wc -l)" -lt 20 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_listen() { # env-value unit-path
FELIS_NANO_LISTEN="$1" NANO_SERVICE="$2" bash -c "$rblock"'
resolve_nano_listen
printf "LISTEN: %s\n" "$FELIS_NANO_LISTEN"'
}
expect "a re-run keeps the unit's listen address" "LISTEN: 127.0.0.1:25580" \
"$(run_listen '' "$sdir/felis-nano.service")"
expect "the operator's address beats the unit's" "LISTEN: 10.0.0.5:8081" \
"$(run_listen 10.0.0.5:8081 "$sdir/felis-nano.service")"
expect "a first install listens on loopback" "LISTEN: 127.0.0.1:8081" \
"$(run_listen '' "$sdir/absent.service")"
expect "a first install takes the operator's address" "LISTEN: 10.0.0.5:8081" \
"$(run_listen 10.0.0.5:8081 "$sdir/absent.service")"
# Whatever the address came from, it reaches the firewall, the summary and the unit as-is.
lblock="$(awk '/^validate_listen\(\) \{/,/^}/' "$BS")"
[ -n "$lblock" ] || { echo "FAIL: no validate_listen found in $BS"; exit 1; }
[ "$(printf '%s\n' "$lblock" | wc -l)" -lt 20 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
check_listen() { # value
bash -c 'die() { printf "DIE: %s\n" "$*"; exit 1; }
'"$lblock"'
validate_listen FELIS_NANO_LISTEN "$1" && echo VALID' _ "$1" 2>&1
}
for v in 8081 127.0.0.1 127.0.0.1:0 127.0.0.1:65536 127.0.0.1:x ::1:8081; do
expect "listen address $v is refused" "DIE: FELIS_NANO_LISTEN" "$(check_listen "$v")"
done
for v in '[::1]:8081' '[::]:8081' :8081 0.0.0.0:8081 127.0.0.1:8081; do
expect "listen address $v is accepted" VALID "$(check_listen "$v")"
done
# The summary hands the operator the URL to paste into the proxy's JVM flags, so it must
# name the address nano actually binds, and the node's only for a wildcard bind.
sblock="$(awk '/^summary_nano\(\) \{/,/^}/' "$BS")"
[ -n "$sblock" ] || { echo "FAIL: no summary_nano found in $BS"; exit 1; }
[ "$(printf '%s\n' "$sblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
kblock="$(awk '/^nano_listen_is_loopback\(\) \{/,/^}/' "$BS")"
[ -n "$kblock" ] || { echo "FAIL: no nano_listen_is_loopback found in $BS"; exit 1; }
[ "$(printf '%s\n' "$kblock" | wc -l)" -lt 10 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_summary() { # listen [proxy-cidr]
FELIS_NANO_LISTEN="$1" FELIS_NANO_PROXY_CIDR="${2:-}" NODE_IP=203.0.113.9 STATE_DIR=/etc/felis bash -c '
ok() { printf "OK: %s\n" "$*"; }
log() { printf "LOG: %s\n" "$*"; }
systemctl() { :; }
'"$kblock"'
'"$sblock"'
summary_nano'
}
expect "a private bind is the address printed" "http://10.0.0.5:8081/session/minecraft/hasJoined" \
"$(run_summary 10.0.0.5:8081)"
expect "an IPv6 loopback bind is printed as bound" "http://[::1]:8081/session/minecraft/hasJoined" \
"$(run_summary '[::1]:8081')"
out="$(run_summary 127.0.0.1:8081)"
expect "a loopback bind is printed as bound" "http://127.0.0.1:8081/session/minecraft/hasJoined" "$out"
expect "a loopback bind keeps its loopback note" "Bound to loopback" "$out"
for v in 0.0.0.0:8081 '[::]:8081' :8081; do
expect "a wildcard $v bind prints the node's address" "http://203.0.113.9:8081/session/minecraft/hasJoined" \
"$(run_summary "$v")"
done
# Loopback is what keeps the firewall shut, and hasJoined takes no token: a default that
# does not classify as loopback turns every fresh nano host into a public auth relay.
run_loopback() { # listen
FELIS_NANO_LISTEN="$1" bash -c "$kblock"'
if nano_listen_is_loopback; then echo LOOPBACK; else echo ROUTABLE; fi'
}
for v in 127.0.0.1:8081 127.0.0.5:8081 localhost:8081 '[::1]:8081'; do
expect "$v is loopback" LOOPBACK "$(run_loopback "$v")"
done
for v in 0.0.0.0:8081 10.0.0.5:8081 '[::]:8081'; do
expect "$v is not loopback" ROUTABLE "$(run_loopback "$v")"
done
ndefault="$(run_listen '' "$sdir/absent.service")"
ndefault="${ndefault#LISTEN: }"
expect "the default listen address (${ndefault:-empty}) is loopback" LOOPBACK "$(run_loopback "$ndefault")"
# --- firewalld admits the proxy alone ---------------------------------------------------
# hasJoined takes no token, so a routable bind is opened only to FELIS_NANO_PROXY_CIDR, never
# to every source, and a re-run closes the port an earlier installer opened to everyone.
cblock="$(awk '/^validate_cidr\(\) \{/,/^}/' "$BS")"
[ -n "$cblock" ] || { echo "FAIL: no validate_cidr found in $BS"; exit 1; }
[ "$(printf '%s\n' "$cblock" | wc -l)" -lt 15 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
check_cidr() { # value
bash -c 'die() { printf "DIE: %s\n" "$*"; exit 1; }
'"$cblock"'
validate_cidr FELIS_NANO_PROXY_CIDR "$1" && echo VALID' _ "$1" 2>&1
}
for v in 10.0.0.7 10.0.0.7/ /32 10.0.0.0/8/9 10.0.0.7/x '10.0.0.7/32 port' '10.0.0.7/32"'; do
expect "proxy CIDR <$v> is refused" "DIE: FELIS_NANO_PROXY_CIDR" "$(check_cidr "$v")"
done
for v in '' 10.0.0.7/32 192.168.0.0/24 fd00::7/128; do
expect "proxy CIDR <$v> is accepted" VALID "$(check_cidr "$v")"
done
fblock="$(awk '/^configure_nano_firewall\(\) \{/,/^}/' "$BS")"
[ -n "$fblock" ] || { echo "FAIL: no configure_nano_firewall found in $BS"; exit 1; }
[ "$(printf '%s\n' "$fblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_fw() { # listen proxy-cidr port-already-open(0|1)
FELIS_NANO_LISTEN="$1" FELIS_NANO_PROXY_CIDR="$2" OPEN="$3" bash -c '
ok() { printf "OK: %s\n" "$*"; }
log() { printf "LOG: %s\n" "$*"; }
warn() { printf "WARN: %s\n" "$*"; }
systemctl() { return 0; }
firewall-cmd() {
case "$*" in *--query-port=*) [ "$OPEN" = 1 ]; return ;; esac
printf "FW: %s\n" "$*"
}
'"$kblock"'
'"$fblock"'
configure_nano_firewall' 2>&1
}
no_blanket_port() { # label output
case "$2" in
*--add-port*) echo "FAIL $1: the port was opened to every source:"; echo "$2"; fails=$((fails + 1)) ;;
*) echo "PASS $1" ;;
esac
}
out="$(run_fw 0.0.0.0:8081 10.0.0.7/32 0)"
expect "a proxy CIDR opens the port to that source alone" \
'FW: --permanent --add-rich-rule=rule family="ipv4" source address="10.0.0.7/32" port port="8081" protocol="tcp" accept' "$out"
no_blanket_port "a proxy CIDR never opens the port to every source" "$out"
expect "an IPv6 proxy CIDR gets an ipv6 rule" 'rule family="ipv6" source address="fd00::7/128"' \
"$(run_fw '[::]:8081' fd00::7/128 0)"
out="$(run_fw 0.0.0.0:8081 '' 0)"
expect "no proxy CIDR says the port stays closed" "WARN: no FELIS_NANO_PROXY_CIDR" "$out"
no_blanket_port "no proxy CIDR opens nothing" "$out"
case "$out" in
*--add-rich-rule*) echo "FAIL no proxy CIDR must add no rule:"; echo "$out"; fails=$((fails + 1)) ;;
*) echo "PASS no proxy CIDR adds no rule" ;;
esac
expect "a re-run closes the port an earlier install opened to everyone" "FW: --permanent --remove-port=8081/tcp" \
"$(run_fw 0.0.0.0:8081 10.0.0.7/32 1)"
case "$(run_fw 127.0.0.1:8081 10.0.0.7/32 1)" in
*FW:*) echo "FAIL a loopback bind must leave firewalld alone"; fails=$((fails + 1)) ;;
*) echo "PASS a loopback bind leaves firewalld alone" ;;
esac
expect "a routable bind with no proxy CIDR is warned about" "WARNING: bound to 10.0.0.5:8081 with no FELIS_NANO_PROXY_CIDR" \
"$(run_summary 10.0.0.5:8081)"
expect "a routable bind with a proxy CIDR names it" "admits 8081/tcp only from" \
"$(run_summary 10.0.0.5:8081 10.0.0.7/32)"
pblock="$(awk '/^prompt_install_mode\(\) \{/,/^}/' "$BS")"
[ -n "$pblock" ] || { echo "FAIL: no prompt_install_mode found in $BS"; exit 1; }
[ "$(printf '%s\n' "$pblock" | wc -l)" -lt 60 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
tblock="$(awk '/^bootstrap_from_tui\(\) \{/,/^}/' "$BS")"
[ -n "$tblock" ] || { echo "FAIL: no bootstrap_from_tui found in $BS"; exit 1; }
[ "$(printf '%s\n' "$tblock" | wc -l)" -lt 5 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
# Only the no-terminal path can run unattended, and with a terminal attached the prompt
# would sit waiting on it. setsid drops the controlling terminal, as cloud-init and CI have.
notty=""
if (: </dev/tty) 2>/dev/null; then
if command -v setsid >/dev/null 2>&1; then notty=setsid; else notty=skip; fi
fi
run_mode() { # unit-path done-marker-path [FELIS_INSTALL_MODE [FELIS_BOOTSTRAP_FROM_TUI]]
INSTALL_MODE="${3:-}" FELIS_BOOTSTRAP_FROM_TUI="${4:-}" NANO_SERVICE="$1" BOOTSTRAP_DONE="$2" \
$notty bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { printf "LOG: %s\n" "$*"; }
'"$tblock"'
'"$pblock"'
prompt_install_mode </dev/null
printf "MODE: %s\n" "$INSTALL_MODE"' 2>&1
}
if [ "$notty" = skip ]; then
echo "SKIP install-mode default: a terminal is attached and there is no setsid to drop it"
else
: > "$sdir/bootstrap.done"
expect "a nano-only host re-runs as nano" "MODE: nano" \
"$(run_mode "$sdir/felis-nano.service" "$sdir/absent.done")"
expect "a host with the full install re-runs as full" "MODE: full" \
"$(run_mode "$sdir/felis-nano.service" "$sdir/bootstrap.done")"
out="$(run_mode "$sdir/absent.service" "$sdir/absent.done")"
expect "a fresh host defaults to full" "MODE: full" "$out"
expect "no controlling terminal takes the no-prompt path" "LOG: no terminal for a prompt" "$out"
# felis setup goes on to need the control plane, so under it nano is refused, and the
# nano-only default above must not apply either.
expect "felis setup refuses FELIS_INSTALL_MODE=nano" "DIE: felis setup installs the full control plane" \
"$(run_mode "$sdir/absent.service" "$sdir/absent.done" nano 1)"
expect "felis setup installs full on a nano-only host" "MODE: full" \
"$(run_mode "$sdir/felis-nano.service" "$sdir/absent.done" "" 1)"
fi
# --- install_go_toolchain checks the tarball before it replaces anything ----------------
# The tarball is unpacked and run as root, so a download that does not hash to the pin is
# refused -- and refused before the working toolchain is removed.
# The function replaces whatever version sits at GOROOT_DIR, so that has to be a directory
# Felis owns, never an operator's /usr/local/go.
case "$(grep '^GOROOT_DIR=' "$BS")" in
'GOROOT_DIR="/opt/felis/'*) echo "PASS the Go toolchain lives under /opt/felis" ;;
*) echo "FAIL the Go toolchain must live under /opt/felis, got: $(grep '^GOROOT_DIR=' "$BS")"; fails=$((fails + 1)) ;;
esac
gblock="$(awk '/^install_go_toolchain\(\) \{/,/^}/' "$BS")"
[ -n "$gblock" ] || { echo "FAIL: no install_go_toolchain found in $BS"; exit 1; }
[ "$(printf '%s\n' "$gblock" | wc -l)" -lt 50 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
gsum="$(printf 'stand-in go toolchain\n' | sha256sum | cut -d' ' -f1)"
groot="$sdir/go"
run_go() { # FELIS_GO_VERSION pinned-amd64-digest [FELIS_GO_SHA256]
FELIS_GO_VERSION="$1" GO_PINNED_VERSION=1.26.4 GO_PINNED_SHA256_AMD64="$2" \
GO_PINNED_SHA256_ARM64=unused FELIS_GO_SHA256="${3:-}" GOROOT_DIR="$groot" TMPDIR="$sdir" bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { printf "LOG: %s\n" "$*"; }
ok() { printf "OK: %s\n" "$*"; }
remember_temp() { printf "TEMP: %s\n" "$1"; }
uname() { echo x86_64; }
curl() { while [ "$#" -gt 1 ] && [ "$1" != "-o" ]; do shift; done
printf "stand-in go toolchain\n" > "$2"; printf "CURL: %s\n" "$2"; }
tar() { printf "TAR: %s\n" "$*"; }
'"$gblock"'
install_go_toolchain'
}
mkdir -p "$groot" && : > "$groot/KEEP"
out="$(run_go 1.26.4 deadbeef)"
expect "a Go download that does not match the pin is refused" \
"DIE: Go 1.26.4 (amd64) checksum mismatch: got ${gsum}, expected deadbeef" "$out"
case "$out" in
*TAR:*) echo "FAIL a refused Go download must not be unpacked"; fails=$((fails + 1)) ;;
*) echo "PASS a refused Go download is not unpacked" ;;
esac
if [ -e "$groot/KEEP" ]; then
echo "PASS a refused Go download leaves the old toolchain in place"
else
echo "FAIL a refused Go download must not remove the old toolchain"; fails=$((fails + 1))
fi
expect "an unpinned FELIS_GO_VERSION without a digest is refused" "DIE: no pinned sha256 for Go 1.99.0" \
"$(run_go 1.99.0 "$gsum")"
expect "an unpinned FELIS_GO_VERSION installs with its own FELIS_GO_SHA256" "TAR: " \
"$(run_go 1.99.0 deadbeef "$gsum")"
out="$(run_go 1.26.4 "$gsum")"
expect "a Go download matching the pin is unpacked" "TAR: " "$out"
gtmp="$(printf '%s\n' "$out" | sed -n 's/^TEMP: //p')"
expect "the Go download is staged in a directory the cleanup removes" \
"CURL: ${gtmp:-<none>}/go1.26.4.linux-amd64.tar.gz" "$out"
# --- a private repo without a token fails with the hint instead of prompting -------------
# git asks for credentials on /dev/tty, where a piped install would sit waiting. Every
# network git call goes through git_auth, so the switch belongs there.
gablock="$(awk '/^git_auth\(\) \{/,/^}/' "$BS")"
[ -n "$gablock" ] || { echo "FAIL: no git_auth found in $BS"; exit 1; }
[ "$(printf '%s\n' "$gablock" | wc -l)" -lt 15 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_git_auth() { # token
FELIS_GITHUB_TOKEN="$1" bash -c '
unset GIT_TERMINAL_PROMPT # whatever runs this harness may have set it already
git() { printf "GIT: prompt=%s\n" "${GIT_TERMINAL_PROMPT:-<unset>}"; }
'"$gablock"'
git_auth clone https://example.invalid/felis.git'
}
expect "git never prompts without a token" "GIT: prompt=0" "$(run_git_auth '')"
expect "git never prompts with a token" "GIT: prompt=0" "$(run_git_auth ghp_example)"
fblock="$(awk '/^fetch_source\(\) \{/,/^}/' "$BS")"
[ -n "$fblock" ] || { echo "FAIL: no fetch_source found in $BS"; exit 1; }
[ "$(printf '%s\n' "$fblock" | wc -l)" -lt 40 ] \
|| { echo "FAIL: the extracted block is not the function -- did its closing brace move?"; exit 1; }
run_fetch() { # src-dir
SRC_DIR="$1" FELIS_REF=main FELIS_REPO_URL=https://example.invalid/felis.git bash -c '
die() { printf "DIE: %s\n" "$*"; exit 1; }
log() { :; }
ok() { :; }
resolve_install_ref() { :; }
stamp_version() { :; }
git_auth() { return 128; }
git() { :; }
'"$fblock"'
fetch_source'
}
expect "a failed clone names the token" "set FELIS_GITHUB_TOKEN" "$(run_fetch "$sdir/src")"
mkdir -p "$sdir/src/.git"
expect "a failed fetch into an existing checkout names the token" "set FELIS_GITHUB_TOKEN" \
"$(run_fetch "$sdir/src")"
# ---------------------------------------------------------------------------------------
if [ "$fails" -eq 0 ]; then
echo "ALL PASS"
+18 -1
View File
@@ -13,7 +13,8 @@
# 1. Prebuilt tars at deploy/images/felis-limbo.tar + felis-lobby.tar (imported as-is).
# 2. Otherwise built on this host with docker, resolving the LOOHP/Limbo CI jar and
# the latest stable Paper jar automatically. Override any of:
# LIMBO_JAR_URL LIMBO_SCHEM_URL LIMBO_VERSION PAPER_JAR_URL PAPER_MC_VERSION
# LIMBO_JAR_URL LIMBO_SCHEM_URL LIMBO_VERSION PAPER_JAR_URL PAPER_JAR_SHA256
# PAPER_MC_VERSION
#
# Toggles: SKIP_BOOTSTRAP=1 (base already up), SKIP_SETUP=1 (stop before the TUI).
set -Eeuo pipefail
@@ -73,9 +74,24 @@ else
: "${PAPER_MC_VERSION:=1.21.8}"
: "${PAPER_JAR_URL:=$(curl -fsSL --max-time 30 "https://fill.papermc.io/v3/projects/paper/versions/${PAPER_MC_VERSION}/builds/latest" | grep -oE 'https://fill-data\.papermc\.io/[^"]+\.jar' | head -1)}"
[ -n "$PAPER_JAR_URL" ] || die "could not resolve the Paper jar; set PAPER_JAR_URL"
# Both Dockerfiles require the jar's digest. The fill-data URL is content-addressed
# (the objects/ path segment IS the sha256), so it is derived rather than asked for;
# a mirror override carries no such segment and must bring its own digest.
if [ -z "${PAPER_JAR_SHA256:-}" ]; then
sha="${PAPER_JAR_URL#*/objects/}"
sha="${sha%%/*}"
case "$sha" in
*[!0-9a-f]*|"") sha="" ;;
esac
if [ "${#sha}" -ne 64 ]; then
die "cannot derive the Paper jar sha256 from PAPER_JAR_URL (not a content-addressed fill-data URL); set PAPER_JAR_SHA256"
fi
PAPER_JAR_SHA256="$sha"
fi
log "building $LOBBY_IMAGE (Paper $PAPER_MC_VERSION)"
docker build -f "$SRC_DIR/deploy/lobby/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
--build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
-t "$LOBBY_IMAGE" "$SRC_DIR"
docker save "$LOBBY_IMAGE" | "$K3S" ctr images import -
@@ -84,6 +100,7 @@ else
log "building $PAPER_IMAGE (plain Paper $PAPER_MC_VERSION, forwarding via the operator initContainer)"
docker build -f "$SRC_DIR/deploy/paper/Dockerfile" \
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
--build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
-t "$PAPER_IMAGE" "$SRC_DIR"
docker save "$PAPER_IMAGE" | "$K3S" ctr images import -
fi
+9
View File
@@ -9,6 +9,7 @@
# Fill v3 API — api.papermc.io v2 has returned HTTP 410 since 2026-07-01):
# docker build -f deploy/lobby/Dockerfile \
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-26.2-<build>.jar \
# --build-arg PAPER_JAR_SHA256=<that same sha — the objects/ path segment> \
# --build-arg LUCKPERMS_JAR_URL="$(curl -fsSL https://metadata.luckperms.net/data/all \
# | grep -o 'https://download.luckperms.net/[^"]*/bukkit/loader/[^"]*\.jar')" \
# -t felis-lobby:demo .
@@ -42,6 +43,10 @@ RUN cd plugins/paper \
# also runs the plugin's Java-21 bytecode, so only the runtime moves.
FROM eclipse-temurin:25-jre
ARG PAPER_JAR_URL
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
# that shape at build time. Checking the digest after the download turns a truncated or
# tampered fetch into a failed build instead of a lobby booted on the wrong bytes.
ARG PAPER_JAR_SHA256
# LuckPerms is required, not optional: the panel's whole permission surface
# (internal/api/handlers_access.go) issues `lp user ...` over RCON, so a lobby built
# without it answers every grant with "Unknown command" — a failure the operator only
@@ -55,12 +60,16 @@ RUN set -eu; \
if [ -z "${PAPER_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \
fi; \
if [ -z "${PAPER_JAR_SHA256:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_SHA256=<paper jar sha256> is required" >&2; exit 1; \
fi; \
if [ -z "${LUCKPERMS_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg LUCKPERMS_JAR_URL=<luckperms bukkit jar> is required" >&2; exit 1; \
fi; \
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
mkdir -p /paper/plugins; \
curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \
echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \
curl -fSL "$LUCKPERMS_JAR_URL" -o /paper/plugins/LuckPerms.jar; \
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \
echo "eula=true" > /paper/eula.txt
+1
View File
@@ -29,6 +29,7 @@ this at every layer:
```
docker build -f deploy/lobby/Dockerfile \
--build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \
--build-arg PAPER_JAR_SHA256=<sha256 of that jar> \
-t felis-lobby:demo .
docker save felis-lobby:demo | sudo k3s ctr images import -
# felis.toml → [velocity] lobby_image = "felis-lobby:demo"
+1 -1
View File
@@ -93,7 +93,7 @@ else
echo " injects it from the <server>-rcon Secret when spec.rcon.enabled is true." >&2
fi
# ponytail: rewritten whole, not merged. Paper loads this file and fills every key it does
# Rewritten whole, not merged. Paper loads this file and fills every key it does
# not find with the default, then writes the full tree back — so a proxies-only file is a
# complete, stable input, and the lobby's other globals are simply always the defaults.
# That is true of a system server Felis owns end to end; if admins are ever allowed to tune
+9
View File
@@ -18,6 +18,7 @@
# API — the SAME url the lobby build resolves, so this reuses it and adds no new dependency):
# docker build -f deploy/paper/Dockerfile \
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-<ver>-<build>.jar \
# --build-arg PAPER_JAR_SHA256=<that same sha — the objects/ path segment> \
# -t felis-paper:demo .
# docker save felis-paper:demo | sudo k3s ctr images import -
# # felis.toml → recommended via 0019_recommended_paper.sql (no [velocity] key points here)
@@ -30,13 +31,21 @@
# to boot on anything older.
FROM eclipse-temurin:25-jre
ARG PAPER_JAR_URL
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
# that shape at build time. Checking the digest after the download turns a truncated or
# tampered fetch into a failed build instead of a server booted on the wrong bytes.
ARG PAPER_JAR_SHA256
RUN set -eu; \
if [ -z "${PAPER_JAR_URL:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \
fi; \
if [ -z "${PAPER_JAR_SHA256:-}" ]; then \
echo "ERROR: --build-arg PAPER_JAR_SHA256=<paper jar sha256> is required" >&2; exit 1; \
fi; \
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
mkdir -p /paper; \
curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \
echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*
COPY deploy/paper/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
@@ -1,41 +0,0 @@
# Foundational subsystems: the initial Felis import (ledger backfill)
- **Type:** feature (initial import) — retroactive ledger entry
- **Date:** 2026-06-26
- **Area:** `apis/`, `internal/` (naming, rcon, store, config, build, backup, operator,
submit, api, platform), `cmd/felis`, `plugins/`
- **Commits:**
- `7fbebfe` feat(apis): MinecraftServer CRD types (v1alpha1) — the lifecycle source of truth (§1)
- `708cdfc` feat(core): naming, RCON, store (Postgres + embedded migrations), config, image-build libraries
- `43ab921` feat(backup): archive-based world backup/restore + the retention/idle reaper
- `78b8cf6` feat(operator): MinecraftServer controller and reconcilers
- `d39605e` feat(submit): user modpack build + admin-approval pipeline (see [modpack-submission-lane](2026-06-26-modpack-submission-lane.md))
- `b508fcc` feat(api): dual-faced felis-api — permissions/LuckPerms, modpack lane, admin fleet read
- `47fcd90` feat(platform): node orchestration + the `cmd/felis` single-binary entrypoint
- `93f143f` feat(plugins): Velocity proxy + Fabric/Forge/NeoForge/Paper integration mods
- **Tasks:** #23 (permissions), #24 (modpack lane), #25 (fleet read)
## What it did
Stood up the whole backend spine in one build-order sweep: the Kubernetes CRD that is
the lifecycle source of truth, the core libraries (deterministic resource naming, the
RCON client, the Postgres store with embedded SQL migrations, config loading, container
image-build helpers), the backup/restore/reaper subsystems, the operator controller
that drives `MinecraftServer` resources, the user-modpack submit+approval pipeline, the
dual-faced (internal/external) felis-api behind a Zero-Trust guard, the platform
orchestrator that wires it all together under `cmd/felis`, and the server-side
integration plugins.
## Why
This is the project's first functional import — the substrate every later change edits.
It predates the change-ledger convention (established `fad48ff`, 2026-07-06), so it never
got a contemporaneous detail doc; this entry backfills one.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the
> change-ledger's detail-doc axis (§ Convention). This entry deliberately describes only
> what these eight commits **introduced** on 2026-06-26 — the named subsystems have been
> extended and reworked many times since (auth, passkey, metrics, quotas, updates), and
> that later work lives in its own dated detail docs, not here. Not independently
> re-verified for this doc; each subsystem was verified at its original commit and the
> current tree builds green at `9911b8c` (WSL oracle, go1.26.4).
@@ -1,30 +0,0 @@
# Modpack submission lane: build/approval pipeline + storage backends (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-06-26 – 2026-07-02
- **Area:** `internal/submit` (build/approval pipeline, storage backends), `internal/api` (submission endpoints)
- **Commits:**
- `d39605e` feat(submit): user modpack build + approval pipeline — an uploaded modpack stays `pending_review` and is never built until an admin approves; approval is a single-winner compare-and-swap handing off to the image-build Job, keeping the mandatory vulnerability scan in front of any push
- `598f3d3` feat(submit): local + S3 backends for modpack upload contexts, installer-selectable
- **Tasks:** #24 (§8 user-submitted modpack approval lane)
## What it did
Built the user-directed extension over the image-build subsystem: a player uploads a
modpack context, it sits in `pending_review`, and an admin's approval is the single-winner
gate that hands off to the build Job — with the vulnerability scan always ahead of any
registry push. `598f3d3` makes the upload-context store pluggable (local filesystem or S3),
selectable at install time.
## Why
Untrusted user content must never build or push unreviewed, and the compare-and-swap
approval guarantees exactly one build per submission even under a double-click or retry.
The storage-backend choice lets a single-node demo use local disk while a real deployment
uses S3, without a code change.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The approval
> compare-and-swap and endpoints were unit-tested at their commits; the S3 path is
> integration-configurable. The panel-side submission/approval UI is the collaborator's
> frontend work and is tracked only by its INDEX rows. Not independently re-verified for
> this doc; current tree green at `9911b8c`.
@@ -1,38 +0,0 @@
# Cloudflare Tunnel + Access edge (cfsetup) + NodePort fencing (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-01
- **Area:** `internal/cfsetup` (pure core + integration runner), `cmd/felis` (TUI edge flow), edge nftables fence
- **Commits:**
- `53a7664` feat(cfsetup): recommended Cloudflare Tunnel + Access edge (§14) — domain- and IdP-agnostic; the load-bearing `validateFailClosed` allowlist refuses any policy that could be public; fail-shut 404 catch-all; the raw game host is never proxied
- `ba13839` feat(breakglass): optional Tunnel + Access setup in the sudo TUI, an independent peer of Owner provisioning
- `a531f5e` fix(cfsetup): keep the connector install in the host apply layer only (drop the duplicate `StartConnector`)
- `2810fe8` fix(cfsetup): repoint a stale DNS record when routing a tunnel hostname
- `7d3be64` feat(cfsetup): start the tunnel connector as a setup step
- `346ec68` refactor(deploy): rework the cloudflare-edge walkthrough — restructured the edge TUI flow and added a tested `cfsetup` integration-runner path (with TUI height-measure/root tests)
- `e058a64` feat(edge): close the panel NodePort to the public after the tunnel is up — nftables at prerouting `raw` (-300), before kube-proxy's NodePort DNAT, gated on the connector actually serving; loopback accepted first so the connector origin hop is untouched
- **Tasks:** #37 (fence panel NodePort to public after tunnel)
## What it did
Stood up the optional one-click Zero-Trust edge: a Cloudflare Tunnel routing only the web
hostnames plus a fail-closed Access application, provisioned from the sudo TUI against the
operator's own Cloudflare account. `e058a64` then closes the Access-bypass hole where a
direct `https://<node-ip>:<nodeport>/` with the right Host header reached the origin
behind Access, by fencing the NodePort at the nftables raw hook so the packet is caught on
its original destination port — but only once the connector is confirmed serving, so
fencing never severs the only web path to a live origin.
## Why
Access is only a security boundary if the origin cannot be reached around it. The
fail-closed policy guard (`validateFailClosed`) and the NodePort fence are the two
load-bearing safety properties: a policy that could be public aborts the run with nothing
created, and a routable-but-unfenced NodePort would defeat the whole edge.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The policy guard,
> ingress generation, request bodies, gating, and the nftables ruleset shape / conn-count
> gate are unit-tested; the live cloudflared/Cloudflare-API and `nft` calls are
> INTEGRATION-ONLY (need a real account). KNOWN-LIMITATION: the fence targets nftables;
> firewalld-native coordination is deferred. Not independently re-verified for this doc;
> current tree green at `9911b8c`.
@@ -1,32 +0,0 @@
# Console auth: local-password login → passwordless migration (ledger backfill)
- **Type:** feature + refactor — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-04
- **Area:** `internal/api` (auth handlers, sessions), `internal/store` (users schema)
- **Commits:**
- `af14f02` feat(api): local-password authentication backend — login/logout/change-password on `op.console`; HttpOnly+Secure+SameSite=Lax host-only server-side sessions (SHA-256, 12h TTL); anti-enumeration uniform bcrypt; JSON-only credential writes (415 otherwise); fails closed unless `local_auth_enabled`
- `0c1cc59` feat(auth): migrate console login to passwordless
- `3b43f05` refactor(api): drop the dead login concurrency limiter and reconcile passwordless comments
- `c20b12c` refactor(api): drop the dead password-era `ResetMailer`, reconcile passkey-unbind docs
- **Tasks:** #27 (B1 thin thread), #79/#80/#81 (residue sweep + primitive adjudication)
## What it did
Shipped the staff local-password door (`af14f02`) as the primary web login when
Zero Trust is not in front of the API, then migrated the console to passwordless
(`0c1cc59`) once email-OTP + passkey were the intended factors. The two refactors
(`3b43f05`, `c20b12c`) then swept the password-era residue — the now-dead login
concurrency limiter and the `ResetMailer` — so no unused password machinery lingered in
the compile path, and reconciled the stale comments that referenced it.
## Why
`op.console` needs a real login even in deployments without a Cloudflare-Access edge; the
password backend was that. Once the passwordless factors landed, keeping the old password
scaffolding around was a bug farm — the sweep is the closeout evidence that the migration
was complete, not half-done.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. `af14f02` was
> covered by Go unit tests (content-type guard, anti-enumeration, forced-change lockdown)
> at its commit. Not independently re-verified for this doc; current tree green at
> `9911b8c` (WSL oracle, go1.26.4).
@@ -1,38 +0,0 @@
# Deploy: one-line bootstrap installer + demo bring-up (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-03
- **Area:** `deploy/` (bootstrap.sh, Dockerfiles, demo-up.sh), image build context
- **Commits:**
- `58fa4b0` feat(deploy): one-line bootstrap installer + distroless felis image (auto-detects apt/dnf, installs Docker/k3s/PostgreSQL, opens pg_hba to the pod CIDR, runs migrations, applies the control-plane bundle, leaves Web disabled pending `felis setup`)
- `94a3b7b` fix(deploy): harden bootstrap for RHEL-family Linux
- `deaa2f8` feat(deploy): zypper support (openSUSE/SLES)
- `318a724` feat(deploy): pacman support (Arch)
- `e5f1682` refactor(deploy)!: TUI (breaking walkthrough restructure)
- `28c3eee` refactor(deploy): improved TUI walkthrough
- `c14ed17` fix(docker): keep embedded `panel/` and `deploy/` in the image build context
- `d9e866f` fix(deploy): make the lobby image actually build (re-include `plugins/paper`, build on `gradle:8.14-jdk21`)
- `b84debf` feat(deploy): one-shot `demo-up.sh` — bootstrap → build/import limbo+lobby images → wire `[velocity]` image refs → `felis setup`, ending in the interactive Owner TUI
- **Tasks:** #26 (Phase A bootstrap verified end-to-end on the Demo VM)
## What it did
Made a bare Linux box a running Felis with one command. `bootstrap.sh` auto-detects the
host package manager across the four major families (apt/dnf/zypper/pacman), installs
whatever is missing (Docker, k3s, PostgreSQL, cloudflared), builds+imports the distroless
felis image, opens `pg_hba` to the pod CIDR, runs migrations, and applies the rendered
control-plane bundle. `demo-up.sh` wraps that plus the login-limbo/lobby image build and
`felis setup` into a single command, stopping only at the Owner-creation TUI it cannot
automate.
## Why
The spec calls for a self-hostable single-node deployment a SysAdmin can stand up without
a Kubernetes background. The package-manager fan-out and the demo wrapper are what make
"one line" true across real distros rather than only on the author's box.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the
> change-ledger's detail-doc axis. `deploy/` is shell + Dockerfiles (not Go-oracle
> verifiable); `d9e866f` records a real build+boot check (limbo `/healthz` 200, lobby
> reaches "Done"). Not independently re-verified for this doc; current tree green at
> `9911b8c`.
@@ -1,35 +0,0 @@
# felis CLI: break-glass recovery console + first-run setup (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-06-30
- **Area:** `cmd/felis` (break-glass/setup TUI, apply, migrate), `internal/api` (audit, owner store), `deploy/`
- **Commits:**
- `e108a37` feat(cli): break-glass emergency console TUI — root-only (`euid==0`), provisions/resets the Owner directly against Postgres, enables local login, prints a durable one-time-password summary
- `2d0bbb0` feat(cli): attribute break-glass recovery to the SysAdmin who runs it — bootstrap / recovery (bcrypt) / root-override, each audited with an honest `verified` flag and payload
- `a94b001` feat(deploy): break-glass Operator account provisioning
- `eb5875a` feat(felis): Operator break-glass op behind an operation menu
- `f5d00f3` feat(cli): `felis apply` for direct CRD creation
- `9c46632` feat(cli): `felis setup` first-run console (shared `runConsoleTUI` model, reclaim protection, cfsetup idempotency, `[auth].admin_hostname` respect)
- `7d91373` fix(migrate): honor `-config` placed after the `up` verb (flag.Parse stops at the first non-flag token)
- **Tasks:** #27 (B1 login→change-pw→TUI reset)
## What it did
Built the local-root recovery and first-run surface that bypasses web Zero Trust by
design. `felis breakGlass` mints or resets the Owner when the web login is unreachable;
`2d0bbb0` makes it accountable by recording *which* SysAdmin broke the glass across three
audited modes. `felis setup` is the non-emergency first-run twin sharing the same console
model. `felis apply` writes a `MinecraftServer` CRD directly, and `7d91373` fixes the
`migrate` flag parse so a configured DB path after `up` is honored.
## Why
An operator with root on the node and a kubeconfig must always be able to recover the
platform — that is break-glass's whole job, so it never refuses. Attribution
(`2d0bbb0`) closes the gap that root is machine authority, not a human identity: the root
gate is necessary but not sufficient for the audit trail.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The core logic was
> covered by Go unit tests over a fake owner store at each commit (auth match/non-match,
> the three audit modes, headless TUI drive). The bubbletea TUI glue is untested by house
> convention. Not independently re-verified for this doc; current tree green at `9911b8c`.
@@ -1,36 +0,0 @@
# Player onboarding data layer §B2: email-OTP, account-link, QR, Bind-Code (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-03
- **Area:** `internal/api` (onboarding/auth-bind handlers), `internal/store` (migrations 0004–0006)
- **Commits:**
- `dbe34a1` feat(api): player email-OTP verification (§B2) — `POST /account/email/{start,verify}`; 6-digit code, SHA-256-at-rest, 10-min TTL, 5-attempt cap enforced in the repo
- `1f8b9bb` feat(api): record account-link auth source (`mojang|thirdparty`) for the dual-Yggdrasil split (§10)
- `116595f` feat(api): QR scan-login completion poll on the internal face (`GET /internal/account/link/status/{mc_uuid}`) — read-only, reuses `UserByMCUUID`, no migration
- `fe2ece0` feat(api): public Bind-Code onboarding (`POST /auth/bind`) — the one pre-account entrypoint of `console.<root_domain>`; refuses a staff-UUID code with 403 without consuming it, so the public door provably never yields an admin principal
- `55592ed` feat(auth): public auth-bind endpoint wiring
- `6c3999a` fix(api): rate-limit email-OTP sends to close the email-bomb vector
- `879b177` fix(api): make OTP-start throttle atomic to close the concurrent-burst bypass
- **Tasks:** #29 (B2 data layer), #32 (OTP rate-limit), #35 (atomic throttle), #39 (console access model)
## What it did
Built the Go-verifiable data layer of forced web onboarding: prove control of an email
(OTP), record which Yggdrasil authenticated an in-game UUID, let a phone already signed in
to the panel complete a QR device-code link, and let an account-less player redeem a
one-time Bind Code minted in the Login Lobby to create+link+session in one public step.
The two fixes bound the OTP abuse surface — a per-target send rate limit and an atomic
reserve that closes the check-then-act race on the attempt counter.
## Why
The spec forces onboarding through the web so every account is provably email-controlled
and UUID-linked before it can operate anything. The `op.console` redline in `fe2ece0` — a
staff-UUID code is refused without being consumed — is what keeps the public console door
from ever minting an admin principal.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The account/session
> logic, single-use codes, and the op.console redline were covered by handler tests + the
> OpenAPI parity gate at each commit; the identity guarantee behind a Bind Code lives in
> velocity/Java (CODE-ONLY) and is not verifiable from this repo. Not independently
> re-verified for this doc; current tree green at `9911b8c`.
@@ -1,32 +0,0 @@
# §B3 username-collision reclaim + account migration (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-06-27 – 2026-07-05
- **Area:** `internal/api` (internal-face reclaim/blacklist, account migrate), `internal/store` (migration 0006)
- **Commits:**
- `a29571d` feat(api): reclaim squatted usernames for Mojang-priority players (§B3, 正版优先) — `POST /internal/player/reclaim` bars the squatter UUID + stashes its data (30-day hold) in one transaction, idempotent, returns the *first* reclaim's expiry; `GET /internal/player/blacklist/{mc_uuid}` is the login-gate check
- `fdb6efb` feat(account): migrate a live account's owned servers to a new account (§B3 inherit)
- **Tasks:** #30 (B3 game-login + username-collision reclaim)
## What it did
Built the data layer of the Mojang-priority collision flow: when the configured
third-party Yggdrasil and official Mojang issue the same username under different UUIDs,
the non-genuine squatter is displaced in favour of the real Mojang owner. Both tables are
keyed by `mc_uuid`, so the genuine player — identical username, *different* UUID — is
never caught by the bar. `fdb6efb` adds the inherit half: migrating an existing account's
owned servers onto a new account.
## Why
Two players cannot hold one username across two Yggdrasils; the spec resolves it in the
genuine Mojang owner's favour with a 30-day data hold for the displaced squatter, told the
truth about how long their data is kept (the first hold's window, never a fresh `now()+30d`
on retry).
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Handlers + the
> in-memory repo contract were unit-tested at each commit; the Postgres SQL path is
> integration-only, and the velocity collision-routing / limbo prompt / authlib
> dual-backend are code-only (Java) and out of this data-layer slice. Not independently
> re-verified for this doc; current tree green at `9911b8c`. Related: the operator-facing
> `/felis migrate` command has its own doc ([felis-migrate-command](2026-07-05-felis-migrate-command.md)).
@@ -1,39 +0,0 @@
# felis-api security + robustness hardening (audit sweep) (ledger backfill)
- **Type:** fix — retroactive ledger entry
- **Date:** 2026-06-30 – 2026-07-01
- **Area:** `internal/api` (login, request-id, listeners, SSE relays, quota/claim, MyServers), `internal/operator`
- **Commits:**
- `7a51c1d` fix(api): bound concurrent login bcrypt to shed CPU-pin floods (429 `auth_busy` before the compare; a cap, not a per-account lockout) — *audit #2*
- `164ac44` fix(api): validate inbound `X-Request-Id` before echo + audit persist (≤64 bytes, log-safe charset) — *audit-integrity*
- `c6c0772` fix(api): read/idle timeouts on all three listeners via a `newAPIServer` factory (closes Slowloris via `ReadHeaderTimeout`; `WriteTimeout` left unset so SSE isn't severed) — *audit #3*
- `3c1d647` fix(api): per-principal SSE stream cap (429 `too_many_streams`) — *audit #1, blast-radius bound*
- `d6e3189` fix(api): per-write deadline on SSE relay to sever a stalled reader (the real leak close behind the cap) — *audit #1*
- `8f41a00` fix(api): clear the SSE write deadline on return so it can't leak onto a reused keep-alive connection — *audit #1*
- `6368ab1` fix(api): `COALESCE` the MyServers `owned` flag so an ownerless row doesn't 500 the listing
- `2a4a81b` fix(api): don't burn the wake cooldown when refused at capacity
- `9873904` fix(operator): populate `Status.Players` from an RCON `list` probe (so the panel doesn't report 0/0)
- **Tasks:** #33 (wake cooldown), #34 (Status.Players), #41–#46 (audit #1–#4)
## What it did
A hardening sweep across the API's abuse and robustness surface: bound the two unbounded
CPU/goroutine amplifiers (concurrent bcrypt, per-principal SSE streams), close the SSE
relay's real stalled-reader leak with a per-write deadline (and clear it so it can't leak
onto a pooled connection), validate the caller-supplied request id before it reaches the
audit trail, set listener timeouts to close Slowloris, and fix two functional bugs — the
ownerless-row 500 and the wake cooldown burned on a capacity refusal.
## Why
Each is a specific, demonstrated failure mode: a login flood pins every core in bcrypt; a
stalled SSE reader leaks a relay goroutine + its upstream kube-apiserver follow *for the
life of the process*; an unvalidated `X-Request-Id` is a CR/LF log-forgery vector. The
`WriteTimeout`-left-unset detail is load-bearing — a blanket write timeout would sever the
healthy long-lived console/build-log streams the platform depends on.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Each fix shipped a
> targeted test at its commit — notably `d6e3189`/`8f41a00` use a deadline-aware
> `ResponseWriter` that fails closed if the guard is removed. The quota-claim TOCTOU
> (audit #4) is a documented KNOWN-LIMITATION (`2c56d17`), closeable only against a real
> Postgres. Not independently re-verified for this doc; current tree green at `9911b8c`.
-25
View File
@@ -1,25 +0,0 @@
# felis_* Prometheus metrics (§23) (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-06-30
- **Area:** `internal/metrics` + the emit sites in build, platform/fleet, and the start lifecycle
- **Commits:**
- `75642d9` feat(metrics): named `felis_*` Prometheus collectors
- `2a93a9e` feat(metrics): record `felis_image_build_failures_total` on failed builds
- `79eae7f` feat(metrics): publish `felis_servers_total` from a fleet snapshot
- `8ac5e64` feat(metrics): observe `felis_start_duration_seconds` across the start lifecycle
- **Tasks:** #17 (§23 felis_* metrics decision)
## What it did
Added the named `felis_*` collector set and wired the three emit points that make it
non-empty: a counter incremented on image-build failure, a gauge published from a fleet
snapshot, and a histogram observed across the server start lifecycle.
## Why
§23 calls for first-class operational metrics under a stable `felis_` namespace rather than
ad-hoc logging, so an operator can alert on build failures, fleet size, and start latency.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Not independently
> re-verified for this doc; current tree green at `9911b8c` (WSL oracle, go1.26.4).
@@ -1,37 +0,0 @@
# Auto-update subsystem: decision core + sources + gatherer + window API (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-07-01 – 2026-07-05
- **Area:** `internal/updates` (pure decision core), `internal/updater` (release sources, gatherer), `internal/api` (window admin API)
- **Commits:**
- `c01f133` feat(updates): pure I/O-free decision core — each tracked component is Pinned (Minecraft, left alone), Notify, or Scheduled (apply only inside a SysAdmin window); never force-applied, never a downgrade, never an auto-applied prerelease
- `3673af6` feat(api): admin API for the maintenance window (`GET`/`PUT /updates/window`), stored as JSON under `platform_settings` — API + persistence only, nothing consumes it yet
- `7464fa7` fix(updates): tag `Window` JSON so the persisted window round-trips (the obvious decode is correct by construction; a zero window fails closed to notify-only)
- `96b3cc9` feat(updater): wire `updates.Run` to a caller with PaperMC v3 release discovery
- `7d27640` feat(updater): GitHub Releases source, routing felis-api/k3s/cloudflared
- `7db57b9` feat(updater): `VersionGatherer` extraction core + CLI gather seam
- **Tasks:** #38 (auto-update: Felis/k3s/components/Velocity, pin Minecraft)
## What it did
Built the auto-update spine as a pure decision core plus the release-discovery sources
(PaperMC, GitHub Releases) and the version gatherer, with a SysAdmin-set maintenance
window read/written through an admin API. Version parsing tolerates the real feeds (leading
`v`, k3s `+k3s1` suffix, calendar versions, prerelease tails) and orders by SemVer
precedence.
## Why
The red lines are `不要强制自动更新` (never force auto-update) and `能不动的就别动`
(Minecraft stays pinned). The design encodes them structurally: a component may be applied
*only* inside a window the operator explicitly set, and Minecraft is Pinned so it is never
touched. `7464fa7`'s fail-closed zero-window (decodes to notify-only, never a rogue apply)
is the safety property for the not-yet-built runner.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The load-bearing
> invariants (pinned never changes, no downgrade, no auto-prerelease, apply-only-in-window)
> and the JSON round-trip contract were unit-tested at their commits. This subsystem is
> deliberately **report-only / integration-deferred**: the concrete Notifier/Applier,
> the `felis update` CLI + CronJob, and the current-version producing seams are declared
> but not wired (see `internal/updater/doc.go`, `openapi.yaml`). Not independently
> re-verified for this doc; current tree green at `9911b8c`.
@@ -1,37 +0,0 @@
# Passkey (WebAuthn) enrollment subsystem + hardening (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-07-01 – 2026-07-02
- **Area:** `internal/passkey` (go-webauthn adapter), `internal/api` (enrollment handlers/audit), `internal/store` (migrations 0007–0009)
- **Commits:**
- `f2c916d` feat(api): passkey enrollment persistence layer
- `742f15f` feat(api): passkey enrollment endpoints
- `0261204` feat(passkey): go-webauthn enrollment verifier adapter (Oracle-verified against a virtual authenticator)
- `fce0fce` feat(passkey): wire the enrollment verifier into felis-api
- `7278cd7` feat(passkey): require + record user verification at enrollment (`UserVerification=required`; capture `user_verified`/`backup_eligible`/`backup_state` — migration 0009) — *fix (d)*
- `cdbb5ab` fix(api): record credential id in the passkey-register audit event so bind/unbind are symmetric — *fix (a)*
- `9953275` fix(api): bound `webauthn_challenges` growth by superseding *all* prior rows per (user, purpose) — *fix (b)*
- `20e31fb` fix(store): cascade-delete passkeys + challenges on user removal (recreate both FKs `ON DELETE CASCADE`, scoped to the passkey tables only) — *fix (c)*
- `54bc6ef` fix(api): clear bound passkeys on password change to close a takeover foothold — *fix (e)*
- **Tasks:** #36 (passkey bind with email-OTP fallback), #48–#52 (fixes a–e)
## What it did
Built the WebAuthn *enrollment* half — persistence, the go-webauthn crypto adapter, and
the register-begin/finish endpoints — then hardened it through the five-fix batch (a–e):
symmetric audit, a bounded challenge table, cascade cleanup, enforced+recorded user
verification, and unbinding every passkey on a password reset so a passkey planted through
a transiently-hijacked session cannot survive as a standing login foothold.
## Why
Passkeys are the phishing-resistant factor with email-OTP as the fallback. The hardening
batch closes the seams that make enrollment safe to *rely on*: without UV enforcement a
passkey proves possession but not user; without the password-reset clear, a planted
passkey outlives the very remediation meant to evict an attacker.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The adapter crypto
> was verified against a virtual authenticator (virtualwebauthn), and each fix shipped
> with a targeted test (UV-negative rejection, challenge-growth bound, cascade, symmetric
> audit) at its commit. Not independently re-verified for this doc; current tree green at
> `9911b8c`. The assertion/login half is a separate doc ([passkey-login](2026-07-01-passkey-login.md)).
-36
View File
@@ -1,36 +0,0 @@
# Passkey (WebAuthn) login: assertion, discoverable, clone-detection (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-07-01 – 2026-07-05
- **Area:** `internal/passkey` (assertion crypto), `internal/api` (login/assertion, unbind, UA-guard), `internal/store` (migrations 0013/0014)
- **Commits:**
- `e035142` feat(passkey): WebAuthn login/assertion crypto adapter (BeginLogin/FinishLogin over go-webauthn, Oracle-verified against a virtual authenticator; surfaces the signature counter as a ceremony fact)
- `ec468ba` feat(auth): discoverable (usernameless) passkey login — the from-zero door the username-first assertion couldn't key on
- `0dbd557` fix(store): renumber the discoverable-login migration 0013 → 0014
- `9e1df12` feat(passkey): advance `sign_count`, reject clone-warned assertions
- `4f59d51` feat(auth): owner-tier passkey-unbind remediation endpoint
- `a63f49d` feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey — a backend-only UA interstitial (the SPA is untouched); asset/API/health requests pass through, an `ua_ack` cookie lets a determined user continue
- **Tasks:** #40 (from-zero discoverable login), #67 (WeChat/QQ UA-guard in `internal/panel`)
## What it did
Built the assertion (login) half of the ceremony: the crypto adapter, then discoverable
credentials so a user with no typed identifier can still log in (the enrollment
identifier problem the earlier deferral doc named), clone detection via the advancing
signature counter, and the owner-tier unbind remediation. `a63f49d` guards the flow at the
transport edge — WebAuthn is unusable inside the WeChat/QQ WebViews, so those UAs get a
bilingual "open in your system browser" page instead of the passkey SPA.
## Why
Enrollment without a login path is half a feature. Discoverable credentials resolve the
blocker recorded in the earlier deferral (`users.email` is nullable/non-unique and a
player's username is their Minecraft UUID, so username-first assertion had nothing to key
on). The UA-guard stops the most common real-world dead end: a passkey prompt that can
never succeed inside an in-app browser.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The assertion crypto
> was verified against a virtual authenticator (enrollment→assertion chain, origin-mismatch
> and unbound-credential rejection); `9e1df12`'s clone policy and the UA-guard pass-through
> were unit-tested at their commits. Not independently re-verified for this doc; current
> tree green at `9911b8c`.
@@ -1,37 +0,0 @@
# System servers: login-limbo + lobby (always-on gate) (ledger backfill)
- **Type:** feature — retroactive ledger entry
- **Date:** 2026-07-02
- **Area:** `internal/config`, `internal/naming`, `internal/api` (CRD readiness), `internal/operator`, `internal/platform`, `cmd/felis`, `plugins/limbo`, `deploy/limbo` + `deploy/lobby`
- **Commits:**
- `9bed51b` feat(config): `[velocity] login_image/lobby_image` — setup provisions the always-on system services only when set (empty = fail-loud skip; no official LOOHP/Limbo image exists)
- `9ef817f` feat(naming): reserved system-server names + service-token identifiers (single source of truth for the internal-API credential Secret)
- `159107b` feat(api): HTTP readiness knob on `MinecraftServer` + user-server fallback defaults to the login gate
- `dc23cb5` feat(operator): system-server pod HTTP readiness probe + login-only `FELIS_SERVICE_TOKEN` env (keyed off the reserved name so it can never leak into a user pod; sourced via `secretKeyRef`, never inlined)
- `3fdb3d0` feat(platform): internal-API base-URL helper + single-sourced token Secret
- `f554d52` feat(cli): provision the reaper-exempt login/lobby servers + replicate the service-token Secret into the minecraft namespace
- `241fe21` feat(limbo): felis-limbo in-game login flow (join → blacklist check → mint bind code → open book to `console.<root_domain>` → poll link-status → BungeeCord transfer to lobby; fail-closed)
- `c7315e4` feat(deploy): login-limbo + lobby images with game-port pinning (server-port pinned to GamePort 25565 on every start)
- **Tasks:** #53–#68 (system-server plumbing L1–L4, limbo plugin, operator env injection)
## What it did
Stood up the always-on authentication gate: reserved, reaper-exempt login/lobby
`MinecraftServer`s provisioned by setup, an HTTP readiness path for the RCON-less LOOHP/Limbo
loader (which reports "started" only after the first tick), and the felis-limbo plugin that
runs the whole onboarding *inside* Limbo before transferring an admitted player to the
lobby. A fresh connection always lands on the login gate, never a user backend, so
authentication is always in front.
## Why
The spec requires that a player authenticate before reaching any real server. That needs a
purpose-built always-on front server (Limbo) that speaks to the internal API — hence the
login-only service-token injection (keyed to the reserved name so it can never reach a user
pod) and the HTTP readiness knob for a loader that has no RCON.
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The Go layer
> (config/naming/readiness/operator env/platform) was unit-tested at each commit; the
> felis-limbo plugin is Java verified against a real Limbo jar via podman (#65), and the
> images carry a real build+boot check (#63, limbo `/healthz` 200 on 25565). Not
> independently re-verified for this doc; current tree green at `9911b8c`.
@@ -1,83 +0,0 @@
# Break-glass "halt a running server" op (#31 B4)
- **Type:** feature (addition)
- **Date:** 2026-07-05
- **Area:** `cmd/felis` — break-glass recovery console (Go, oracle-verifiable)
- **Commit:** `c2ee21a` — feat(breakglass): add halt-a-server op to the recovery console (§B4)
- **Task:** #31 Phase B4 (felis TUI break-glass ops)
## What it does
Adds a **"Halt a running server"** operation to the root-gated break-glass console.
The operator picks a server from the live fleet and the console flips that
`MinecraftServer` CRD's `spec.desiredState` to `Stopped`, letting the operator
reconcile it into a graceful shutdown. It is the emergency "stop this now" lever for
when the panel is unreachable but the box still has `root` + a kubeconfig.
## Why
The break-glass console already provisions the Owner and adds Operators, but there
was no local, panel-independent way to **stop** a misbehaving server (runaway,
compromised, resource-pinning). Halting is a reversible state nudge — the safest
possible break-glass power — so it belongs in the same root-gated recovery surface.
## Design decisions
- **CRD write, not pod kill.** The console flips `spec.desiredState=Stopped` with a
**spec-only merge patch** (`client.MergeFrom`), never a full-object `Update`. The
operator writes `status` on the same object continuously; a merge patch of
`spec.desiredState` touches a disjoint field and cannot race/clobber the operator's
status writes. A halt is therefore exactly the CRD write the operator already knows
how to honour.
- **Authority = root + kubeconfig.** The accountable actor is the OS user who
escalated to root (`osUser`), recorded for attribution — not proof. The root gate
plus kubeconfig possession *is* the authority, so (unlike the owner/operator paths)
no credential-minting auth sub-flow is needed for a reversible state change.
- **System servers allowed but named.** Halting the `login`/`lobby` system servers
takes the shared front door down (login has no fallback). Break-glass is deliberately
full power, so the console **warns** rather than forbids: a `⚠ system` tag in the
picker and an explicit `WARNING` line in the post-exit summary.
- **Audit is best-effort.** `performHalt` mirrors `performBreakGlass`: the halt
succeeds even if the audit sink is down (break-glass must work with logging broken);
any audit error rides back in the outcome and is surfaced as a summary `WARNING`.
- **Already-stopped is a no-op** reported distinctly ("was already stopped" vs "is now
stopping"), so the console never claims a stop it didn't perform.
- **Namespace from config.** The target namespace is `cfg.K8s.Namespace`, threaded
through the console constructors — never hardcoded.
## Files
| File | Change |
|---|---|
| `cmd/felis/halt.go` | **new** — pure core (no bubbletea): `listServersForHalt`, `haltServer` (merge patch), `isSystemServer`, `performHalt`, `auditHalt` |
| `cmd/felis/halt_test.go` | **new** — table tests against a controller-runtime **fake client** (applies patches for real): running→stopped persists, already-stopped no-op, missing→error, system flag, list projection + desired-state fallback, audit success, audit-failure-still-halts |
| `cmd/felis/tui_halt.go` | **new** — bubbletea/huh shell mirroring `ownerModel` (load → pick → work → done), empty-fleet guard, `⚠ system` picker labels, outcome card |
| `cmd/felis/tui_menu.go` | `bgHaltServer` enum + "Halt a running server" menu option |
| `cmd/felis/tui_root.go` | `namespace` field; `bgHaltServer` dispatch to `newHaltModel`; `haltResultMsg` terminal handling |
| `cmd/felis/breakglass.go` | halt fields on `breakGlassResult`; `namespace` threaded through `runBreakGlassTUI`/`runSetupTUI`/`runConsoleTUI`; post-exit halt summary (stopping / already-stopped, system + audit warnings, restart hint) |
| `cmd/felis/setup.go` | pass `cfg.K8s.Namespace` into `runSetupTUI` |
| `cmd/felis/tui_root_test.go` | pass `"minecraft"` namespace into `newRootModel` test call |
## Verification
WSL oracle (go1.26.4, FedoraLinux-44), authoritative for Go:
```
go build ./... → BUILD_OK
go vet ./cmd/felis/... → VET_OK
go test ./... → all 20 packages ok, ALL_GREEN
```
The core (`halt.go`) is fully unit-tested against a real `fake.Client`, which applies
the merge patch, so the test asserts the **persisted** `spec.desiredState`, not merely
that `Patch` was called. `tui_halt.go` is thin bubbletea glue (untested by house
convention, mirrors the existing `tui_owner.go`).
## Self-review outcome
- **ponytail (over-engineering):** lean — no one-impl interface, every field consumed,
audit seam justified. Nothing cut.
- **correctness:** caught and fixed a misleading restart hint — the summary originally
pointed at `felis apply`, but that command is **create-only** (errors "already
exists" on an existing server); corrected to "restart from the panel, or set
`spec.desiredState` back to Running."
@@ -1,80 +0,0 @@
# `/felis migrate` in-game command (§B3 inherit, Velocity side)
- **Type:** feature (addition)
- **Date:** 2026-07-05
- **Area:** `plugins/velocity` + `plugins/shared` — Velocity proxy plugin (Java, compile-verified)
- **Commit:** `c1aa38b` — feat(velocity): add /felis migrate to open an account migration (§B3 inherit)
- **Task:** completes the code-only gap named in `internal/api/handlers_account_migrate.go`
## What it does
Adds the in-game `/felis migrate` command that a player runs to **open an account
migration** — the first step of handing their owned servers to another account (spec
§B3 "inherit", scenario A). The command posts the player's Mojang-verified UUID to the
backend, which puts that account into migrate mode (`state=initiated`). The player then
finishes the migration on the web console (prove it's them, name the receiving account,
redeem a one-time code).
The Go backend (`handleMigrateStart` and the web-driven steps 2–4) already existed and
was tested; its header comment explicitly named **"the `/felis migrate` command that
calls handleMigrateStart"** as the code-only gap. This change closes that gap.
## Why
Without the in-game command, the migration flow had no entry point — the backend
handler was reachable only in theory. `/felis migrate` is the trustworthy initiator:
Velocity has already established the caller's online-mode UUID, so the sensitive proof
can be deferred to the web step-up while the in-game command just opens the migration.
## Design decisions
- **Mirrors the existing command suite verbatim.** `doMigrate` follows `doClaim`;
`migrateError` follows `claimError`; `migrateStart` follows `claim`/`opLoginApprove`.
No new imports, types, or idioms — every construct already appears in the same files.
- **Identity-bound + out-of-limbo, but server-independent.** Like `claim`, it requires
a real player past the login limbo (`requirePlayer` + `ensureOutOfLimbo`). Unlike
`claim`, it acts on the caller's *account*, not the server they stand on, so there is
**no** `registry`/current-server check.
- **Expects HTTP 201.** `migrateStart` posts to
`/api/v1/internal/account/migrate/start` and expects **201 Created** (`handleMigrateStart`
returns `StatusCreated`) — not 200 like the other calls. A 201 that does not affirm
`started:true` is treated as a contract breach, not a refusal.
- **Error mapping matches the handler's refusals:** 404 `not_linked` → "Link your
account on the web console before migrating"; 409 `account_retired` → "This account
can't start a migration (already migrated or retired)"; transport (0) and default →
generic retry text.
- **Points the player to the console on success.** The command only *opens* the
migration, so on success it prints the player web console URL
(`https://console.<root_domain>`, derived from config — never a hardcoded domain) and
a one-line description of the remaining steps. A proxy-side `logger.info` records the
initiating username against the UUID (the backend audit only has the UUID).
## Files
| File | Change |
|---|---|
| `plugins/shared/.../link/FelisApiClient.java` | **+`migrateStart(UUID)`** — POST mc_uuid, expect 201, affirm `started:true` |
| `plugins/velocity/.../FelisVelocityPlugin.java` | `migrate` literal in the Brigadier tree; **`doMigrate`** handler; **`migrateError`** mapper; `/felis migrate` help line |
## Verification
Java is not oracle-verifiable via the Go suite, but it **is** compile-verifiable via
the podman gradle toolchain established in #63/#65:
```
podman run --rm -v plugins:/work -w /work/velocity \
docker.io/library/gradle:jdk17 gradle --no-daemon compileJava
→ BUILD SUCCESSFUL in 19s (compiled against real velocity-api:3.3.0-SNAPSHOT)
```
The change compiles clean against the real Velocity API jar (including the shared
`FelisApiClient` compiled straight into the velocity module). The backend contract it
speaks to (`handleMigrateStart`) is covered by `handlers_account_migrate_test.go` on
the Go side.
## Self-review outcome
- **ponytail (over-engineering):** lean — pure mirror of three existing, compiling
methods; no speculative abstraction. Nothing cut.
- **correctness:** the one contract divergence (201 vs 200) was verified against the Go
handler source before writing.
@@ -1,29 +0,0 @@
# Operator: idle auto-stop, quotas, startup/readiness timeouts, /readyz (ledger backfill)
- **Type:** feature + fix — retroactive ledger entry
- **Date:** 2026-07-05
- **Area:** `internal/operator` (idle stop, timeouts), `internal/api` (quotas, /readyz)
- **Commits:**
- `91bfa27` feat(operator): idle auto-stop (§8)
- `e574749` feat(api): enforce CPU/memory/storage quotas (§9.3, §22)
- `7f7e459` fix(operator): enforce startup and readiness timeouts (§5, §8)
- `7becb38` fix(api): implement `/readyz` with real DB + K8s API + CRD checks (§7)
- **Tasks:** §5/§7/§8/§9.3/§22 operator + resource-governance spec items
## What it did
Rounded out the operator's lifecycle governance: stop idle servers automatically, enforce
per-resource CPU/memory/storage quotas at claim/create, bound how long a server may sit in
startup/readiness before the operator gives up, and make `/readyz` a real dependency check
(DB, Kubernetes API, and the CRD) rather than a static 200.
## Why
An orchestrator that never reclaims idle capacity or bounds startup will accumulate stuck
and wasteful workloads; a `/readyz` that always returns 200 tells the load balancer a
broken control plane is healthy. These are the spec's resource-governance and
readiness-correctness requirements (§5/§7/§8/§9.3/§22).
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Covered by Go unit
> tests at each commit. Not independently re-verified for this doc; current tree green at
> `9911b8c` (WSL oracle, go1.26.4).
@@ -1,117 +0,0 @@
# Break-glass "back up a world now" console peer (§B4 "Sync", phase 2b)
- **Type:** feature (addition)
- **Date:** 2026-07-07
- **Area:** `cmd/felis` (Go, oracle-verified); `internal/api` + `docs/openapi.yaml`
(the internal endpoint's `os_user` accountability extension)
- **Commit:** `fc748d3`
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation. Per the user's
**"两者都要"** decision the feature was built in two halves: the felis-api endpoint
that does the real backup-Job orchestration (phase 1 `7a7c0d5` external face, phase
2a `f2fc57c` internal face) and a break-glass menu peer that calls it while the API
is alive. **This change is that peer — phase 2b — the last build step of "Sync".**
(S3 remains open in B4; this does not close the phase.)
## What it does
Adds a **"Back up a world now (Sync)"** operation to the root-gated break-glass
console. The operator picks a server from the live fleet; the peer resolves the
`felis-api-internal` ClusterIP Service and the service token from the control
namespace, then POSTs the internal backup endpoint to snapshot that server's world
while felis-api is alive. It is the on-node counterpart to the halt op: an emergency
"snapshot this now" lever for when the panel is unreachable but the box still has
`root` + a kubeconfig and the API is running.
It also extends the internal backup endpoint (phase 2a) to accept an optional
`{"os_user":"..."}` body so the audit row names the operator at the keyboard rather
than the generic `break-glass`.
## Why
The console cannot render the backup Job itself — it lacks the deployment coordinates
(`FELIS_IMAGE`, `FELIS_BACKUP_PVC`) that only felis-api holds. So, unlike the halt peer
(which writes the `MinecraftServer` CRD directly), the backup peer must go **through**
the API. The internal face exists precisely so an on-node machine caller with the
service token — no browser session, no Cloudflare-Access Principal — can reach that
orchestration. This change is the client that knocks on that door.
## Design decisions
- **Goes through the API, does not orchestrate locally.** Mirrors the phase-2a
rationale: deployment coordinates live only in felis-api. The peer's job is to
resolve the endpoint, authenticate with the token, and translate the HTTP result
into a friendly outcome card — not to build a Job.
- **ClusterIP resolution, not DNS.** `resolveInternalAPI` `Get`s the
`platform.APIInternalServiceName` Service and dials its `Spec.ClusterIP:APIInternalPort`
directly, erroring on an empty or `None` (headless) ClusterIP. The on-node console's
host resolver is not CoreDNS, so the in-cluster Service DNS name would not resolve
from the host; the ClusterIP is routable from the node and is what the `2ba9948`
dedicated ClusterIP Service exists to provide.
- **`os_user` attribution, parity with halt.** The console sends the escalated OS user
in the request body; the endpoint makes it the audit actor. The body is decoded
whenever `ContentLength != 0` — **not** gated on `Content-Type` (`decodeJSON` checks
only for unknown/trailing fields, not the header) — so a console that forgets the
header still records the operator. Absent/blank falls back to `break-glass`. The
console does **not** double-audit: the API audits at the boundary, single-sourced.
- **Stopped-gate stays server-side.** The world PVC is RWO, so the server must be
stopped. The peer does not pre-check this; it lets the API's stopped-gate return
`409 not_stopped` and renders that as a "must be stopped — halt it first" card. The
safety check is single-sourced in the API, never duplicated (and possibly drifting)
in the console.
- **A running pick ends the session with exit 1 — deliberately accepted.** The picker
lists all servers (a running pick is easy to hit), and a `409` surfaces as an error
through `backupResultMsg{err}` → root sets `m.err` → `breakglass.go` prints the
friendly card to stderr and exits non-zero. A dedicated `not_stopped` non-error path
was considered and **rejected**: every break-glass op ends the session anyway (all
`tea.Quit`), so `409`-vs-success differs only in exit code — marginal for an
interactive TUI. No error is swallowed; the friendly card is shown either way. Adding
a soft-landing state machine for one status code is complexity the interactive
surface does not earn.
- **Core/shell split, mirrors halt.** All decision logic (`resolveInternalAPI`,
`requestBackup`, `backupErrorFromResponse`, `performBackupNow`) lives in `backupnow.go`
and is unit-tested against a controller-runtime fake client + `httptest`.
`tui_backupnow.go` is thin bubbletea/huh glue (untested by house convention, mirrors
`tui_halt.go`). The picker **reuses** `listServersForHalt`/`haltableServer` rather
than cloning a second server-listing path.
## Files
| File | Change |
|---|---|
| `cmd/felis/backupnow.go` | **new** — pure core: `resolveInternalAPI` (Service ClusterIP + token secret), `requestBackup` (POST + Bearer + `os_user` body, status→outcome), `backupErrorFromResponse` (409/503/404/error-body mapping), `performBackupNow` |
| `cmd/felis/backupnow_test.go` | **new** — table tests against a fake client + `httptest`: happy resolve, headless/empty-token errors; 202 asserts Bearer + `os_user` + path; 409/503/404 + transport-failure mapping |
| `cmd/felis/tui_backupnow.go` | **new** — bubbletea/huh shell mirroring `tui_halt.go`: load → pick → work → outcome card; empty-fleet guard; friendly error card |
| `cmd/felis/tui_menu.go` | `bgSyncBackup` enum + "Back up a world now (Sync)" menu option after the halt option |
| `cmd/felis/tui_root.go` | `bgSyncBackup` dispatch to `newBackupModel`; `backupResultMsg` terminal handling into `breakGlassResult` |
| `cmd/felis/breakglass.go` | `backedUp`/`backupServer`/`backupStatus` result fields; cancel guard; post-exit backup summary |
| `internal/api/handlers_backups.go` | `handleInternalBackup` decodes the optional `os_user` body (gated on `ContentLength`, not `Content-Type`) and passes it as the audit actor; defaults `break-glass` |
| `internal/api/handlers_backup_now_test.go` | **+subtest** in `TestInternalBackup`: `os_user` body attributes the audit to the operator |
| `docs/openapi.yaml` | document the `internalBackupNow` optional `os_user` request body |
## Verification
WSL oracle (go1.26.4, FedoraLinux-44, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL_GREEN
```
`cmd/felis` and `internal/api` both re-ran (not cached), so the new `backupnow_test.go`
and the added `TestInternalBackup` subtest executed. `TestOpenAPIMatchesServedRoutes`
still passes: the `os_user` body is an addition to an already-documented operation, so
the served⇔documented route match is unchanged. The core is covered against a real
`fake.Client` (resolves the Service/Secret) + `httptest.Server` (asserts the wire
request and maps every status), so the tests exercise persisted/observable behaviour,
not merely that a call was made. `tui_backupnow.go` is thin glue, untested per the
`tui_halt.go` convention.
## Self-review outcome
- **ponytail (over-engineering):** lean — no new abstraction beyond the four core
functions two call sites (test + TUI) already justify; the picker reuses halt's
server-list core rather than cloning it; the deliberate rejection of a `not_stopped`
soft-landing path kept the state machine at four steps. Nothing cut.
- **correctness:** the `os_user` decode is gated on `ContentLength`, matching how
`decodeJSON` actually works (no `Content-Type` check), so a header-less console still
attributes correctly; the stopped-gate and audit stay single-sourced server-side, so
the peer cannot drift from the endpoint on the RWO safety check or the audit record.
@@ -1,82 +0,0 @@
# Separate ClusterIP Service for the felis-api internal face (8081)
- **Type:** bug fix (latent networking gap) + enabling change
- **Date:** 2026-07-07
- **Area:** `internal/platform` — Go (struct render oracle-verified; packet path **not**
verifiable in this environment — see Verification)
- **Commit:** `2ba9948`
- **Task:** #31 Phase B4 — surfaced while wiring the break-glass backup console peer
(phase 2b): the peer needs a routable path to the internal face, and that path was
broken for the login pod too.
## What it does
Renders a **new ClusterIP-only Service `felis-api-internal`** (control namespace)
that fronts the felis-api pod's internal port 8081, and repoints
`InternalAPIBaseURL` (the URL baked into the login pod's `FELIS_API_BASE_URL`) at
that Service name. Adds exported `APIInternalServiceName` / `APIInternalPort` so the
on-node break-glass console can resolve the Service's ClusterIP and dial it.
## Why (the latent bug)
The login limbo pod is configured with
`FELIS_API_BASE_URL = http://felis-api.<ns>.svc.cluster.local:8081` (setup.go) and
dials the internal face with the service token to mint bind codes and poll link
status. But the only Service named `felis-api` is the **external** face: a NodePort
Service that declares **only** port 443. A Service answers only on its declared
ports, so `felis-api:8081` had no backend — **every login-pod call to the internal
API silently failed to connect.** `deploy/limbo/README.md` even documented the
"login-pod → felis-api internal-port (8081) path" as reachable; it was not.
## Design decisions
- **A separate Service, not a second port on `felis-api`.** A `Type: NodePort`
Service allocates a node port for **every** declared port, with no per-port
opt-out. Folding 8081 into the NodePort `felis-api` Service would therefore publish
the internal face — which is service-token-only, explicitly **no Zero Trust** — on
every node's external IP. That violates the two-face security posture. A distinct
`ClusterIP` Service exposes 8081 **in-cluster only**: reachable by the login pod via
cross-namespace DNS, and by the on-node console via the ClusterIP (kube-proxy
programs ClusterIPs into the node's routing).
- **Repoint `InternalAPIBaseURL` to the new Service name.** The helper single-sources
the name the login pod is told to call; pointing it at `felis-api-internal` keeps
the login pod and the Service in agreement by construction.
- **Export the name + port for the console.** The break-glass backup peer (phase 2b)
resolves `APIInternalServiceName`'s ClusterIP at runtime and dials
`http://<clusterIP>:APIInternalPort` — it cannot use the cluster-DNS form because
the host's resolver is not CoreDNS.
## Files
| File | Change |
|---|---|
| `internal/platform/workloads.go` | **+`apiInternalService`** (ClusterIP, 8081→`internal`), wired into `Workloads()`; **+exported `APIInternalServiceName`/`APIInternalPort`**; `InternalAPIBaseURL` repointed at the internal Service, comment corrected |
| `internal/platform/workloads_test.go` | **+`TestAPIInternalService_ClusterIP`** — ClusterIP (never NodePort), 8081→`internal`, no nodePort, selects the api pods, name distinct from `felis-api` |
| `deploy/limbo/README.md` | document the `felis-api-internal` Service; correct the reachability note |
| `docs/troubleshooting.md` | §6 note: internal calls reached via `felis-api-internal`; a *connect* failure (not 401) points at that Service |
## Verification
WSL oracle (go1.26.4, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL GREEN
```
`TestAPIInternalService_ClusterIP` freezes the Service's shape. **This is a
code-level fix only.** `go build/vet/test` verifies the Service *struct* renders
correctly; it verifies **nothing** about packets flowing — not the login pod's
in-cluster call, not the console's host→ClusterIP dial (which relies on kube-proxy's
OUTPUT-chain DNAT, present on k3s but unverified here), not that 8081 is programmed
on a live cluster. Per the project's "Java/K8s code-only" reality, the runtime path
is **pending real-cluster verification**; the manifest-level defect (a DNS name with
no backing port) is fixed and asserted.
## Self-review outcome
- **ponytail (over-engineering):** one Service + two exported identifiers, all
load-bearing (the login pod and the console both need the routable 8081). No new
abstraction; `apiInternalService` mirrors `apiService`/`registryService`.
- **correctness / security:** the ClusterIP-not-NodePort choice is the crux — it keeps
the no-Zero-Trust internal face off every node's external interface, which a second
port on the NodePort Service could not.
@@ -1,83 +0,0 @@
# Internal-face break-glass world backup endpoint (§B4 "Sync", phase 2a)
- **Type:** feature (addition)
- **Date:** 2026-07-07
- **Area:** `internal/api`, `docs/openapi.yaml` — Go, oracle-verified
- **Commit:** `f2fc57c`
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation. Per the user's
**"两者都要"** decision the feature is built in two halves: the felis-api endpoint
that does the real backup-Job orchestration (phase 1, `7a7c0d5`) and a break-glass
menu peer that calls it while the API is alive (phase 2b, follow-up). **This change
is phase 2a: the second, internal face of that endpoint** — the door the console
peer will knock on.
## What it does
Adds `POST /api/v1/internal/servers/{name}/backup`, an **internal-face** twin of the
external `POST /api/v1/servers/{name}/backup`. The on-node break-glass console (root
on the host, holding the service token) POSTs here to snapshot a stopped world while
felis-api is alive. Same 202 `backing_up` / 409 `not_stopped` / 503
`backup_unavailable` / 404 / 400 `bad_name` surface as the external face.
## Why
The console cannot render the backup Job itself: it lacks the deployment coordinates
(`FELIS_IMAGE`, `FELIS_BACKUP_PVC`) that only felis-api holds — the same reason the
endpoint exists at all (phase 1). But the external face requires a Cloudflare-Access
Principal the console does not have. The internal face authenticates with the service
token (a trusted machine caller, no Principal), so the console can reach the same
orchestration without a browser session.
## Design decisions
- **No owner gate on the internal face.** The external handler enforces owner-or-admin
from the Principal; the internal handler has none — the service token IS the
authorization (the operator already has root on the node), so a server owned by
someone else still backs up. This mirrors how the other internal-face handlers
(op-login approve, QR poll) trust the token rather than a Principal.
- **Shared `enqueueBackup` tail.** The RWO stopped-gate, the optional-Backuper 503,
the async hand-off, and the audit+202 were refactored out of `handleBackupNow` into
a single `enqueueBackup(w, r, name, rec, actor, source)` that both faces call. The
two faces differ **only** in how the caller is authorized and in the audit
actor/source — the security-critical stopped-gate is single-sourced so the faces
cannot drift apart.
- **Audit attributed to break-glass/internal.** The internal handler audits directly
via `Repo.Audit` with `Actor:"break-glass", Source:"internal"` (the `a.audit`
helper hardcodes `Source:"external"`), so a console-initiated backup is
distinguishable in the audit log from an owner's self-service one.
- **Console does not double-audit.** Unlike the halt peer — which writes the CRD
directly and audits locally — the backup peer goes through the API, and the API
audits at the boundary. Auditing is single-sourced there; the console will not
emit its own row.
## Files
| File | Change |
|---|---|
| `internal/api/handlers_backups.go` | **+`handleInternalBackup`**, **+`enqueueBackup`**; `handleBackupNow` tail now calls `enqueueBackup(..., p.Email, "external")` |
| `internal/api/api.go` | register `POST /api/v1/internal/servers/{name}/backup` on the internal-face route table |
| `docs/openapi.yaml` | document the `internalBackupNow` operation (`x-felis-face: [internal]`, `serviceToken` security) |
| `internal/api/handlers_backup_now_test.go` | **+`TestInternalBackup`** — no-owner-gate, break-glass/internal audit, stopped-gate/503/404/400 |
## Verification
WSL oracle (go1.26.4, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL GREEN
```
`TestOpenAPIMatchesServedRoutes` gates the new route against `docs/openapi.yaml` in
both directions (served⇔documented) and passes. `TestInternalBackup` (5 subtests) and
the existing `TestBackupNow` (10) both pass — the external refactor is
behaviour-preserving (same audit actor `p.Email`/source `external`).
## Self-review outcome
- **ponytail (over-engineering):** the internal face is not a copy of the external
handler — the shared tail (`enqueueBackup`) collapses the duplication, and the two
handlers hold only their distinct auth + audit-attribution. No new abstraction
beyond the one shared function two callers already justify.
- **correctness:** the no-owner-gate difference is deliberate and matches the other
service-token handlers; the stopped-gate is unchanged and now single-sourced, so the
external and internal faces cannot diverge on the RWO safety check.
@@ -1,115 +0,0 @@
# On-demand world backup (§B4 break-glass "Sync"; felis-api endpoint + Job executor)
- **Type:** feature (addition)
- **Date:** 2026-07-07
- **Area:** `internal/backupjob` (new pkg), `internal/api`, `cmd/felis`, `docs/openapi.yaml` — Go, oracle-verified
- **Commit:** `7a7c0d5`
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation, resolved with the user as **immediate/on-demand world backup**. Per the user's "两者都要" decision this is built in two halves: **(this change) the felis-api endpoint that does the real backup-Job orchestration**, and (a follow-up) a break-glass menu peer that calls it while the API is alive.
## What it does
Adds `POST /api/v1/servers/{name}/backup`: an owner or admin snapshots a **stopped**
server's world into the archive store on demand, recorded as a first-class
`world_backups` row (reason `manual`) — restorable later by the existing restore path
and expired by the reaper's retention pass, so it never leaks as an orphan archive.
The backup runs asynchronously as a one-shot Kubernetes Job (the new
`internal/backupjob` package), mirroring how restore and image builds hand off to
Jobs. The handler answers **202 `backing_up`**.
## Why
felis-api cannot archive a world in-process: the world PVC is **RWO** and owned by the
operator's StatefulSet, so the API has nothing to mount at request time — the same
constraint that already makes `internal/restore` a Job. The break-glass console (which
runs direct-to-Postgres) likewise lacks the deployment coordinates (`FELIS_IMAGE`,
`FELIS_BACKUP_PVC`) needed to render the Job. Both point to the same home: the
orchestration belongs in felis-api, which holds those coordinates; other callers
invoke the endpoint.
## Design decisions
- **Backup Job self-records its `world_backups` row.** Unlike the restore Job — which
is deliberately DB-blind because it processes a potentially poisoned archive — the
backup Job **does** mount the felis config Secret and inserts its own backup row,
exactly like the reaper (the only other component holding both a world mount and the
database). This avoids the archive-then-async-record split that would otherwise leak
orphan archives on a crash. The security review for that one departure lives in
`internal/backupjob/jobspec.go` and is frozen by `jobspec_test.go`. Rationale: a
backup only **reads** a world the operator already owns and tars it (bytes, never
executed), so restore's poisoned-input threat does not apply; its blast radius (DB +
two PVCs) is a strict subset of the reaper's, and it never deletes a PVC nor calls
the K8s API (SA token stays un-mounted).
- **World mounted read-only, backup PVC read-write** — the mirror image of restore.
- **Stopped-gate (409 `not_stopped`).** The world PVC is RWO and held by a running
server, so a backup Job cannot double-mount it; the handler refuses unless the server
is fully stopped (`info.Ready || DesiredState != Stopped`). This also guarantees a
quiescent, non-torn archive. Mirrors `handleRestoreBackup`'s gate.
- **Authorization is restore's front half, minus the former-owner match.** Backup is
initiated by the **current** owner and records **their** ownership, so there is no
prior owner's data to leak — the leak guard that restore needs does not apply here.
An admin may back up an unowned (released) world; the recorded former owner is then
empty, exactly as the reaper records for an unowned reap.
- **Unique Job name per request.** Each backup Job is named `backup-<server>-<rand>`,
not a deterministic `backup-<server>`. A deterministic name would collide with a
just-finished Job still inside its `TTLSecondsAfterFinished` window (10m), and the
`AlreadyExists → 202` path would then silently produce **no** archive — the exact
window a user (or the console "立即备份" button) retries in. Unique names make every
request produce its own archive; `ErrAlreadyExists` remains only as a defensive
no-op on the ~impossible suffix collision. Ceiling (documented in `backup.go`): two
truly simultaneous taps may schedule two backup Pods — both mount the world PVC
read-only, so neither corrupts anything; single-flight-on-running is the upgrade
path if a double-tap storm ever appears.
- **One retention clock.** The entrypoint reuses the reaper's `reaperConfig` derivation
so a manual backup expires on the same schedule as an inactivity backup — one policy,
not two. The `"bk-"+hex` id scheme also matches, so manual and inactivity backups are
indistinguishable downstream.
- **Fail-safe on record failure.** If the row insert fails, the entrypoint deletes the
just-written archive so a failed backup leaves no unrecorded bytes.
- **Optional executor, honest 503.** Wired only when `FELIS_IMAGE` + `FELIS_BACKUP_PVC`
are supplied (same gate as restore); otherwise `API.Backuper` is nil and the endpoint
returns 503 `backup_unavailable`, so the authorization boundary is exercised before
the Job executor is deployable.
## Files
| File | Change |
|---|---|
| `internal/backupjob/jobspec.go` | **new** — `BackupJob` renderer + `BackupJobName`; weak SA, token off, hardened container, world RO / backup RW, config-Secret mount |
| `internal/backupjob/backup.go` | **new** — `Backuper` (idempotent enqueue) + `Config`/`withDefaults` |
| `internal/backupjob/k8sjobs.go` | **new** — controller-runtime `CreateBackupJob` (AlreadyExists → idempotent) |
| `internal/backupjob/jobspec_test.go` | **new** — freezes the Job's security shape incl. the deliberate config-Secret mount |
| `internal/backupjob/backup_test.go` | **new** — asserts each `Backup` call mints a unique Job name (repeat-tap must not silently no-op) |
| `cmd/felis/backup.go` | **new** — `felis backup` in-Pod entrypoint: archive + self-record + orphan-cleanup |
| `cmd/felis/run.go` | dispatch `case "backup"` + usage line |
| `internal/api/backuper.go` | **new** — the narrow `Backuper` port |
| `internal/api/handlers_backups.go` | **+`handleBackupNow`** |
| `internal/api/api.go` | `Backuper` field + `POST /servers/{name}/backup` route |
| `internal/api/handlers_backup_now_test.go` | **new** — `fakeBackuper` + handler subtests |
| `internal/api/backuper_wire_test.go` | **new** — compile-time `Backuper = (*backupjob.Backuper)(nil)` |
| `cmd/felis/api.go` | wire `backuper` under the `FELIS_IMAGE`+`FELIS_BACKUP_PVC` gate; `backupConfig` helper |
| `docs/openapi.yaml` | document the `backupNow` operation |
## Verification
WSL oracle (go1.26.4, authoritative for Go):
```
go build ./... && go vet ./... && go test ./... → ALL GREEN
```
The `internal/api` OpenAPI served-route contract test (`TestOpenAPIMatchesServedRoutes`)
initially failed — the new route was served but undocumented — and passes after adding
the `backupNow` operation to `docs/openapi.yaml`. `internal/backupjob` and the new
handler subtests pass. The controller-runtime `K8sJobs` binding is integration-only
(needs a live cluster) and is exercised only by the interface conformance test.
## Self-review outcome
- **ponytail (over-engineering):** the backup Job is a near-mirror of the restore Job,
not a shared parameterization — deliberate, because its security shape differs (it
holds DB creds) and must be asserted independently, not hidden behind a shared knob.
No speculative config; `Config.withDefaults` fills only real deployment values.
- **correctness:** the RWO stopped-gate and the self-recording atomicity were traced to
the reaper and restore before writing; the former-owner asymmetry vs restore is
justified above.
@@ -1,155 +0,0 @@
# Test-quality integrity audit — do the verifications verify FUNCTION, or just go green?
- **Type:** audit / verification evidence (no code changed)
- **Date:** 2026-07-07
- **Method:** mutation testing on the WSL oracle (go1.26.4) + per-function coverage backbone
- **Scope:** the load-bearing safety invariants the backfilled change ledger *claims* were tested
- **Tree state:** every mutation reverted; authoritative Windows-git working tree clean at `5a7cd5a`
- **Point verified:** each gate is broken at `HEAD` (`5a7cd5a`), not per-commit — this is the
right reading of "does the verification verify the FUNCTION": the current test pins the
current implementation. A per-commit sweep would audit history hygiene, a different question.
## Why this audit exists
The ledger backfill asserts, per subsystem, that a set of load-bearing safety
properties are "unit-tested". A passing suite proves the tests are GREEN; it does
not prove they would go RED if the behaviour broke. Those are different claims —
"passing ≠ verifying". This audit closes that gap the only way that earns the word
*verified*: **break the implementation, confirm the specific test turns red.** A
subagent (or a human) *reading* a test and judging it "looks thorough" reproduces
the exact error being audited (looks-right ≠ verifies), so reading was used only to
locate the gate line; the verdict is always the mutation result.
## Result: 18 / 18 crown-jewel invariants mutation-verified
Each row is a one-line break of the implementation, run against its own package on
the oracle. **CAUGHT = the suite went red** = the test genuinely pins the behaviour.
| # | Invariant (claimed tested) | Impl gate mutated | Verdict |
|---|---|---|---|
| 1 | Pinned component is NEVER changed | `plan.go` pin branch → fall through | CAUGHT |
| 2 | A downgrade is NEVER proposed | `plan.go` `lv.After(current)` → `true` | CAUGHT |
| 3 | A prerelease is NEVER auto-applied | `plan.go` `!lv.IsPrerelease()` → `true` | CAUGHT |
| 4 | Apply ONLY inside the SysAdmin window | `plan.go` `Window.Contains(now)` → `true` | CAUGHT |
| 5 | `After` is strict (no equal-version churn) | `version.go` `> 0` → `>= 0` | CAUGHT |
| 6 | Clone-warned assertion refused fail-closed | `handlers_passkey.go` `if va.CloneWarning` → `if false` | CAUGHT |
| 7 | Approval CAS builds exactly once | `submit.go` `if !won` → `if false` | CAUGHT |
| 8 | An `everyone` base is not fail-open | `cfsetup.go` `if !includeHasEveryone` → `if true` | CAUGHT |
| 9 | Scoped-identity recognition actually admits | `cfsetup.go` `scoped = true` → `scoped = false` | CAUGHT |
| 10 | SSE per-principal stream cap holds | `api.go` cap-disable threshold | CAUGHT |
| 11 | OTP atomic reserve → one winner per burst | `api.go` `Sub(last) < window` → `< 0` | CAUGHT |
| 12 | Idle server is auto-stopped | `reconciler.go` `AutoStopEnabled &&` → `false &&` | CAUGHT |
| 13 | Startup/readiness timeout fires | `reconciler.go` `>= timeout` → `>= timeout + 1h` | CAUGHT |
| 14 | `/readyz` 503s when DB/K8s is down | `handlers_internal.go` dep-check `err != nil` → `false` | CAUGHT |
| 15 | NodePort fence only fires once connector serves | `tui_edge_apply.go` `connectorConnCount` parse-fail `return 0` → `1` | CAUGHT |
| 16 | CRITICAL-CVE build is NEVER admitted | `build.go` scan-gate `JobFailed`→`StatusFailed` → `StatusSucceeded` | CAUGHT |
| 17 | A user can NEVER claim a reserved system name | `naming.go` `reserved[name]` → `reserved["__nomatch__"]` | CAUGHT |
| 18 | Service token reaches ONLY the login pod | `builders.go` `Name == SystemLoginServer` → `true` | CAUGHT |
Rows 15–18 close the gap a review of this audit surfaced: the first pass verified a
*subset* and worded the verdict as the whole set. They are the four remaining
load-bearing safety properties the ledger docs name as "unit-tested" (§ *Documented-tested
claim reconciliation* below). Each mutation produced a real `--- FAIL` on the specifically
named test — e.g. #15 reddened `TestConnectorConnCount/garbage_is_not_a_healthy_tunnel`,
#16 `TestSyncFailedDoesNotAdmitImage`, #18 `TestBuildEnvWithholdsServiceTokenFromUserServers`
— i.e. an assertion failure, not a compile break.
Not one crown-jewel test was vacuous. The `cfsetup` fail-closed test additionally
feeds five distinct *violating* policies (bare-everyone, everyone-OR-identity,
unrecognized `ip` type, empty rule, wrong decision) and asserts each is rejected —
strong negative-path coverage, confirmed by mutations #8–#9.
## Coverage backbone — what no oracle test executes (failure-mode B)
Coverage triages code that no test even runs (so it cannot be verified). It does NOT
itself earn "verified" — high coverage with weak asserts is the same green-number
trap. Per-function scan of the security packages:
**Integration-only by design (0% on the oracle — honest, NOT a gap).** The real
adapters run only against live infra; unit tests exercise the ports through fakes:
- `pgrepo.go` — all SQL, **including `QuotaAvailable`/`QuotaCheck` (the quota TOCTOU
atomic claim)**. This matches task #45's own "ENV-blocked" note: the atomic claim
is a Postgres `INSERT … WHERE`, verifiable only against a real DB.
- `k8scluster.go`, the K8s console/log-stream adapters — real Kubernetes/RCON I/O.
- `tui_edge_apply.go` `verifyConnectorServing` + the nftables fence apply — shell out to
live `cloudflared`/`nft`. **Correction from the first pass:** the doc splits this from the
*pure* `connectorConnCount` decision gate, which IS unit-tested and is now mutation-proven
(#15). The first pass wrongly folded the whole fence into "integration-only"; only the live
calls are. The gate that decides *whether* to fence is verified.
**Genuine coverage gap (untested at the HTTP layer — "not verified").** These are
*missing* tests, not fake-passing ones:
- `handlers_users.go` — the P5 SysAdmin account-management suite: `handleCreateUser`,
`handleGetQuotas`, `handleSetQuotas`, `handleListUsers`, `handleGetUser`,
`handlePatchUser`, `handleDeleteUser`, `handleDisableUser`, `handleLinkAccount`,
`handleUnlinkAccount`, `handleListUserSessions`, `handleRevokeUserSessions`, and
`validateUsername`. All 0%; no `handlers_users*_test.go` exists. (The adjacent
`DELETE …/passkeys` remediation handler *is* tested by `TestUnbindUserPasskeys`.)
- `handleReady` — the internal-face "server is up" push (distinct from the tested
`handleReadyz`); 0%.
These handlers are owner/operator-role-gated, so the blast radius is bounded, but
`validateUsername` is load-bearing input validation and is the highest-value target
for a follow-up test. **Recommendation:** add an `handlers_users_test.go` covering
create/quota/link + `validateUsername` negative paths. Filed as a proposed change,
not made here (this is a read-and-verify audit — no test/impl was modified).
## Cheap tells (static pre-pass)
- 3 `t.Skip` sites, all benign: RNG-collision reruns (a 1-in-10^6 OTP code clash),
not coverage-gating skips.
- No test file falls below 2 assertions per test function.
## Documented-tested claim reconciliation
To avoid the subset-verified/whole-worded trap a second time, every "unit-tested"
string in the ledger docs was enumerated (`grep -niE "unit-tested" docs/changes/*.md`)
and mapped to a verdict — verified fail-open gates get a mutation; behavioural/contract
claims are scoped, not silently dropped:
| Doc claim | Verdict |
|---|---|
| modpack: approval CAS builds once | mutation #7 |
| modpack: **scan in front of any push** | mutation #16 |
| cloudflare: `validateFailClosed` refuses public policy | mutations #8–#9 |
| cloudflare: **conn-count fence gate** | mutation #15 |
| system-servers: **naming reservation** | mutation #17 |
| system-servers: **service-token → login pod only** | mutation #18 |
| operator: idle stop / startup+readiness timeout / `/readyz` | mutations #12 / #13 / #14 |
| auto-update: pin / no-downgrade / no-prerelease / window / strict-`After` | mutations #1–#5 |
| passkey: clone-warned assertion refused | mutation #6 |
**Scoped, NOT individually mutation-proven** (behavioural/contract-level, not fail-open
safety gates — they rest on the green suite + the coverage backbone, and are called out here
rather than folded into the verdict):
- console-auth: content-type guard, anti-enumeration, forced-change lockdown. Anti-enumeration
is the one with security weight; the current public login door is email-OTP/passkey, and its
anti-enumeration behaviour is a candidate for a future mutation pass.
- break-glass setup: owner-auth match/non-match over the fake store.
- username-reclaim + auto-update JSON round-trip: in-memory repo contract / serialization.
## Toolchain honesty
Only Go runs on the oracle. The felis-limbo plugin (Java) is podman-verified against
a real Limbo jar (#65); the limbo/lobby images carry a build+boot check (#63). Those
completions were never a green-Go-tests claim and are not audited as if they were.
## Verdict
The commit history's verification claims are **accurate**: all 18 load-bearing *fail-open
safety gates* the ledger names as tested — spanning every subsystem, reconciled one-for-one
against the docs' "unit-tested" claims above — are mutation-proven to pin behaviour, not
merely to pass. No crown-jewel test was vacuous. The shortfalls are (a) integration seams
unrunnable on the oracle by design (honestly classified — including the live
`cloudflared`/`nft` fence-apply, whose *decision* gate is nonetheless verified); (b) one
untested cluster of admin user-management handlers — a missing test, not a false green; and
(c) a residue of behavioural/contract-level "unit-tested" claims (console anti-enumeration,
break-glass owner-auth, repo/JSON contracts) that rest on the green suite plus coverage and
are scoped above rather than individually mutation-proven — the honest boundary of this pass.
**Method note.** The first pass mutation-verified 14 gates but worded its verdict as "every"
invariant; a review caught that 4 documented safety gates (fence, scan, naming, service-token)
were named-as-tested yet unverified, and one (the fence) was mis-classified as integration-only.
Those four are now mutation-proven (#15–#18) and the classification corrected. The lesson is
the audit's own thesis turned on itself: *reading a scope and judging it complete* reproduces
the *looks-right ≠ verifies* error — only the enumerate-and-mutate reconciliation earns the word.
@@ -1,129 +0,0 @@
# Adversarial input-validation audit — every dangerous sink fails closed
- **Type:** negative-path / input-validation audit (no production code change)
- **Date:** 2026-07-08
- **Area:** `internal/api`, `internal/naming`, `internal/submit`, `internal/rcon`
- **Task:** a different question from the #82 / round-2 mutation audits. Those asked
*do the TESTS catch a gate regression?* This asks *does the CODE reject hostile
INPUT, or does bad data PASS?* — feed the real endpoints malformed, boundary, and
hostile bodies (故意加错误数据) and confirm they fail closed (4xx) rather than letting
the garbage reach a sink.
## Method — sink-first, not fuzz-everything
The low-hanging garbage (oversized body, unknown field, wrong content-type) is already
caught by the universal body guards, so a blanket "fuzz all ~80 handlers" would burn
effort where the answer is known. The real "can bad data PASS?" risk lives at the
**sinks** — the few places a request string is concatenated into an RCON command, used
as a K8s object name, joined into a filesystem/archive path, put in a SQL query, or
accepted as an enum/quantity **without a validator in front**. So the audit traces each
dangerous sink class from its handler entry to the sink, and for the crown-jewel class
(text → RCON) **mutation-verifies** the guard is non-vacuously pinned: loosen the guard
in source, run the package tests, confirm the specifically-named negative test reddens
(`--- FAIL: <subtest>`), then revert. Oracle: WSL Fedora-44, go1.26.4.
## The universal belt (caught before any sink)
`decodeJSON` (`internal/api/util.go`) wraps every body in
`http.MaxBytesReader(w, r.Body, 1<<20)` (1 MiB cap), sets `DisallowUnknownFields()`, and
rejects trailing data after the first JSON value — all → `400 bad_request`.
`requireJSONContentType` returns `415 unsupported_media_type` on credential writes
(a CSRF belt). So oversized, unknown-field, multi-document, and wrong-type bodies never
reach a handler body at all.
## The dangerous sinks — each traced fail-closed
### 1. Text → RCON (the lead). Two vectors, both fenced.
**(a) Structured commands** — `handlers_access.go` (LuckPerms permission/group,
whitelist/ban/kick). Every operand is validated against an anchored allow-list charset
*before* it is concatenated: `mcNameRe = ^[A-Za-z0-9_]{1,16}$` (player),
`lpNodeRe = ^[A-Za-z0-9_.*-]{1,64}$` (node), `lpCtxRe = ^[A-Za-z0-9_-]{1,48}$`
(world/group). No space, separator, or control character can appear in a validated
operand, and **there is no free-text field anywhere** — a ban/kick deliberately carries
no reason string (that would be the one splice vector). Go's `$` is `\z` (absolute end,
not `\Z`), so even a single trailing `\n` is rejected.
*Mutation-verified:* loosening `mcNameRe` to admit a space
(`^[A-Za-z0-9_ ]{1,16}$`) reddens
`TestAccessInjectionRejected/{whitelist,ban,kick,permission}_player_space` — the guard
is real, not vacuous.
**(b) Free-text passthrough** — `handlers_console.go` `handleCommand`, POST
`/servers/{name}/command`. This is the *one deliberate* free-text → RCON vector, and it
is **owner/admin-gated** (403 for a stranger, 404 for an unknown server). Its input
fence: trim + strip a single leading `/`, reject empty, cap at 1024 bytes, and
`strings.IndexFunc(command, func(c rune) bool { return c < 0x20 }) >= 0 → 400` — every
C0 control (incl. `\n`) is rejected so one request cannot splice a second command.
*Mutation-verified:* disabling the scan (`c < 0x20` → `c < 0x00`) reddens
`TestConsoleCommand/control_character_(newline)_->_400,_no_RCON_call`, whose input is
literally `{"command":"say hi\nop attacker"}` and whose assertion is 400 **and**
`console.calls == 0`. (Severity note: because this vector is owner-gated by design, the
scan is an audit-integrity measure — one request = one command — not a privilege
boundary; a splice on your *own* server escalates nothing, since the owner may already
run any RCON command. The fence exists regardless.)
### 2. Break-glass / internal-face (the newest code — scrutinised specifically)
This surface runs under a "service-token-authed / local-root, inputs trusted" posture,
the classic place a field-level guard gets skipped. Traced end to end:
- **Server name** — every internal handler (`handleReady`, `handleJoinEvent`,
`handleInternalWake`, `handleInternalClaim`, `handleInternalMenuStatus`) validates the
path name with `naming.ValidateServerName` before use.
- **`mc_uuid`** (join/wake/claim bodies) — checked non-empty, then flows *only* to
DB-parameterized calls (`RecordJoin`, `UserByMCUUID`, `UUIDInAllowlist`). The
"allowlist" is a **DB table**, not a live RCON `whitelist add` — there is no
`mc_uuid` → RCON path.
- **Break-glass "OP-create"** (`performAddOperator` → `provisionOperator` →
`InsertOperator(ctx, id, username, email)`) is a **parameterized DB INSERT** creating a
*panel staff account*, **not** a Minecraft `op` RCON command. The hypothesised
name → RCON `op` sink was checked and **does not exist** in this shape; the username is
`TrimSpace`d and reaches only `$N`-parameterized SQL, from a local-root caller.
- **`os_user`** (internal backup attribution) — `TrimSpace`d, sets only the audit actor
(a DB row); shown non-vacuous in the round-2 backup/restore audit (`b7b4a3b`).
### 3. SQL injection — dismissed.
`pgrepo.go` uses uniform `$1/$2/$3` parameterization throughout
(`QueryRowContext`/`ExecContext(ctx, q, args…)`); no request string is `Sprintf`'d into a
query.
### 4. Path traversal (submit) — dismissed.
`internal/submit` validates the submission id (rejects `..`, path separators, uppercase,
space, empty), and the on-disk blob name is a **fixed** constant (`contextBlobName`) — no
attacker-supplied filename is ever joined. The hostile-id matrix
`{"../evil","sub/../../etc","SUB-UPPER","has space","","a/b"}` is test-pinned in both the
local and S3 backends. The one free-form field a submission carries (`DisplayName`) is
charset-constrained by `displayNameRE` and rejects control chars
(`submit_test.go` "control chars" case).
### 5. K8s object names — validated at every cluster write.
`ValidateServerName` / `ValidateSystemServerName` (`^[a-z0-9-]{3,32}$`, no
leading/trailing dash, reserved-name set) and `ValidateHostname`
(`dnsLabelRE`, single label under the configured root domain) gate every create/patch.
`cluster.go` documents the invariant: "there is no free-form YAML path — every field is a
typed, validated value," so no raw CRD field can be smuggled through a create/patch body.
### 6. Numeric / enum — fail-closed.
`resolveResources` routes **every** quantity (memory, `resources.cpu/memory`, requests)
through `parsePositiveQuantity`, which rejects `q.Sign() <= 0` (negative *and* zero) with
a field-named 400, plus a request>limit guard; `parseStorageSize` carries the same guard.
Enums are closed sets: `parseAutostartPolicy` (ownerOnly/public/allowlist), the
access-action switch, and the image `ImageAdmitted` allow-list (no free image string).
## Verdict
**Bad data does not pass.** Every dangerous sink is fail-closed — including the
break-glass / internal-face surface, where the hypothesised `mc_uuid`/OP-create → RCON
paths were traced and found not to exist (parameterized DB, not RCON). Both text → RCON
vectors — structured (`handlers_access`) and free-text (`handleCommand`) — are
mutation-pinned by their named negative tests. No gap was found and no production code
changed; the honest result of "故意加错误数据" is that the input surface rejects it.
Scope is deliberately bounded to the dangerous sinks and the newest (break-glass) code,
not an exhaustive fuzz of all ~80 handlers — the claim proven is "every place user input
reaches a dangerous sink validates before the sink," by trace plus two mutations, not
"every handler was fuzzed."
@@ -1,57 +0,0 @@
# §B4 phase close — S3 archive backend deferred by design (decision record)
- **Type:** decision / scope record (no code changed)
- **Date:** 2026-07-08
- **Area:** `internal/config`, `internal/backup` — the archive-store backend selection
- **Task:** #31 Phase B4 break-glass ops — **closes the phase**, superseding the phase-2b
note in [break-glass-backup-peer](2026-07-07-break-glass-backup-peer.md) ("S3 remains
open in B4; this does not close the phase").
## Decision
The three operational B4 break-glass ops are built, wired into the recovery console menu,
and oracle-verified:
| Op | Menu enum | Commit |
|---|---|---|
| Provision/reset Owner | `bgProvisionOwner` | (Phase B1 lineage) |
| Add Operator ("OP create") | `bgAddOperator` | recovery-console menu |
| Halt a running server | `bgHaltServer` | `c2ee21a` |
| Back up a world now ("Sync") | `bgSyncBackup` | `7a7c0d5` / `f2fc57c` / `fc748d3` |
The fourth B4 line item — **S3 archive backend (`tarS3`)** — is **deferred by design**, not
left as a silent gap. It is closed as a documented deferral and **#31 is done**.
## Why deferring is safe (not a loose end)
- **Fail-closed at config load, frozen by a test.** `config.Load` rejects
`store = "tarS3"` (and `volumeSnapshot`, `longhorn`) with an error that points the
operator at the `tarLocal` remediation. `TestLoadRejectsUnimplementedArchiveStore`
freezes exactly this: a config naming an unimplemented backend fails at load, so
felis-api can never boot green while the reaper CronJob fails every run and restore
silently 503s. tarS3 cannot be selected into a broken state.
- **Peer to two other deferred backends.** `tarS3` sits beside `volumeSnapshot` and
`longhorn` as recognized-but-unimplemented store names. The spec's own phasing is
tarLocal-first ("起步 `tarLocal` … 要异地/跨集群 → `tarS3`"): the baseline single-node
path is `tarLocal` (tar → backup PVC), which is implemented, tested, and the backend
every built backup/restore path uses today.
- **Offsite/cross-cluster DR is the only capability gap**, and it is opt-in future work,
not a correctness hole in the shipped baseline.
## The build path, when offsite DR is wanted
`minio-go/v7` is already vendored (the modpack upload lane's `internal/submit/s3store.go`),
so tarS3 adds no dependency. A future build is bounded:
1. `internal/backup/tars3.go` — a `WorldArchiver` reusing the existing package-level
`writeTarGz`/`readTarGz`/`pruneToManifest`, streaming the tar to an object via
`PutObject` (size −1, multipart) and reading it back via `GetObject`, mirroring
`s3store.go`'s fakeable-client testability.
2. Store-selection factory in the backup/reaper entrypoint (`store = "tarS3"` → construct
the minio-backed archiver) + remove tarS3 from the config fail-closed list (update
`TestLoadRejectsUnimplementedArchiveStore` to keep only volumeSnapshot/longhorn).
3. Inject the S3 Secret into the backup/reaper Job Pods (integration wiring, like the
Kaniko S3-context credential path).
The live-S3 upload + Secret-into-Pod would be integration-only verified, exactly as the
modpack S3 backend and the pgrepo SQL are.
@@ -1,93 +0,0 @@
# Backup/restore data-safety mutation audit (round 2) — 7 gates pinned, 1 gap closed
- **Type:** test-quality audit + one test added (code change: `internal/api/handlers_backups_test.go`)
- **Date:** 2026-07-08
- **Area:** `internal/api` (`handlers_backups.go`), `internal/backupjob`
- **Task:** continues the #82 test-quality integrity audit onto the on-demand
backup/restore surface, which postdates the 18-gate round-1 audit
([mutation audit `4626ab5`](2026-07-07-test-quality-mutation-audit.md)). The
backup/restore endpoints (`7a7c0d5` / `f2fc57c` / `fc748d3`) were not in that pass.
## Method
Same as round 1: apply a one-line mutation to a fail-open gate in the source, run the
package tests, confirm the **specifically-named** test reddens with an *assertion*
failure (`--- FAIL: <subtest>`), then revert. A build break (`declared and not used`,
`undefined`) is not a valid verdict, so mutations are operator-flips that keep every
operand referenced (`!=`→`==`, drop a `!`, a literal→`true`, or `if false && <orig>` to
disable a gate without orphaning its variables). Airtightness: each mutation is re-run
with `-run` scoped to the intended subtest and `grep -- "--- FAIL: <subtest>"`, so a
reddening sibling can't be mistaken for the gate under test. Oracle: WSL Fedora-44,
go1.26.4.
## Fail-open gates mutation-verified (all CAUGHT at the named subtest)
| # | Gate (file:line) | What it guards | Mutation | Subtest that reddened |
|---|---|---|---|---|
| A | `handlers_backups.go:282` enqueueBackup stopped-gate | RWO double-mount / torn archive while the world is up | `!=`→`==` | `TestBackupNow/starting_server_->_409_not_stopped` |
| B | `:151` restore stopped-gate | restore Job can't mount a live world's RWO PVC | `!=`→`==` | `TestRestoreBackup/starting_server_->_409_not_stopped` |
| C | `:136` restore former-owner match | a fresh claimant resurrecting the previous owner's world | `!=`→`==` | `TestRestoreBackup/current_owner_who_is_not_former_owner_->_403` |
| D | `:116` restore cross-server guard | restoring server A's backup onto server B | `!=`→`==` | `TestRestoreBackup/restore_by_backup_id_cross-server_->_403` |
| E | `:214` backup owner-or-admin authz | a stranger backing up someone else's world | drop `!` | `TestBackupNow/non-owner_->_403,_no_backup` |
| F | `:26` list cross-user scope | a user seeing other tenants' backups | `p.IsAdmin()`→`true` | `TestListBackups/user_sees_only_own_former-owned_present_backups` |
## The gap this audit found — and closed
**Restore's owner-or-admin gate (`handlers_backups.go:85`) was not pinned by any test.**
It is the twin of gate E, but the two are *not* symmetric. Disabling it
(`if false && !a.isOwnerOrAdmin(p, rec)`, which keeps `rec` referenced so the package
still builds) reddened **nothing** — `go test ./internal/api/` stayed `ok`. The same
disable applied to backup's L214 (gate E) reddened `non-owner` immediately, proving the
technique valid and the asymmetry real.
Root cause: the former-owner gate at L136 backstops every non-owner case the suite
exercised (a stranger and a wrong-backup current owner both fail L136 *and* L85, so
L136's 403 masks a broken L85). The one case only L85 catches went untested: a
**superseded former owner** — a user who owned a server, took this backup
(`FormerOwner=them`), then released it to a *new* owner. They still pass L136 (they *are*
the former owner) but must be stopped by L85, or they could roll the new owner's live
server back onto their old world (cross-tenant clobber). The handler comment names this
the "must re-claim first" rule (`handlers_backups.go:77-79`).
**Fix (code):** added `TestRestoreBackup/former owner after release -> 403, no restore`,
the mirror of the existing L136 test. Verified both directions: green on the clean tree,
and it is the sole subtest that reddens when L85 is disabled — so it now pins the owner
gate specifically, not L136. No production code changed; `handlers_backups.go` is a pure
test addition away from where it was.
## Enumeration — covered vs. scoped (so "the gates" means all of them)
- **Fail-open data-safety gates — all pinned:** A–F above, plus L85 (now closed). 7/7.
- **Accountability, not fail-open (verified non-vacuous):** the `os_user` attribution at
`handlers_backups.go:254` — a supplied operator name overrides the default `break-glass`
audit actor. Mutating `u != ""`→`u == ""` reddens
`TestInternalBackup/os_user_body_attributes_the_audit_to_the_operator`, so the
attribution test isn't vacuous. A failure here degrades the audit actor; it grants no
bypass, so it is out of the fail-open bucket.
- **Contract/behavioral (tested, out of mutation scope):** nil `Backuper`/`Restorer` → 503;
a failed backup/restore → 500 **not** audited; `backup_ref` never serialized to the wire;
the internal-face actor defaults to `break-glass`. Each has a direct test; none is a
fail-open safety gate.
- **`internal/backupjob` (glanced, not mutated):** orchestration only — each backup gets a
unique Job name (`BackupJobName` + random suffix) so a repeat "立即备份" tap can't collide
with a just-finished Job still inside its TTL; `ErrAlreadyExists` is a defensive no-op;
`Backup` returns once the Job is created (the async 202 is honest). Unit-tested against a
fake `Jobs`; the controller-runtime `k8sjobs.go` and the `jobspec.go` Pod shape (weak SA
with its token un-mounted, config Secret mounted for the self-recorded row, read-only
world mount) are integration-verified per the package doc — not fail-open API gates.
## Coverage nuance (documented, not a gate failure)
The stopped-gate is `if info.Ready || info.DesiredState != DesiredStopped`. The `!=`→`==`
mutation pins the `DesiredState` operand (both A and B reddened), but `info.Ready` is not
*independently* pinned: no test sets `Ready=true` together with `DesiredState=Stopped` —
the stopping-but-still-up race. Low risk because in practice `Ready` drops as
`DesiredState` leaves `Stopped`, but the belt-and-suspenders `Ready` operand rides on
coverage of the operand beside it rather than its own case.
## Verdict
The backup/restore data-safety surface is a coherent unit, and this closes it: **7/7
fail-open gates pinned** (6 pre-existing, 1 added this round), one accountability gate
shown non-vacuous, one coverage edge documented. Not extended to every handler — that
would be an unbounded "continue the audit."
@@ -1,110 +0,0 @@
# Felis-nano — the federating `hasJoined` multiplexer (step 1: the Go resolver)
- **Type:** feature (new endpoint) — the verifiable "brain" of Felis-nano
- **Date:** 2026-07-12
- **Area:** `internal/api` (`handlers_hasjoined.go` + test), `internal/api/api.go`
(route + `AuthSources` field), `docs/openapi.yaml`, `go.mod`
- **Task:** Felis-nano provides a MultiLogin-like capability — one Velocity proxy that
accepts logins verified by **several** Yggdrasil auth servers at once (Mojang + N
third-party roots), 正版优先 (Mojang-first). This change builds **step 1**: the Go
`hasJoined` multiplexer that does the federating verification. It is the only part of
the plan that produces immediate verifiable hard evidence (a unit-tested HTTP endpoint);
the two delivery shells that point Velocity at it (a JVM `-Dmojang.sessionserver` flag,
and a thin reflection-hook plugin for third-party servers) are later steps.
## What Velocity asks for, and what this answers
On a Minecraft login Velocity's authlib computes the `serverId` hash and issues
`GET /session/minecraft/hasJoined?username=<name>&serverId=<hash>[&ip=<ip>]` against
whatever URL its `mojang.sessionserver` system property names. A 200 with a game profile
means "verified"; a 204 means "not verified" and authlib rejects the login. Vanilla points
this at Mojang alone. Felis-nano points it **here**, and this endpoint fans the same query
out to the configured Yggdrasil roots **in priority order**, returning the first source
that validates. Each upstream Yggdrasil runs its own `serverId`-hash check — the
multiplexer only relays, it computes no hashes.
## The one non-negotiable transform — per-source UUID namespacing
A third-party Yggdrasil's UUIDs are **self-asserted**: nothing stops a malicious source
from answering with a *genuine Mojang player's* UUID. If that UUID were emitted as-is, the
third-party could impersonate any Mojang player with full UUID fidelity — and the reclaim/
blacklist layer could never catch it, because its whole invariant is "the genuine Mojang
player has a **different** UUID from any squatter." That invariant would simply be false.
So the resolver rewrites every non-identity source's profile into a per-source namespace
**before it leaves the resolver** — the single entry point every login crosses:
```
canonical = UUIDv3(felisAuthNS, tag + ":" + nativeID) // third-party
canonical = the source's UUID verbatim // Mojang (Identity: true)
```
MD5 (UUIDv3) preimage resistance means no third-party can mint a value inside Mojang's
UUID space; the per-`tag` prefix means two sources can't collide onto one identity. Every
downstream key — `account_links`, `username_blacklist`, owner checks — then sees exactly
one canonical UUID per real identity, so the reclaim invariant is true **by construction**,
not by assumption.
## Fail-closed details that bite if wrong
- **Bar gate at the chokepoint.** The canonical UUID is checked against
`Repo.IsUsernameBlacklisted` *before* the profile is returned, so a reclaimed squatter
stays out even on a consumer that has no limbo plugin. Keyed on the **dashed** canonical
(`.String()`) — the exact form `Repo.ReclaimUsername` stores. A DB error there fails
closed (non-200 → authlib rejects), matching the existing `handleCheckBlacklist` pattern.
- **Emit undashed.** authlib's `GameProfile` expects the 32-hex undashed `id`
(`hex.EncodeToString(u[:])`); the DB/reclaim/blacklist keys are dashed. The resolver
**checks** on the dashed string and **emits** the undashed one. Mixing the two forms is a
silent gate miss — pinned by the tests below.
- **`properties` relayed verbatim** (`[]json.RawMessage`) so a source's signed textures
survive the multiplexer untouched.
- **Inert by default.** `AuthSources` is nil until `cmd/felis` wires configured sources,
so the endpoint 204s every login until deliberately configured — it ships off.
- **Public internal-face route.** authlib sends no service token, so the route is mounted
`Public: true` on the internal face (like `/healthz`); no third face is introduced. The
OpenAPI parity test enforces `x-felis-face: [internal]` + `x-felis-tier: public`.
## Files
| File | Change |
|---|---|
| `internal/api/handlers_hasjoined.go` | **new** — `handleHasJoined` + `resolveHasJoined` + `AuthSource`/`sessionProfile` types + `felisAuthNS` |
| `internal/api/handlers_hasjoined_test.go` | **new** — `TestHasJoined`, 7 subtests over `httptest` fake Yggdrasil roots |
| `internal/api/api.go` | `AuthSources []AuthSource` field (nil = inert) + `GET /session/minecraft/hasJoined` `Public` internal route |
| `docs/openapi.yaml` | `/session/minecraft/hasJoined` path — `x-felis-face: [internal]`, `x-felis-tier: public`, `security: []` |
| `go.mod` | promote `github.com/google/uuid` indirect→direct (first direct importer) |
## Verification evidence
Oracle: WSL Fedora-44, go1.26.4. `go build ./...` → `BUILD-OK`. Full `internal/api`
package green (`ok felis.lolicon.best/internal/api`), `go vet ./internal/api/` clean. The
full package (not a `-run` filter) was run because this change edits two shared surfaces —
the `API` struct and the `internalAPIRoutes()` table — where a route that isn't under
`/api/v1/` is exactly what a route-table-driven invariant test would trip; nothing
reddened.
`TestHasJoined` — 7 subtests, all PASS:
1. `mojang identity passthrough` — Mojang UUID unchanged, `properties` relayed.
2. **`thirdparty UUID rewritten, never emitted as-is`** — the security invariant: an evil
source returns real Notch's Mojang UUID; the resolver emits neither that UUID nor any
Mojang-space value, but the deterministic `UUIDv3(felisAuthNS, "littleskin:"+id)`.
3. `mojang priority wins over thirdparty` — Mojang-first ordering.
4. `fallthrough to thirdparty when mojang 204s` — priority scan continues past a 204.
5. `no source validates -> 204`.
6. `barred canonical UUID -> 204` — the reused reclaim bar gate holds at the resolver.
7. `missing username -> 204` — no source touched on a malformed query.
`TestOpenAPIMatchesServedRoutes` PASS — the new route's served facets match its
`docs/openapi.yaml` entry in both directions.
## Deferred (not in this change)
- **Source configuration** (step 2): a `tag`/`type`/`url`/`priority` schema and
`cmd/felis` wiring that populates `AuthSources`. Until then the endpoint is inert.
- **Delivery shells** (step 3): Shell 1 = the `-Dmojang.sessionserver` JVM flag on
Felis-managed proxies; Shell 2 = the thin reflection-hook Velocity plugin for
third-party servers, plus a Velocity verification runbook. The user has a real
server to test Shell 2 against.
- **Name-match / textures-signature enforcement** — deliberately out of scope: identity
is `tag:nativeID`, not the name, and textures are the cosmetic bucket relayed verbatim.
-255
View File
@@ -1,255 +0,0 @@
# Felis change ledger
The index of every functional change to Felis — what it did and which commit records
it. This is the durable, in-repo map that `git log` alone doesn't give: it links
substantial changes to their detail docs and flags work that is built and verified but
not yet committed.
## Convention
- **Every functional change** (a feature addition, a behaviour change, a bug fix) gets:
1. a dated detail doc in this directory — `docs/changes/YYYY-MM-DD-<slug>.md`, covering
_what it did, why, the files touched, and the verification evidence_; and
2. a row in the ledger below, carrying its **commit record** (the short SHA).
- A change that is **built and verified but not yet committed** (e.g. while PGP signing
is locked) sits in **Pending** with `commit: pending`, and moves into the ledger with
its real SHA once committed.
- Pure-cosmetic or non-functional commits (docs, style) still appear in the ledger table
for completeness, but do not require a dedicated detail doc.
- The ledger table is generated losslessly from git history and can be regenerated:
```
git log --reverse --pretty=format:'| %h | %ad | %s |' --date=short
```
## Pending (built + verified, not yet committed)
_None._
## Detail docs
Depth docs for substantial changes, keyed to the commit(s) they cover. The committed
ledger table below stays a lossless mirror of `git log` (so it can be regenerated); this
section is where a row's detail doc, when it has one, is found. Most rows — panel/UI,
docs, chore, style — have no detail doc by convention and are recorded by their table row
alone. Entries marked *(backfill)* were reconstructed retroactively on 2026-07-07 from git
history to close the ledger's detail-doc axis for the pre-convention functional commits;
each carries a backfill note stating it was not independently re-verified. Frontend/`panel`
commits are the collaborator's UI work and are not given detail docs here.
| Detail doc | Commit(s) | Scope |
|---|---|---|
| [foundational-subsystems](2026-06-26-foundational-subsystems.md) *(backfill)* | `7fbebfe` `708cdfc` `43ab921` `78b8cf6` `d39605e` `b508fcc` `47fcd90` `93f143f` | initial import: CRD, core libs, backup, operator, submit, api, platform, plugins |
| [modpack-submission-lane](2026-06-26-modpack-submission-lane.md) *(backfill)* | `d39605e` `598f3d3` | §8 modpack build/approval pipeline + local/S3 backends |
| [deploy-bootstrap-installer](2026-06-27-deploy-bootstrap-installer.md) *(backfill)* | `58fa4b0` `94a3b7b` `deaa2f8` `318a724` `e5f1682` `28c3eee` `c14ed17` `d9e866f` `b84debf` | one-line bootstrap installer + demo bring-up |
| [console-auth-passwordless](2026-06-27-console-auth-passwordless.md) *(backfill)* | `af14f02` `0c1cc59` `3b43f05` `c20b12c` | local-password login → passwordless migration + residue sweep |
| [felis-cli-break-glass-setup](2026-06-27-felis-cli-break-glass-setup.md) *(backfill)* | `e108a37` `2d0bbb0` `a94b001` `eb5875a` `f5d00f3` `9c46632` `7d91373` | break-glass recovery console + first-run setup + apply/migrate |
| [cloudflare-tunnel-access-edge](2026-06-27-cloudflare-tunnel-access-edge.md) *(backfill)* | `53a7664` `ba13839` `a531f5e` `2810fe8` `7d3be64` `346ec68` `e058a64` | §14 Tunnel + fail-closed Access edge + NodePort fence |
| [player-onboarding-b2](2026-06-27-player-onboarding-b2.md) *(backfill)* | `dbe34a1` `1f8b9bb` `116595f` `fe2ece0` `55592ed` `6c3999a` `879b177` | §B2 email-OTP, account-link, QR, Bind-Code + OTP throttle |
| [username-reclaim-b3](2026-06-27-username-reclaim-b3.md) *(backfill)* | `a29571d` `fdb6efb` | §B3 Mojang-priority reclaim + account migration |
| [felis-metrics](2026-06-30-felis-metrics.md) *(backfill)* | `75642d9` `2a93a9e` `79eae7f` `8ac5e64` | §23 felis_* Prometheus collectors |
| [felis-api-hardening](2026-06-30-felis-api-hardening.md) *(backfill)* | `7a51c1d` `164ac44` `c6c0772` `3c1d647` `d6e3189` `8f41a00` `6368ab1` `2a4a81b` `9873904` | audit #1–#3 + robustness fixes |
| [passkey-enrollment](2026-07-01-passkey-enrollment.md) *(backfill)* | `f2c916d` `742f15f` `0261204` `fce0fce` `7278cd7` `cdbb5ab` `9953275` `20e31fb` `54bc6ef` | WebAuthn enrollment + hardening a–e |
| [passkey-login](2026-07-01-passkey-login.md) *(backfill)* | `e035142` `ec468ba` `0dbd557` `9e1df12` `4f59d51` `a63f49d` | WebAuthn assertion/discoverable login + UA-guard |
| [auto-update-subsystem](2026-07-01-auto-update-subsystem.md) *(backfill)* | `c01f133` `3673af6` `7464fa7` `96b3cc9` `7d27640` `7db57b9` | update decision core + sources + gatherer + window API (report-only) |
| [system-servers-login-limbo-lobby](2026-07-02-system-servers-login-limbo-lobby.md) *(backfill)* | `9bed51b` `9ef817f` `159107b` `dc23cb5` `3fdb3d0` `f554d52` `241fe21` `c7315e4` | always-on login-limbo + lobby auth gate |
| [operator-idle-quota-readiness](2026-07-05-operator-idle-quota-readiness.md) *(backfill)* | `91bfa27` `e574749` `7f7e459` `7becb38` | idle auto-stop, quotas, timeouts, /readyz |
| [break-glass-halt](2026-07-05-break-glass-halt.md) | `c2ee21a` | §B4 break-glass halt-a-server op |
| [felis-migrate-command](2026-07-05-felis-migrate-command.md) | `c1aa38b` | §B3 `/felis migrate` account migration |
| [on-demand-world-backup](2026-07-07-on-demand-world-backup.md) | `7a7c0d5` | §B4 Sync phase 1 — external backup endpoint + Job |
| [internal-backup-endpoint](2026-07-07-internal-backup-endpoint.md) | `f2fc57c` | §B4 Sync phase 2a — internal-face backup endpoint |
| [internal-api-clusterip-service](2026-07-07-internal-api-clusterip-service.md) | `2ba9948` | felis-api internal-face ClusterIP Service |
| [break-glass-backup-peer](2026-07-07-break-glass-backup-peer.md) | `fc748d3` | §B4 Sync phase 2b — console backup peer |
| [adversarial-input-audit](2026-07-08-adversarial-input-audit.md) | `667c6d3` | adversarial input-validation audit — sink-first negative-path, two text→RCON guards mutation-pinned, no code change |
| [felis-nano-hasjoined-resolver](2026-07-12-felis-nano-hasjoined-resolver.md) | `ff550c4` | Felis-nano §B3 — federating hasJoined multiplexer + per-source UUID namespacing |
## Committed change ledger
Oldest first (project build order). Commit = short SHA on `main`. Frontend/`panel`
commits are the collaborator's UI work; backend (Go/Java/K8s) is tracked here as the
primary record.
| Commit | Date | Change |
|---|---|---|
| 5b7b38d | 2026-06-26 | chore: add Go module manifest and ignore rules |
| 5a30aa5 | 2026-06-26 | docs: add OpenAPI 3.1 served-route contract |
| 7fbebfe | 2026-06-26 | feat(apis): add MinecraftServer CRD types (v1alpha1) |
| 708cdfc | 2026-06-26 | feat(core): add naming, RCON, store, config, and image-build libraries |
| 43ab921 | 2026-06-26 | feat(backup): add backup, restore, and reaper subsystems |
| 78b8cf6 | 2026-06-26 | feat(operator): add MinecraftServer controller and reconcilers |
| d39605e | 2026-06-26 | feat(submit): add user modpack build and approval pipeline |
| b508fcc | 2026-06-26 | feat(api): add felis-api service with permissions, modpack lane, and fleet read |
| 47fcd90 | 2026-06-26 | feat(platform): add node orchestration and the felis entrypoint |
| eee00c2 | 2026-06-26 | feat(panel): add three-sided web console (User, Admin, SysAdmin) |
| 93f143f | 2026-06-26 | feat(plugins): add Velocity proxy and Fabric/Forge/NeoForge/Paper integration mods |
| ce0ba76 | 2026-06-27 | chore(api): add kubebuilder object-generation markers to v1alpha1 |
| 7d91373 | 2026-06-27 | fix(migrate): honor -config flag placed after the up verb |
| 99de43f | 2026-06-27 | chore: ignore plugin build artifacts and editor config |
| 58fa4b0 | 2026-06-27 | feat(deploy): add one-line bootstrap installer and container image |
| af14f02 | 2026-06-27 | feat(api): local-password authentication backend |
| e108a37 | 2026-06-27 | feat(cli): break-glass emergency console TUI |
| 885c4a9 | 2026-06-27 | feat(panel): local-password login and forced password change |
| 2d0bbb0 | 2026-06-27 | feat(cli): attribute break-glass recovery to the SysAdmin who runs it |
| dbe34a1 | 2026-06-27 | feat(api): add player email OTP verification (spec §B2 onboarding) |
| 1f8b9bb | 2026-06-27 | feat(api): record account-link auth source (mojang/thirdparty) |
| a29571d | 2026-06-27 | feat(api): reclaim squatted usernames for Mojang-priority players (spec §B3) |
| 53a7664 | 2026-06-27 | feat(cfsetup): recommended Cloudflare Tunnel + Access edge setup |
| ba13839 | 2026-06-27 | feat(breakglass): optional Cloudflare Tunnel + Access setup in the TUI |
| a5a6482 | 2026-06-27 | feat: dev mock |
| dd2fc6f | 2026-06-27 | feat(panel): i18n |
| 7b458da | 2026-06-27 | feat(panel): light/dark theme |
| 51c9eab | 2026-06-27 | docs: CONTRIBUTOR.md |
| 74e7e6e | 2026-06-27 | docs: CONTRIBUTING.md |
| b81b334 | 2026-06-27 | Merge branch 'main' of https://github.com/MliroLirrorsIngenuity/Felis |
| 9d13787 | 2026-06-27 | refactor(panel): dashboard |
| f3521e3 | 2026-06-27 | refactor(panel): uniform margins |
| 18b4be0 | 2026-06-27 | refactor(panel): uniform title icon styles |
| 665841b | 2026-06-27 | fix(panel): remove internal spec references from user-facing text |
| cf88bcc | 2026-06-27 | fix(panel): extract hardcoded security note into i18n keys |
| 061482d | 2026-06-27 | fix(panel): extract hardcoded Chinese text to i18n keys |
| cfe126f | 2026-06-27 | feat(panel): add RCON command input to server console |
| ae91133 | 2026-06-27 | style(panel): refine button styles with shadow, active scale, toned-down colors |
| e189265 | 2026-06-27 | refactor(panel): compact server card layout, denser grid |
| f4df3e2 | 2026-06-27 | feat(panel): pagination for server lists |
| b512a18 | 2026-06-27 | refactor(panel): adjust margins |
| a49443c | 2026-06-27 | refactor(panel): simplify sidebar |
| 86f2ae4 | 2026-06-28 | feat(panel): sidebar foot shows current account + sign-out; reorder Account cards |
| f5d00f3 | 2026-06-28 | feat(cli): implement felis apply command for direct CRD creation |
| 832b200 | 2026-06-28 | fix(panel): reactive system theme detection |
| c9cd9dc | 2026-06-28 | style(panel): unify dialog animation to fade and scale from center |
| 94a3b7b | 2026-06-28 | fix(deploy): harden bootstrap for RHEL-family Linux |
| 9c46632 | 2026-06-28 | feat(cli): add felis setup first-run console with reclaim protection and cfsetup idempotency |
| 5450c26 | 2026-06-28 | chore: normalize line endings and apply formatting |
| deaa2f8 | 2026-06-28 | feat(deploy): add zypper support for openSUSE/SLES |
| 318a724 | 2026-06-28 | feat(deploy): add pacman support for Arch Linux |
| e5f1682 | 2026-06-29 | refactor(deploy)!: TUI |
| 28c3eee | 2026-06-30 | refactor(deploy): improved TUI walkthrough |
| 116595f | 2026-06-30 | feat(api): add QR scan-login completion poll on the internal face |
| a94b001 | 2026-06-30 | feat(deploy): add break-glass Operator account provisioning |
| 346ec68 | 2026-06-30 | refactor(deploy): improved cloudflare walkthrough |
| 563041a | 2026-06-30 | feat(panel): add fail-closed role-switcher view-mode logic |
| 50b8487 | 2026-06-30 | feat(panel): wire role-switcher into the app shell |
| 75642d9 | 2026-06-30 | feat(metrics): add named felis_* Prometheus collectors |
| 2a93a9e | 2026-06-30 | feat(metrics): record felis_image_build_failures_total on failed builds |
| 79eae7f | 2026-06-30 | feat(metrics): publish felis_servers_total from a fleet snapshot |
| 8ac5e64 | 2026-06-30 | feat(metrics): observe felis_start_duration_seconds across the start lifecycle |
| 676407d | 2026-06-30 | docs(diagrams): align §28 sequence diagrams with implemented routes |
| ac02c69 | 2026-06-30 | docs(troubleshooting): add operator failure-mode checklist |
| eb5875a | 2026-06-30 | feat(felis): add Operator break-glass op behind an operation menu |
| c14ed17 | 2026-06-30 | fix(docker): keep embedded panel/ and deploy/ in the image build context |
| 2a4a81b | 2026-06-30 | fix(api): don't burn wake cooldown when refused at capacity |
| 6c3999a | 2026-06-30 | fix(api): rate-limit email-OTP sends to close the email-bomb vector |
| 9873904 | 2026-06-30 | fix(operator): populate Status.Players from an RCON list probe |
| 879b177 | 2026-06-30 | fix(api): make OTP-start throttle atomic to close concurrent-burst bypass |
| 29f5341 | 2026-07-01 | docs(api): correct cooldownLimiter doc for its OTP reuse |
| 7507cfa | 2026-07-01 | Revert "feat(panel): wire role-switcher into the app shell" |
| f2c916d | 2026-07-01 | feat(api): add passkey enrollment persistence layer |
| 742f15f | 2026-07-01 | feat(api): add passkey enrollment endpoints |
| d2de11a | 2026-07-01 | feat(panel): fleet |
| 0261204 | 2026-07-01 | feat(passkey): add go-webauthn enrollment verifier adapter |
| fce0fce | 2026-07-01 | feat(passkey): wire enrollment verifier into felis-api |
| 2810fe8 | 2026-07-01 | fix(cfsetup): repoint stale DNS record when routing a tunnel hostname |
| 7d3be64 | 2026-07-01 | feat(cfsetup): start the tunnel connector as a setup step |
| a531f5e | 2026-07-01 | fix(cfsetup): keep connector install in the host apply layer only |
| e058a64 | 2026-07-01 | feat(edge): close the panel NodePort to the public after the tunnel is up |
| c01f133 | 2026-07-01 | feat(updates): add pure decision core for component self-update |
| fe2ece0 | 2026-07-01 | feat(api): add public Bind-Code onboarding for the player console |
| 3673af6 | 2026-07-01 | feat(api): add admin API for the SysAdmin-set auto-update maintenance window |
| 7464fa7 | 2026-07-01 | fix(updates): tag Window JSON so the persisted maintenance window round-trips |
| e035142 | 2026-07-01 | feat(passkey): add WebAuthn login/assertion crypto adapter |
| f34711c | 2026-07-01 | docs(api): record passkey login-handler deferral rationale |
| 7a51c1d | 2026-07-01 | fix(api): bound concurrent login bcrypt to shed CPU-pin floods |
| 164ac44 | 2026-07-01 | fix(api): validate inbound X-Request-Id before echo and audit persist |
| c6c0772 | 2026-07-01 | fix(api): set read/idle timeouts on the felis-api listeners |
| 3c1d647 | 2026-07-01 | fix(api): cap concurrent SSE streams per principal |
| d6e3189 | 2026-07-01 | fix(api): bound SSE relay writes with a deadline to sever stalled readers |
| 2c56d17 | 2026-07-01 | docs(api): record the quota-claim TOCTOU as a KNOWN-LIMITATION (audit #4) |
| 8f41a00 | 2026-07-01 | fix(api): clear the SSE write deadline on return so it can't leak to a reused connection |
| 15c58d9 | 2026-07-01 | feat(panel): player management |
| 8ae65ae | 2026-07-02 | feat(panel): backup management |
| a15ff55 | 2026-07-02 | refactor(panel): optimize player list layout and horizontal operations |
| 149f01a | 2026-07-02 | fix(panel): change console button to outline variant on my servers page |
| 4ecaf3c | 2026-07-02 | refactor(panel): set defaultOpen parameter of whitelist card to false |
| a9dbc8b | 2026-07-02 | feat(panel): add search and status filtering to my servers page |
| fa7bab5 | 2026-07-02 | style(panel): refine search and filter layout to align with header |
| 70d17a0 | 2026-07-02 | feat(panel): align my servers page search layout with fleet table |
| 5a8eff1 | 2026-07-02 | feat(panel): remove developer comment footer cards from my servers and server admin pages |
| 92770ea | 2026-07-02 | style(panel): adjust pagination padding to pt-3 for balanced spacing |
| 6e43a46 | 2026-07-02 | fix(panel): pin sidebar navigation and enable independent content scroll |
| 3b4298d | 2026-07-02 | refactor(panel): unify servers cockpit layout, resolve duplicate pages and adjust spacing |
| c0d333b | 2026-07-02 | feat(panel): support full server config edit dialog with status prefilling |
| 6368ab1 | 2026-07-02 | fix(api): coalesce MyServers owned flag so ownerless rows do not 500 |
| cdbb5ab | 2026-07-02 | fix(api): record credential id in passkey-register audit event |
| 9953275 | 2026-07-02 | fix(api): bound webauthn_challenges growth by superseding all prior rows |
| 20e31fb | 2026-07-02 | fix(store): cascade-delete passkeys and challenges on user removal |
| 7278cd7 | 2026-07-02 | feat(passkey): require and record user verification at enrollment |
| 54bc6ef | 2026-07-02 | fix(api): clear bound passkeys on password change to close a takeover foothold |
| 19f500b | 2026-07-02 | style(panel): update destructive red color and rename wake to start |
| e0bc288 | 2026-07-02 | feat(panel): implement image build pipeline and admin whitelist with mock dev api |
| 8594622 | 2026-07-02 | feat(panel): implement email OTP verification and passkey registration management |
| 9bed51b | 2026-07-02 | feat(config): add [velocity] login_image/lobby_image for system servers |
| 9ef817f | 2026-07-02 | feat(naming): system-server names, validation, and service-token identifiers |
| 159107b | 2026-07-02 | feat(api): HTTP readiness knob on MinecraftServer and login-gate fallback default |
| dc23cb5 | 2026-07-02 | feat(operator): system-server pod readiness probe and login service-token env |
| 3fdb3d0 | 2026-07-02 | feat(platform): internal API base-URL helper and single-sourced token secret |
| f554d52 | 2026-07-02 | feat(cli): provision login/lobby system servers with login env and token replica |
| a63f49d | 2026-07-02 | feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey |
| 241fe21 | 2026-07-02 | feat(limbo): felis-limbo in-game login flow over the shared account-link client |
| c7315e4 | 2026-07-02 | feat(deploy): login-limbo and lobby images with game-port pinning |
| 191640c | 2026-07-02 | feat(panel): implement admin submission approval and reject queue |
| 598f3d3 | 2026-07-02 | feat(submit): local + S3 backends for modpack upload contexts, installer-selectable |
| adf0d99 | 2026-07-02 | feat(panel): implement user-side modpack submissions with drag & drop context upload |
| d9e866f | 2026-07-03 | fix(deploy): make the lobby image actually build |
| b84debf | 2026-07-03 | feat(deploy): one-shot demo bring-up wrapper |
| 73d6ec1 | 2026-07-03 | feat(mock): add mock submissions for owner account |
| 5427bc7 | 2026-07-03 | feat(servers): support claiming servers directly from ServersPage list |
| 55592ed | 2026-07-03 | feat(auth): support public auth bind endpoint |
| 804459c | 2026-07-03 | feat(panel): implement admin maintenance window settings page |
| dd7dff6 | 2026-07-03 | feat(panel): support importing parameters from submission with owner-restricted unapproved entries |
| a8701c1 | 2026-07-03 | fix(panel): prevent automatic wake during server claim in mock api |
| f749c2c | 2026-07-03 | style(panel): resolve double borders and uneven padding in server console |
| 5f402b1 | 2026-07-03 | fix(panel): force dark mode and pure black bg on server console card |
| b7d8000 | 2026-07-03 | feat(panel): implement dedicated LuckPerms permissions and groups management sub-page |
| e60784b | 2026-07-04 | fix(panel): eliminate page collapse and scroll shifts during LuckPerms query reload |
| 439f19e | 2026-07-04 | fix(panel): prevent page collapse and scroll shifts in players and bans management sections during reload |
| 83e57b4 | 2026-07-04 | fix(panel): implement two-step confirmation for claiming a server to prevent accidental operations |
| 3347cc0 | 2026-07-04 | feat(panel): implement user management administration panel with sessions and minecraft link support |
| 67b4e19 | 2026-07-04 | fix(panel): override generic already_exists error message during user creation and profile editing |
| 2c95da8 | 2026-07-04 | fix(panel/i18n): add missing users_col_user key to translation files |
| 627883e | 2026-07-04 | fix(panel): refine reset password messages and fix empty email placeholder in mock api response |
| 0c1cc59 | 2026-07-04 | feat(auth): migrate console login to passwordless |
| 3b43f05 | 2026-07-04 | refactor(api): drop dead login concurrency limiter and reconcile passwordless comments |
| 4f59d51 | 2026-07-04 | feat(auth): add owner-tier passkey-unbind remediation endpoint |
| c20b12c | 2026-07-04 | refactor(api): drop dead password-era ResetMailer, reconcile passkey-unbind docs |
| 0a2accd | 2026-07-04 | chore: stop tracking Autohand-generated AGENTS.md |
| 96b3cc9 | 2026-07-04 | feat(updater): wire updates.Run to a caller with PaperMC v3 release discovery |
| 9896fe1 | 2026-07-05 | docs(updater): correct PaperMC UA/fixture overclaims, re-tier the boundary |
| 7d27640 | 2026-07-05 | feat(updater): add GitHub Releases source and route felis-api/k3s/cloudflared |
| bd49313 | 2026-07-05 | refactor(panel): 抽取 10 个公共组件,消除 ~150 处重复代码 |
| 91bfa27 | 2026-07-05 | feat(operator): implement idle auto-stop (spec §8) |
| e574749 | 2026-07-05 | feat(api): enforce CPU/memory/storage quotas (spec §9.3, §22) |
| 7db57b9 | 2026-07-05 | feat(updater): add VersionGatherer extraction core and CLI gather seam |
| ec468ba | 2026-07-05 | feat(auth): add discoverable (usernameless) passkey login |
| 154002e | 2026-07-05 | docs(auth): cite MultiLogin reference for UUID-keyed reclaim split |
| 0dbd557 | 2026-07-05 | fix(store): renumber discoverable-login migration 0013 -> 0014 |
| 9e1df12 | 2026-07-05 | feat(passkey): advance sign_count, reject clone-warned assertions |
| 7f7e459 | 2026-07-05 | fix(operator): enforce startup and readiness timeouts (§5, §8) |
| 7becb38 | 2026-07-05 | fix(api): implement /readyz with real DB + K8s API + CRD checks (§7) |
| 9079a2c | 2026-07-05 | feat(panel): implement email otp and passkey login interface |
| cfe68ae | 2026-07-05 | fix(panel): align status distribution order to put Stopped at the end |
| bbcfaeb | 2026-07-05 | refactor(panel): remove redundant voxel network topology description subtitle |
| fdb6efb | 2026-07-05 | feat(account): migrate a live account's owned servers to a new account (§B3 inherit) |
| abad137 | 2026-07-06 | style(panel): unify vertical spacing below PageHeader across pages |
| c2ee21a | 2026-07-06 | feat(breakglass): add halt-a-server op to the recovery console (§B4) |
| c1aa38b | 2026-07-06 | feat(velocity): add /felis migrate to open an account migration (§B3 inherit) |
| 7a7c0d5 | 2026-07-07 | feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync) |
| f2fc57c | 2026-07-07 | feat(api): add internal-face break-glass world backup endpoint (§B4 Sync) |
| 2ba9948 | 2026-07-07 | fix(platform): front the felis-api internal face on its own ClusterIP Service |
| fc748d3 | 2026-07-07 | feat(breakglass): add "back up a world now" console peer (§B4 Sync) |
| 9911b8c | 2026-07-07 | docs(changes): record the break-glass backup console peer (§B4 Sync phase 2b) |
| 096d597 | 2026-07-07 | docs(changes): backfill detail docs for pre-ledger functional commits |
| 5a7cd5a | 2026-07-07 | docs(changes): fold 346ec68 cloudflare-edge walkthrough into its detail doc |
| 4626ab5 | 2026-07-07 | docs(changes): mutation-audit the ledger's "unit-tested" safety claims |
| 729bd7b | 2026-07-07 | docs(changes): index the mutation audit and two lagging ledger rows |
| c67a4d3 | 2026-07-08 | docs(changes): close §B4 with the S3 archive backend deferred by design |
| 85b8a92 | 2026-07-08 | test(api): pin restore's owner gate against a superseded former owner |
| b7b4a3b | 2026-07-08 | docs(changes): record the round-2 backup/restore mutation audit and index the owner-gate test |
+36 -12
View File
@@ -460,20 +460,21 @@ paths:
operationId: hasJoined
summary: Multi-source session verifier (Felis-nano hasJoined multiplexer).
description: >-
Velocity's authlib is pointed here via -Dmojang.sessionserver or a thin login
hook. Unauthenticated — the vanilla sessionserver protocol carries no token. The
query is fanned out to the configured Yggdrasil roots in priority order (the
Mojang identity source first); the first source to validate the serverId hash
wins. A non-identity source's self-asserted UUID is rewritten into a per-source
namespace (UUIDv3) before return, so it can never land in Mojang's UUID space.
A rejected or barred login is 204, which authlib maps to a verify failure.
Velocity is pointed here with -Dmojang.sessionserver and sends the request itself.
Unauthenticated — the vanilla sessionserver protocol carries no token. The query
is fanned out to the configured Yggdrasil roots in priority order (the Mojang
identity source first); the first source to validate the serverId hash wins. A
non-identity source's self-asserted UUID is rewritten into a per-source namespace
(UUIDv3) before return, so it can never land in Mojang's UUID space. A rejected
or barred login is 204, which Velocity answers with its online-mode-only kick.
Any other non-200 status makes Velocity report the auth servers as down.
x-felis-face: [internal]
x-felis-tier: public
security: []
parameters:
- { name: username, in: query, required: true, schema: { type: string } }
- { name: serverId, in: query, required: true, schema: { type: string } }
- { name: ip, in: query, required: false, schema: { type: string } }
- { name: username, in: query, required: true, schema: { type: string, maxLength: 64 } }
- { name: serverId, in: query, required: true, schema: { type: string, maxLength: 64 } }
- { name: ip, in: query, required: false, schema: { type: string, maxLength: 64 } }
responses:
'200':
description: A source validated the session; the canonical game profile.
@@ -484,10 +485,33 @@ paths:
required: [id, name]
properties:
id: { type: string, description: Canonical UUID, undashed 32-hex. }
name: { type: string }
name:
type: string
description: >-
The name the source returned. A third-party player whose name is
registered to a Mojang account gets it back as PREFIX_name, cut to
16 characters.
properties: { type: array, items: { type: object } }
'204':
description: No source validated the session, or the resolved UUID is barred.
description: >-
Not admitted, with no source asked when username or serverId is missing or a
parameter is over 64 bytes. Otherwise no source validated the session, the
canonical UUID is barred, a third-party source returned a name that is not a
legal Minecraft username, or the identity source returned an unparseable id.
'400':
description: >-
The request declared a body. No body is sent back, and the connection is
closed.
'500':
description: The bar-list lookup failed, so the login is not admitted.
content:
application/json:
schema: { $ref: '#/components/schemas/Error' }
'503':
description: >-
No source validated the session and at least one source failed (transport
error, redirect, unexpected status, or a 200 without a usable profile). Its
player may be the one logging in, so this is not answered as a 204. No body.
# -------------------------------------------------- internal: servers ------
/api/v1/servers:
+4 -4
View File
@@ -12,7 +12,7 @@ require (
github.com/go-webauthn/webauthn v0.17.4
github.com/golang-jwt/jwt/v5 v5.3.1
github.com/google/uuid v1.6.0
github.com/jackc/pgx/v5 v5.7.1
github.com/jackc/pgx/v5 v5.9.2
github.com/minio/minio-go/v7 v7.2.1
github.com/prometheus/client_golang v1.19.1
github.com/prometheus/client_model v0.6.1
@@ -92,12 +92,12 @@ require (
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/crypto v0.52.0 // indirect
golang.org/x/exp v0.0.0-20231006140011-7918f672742d // indirect
golang.org/x/net v0.54.0 // indirect
golang.org/x/net v0.55.0 // indirect
golang.org/x/oauth2 v0.21.0 // indirect
golang.org/x/sync v0.20.0 // indirect
golang.org/x/sync v0.21.0 // indirect
golang.org/x/sys v0.45.0 // indirect
golang.org/x/term v0.43.0 // indirect
golang.org/x/text v0.37.0 // indirect
golang.org/x/text v0.39.0 // indirect
golang.org/x/time v0.3.0 // indirect
gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect
google.golang.org/protobuf v1.36.10 // indirect
+10 -10
View File
@@ -116,8 +116,8 @@ github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsI
github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM=
github.com/jackc/pgx/v5 v5.7.1 h1:x7SYsPBYDkHDksogeSmZZ5xzThcTgRz++I5E+ePFUcs=
github.com/jackc/pgx/v5 v5.7.1/go.mod h1:e7O26IywZZ+naJtWWos6i6fvWK+29etgITqrqHLfoZA=
github.com/jackc/pgx/v5 v5.9.2 h1:3ZhOzMWnR4yJ+RW1XImIPsD1aNSz4T4fyP7zlQb56hw=
github.com/jackc/pgx/v5 v5.9.2/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/josharian/intern v1.0.0 h1:vlS4z54oSdjm0bgjRigI+G1HpF+tI+9rE5LLzOg8HmY=
@@ -245,15 +245,15 @@ golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn
golang.org/x/net v0.0.0-20190620200207-3b0461eec859/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s=
golang.org/x/net v0.0.0-20200226121028-0de0cce0169b/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s=
golang.org/x/net v0.0.0-20201021035429-f5854403a974/go.mod h1:sp8m0HH+o8qH0wwXwYZr8TS3Oi6o0r6Gce1SSxlDquU=
golang.org/x/net v0.54.0 h1:2zJIZAxAHV/OHCDTCOHAYehQzLfSXuf/5SoL/Dv6w/w=
golang.org/x/net v0.54.0/go.mod h1:Sj4oj8jK6XmHpBZU/zWHw3BV3abl4Kvi+Ut7cQcY+cQ=
golang.org/x/net v0.55.0 h1:bcvxaJn3e1U6InsFWt1JUq1aSjnRxLzT2rtD2KfkDF8=
golang.org/x/net v0.55.0/go.mod h1:L5U2KuzuOe1lY7Z+aWVIKK6qEeJXnXV9yzGA+WCHJww=
golang.org/x/oauth2 v0.21.0 h1:tsimM75w1tF/uws5rbeHzIWxEqElMehnc+iW793zsZs=
golang.org/x/oauth2 v0.21.0/go.mod h1:XYTD2NtWslqkgxebSiOHnXEap4TF09sJSc7H1sXbhtI=
golang.org/x/sync v0.0.0-20190423024810-112230192c58/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20190911185100-cd5d95a43a6e/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20201020160332-67f06af15bc9/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.20.0 h1:e0PTpb7pjO8GAtTs2dQ6jYa5BWYlMuX047Dco/pItO4=
golang.org/x/sync v0.20.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM=
golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20200930185726-fdedc70b468f/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
@@ -265,16 +265,16 @@ golang.org/x/term v0.43.0 h1:S4RLU2sB31O/NCl+zFN9Aru9A/Cq2aqKpTZJ6B+DwT4=
golang.org/x/term v0.43.0/go.mod h1:lrhlHNdQJHO+1qVYiHfFKVuVioJIheAc3fBSMFYEIsk=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.3/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.37.0 h1:Cqjiwd9eSg8e0QAkyCaQTNHFIIzWtidPahFWR83rTrc=
golang.org/x/text v0.37.0/go.mod h1:a5sjxXGs9hsn/AJVwuElvCAo9v8QYLzvavO5z2PiM38=
golang.org/x/text v0.39.0 h1:UbZz4pLOvn600D6Oh6GGEI6VAmndrEBLv8/6BEXzyus=
golang.org/x/text v0.39.0/go.mod h1:3UwRclnC2g0TU9x8PZiyfOajCd1zaUNHF9cvqcQZ+ZM=
golang.org/x/time v0.3.0 h1:rg5rLMjNzMS1RkNLzCG38eapWhnYLFYXDXj2gOlr8j4=
golang.org/x/time v0.3.0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
golang.org/x/tools v0.0.0-20191119224855-298f0cb1881e/go.mod h1:b+2E5dAYhXwXZwtnZ6UAqBI28+e2cm9otk0dWdXHAEo=
golang.org/x/tools v0.0.0-20200619180055-7c47624df98f/go.mod h1:EkVYQZoAsY45+roYkvgYkIh4xh/qjgUK9TdY2XT94GE=
golang.org/x/tools v0.0.0-20210106214847-113979e3529a/go.mod h1:emZCQorbCU4vsT4fOWvOPXz4eW1wZW4PmDk9uLelYpA=
golang.org/x/tools v0.44.0 h1:UP4ajHPIcuMjT1GqzDWRlalUEoY+uzoZKnhOjbIPD2c=
golang.org/x/tools v0.44.0/go.mod h1:KA0AfVErSdxRZIsOVipbv3rQhVXTnlU6UhKxHd1seDI=
golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191011141410-1b5146add898/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
+12 -13
View File
@@ -139,9 +139,9 @@ type API struct {
// AuthSources is the Felis-nano multi-source hasJoined multiplexer's upstream
// Yggdrasil list, in priority order (config order; the Mojang Identity source
// first for 正版优先). Nil — the default — makes the session verifier reject every
// login (204), so the endpoint ships inert until cmd/felis wires configured
// sources. Consumed by handleHasJoined (handlers_hasjoined.go).
// first for 正版优先). Nil makes the session verifier reject every login (204);
// cmd/felis always wires at least the Mojang source through authSourcesFromConfig.
// Consumed by handleHasJoined (handlers_hasjoined.go).
AuthSources []AuthSource
// Now is the clock, injectable for tests. Defaults to time.Now.
@@ -303,12 +303,12 @@ func (a *API) internalAPIRoutes() []apiRoute {
// Mojang player (same name, different UUID) always passes.
{Method: "POST", Pattern: "/api/v1/internal/player/reclaim", h: a.handleReclaimUsername},
{Method: "GET", Pattern: "/api/v1/internal/player/blacklist/{mc_uuid}", h: a.handleCheckBlacklist},
// Felis-nano multi-source session verifier (spec §B3 player game-login).
// Velocity's authlib is pointed here (-Dmojang.sessionserver or a thin login
// hook); it speaks the vanilla sessionserver protocol and carries no token, so
// this is Public. It fans hasJoined out to the configured Yggdrasil roots
// (Mojang-first) and rewrites third-party UUIDs into a per-source namespace
// before returning the canonical profile (handlers_hasjoined.go).
// Felis-nano multi-source session verifier, behind player game-login. Velocity is
// pointed here with -Dmojang.sessionserver and issues the request itself; it speaks
// the vanilla sessionserver protocol and carries no token, so this is Public. It
// fans hasJoined out to the configured Yggdrasil roots (Mojang-first) and rewrites
// third-party UUIDs into a per-source namespace before returning the canonical
// profile (handlers_hasjoined.go).
{Method: "GET", Pattern: "/session/minecraft/hasJoined", Public: true, h: a.handleHasJoined},
// Op-login (passwordless op.console login): a staff member starts the login
// on the web, and an ONLINE in-game admin vouches for it via velocity's
@@ -674,10 +674,9 @@ func principalFromContext(ctx context.Context) *Principal {
// #35); cross-replica bounding would need a shared store (out of scope for the
// single-replica demo).
type cooldownLimiter struct {
mu sync.Mutex
now func() time.Time
last map[string]time.Time
window time.Duration
mu sync.Mutex
now func() time.Time
last map[string]time.Time
}
// allowed reports whether name may wake now WITHOUT recording the attempt. A
+2 -2
View File
@@ -242,7 +242,7 @@ func (f *fakeRepo) QuotaCheck(_ context.Context, userID string, _ string, _ Reso
// For hermetic tests, QuotaCheck delegates to the same QuotaAvailable
// store — tests that care about per-dimension checks should use
// fakeQuotas with direct inspection.
return f.QuotaAvailable(nil, userID)
return f.QuotaAvailable(context.TODO(), userID)
}
func (f *fakeRepo) UpdateServerResources(_ context.Context, _ string, _, _, _ int) error { return nil }
@@ -1335,7 +1335,7 @@ func (c *fakeCluster) GetBySubdomain(_ context.Context, s string) (*ServerInfo,
return nil, ErrNotFound
}
func (c *fakeCluster) ListServers(_ context.Context) ([]ServerInfo, error) { return c.list, nil }
func (c *fakeCluster) Ping(_ context.Context) error { return c.pingErr }
func (c *fakeCluster) Ping(_ context.Context) error { return c.pingErr }
func (c *fakeCluster) SetDesiredState(_ context.Context, n string, s v1alpha1.DesiredState) error {
c.desired[n] = s
return nil
+1 -1
View File
@@ -416,7 +416,7 @@ type lpPermissionView struct {
// maxLPInfoPages bounds how many "permission info" pages the read projector
// chases per request. LuckPerms paginates its reply, so one command shows only
// the first page; we follow the header's page count up to this cap.
// ponytail: 10 pages ≈ 150 entries — raise if a real user outgrows it.
// 10 pages ≈ 150 entries — raise if a real user outgrows it.
const maxLPInfoPages = 10
// handleAccessLuckPermsInfo is the read projector for a player's LuckPerms
+110 -44
View File
@@ -6,6 +6,7 @@ import (
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"net/url"
"regexp"
@@ -16,14 +17,14 @@ import (
"github.com/google/uuid"
)
// Felis-nano: the multi-source hasJoined multiplexer (spec §B3 player game-login).
// Felis-nano: the multi-source hasJoined multiplexer behind player game-login.
//
// Velocity's session verifier (authlib) is pointed here — via -Dmojang.sessionserver
// on Felis-managed proxies, or a thin login-pipeline hook on third-party servers. On
// login Velocity computes the serverId hash and GETs hasJoined; this endpoint fans that
// query out to the configured Yggdrasil roots in priority order (Mojang first, 正版优先)
// and returns the first source that validates. Each upstream Yggdrasil runs its own
// serverId-hash check — the multiplexer only relays, it computes no hashes.
// Velocity is pointed here with -Dmojang.sessionserver and issues the request itself, not
// through authlib. On login it computes the serverId hash and GETs hasJoined; this
// endpoint fans that query out to the configured Yggdrasil roots in priority order
// (Mojang first, 正版优先) and returns the first source that validates. Each upstream
// Yggdrasil runs its own serverId-hash check — the multiplexer only relays, it computes
// no hashes.
//
// The one non-negotiable transform: a non-identity (third-party) source's UUID is
// self-asserted, so its profile is rewritten into a per-source namespace
@@ -33,6 +34,10 @@ import (
// preimage resistance means no third-party source can mint a Mojang-space UUID, and the
// per-tag namespace means two sources cannot collide onto one identity. Every downstream
// key (account_links, username_blacklist, owner checks) then sees one canonical UUID.
//
// One consequence a backend operator meets: a chat-session key a third-party source signed
// over its native UUID cannot verify against the canonical one, even on a backend that
// trusts that source's key. Such players' chat can only be accepted unsigned.
// felisAuthNS is the fixed UUIDv3 namespace every third-party profile is rewritten
// under (see the rewrite rationale above). Derived from the project name, not a magic
@@ -41,9 +46,27 @@ var felisAuthNS = uuid.NewSHA1(uuid.NameSpaceURL, []byte("nano.felis.lolicon.bes
// authHTTPClient calls the upstream Yggdrasil roots. The timeout bounds one login
// against a hung source; the resolver moves on to the next source on any failure.
// ponytail: one shared client, sequential priority scan — a third-party login costs one
// One shared client, sequential priority scan — a third-party login costs one
// wasted Mojang round-trip; add parallel fan-out only if login latency bites.
var authHTTPClient = &http.Client{Timeout: 5 * time.Second}
var authHTTPClient = &http.Client{
Timeout: 5 * time.Second,
Transport: upstreamTransport,
// A redirect is not a hasJoined answer. Following one would let a configured root point
// this host at any URL it can reach — this listener included, where each hop re-runs the
// whole source scan inside the same login's timeout. The 3xx is returned as-is and the
// resolver skips that source like any other non-200.
CheckRedirect: func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse },
}
// upstreamTransport caps response headers, which the 64 KiB body limit does not cover. The
// default allows 1 MiB, so a root that sends that much and then stalls the body pins a few
// MiB per in-flight login for the whole timeout, and enough parallel logins OOM the host
// for every source. Real roots answer in well under 1 KiB of headers.
var upstreamTransport = func() *http.Transport {
t := http.DefaultTransport.(*http.Transport).Clone()
t.MaxResponseHeaderBytes = 16 << 10
return t
}()
// AuthSource is one upstream Yggdrasil root in the multiplexer's priority list (config
// order = priority). URL is the full hasJoined endpoint the query string is appended to.
@@ -59,11 +82,14 @@ type AuthSource struct {
}
// sessionProfile is the Mojang hasJoined contract. properties is relayed verbatim
// (json.RawMessage) so a source's signed textures survive the multiplexer untouched.
// (json.RawMessage) so a source's signed textures survive the multiplexer untouched, and
// it is always emitted as an array: Velocity's GameProfile parser throws on a missing or
// null properties key, while a Yggdrasil root may legitimately send [] or omit it for a
// player with no skin.
type sessionProfile struct {
ID string `json:"id"`
Name string `json:"name"`
Properties []json.RawMessage `json:"properties,omitempty"`
Properties []json.RawMessage `json:"properties"`
}
// HasJoinedHandler returns an http.Handler serving only the Felis-nano hasJoined
@@ -80,19 +106,36 @@ func HasJoinedHandler(sources []AuthSource, repo Repo) http.Handler {
}
// handleHasJoined is the multi-source session verifier (Felis-nano). It is a Public
// internal-face route: authlib speaks the vanilla sessionserver protocol and sends no
// internal-face route: Velocity speaks the vanilla sessionserver protocol and sends no
// service token. A rejected login is 204 No Content — exactly what Mojang returns for an
// invalid session, which authlib maps to "failed to verify username".
// invalid session, which Velocity answers with its online-mode-only kick.
func (a *API) handleHasJoined(w http.ResponseWriter, r *http.Request) {
// Velocity sends no body. When a request declares one anyway, net/http tries to drain it
// before writing any answer, so one that never arrives holds the connection with no
// timeout: ReadHeaderTimeout stopped at the headers. Marking the reply as the last one on
// this connection skips the drain.
if r.ContentLength != 0 {
w.Header().Set("Connection", "close")
w.WriteHeader(http.StatusBadRequest)
return
}
q := r.URL.Query()
username, serverID := q.Get("username"), q.Get("serverId")
if username == "" || serverID == "" {
username, serverID, ip := q.Get("username"), q.Get("serverId"), q.Get("ip")
if username == "" || serverID == "" ||
len(username) > maxHasJoinedParam || len(serverID) > maxHasJoinedParam || len(ip) > maxHasJoinedParam {
w.WriteHeader(http.StatusNoContent)
return
}
prof, src := a.resolveHasJoined(r.Context(), username, serverID, q.Get("ip"))
prof, src, failed := a.resolveHasJoined(r.Context(), username, serverID, ip)
if prof == nil {
// With a source down, "nobody knows this player" is not established: its player may
// be the one logging in. 503 makes Velocity report the auth servers as down and log
// the status, where a 204 would tell that player their account is offline-mode.
if failed {
w.WriteHeader(http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusNoContent)
return
}
@@ -142,11 +185,19 @@ func (a *API) handleHasJoined(w http.ResponseWriter, r *http.Request) {
return
}
// Emit the canonical UUID undashed — the 32-hex form authlib's GameProfile expects.
// Emit the canonical UUID undashed — the 32-hex form Velocity's GameProfile expects.
prof.ID = hex.EncodeToString(canonical[:])
if prof.Properties == nil {
prof.Properties = []json.RawMessage{} // a nil slice would marshal as null
}
writeJSON(w, http.StatusOK, prof)
}
// maxHasJoinedParam bounds each query value before it is forwarded to every source. What
// Velocity sends fits with room to spare: a login name of at most 16 characters, a signed
// SHA-1 hex serverId of at most 41, a textual IP address. Only a direct caller sends more.
const maxHasJoinedParam = 64
// mcUsernameRe is Minecraft's username charset — the trust boundary on a third-party
// source's self-asserted profile name.
var mcUsernameRe = regexp.MustCompile(`^[A-Za-z0-9_]{3,16}$`)
@@ -158,7 +209,7 @@ const mcUsernameMax = 16
// TRUNCATED to fit rather than the rename being skipped when it would not fit — skipping is
// what would silently hand a 14-character premium name back to the squatter.
//
// ponytail: two players of one source whose names agree on their first mcUsernameMax-len(prefix)-1
// Two players of one source whose names agree on their first mcUsernameMax-len(prefix)-1
// characters truncate onto the same in-game name, as does a prefixed name that happens to be
// a premium name itself. Both cost an "already connected" bounce, not an identity: the UUID
// rewrite is what keeps players apart, and it does not depend on the name at all. Add a
@@ -179,15 +230,15 @@ var mojangProfileAPI = "https://api.mojang.com/users/profiles/minecraft/"
// profileHTTPClient is deliberately more impatient than authHTTPClient: the name lookup is a
// SECOND Mojang round-trip on a third-party login (the identity leg already spent one), and
// api.mojang.com is exactly what is unreliable from the networks these servers sit on. A
// slow answer falls back to the cache instead of holding the login open.
var profileHTTPClient = &http.Client{Timeout: 2 * time.Second}
// slow answer counts as taken instead of holding the login open.
var profileHTTPClient = &http.Client{Timeout: 2 * time.Second, Transport: upstreamTransport}
// A name's premium status changes on human timescales, not per login, so it is cached — but
// asymmetrically, because the two directions have very different costs. "Taken" is nearly
// permanent (Mojang does not recycle names), while "free" can stop being true the moment
// someone buys that name, and a stale "free" is the dangerous one: it leaves a squatter
// holding a name its real owner has just bought. So a "free" answer is trusted for minutes
// and a "taken" answer for a day.
// asymmetrically, because the two directions have very different costs. "Taken" changes
// only when its owner renames away, and a stale "taken" costs a third-party player nothing
// but a prefix. "Free" can stop being true the moment someone buys that name, and a stale
// "free" is the dangerous one: it leaves a squatter holding a name its real owner has just
// bought. So a "free" answer is trusted for minutes and a "taken" answer for a day.
const (
premiumTakenTTL = 24 * time.Hour
premiumFreeTTL = 10 * time.Minute
@@ -205,11 +256,11 @@ var premiumNames = struct {
}{m: make(map[string]premiumEntry)}
// isPremiumName reports whether username belongs to a real Mojang account — which is what
// makes a third-party player holding it a squatter. On a lookup failure it prefers a stale
// cached answer, and with nothing cached it fails CLOSED (assume premium → rename the
// third-party player): a Mojang outage must not let a squatter keep a name the real owner is
// about to log in with. Being wrong that way costs a cosmetic prefix; being wrong the other
// way bounces the name's actual owner off the proxy.
// makes a third-party player holding it a squatter. A lookup failure fails CLOSED (assume
// premium → rename the third-party player), even over an expired "free": the name may have
// been bought since, and a Mojang outage must not let a squatter keep it. Being wrong that
// way costs a cosmetic prefix; being wrong the other way bounces the name's actual owner off
// the proxy.
func isPremiumName(ctx context.Context, username string) bool {
key := strings.ToLower(username)
@@ -222,16 +273,14 @@ func isPremiumName(ctx context.Context, username string) bool {
taken, err := lookupPremiumName(ctx, username)
if err != nil {
if hit {
return cached.taken
}
return true
}
premiumNames.Lock()
// ponytail: bounded by dropping the whole map rather than evicting LRU — entries are
// only minted by players who actually authenticated somewhere, so this is a backstop
// against an unbounded map, not a cache policy worth tuning.
// Bounded by dropping the whole map rather than evicting LRU. Any third-party source that
// validates a login mints an entry, so a hostile one can force clears; that costs repeat
// lookups, or a fail-closed prefix while Mojang is unreachable, never an identity. This
// is a backstop against an unbounded map, not a cache policy worth tuning.
if len(premiumNames.m) >= premiumCacheMax {
clear(premiumNames.m)
}
@@ -270,9 +319,11 @@ func lookupPremiumName(ctx context.Context, username string) (bool, error) {
}
// resolveHasJoined queries each configured source in priority order and returns the
// first that validates the session (200 with a profile). A source that is down, answers
// non-200 (204 = "not my player"), or returns garbage is skipped.
func (a *API) resolveHasJoined(ctx context.Context, username, serverID, ip string) (*sessionProfile, AuthSource) {
// first that validates the session (200 with a profile). 204 is "not my player". Any other
// outcome (unreachable, another status, a body that is not a profile) skips the source too,
// but is logged with its tag and reported as failed: otherwise a dead or mistyped source
// looks exactly like a player it does not know, and nobody finds out.
func (a *API) resolveHasJoined(ctx context.Context, username, serverID, ip string) (prof *sessionProfile, src AuthSource, failed bool) {
for _, src := range a.AuthSources {
u := src.URL + "?username=" + url.QueryEscape(username) + "&serverId=" + url.QueryEscape(serverID)
if ip != "" {
@@ -280,23 +331,38 @@ func (a *API) resolveHasJoined(ctx context.Context, username, serverID, ip strin
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
log.Printf("hasJoined: source %q: %v", src.Tag, err)
failed = true
continue
}
resp, err := authHTTPClient.Do(req)
if err != nil {
log.Printf("hasJoined: source %q: %v", src.Tag, err)
failed = true
continue
}
if resp.StatusCode != http.StatusOK {
resp.Body.Close()
switch {
case resp.StatusCode == http.StatusNoContent:
case resp.StatusCode >= 300 && resp.StatusCode < 400:
log.Printf("hasJoined: source %q answered %s with Location %q; redirects are not followed, so set its url to the final endpoint", src.Tag, resp.Status, resp.Header.Get("Location"))
failed = true
default:
log.Printf("hasJoined: source %q answered %s", src.Tag, resp.Status)
failed = true
}
continue
}
var prof sessionProfile
err = json.NewDecoder(io.LimitReader(resp.Body, 1<<16)).Decode(&prof)
var p sessionProfile
err = json.NewDecoder(io.LimitReader(resp.Body, 1<<16)).Decode(&p)
resp.Body.Close()
if err != nil || prof.ID == "" {
if err != nil || p.ID == "" {
log.Printf("hasJoined: source %q answered 200 without a usable profile (err=%v)", src.Tag, err)
failed = true
continue
}
return &prof, src
return &p, src, failed
}
return nil, AuthSource{}
return nil, AuthSource{}, failed
}
+299 -12
View File
@@ -1,13 +1,21 @@
package api
import (
"bufio"
"context"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"net"
"net/http"
"net/http/httptest"
"path"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/google/uuid"
)
@@ -68,6 +76,13 @@ func fakeYgg(t *testing.T, id, name string) *httptest.Server {
return srv
}
// blacklistDownRepo is a store whose bar list cannot be read.
type blacklistDownRepo struct{ *fakeRepo }
func (blacklistDownRepo) IsUsernameBlacklisted(context.Context, string) (bool, error) {
return false, errors.New("bar list unreachable")
}
func getHasJoined(h http.Handler, username, serverID string) *httptest.ResponseRecorder {
return do(h, "GET", "/session/minecraft/hasJoined?username="+username+"&serverId="+serverID, "", nil)
}
@@ -81,13 +96,12 @@ func profileOf(t *testing.T, w *httptest.ResponseRecorder) sessionProfile {
return p
}
// undashed is the 32-hex form the resolver must emit (authlib's GameProfile format).
// undashed is the 32-hex form the resolver must emit (Velocity's GameProfile format).
func undashed(u uuid.UUID) string { return hex.EncodeToString(u[:]) }
const notchMojangID = "069a79f444e94726a5befca90e38aaf5" // a real Mojang-space UUID, undashed
func TestHasJoined(t *testing.T) {
// A trusted (Mojang) source passes its UUID through byte-for-byte.
t.Run("mojang identity passthrough", func(t *testing.T) {
mojang := fakeYgg(t, notchMojangID, "Notch")
api := newTestAPI(newFakeRepo(), newFakeCluster())
@@ -129,9 +143,14 @@ func TestHasJoined(t *testing.T) {
if got != want {
t.Fatalf("rewrite = %q, want deterministic UUIDv3 %q", got, want)
}
// Also pinned as literals: every third-party player's UUID, and every ban and account
// link keyed on one, depends on these exact bytes, and a change to the namespace seed
// moves both sides of the derived comparison above at once.
if ns := felisAuthNS.String(); ns != "07228eae-77f6-500e-9dc0-436afbc87c27" || got != "b63bcc1c611432eeb7b3af3a15012e48" {
t.Fatalf("namespace %s / rewrite %s changed; every existing third-party player would get a new UUID", ns, got)
}
})
// Mojang is priority-first: when both would validate the same name, Mojang wins.
t.Run("mojang priority wins over thirdparty", func(t *testing.T) {
stubMojangNames(t, "Notch")
mojang := fakeYgg(t, notchMojangID, "Notch")
@@ -153,8 +172,6 @@ func TestHasJoined(t *testing.T) {
}
})
// Mojang doesn't know the player (204) → fall through to the third-party source,
// whose profile is returned rewritten.
t.Run("fallthrough to thirdparty when mojang 204s", func(t *testing.T) {
stubMojangNames(t)
mojang := fakeYgg(t, "", "") // 204: not my player
@@ -174,7 +191,6 @@ func TestHasJoined(t *testing.T) {
}
})
// No source validates → 204 (authlib maps this to a verify failure).
t.Run("no source validates -> 204", func(t *testing.T) {
a := fakeYgg(t, "", "")
b := fakeYgg(t, "", "")
@@ -188,6 +204,111 @@ func TestHasJoined(t *testing.T) {
}
})
// A source that could not answer has not said no. With nobody validating, the login is
// an outage whether that source errored, redirected or was unreachable.
t.Run("failing source and no validator -> 503", func(t *testing.T) {
down := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusServiceUnavailable)
}))
t.Cleanup(down.Close)
redirector := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, "https://elsewhere.example/hasJoined", http.StatusMovedPermanently)
}))
t.Cleanup(redirector.Close)
nobody := fakeYgg(t, "", "")
for _, failing := range []string{down.URL, redirector.URL, "http://127.0.0.1:1"} {
api := newTestAPI(newFakeRepo(), newFakeCluster())
api.AuthSources = []AuthSource{
{Tag: "mojang", URL: nobody.URL, Identity: true},
{Tag: "littleskin", Prefix: "LS", URL: failing},
}
if w := getHasJoined(api.InternalHandler(), "Ghost", "abc"); w.Code != http.StatusServiceUnavailable {
t.Fatalf("source %s: code = %d, want 503", failing, w.Code)
}
}
})
// A skinless player's profile may come back with properties [], null or absent. The
// relay must still send an array: Velocity's GameProfile parser throws on a missing or
// null key and the login hangs, where the same answer sent straight to Velocity works.
t.Run("properties always emitted as an array", func(t *testing.T) {
for _, upstream := range []string{
`{"id":"` + notchMojangID + `","name":"Notch","properties":[]}`,
`{"id":"` + notchMojangID + `","name":"Notch","properties":null}`,
`{"id":"` + notchMojangID + `","name":"Notch"}`,
} {
src := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = w.Write([]byte(upstream))
}))
api := newTestAPI(newFakeRepo(), newFakeCluster())
api.AuthSources = []AuthSource{{Tag: "mojang", URL: src.URL, Identity: true}}
w := getHasJoined(api.InternalHandler(), "Notch", "abc")
src.Close()
var body map[string]json.RawMessage
if err := json.Unmarshal(w.Body.Bytes(), &body); err != nil || w.Code != http.StatusOK {
t.Fatalf("upstream %s: code = %d body = %q", upstream, w.Code, w.Body.String())
}
if got := string(body["properties"]); got != "[]" {
t.Errorf("upstream %s: properties = %q, want []", upstream, got)
}
}
})
// A root that answers with a redirect is skipped, not followed: following it lets that
// root aim this host at arbitrary URLs, including its own hasJoined route, which re-enters
// the scan and multiplies the upstream traffic one login causes.
t.Run("redirecting source is skipped, not followed", func(t *testing.T) {
stubMojangNames(t)
var followed atomic.Bool
target := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
followed.Store(true)
_ = json.NewEncoder(w).Encode(map[string]any{"id": notchMojangID, "name": "Notch"})
}))
t.Cleanup(target.Close)
redirector := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, target.URL+"/hasJoined?"+r.URL.RawQuery, http.StatusFound)
}))
t.Cleanup(redirector.Close)
honest := fakeYgg(t, "0123456789abcdef0123456789abcdef", "Steve0")
api := newTestAPI(newFakeRepo(), newFakeCluster())
api.AuthSources = []AuthSource{
{Tag: "evil", Prefix: "EV", URL: redirector.URL},
{Tag: "littleskin", Prefix: "LS", URL: honest.URL},
}
w := getHasJoined(api.InternalHandler(), "Steve0", "abc")
if followed.Load() {
t.Fatal("the redirect was followed")
}
if w.Code != http.StatusOK || profileOf(t, w).Name != "Steve0" {
t.Fatalf("code = %d body = %q, want the next source's player", w.Code, w.Body.String())
}
})
// A root whose headers blow past the cap is dropped like any failed source, even when the
// body behind them is a well-formed profile.
t.Run("oversized response headers skip the source", func(t *testing.T) {
stubMojangNames(t)
bloated := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("X-Padding", strings.Repeat("a", 64<<10))
_ = json.NewEncoder(w).Encode(map[string]any{"id": notchMojangID, "name": "Steve0"})
}))
t.Cleanup(bloated.Close)
honest := fakeYgg(t, "0123456789abcdef0123456789abcdef", "Steve0")
api := newTestAPI(newFakeRepo(), newFakeCluster())
api.AuthSources = []AuthSource{
{Tag: "evil", Prefix: "EV", URL: bloated.URL},
{Tag: "littleskin", Prefix: "LS", URL: honest.URL},
}
w := getHasJoined(api.InternalHandler(), "Steve0", "abc")
want := undashed(uuid.NewMD5(felisAuthNS, []byte("littleskin:0123456789abcdef0123456789abcdef")))
if w.Code != http.StatusOK || profileOf(t, w).ID != want {
t.Fatalf("code = %d body = %q, want the next source's player", w.Code, w.Body.String())
}
})
// The reused bar gate: a barred CANONICAL UUID is rejected at the resolver, so a
// reclaimed squatter stays out even on a consumer with no limbo plugin. Keyed on the
// dashed canonical (post-rewrite), the same form Repo.ReclaimUsername stores.
@@ -205,13 +326,179 @@ func TestHasJoined(t *testing.T) {
}
})
// Missing query fields → 204 without touching any source.
t.Run("missing username -> 204", func(t *testing.T) {
// With the bar list unreadable, nobody can say the player is not barred; the login must
// not go through. (Velocity reports the 500 as the auth servers being down.)
t.Run("bar list lookup error -> not admitted", func(t *testing.T) {
mojang := fakeYgg(t, notchMojangID, "Notch")
api := newTestAPI(blacklistDownRepo{newFakeRepo()}, newFakeCluster())
api.AuthSources = []AuthSource{{Tag: "mojang", URL: mojang.URL, Identity: true}}
if w := getHasJoined(api.InternalHandler(), "Notch", "abc"); w.Code == http.StatusOK {
t.Fatalf("admitted with the bar list unreadable (%q)", w.Body.String())
}
})
// Mojang is trusted for its UUIDs, which is exactly why one that does not parse must not
// be emitted as some default: every such login would share the nil UUID.
t.Run("identity source with an unparseable id -> 204", func(t *testing.T) {
mojang := fakeYgg(t, "not-a-uuid", "Notch")
api := newTestAPI(newFakeRepo(), newFakeCluster())
api.AuthSources = []AuthSource{{Tag: "mojang", URL: "http://127.0.0.1:0", Identity: true}}
w := do(api.InternalHandler(), "GET", "/session/minecraft/hasJoined?serverId=abc", "", nil)
if w.Code != http.StatusNoContent {
t.Fatalf("code = %d, want 204", w.Code)
api.AuthSources = []AuthSource{{Tag: "mojang", URL: mojang.URL, Identity: true}}
if w := getHasJoined(api.InternalHandler(), "Notch", "abc"); w.Code != http.StatusNoContent {
t.Fatalf("code = %d, want 204 (%q)", w.Code, w.Body.String())
}
})
// The player's address is what lets a source refuse a session relayed from another IP
// (prevent-proxy-connections); it has to reach the source unchanged.
t.Run("ip is forwarded to the source", func(t *testing.T) {
got := make(chan string, 1)
src := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
got <- r.URL.Query().Get("ip")
w.WriteHeader(http.StatusNoContent)
}))
t.Cleanup(src.Close)
api := newTestAPI(newFakeRepo(), newFakeCluster())
api.AuthSources = []AuthSource{{Tag: "mojang", URL: src.URL, Identity: true}}
do(api.InternalHandler(), "GET", "/session/minecraft/hasJoined?username=Notch&serverId=abc&ip=203.0.113.9", "", nil)
if ip := <-got; ip != "203.0.113.9" {
t.Fatalf("source saw ip %q, want 203.0.113.9", ip)
}
})
// A GET that declares a body it never sends must still be answered and lose its
// connection; otherwise each such socket stays open for as long as the client likes.
t.Run("request declaring a body is refused and closed", func(t *testing.T) {
api := newTestAPI(newFakeRepo(), newFakeCluster())
srv := httptest.NewServer(api.InternalHandler())
t.Cleanup(srv.Close)
conn, err := net.Dial("tcp", srv.Listener.Addr().String())
if err != nil {
t.Fatal(err)
}
defer conn.Close()
_, _ = io.WriteString(conn, "GET /session/minecraft/hasJoined?username=a&serverId=b HTTP/1.1\r\nHost: x\r\nContent-Length: 1000\r\n\r\n")
_ = conn.SetReadDeadline(time.Now().Add(3 * time.Second))
resp, err := http.ReadResponse(bufio.NewReader(conn), nil)
if err != nil {
t.Fatalf("no answer while the declared body never arrives: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest || !resp.Close {
t.Fatalf("code = %d close = %v, want 400 with Connection: close", resp.StatusCode, resp.Close)
}
})
// A missing or oversized field is answered 204 before any source sees it. The source
// here validates anything it is asked, so only a request that never reaches it is a 204.
t.Run("missing or oversized query field -> 204 without asking a source", func(t *testing.T) {
var hits atomic.Int32
src := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
hits.Add(1)
_ = json.NewEncoder(w).Encode(map[string]any{"id": notchMojangID, "name": "Notch"})
}))
t.Cleanup(src.Close)
api := newTestAPI(newFakeRepo(), newFakeCluster())
api.AuthSources = []AuthSource{{Tag: "mojang", URL: src.URL, Identity: true}}
long := strings.Repeat("a", maxHasJoinedParam+1)
for _, query := range []string{
"serverId=abc",
"username=Notch",
"username=" + long + "&serverId=abc",
"username=Notch&serverId=" + long,
"username=Notch&serverId=abc&ip=" + long,
} {
w := do(api.InternalHandler(), "GET", "/session/minecraft/hasJoined?"+query, "", nil)
if w.Code != http.StatusNoContent || hits.Load() != 0 {
t.Fatalf("%.40s: code = %d, source asked %d times; want 204 and 0", query, w.Code, hits.Load())
}
}
if w := getHasJoined(api.InternalHandler(), "Notch", "abc"); w.Code != http.StatusOK {
t.Fatalf("a well-formed login: code = %d, want 200 (the source is live)", w.Code)
}
})
}
// profileAPIAnswering stands in for api.mojang.com answering every lookup with status and
// counts how often it is asked. 200 means taken, 404 free, anything else is an outage.
func profileAPIAnswering(t *testing.T, status int) *atomic.Int32 {
t.Helper()
var n atomic.Int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
n.Add(1)
w.WriteHeader(status)
}))
t.Cleanup(srv.Close)
setProfileAPI(t, srv.URL+"/")
return &n
}
func seedPremium(name string, taken bool, age time.Duration) {
premiumNames.Lock()
premiumNames.m[strings.ToLower(name)] = premiumEntry{taken: taken, at: time.Now().Add(-age)}
premiumNames.Unlock()
}
// The premium-name cache decides on each third-party login whether the player keeps their
// name. Each case pins one rule that an innocent-looking edit would break unnoticed.
func TestPremiumNameCache(t *testing.T) {
ctx := context.Background()
for _, status := range []int{http.StatusTooManyRequests, http.StatusServiceUnavailable} {
t.Run(fmt.Sprintf("mojang %d is an outage, not a free name", status), func(t *testing.T) {
profileAPIAnswering(t, status)
if !isPremiumName(ctx, "Notch") {
t.Fatal("name treated as free; a squatter would keep it through the outage")
}
})
}
t.Run("a free answer is asked again once it expires", func(t *testing.T) {
n := profileAPIAnswering(t, http.StatusNotFound)
seedPremium("Steve0", false, premiumFreeTTL+time.Minute)
isPremiumName(ctx, "Steve0")
if n.Load() != 1 {
t.Fatalf("expired free answer: %d lookups, want 1", n.Load())
}
})
t.Run("answers within their TTL come from the cache", func(t *testing.T) {
n := profileAPIAnswering(t, http.StatusNotFound)
seedPremium("Steve0", false, premiumFreeTTL-time.Minute)
seedPremium("Notch", true, premiumTakenTTL-time.Hour)
if isPremiumName(ctx, "Steve0") || !isPremiumName(ctx, "Notch") || n.Load() != 0 {
t.Fatalf("cached answers not served as cached (%d lookups)", n.Load())
}
})
t.Run("an expired taken answer outlives an outage", func(t *testing.T) {
profileAPIAnswering(t, http.StatusServiceUnavailable)
seedPremium("Notch", true, 2*premiumTakenTTL)
if !isPremiumName(ctx, "Notch") {
t.Fatal("known premium name treated as free during an outage")
}
})
// The name may have been bought since it was last seen free, and its new owner is who a
// stale "free" would lock out for as long as the lookups keep failing.
t.Run("an expired free answer does not outlive an outage", func(t *testing.T) {
profileAPIAnswering(t, http.StatusTooManyRequests)
seedPremium("Steve0", false, premiumFreeTTL+time.Minute)
if !isPremiumName(ctx, "Steve0") {
t.Fatal("stale free answer trusted during an outage; a squatter keeps a just-bought name")
}
})
t.Run("the cache is cleared rather than grown past its bound", func(t *testing.T) {
profileAPIAnswering(t, http.StatusNotFound)
for i := range premiumCacheMax {
seedPremium(fmt.Sprintf("n%d", i), false, 0)
}
isPremiumName(ctx, "Steve0")
premiumNames.Lock()
size := len(premiumNames.m)
premiumNames.Unlock()
if size != 1 {
t.Fatalf("cache holds %d entries after passing its bound, want 1", size)
}
})
}
+1 -1
View File
@@ -149,7 +149,7 @@ func (a *API) handleInternalWake(w http.ResponseWriter, r *http.Request) {
}
if !ok {
writeError(w, r, newError(http.StatusServiceUnavailable, "at_capacity",
"the cluster is at its running-server cap (spec §9.1); retry once a server stops"))
"the cluster is at its running-server cap; retry once a server stops"))
return
}
+3 -3
View File
@@ -40,16 +40,16 @@ func TestSetupNoSMTPFlow(t *testing.T) {
}
// 2. Record the email — NO OTP. The row is written but email_verified stays false.
w := do(h, "POST", "/api/v1/account/email", `{"email":"[email protected]"}`, jsonHeader)
w := do(h, "POST", "/api/v1/account/email", `{"email":"[email protected]"}`, jsonHeader)
if w.Code != http.StatusOK {
t.Fatalf("set-email code = %d, want 200 (%s)", w.Code, w.Body.String())
}
if u := repo.staff["owner"]; u.Email != "[email protected]" || u.EmailVerified {
if u := repo.staff["owner"]; u.Email != "[email protected]" || u.EmailVerified {
t.Fatalf("after record: email=%q verified=%v, want the address recorded and UNVERIFIED", u.Email, u.EmailVerified)
}
// 3. Email recorded but no passkey → STILL required (email verification is not the gate).
if s := status(t); s["setup_required"] != true || s["email"] != "[email protected]" {
if s := status(t); s["setup_required"] != true || s["email"] != "[email protected]" {
t.Fatalf("email-only status = %v, want setup_required=true (passkey still missing)", s)
}
+4 -4
View File
@@ -52,7 +52,7 @@ func (a *API) handleWake(w http.ResponseWriter, r *http.Request) {
}
if !ok {
writeError(w, r, newError(http.StatusServiceUnavailable, "at_capacity",
"the cluster is at its running-server cap (spec §9.1); retry once a server stops"))
"the cluster is at its running-server cap; retry once a server stops"))
return
}
@@ -526,7 +526,7 @@ func resolveResources(memory string, rr *resourceRequest) (string, corev1.Resour
memLim, ok := limits[corev1.ResourceMemory]
if !ok || memLim.IsZero() {
return "", corev1.ResourceRequirements{}, newError(http.StatusInternalServerError, "internal",
"refusing to create a server without a memory ceiling (§22)")
"refusing to create a server without a memory ceiling")
}
return deriveJavaHeap(memLim), corev1.ResourceRequirements{Limits: limits, Requests: requests}, nil
@@ -705,8 +705,8 @@ func (a *API) handlePatchServer(w http.ResponseWriter, r *http.Request) {
// base ceiling to widen (this endpoint does not read the current spec back), so
// it is rejected rather than guessed.
var (
newResources corev1.ResourceRequirements
resUpdated bool
newResources corev1.ResourceRequirements
resUpdated bool
)
if body.Memory != nil {
javaMemory, resources, err := resolveResources(*body.Memory, body.Resources)
+2 -10
View File
@@ -88,11 +88,7 @@ func (a *API) handleCreateUser(w http.ResponseWriter, r *http.Request) {
return
}
u, err := a.Repo.CreateUser(r.Context(), CreateUserInput{
Username: body.Username,
Email: body.Email,
Role: body.Role,
}, p.Email)
u, err := a.Repo.CreateUser(r.Context(), CreateUserInput(body), p.Email)
if err != nil {
if errors.Is(err, ErrConflict) {
writeError(w, r, newError(http.StatusConflict, "already_exists",
@@ -156,11 +152,7 @@ func (a *API) handlePatchUser(w http.ResponseWriter, r *http.Request) {
return
}
u, err := a.Repo.UpdateUser(r.Context(), id, UpdateUserInput{
Username: body.Username,
Email: body.Email,
Role: body.Role,
}, p.Email)
u, err := a.Repo.UpdateUser(r.Context(), id, UpdateUserInput(body), p.Email)
if err != nil {
if errors.Is(err, ErrNotFound) {
writeError(w, r, newError(http.StatusNotFound, "not_found", "user not found"))
+15
View File
@@ -0,0 +1,15 @@
package api
import (
"os"
"testing"
)
// TestMain points the premium-name lookup at an address nothing listens on before any test
// runs. A test that forgets stubMojangNames then fails the same way everywhere (lookup
// error, fail closed, rename) instead of asking the real api.mojang.com, whose answer
// changes the day someone buys the name and which CI may not reach at all.
func TestMain(m *testing.M) {
mojangProfileAPI = "http://127.0.0.1:1/"
os.Exit(m.Run())
}
+77 -28
View File
@@ -146,13 +146,17 @@ func (p *PGRepo) RedeemPlayerBindCode(ctx context.Context, newUserID, code strin
var mcUUID, authSource string
switch err := tx.QueryRowContext(ctx,
`SELECT mc_uuid, auth_source FROM account_link_codes WHERE code = $1 AND expires_at > $2`,
`SELECT mc_uuid, auth_source FROM account_link_codes WHERE code = $1 AND expires_at > $2
FOR UPDATE`,
code, now).Scan(&mcUUID, &authSource); {
case errors.Is(err, sql.ErrNoRows):
return "", "", "", ErrLinkCodeInvalid
case err != nil:
return "", "", "", err
}
// The lock above serialises redeemers of ONE code; the ON CONFLICT arms below
// cover the rarer cross-code race (two live codes for the same UUID redeemed
// together), where both transactions reach the inserts before either commits.
// Create-or-fetch keyed on the verified UUID. An already-linked role='user' player
// is fetched (idempotent "log in via the game"); a role='admin' STAFF account is
@@ -165,16 +169,27 @@ func (p *PGRepo) RedeemPlayerBindCode(ctx context.Context, newUserID, code strin
mcUUID).Scan(&userID, &existingRole); {
case errors.Is(err, sql.ErrNoRows):
if _, err := tx.ExecContext(ctx,
`INSERT INTO users (id, username, role) VALUES ($1, $2, 'user')`,
`INSERT INTO users (id, username, role) VALUES ($1, $2, 'user')
ON CONFLICT (username) DO NOTHING`,
newUserID, mcUUID); err != nil {
return "", "", "", fmt.Errorf("create player: %w", err)
}
// Re-read by username so a cross-code race converges on the winner's row
// (our id was discarded by DO NOTHING) instead of a bare 500.
var role string
if err := tx.QueryRowContext(ctx,
`SELECT id, role::text FROM users WHERE username = $1`, mcUUID).Scan(&userID, &role); err != nil {
return "", "", "", fmt.Errorf("create player: %w", err)
}
if role != "user" {
return "", "", "", ErrPlayerBindForbidden
}
if _, err := tx.ExecContext(ctx,
`INSERT INTO account_links (user_id, mc_uuid, auth_source) VALUES ($1, $2, $3)`,
newUserID, mcUUID, authSource); err != nil {
`INSERT INTO account_links (user_id, mc_uuid, auth_source) VALUES ($1, $2, $3)
ON CONFLICT (mc_uuid) DO NOTHING`,
userID, mcUUID, authSource); err != nil {
return "", "", "", fmt.Errorf("write account link: %w", err)
}
userID = newUserID
case err != nil:
return "", "", "", err
default:
@@ -1401,7 +1416,7 @@ func (p *PGRepo) UpdateUser(ctx context.Context, userID string, patch UpdateUser
argn++
args = append(args, userID)
q := `UPDATE users SET ` + fmt.Sprintf("%s", sets[0])
q := "UPDATE users SET " + sets[0]
for _, s := range sets[1:] {
q += ", " + s
}
@@ -1878,29 +1893,61 @@ func (p *PGRepo) UserByEmail(ctx context.Context, email string) (*StaffUser, err
return &u, nil
}
// ConsumeLoginEmailOTP redeems a live code for the PRE-SESSION email login door.
// Unlike VerifyEmailOTP it has no identity side-effects: it neither writes
// users.email nor runs the verified-email uniqueness guard — login already
// resolved the userID via UserByEmail, which requires email_verified, so the
// address is settled. Zero rows affected (no live code, expired, consumed, or
// hash mismatch) → ErrNotFound.
// ConsumeLoginEmailOTP redeems the live code for the PRE-SESSION email login door
// with the SAME lifecycle as VerifyEmailOTP (FOR UPDATE, expiry + attempt cap
// before the hash compare, a mismatch charges one attempt without consuming) but
// with NO identity side-effects: it neither writes users.email nor runs the
// verified-email uniqueness guard — login already resolved the userID via
// UserByEmail, which requires email_verified, so the address is settled. Errors
// are exactly ErrOTPInvalid / ErrOTPLocked (ErrEmailTaken is structurally
// impossible here).
func (p *PGRepo) ConsumeLoginEmailOTP(ctx context.Context, userID, purpose, codeHash string, now time.Time) error {
res, err := p.db.ExecContext(ctx,
`UPDATE email_otps SET consumed_at = $4
WHERE user_id = $1 AND purpose = $2 AND code_hash = $3
AND consumed_at IS NULL AND expires_at > $4`,
userID, purpose, codeHash, now)
tx, err := p.db.BeginTx(ctx, nil)
if err != nil {
return err
}
n, err := res.RowsAffected()
if err != nil {
defer tx.Rollback() //nolint:errcheck // no-op after commit
var (
id string
storedHash string
attempts int
expiresAt time.Time
)
switch err := tx.QueryRowContext(ctx,
`SELECT id, code_hash, attempts, expires_at FROM email_otps
WHERE user_id = $1 AND purpose = $2 AND consumed_at IS NULL
ORDER BY created_at DESC LIMIT 1 FOR UPDATE`,
userID, purpose).Scan(&id, &storedHash, &attempts, &expiresAt); {
case errors.Is(err, sql.ErrNoRows):
// Nothing live: never minted, already consumed, or superseded.
return ErrOTPInvalid
case err != nil:
return err
}
if n == 0 {
return ErrNotFound
if !expiresAt.After(now) {
return ErrOTPInvalid
}
return nil
if attempts >= otpMaxAttempts {
return ErrOTPLocked
}
if storedHash != codeHash {
if _, err := tx.ExecContext(ctx,
`UPDATE email_otps SET attempts = attempts + 1 WHERE id = $1`, id); err != nil {
return fmt.Errorf("record otp attempt: %w", err)
}
if err := tx.Commit(); err != nil {
return err
}
return ErrOTPInvalid
}
if _, err := tx.ExecContext(ctx,
`UPDATE email_otps SET consumed_at = $2 WHERE id = $1`, id, now); err != nil {
return fmt.Errorf("consume otp: %w", err)
}
return tx.Commit()
}
// ---- op.console staff login: in-game approval state machine (spec §B op-login) ----
@@ -1949,12 +1996,14 @@ func (p *PGRepo) OpLoginRequestByID(ctx context.Context, id string) (*OpLoginReq
// ListPendingOpLogins returns the live (pending, unconsumed, unexpired at now)
// requests oldest-first, for the in-game admin's approval prompt. A resolved or
// expired request drops out of the list, so an admin only ever sees actionable
// attempts.
// attempts. The username is joined because the approval prompt names the staff
// account; created_at orders the list and lets the prompt show how long a request
// has been waiting.
func (p *PGRepo) ListPendingOpLogins(ctx context.Context, now time.Time) ([]OpLoginRequest, error) {
const q = `SELECT id, user_id, email, expires_at
FROM op_login_requests
WHERE consumed_at IS NULL AND approved_at IS NULL AND expires_at > $1
ORDER BY created_at`
const q = `SELECT r.id, r.user_id, u.username, r.email, r.expires_at, r.created_at
FROM op_login_requests r JOIN users u ON u.id = r.user_id
WHERE r.consumed_at IS NULL AND r.approved_at IS NULL AND r.expires_at > $1
ORDER BY r.created_at`
rows, err := p.db.QueryContext(ctx, q, now)
if err != nil {
return nil, err
@@ -1963,7 +2012,7 @@ func (p *PGRepo) ListPendingOpLogins(ctx context.Context, now time.Time) ([]OpLo
var out []OpLoginRequest
for rows.Next() {
var r OpLoginRequest
if err := rows.Scan(&r.ID, &r.UserID, &r.Email, &r.ExpiresAt); err != nil {
if err := rows.Scan(&r.ID, &r.UserID, &r.Username, &r.Email, &r.ExpiresAt, &r.CreatedAt); err != nil {
return nil, err
}
r.Status = "pending"
-5
View File
@@ -7,7 +7,6 @@ import (
"encoding/base64"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"net"
"net/http"
@@ -185,7 +184,3 @@ func localAuthEnabled(ctx context.Context, repo Repo) bool {
// ensure SessionAuth satisfies ExternalAuth at compile time.
var _ ExternalAuth = SessionAuth{}
// errIsNotFound is a small helper so handlers can branch on the repo's sentinel
// without importing errors at every call site.
func errIsNotFound(err error) bool { return errors.Is(err, ErrNotFound) }
+2 -2
View File
@@ -180,7 +180,7 @@ type Backuper struct {
// just-finished Job still inside its TTL window. ErrAlreadyExists is kept only as a
// defensive no-op against the astronomically unlikely suffix collision.
//
// ponytail: unique names mean two truly simultaneous taps can schedule two backup
// Unique names mean two truly simultaneous taps can schedule two backup
// Pods; both mount the world PVC read-only so neither corrupts anything, and if they
// land on different nodes the RWO attach fails one cleanly. Add single-flight-on-
// running only if a real double-tap storm ever shows up.
@@ -200,7 +200,7 @@ func (b *Backuper) Backup(ctx context.Context, serverName, formerOwner string) e
func jobNameSuffix() string {
var b [4]byte
if _, err := rand.Read(b[:]); err != nil {
// ponytail: crypto/rand only fails if the OS RNG is gone — unrecoverable.
// crypto/rand only fails if the OS RNG is gone — unrecoverable.
panic("backupjob: crypto/rand: " + err.Error())
}
return hex.EncodeToString(b[:])
+2 -2
View File
@@ -181,8 +181,8 @@ func TestBackupJobArgsCarryServerAndOwner(t *testing.T) {
func TestBackupJobRejectsMissingInputs(t *testing.T) {
for _, tc := range []struct {
name string
mut func(*JobParams)
name string
mut func(*JobParams)
}{
{"no image", func(p *JobParams) { p.Image = "" }},
{"no world pvc", func(p *JobParams) { p.WorldPVC = "" }},
+1 -1
View File
@@ -130,7 +130,7 @@ func (r *ExecRunner) CreateTunnel(ctx context.Context, name string) (string, str
// the file (authenticating with cert.pem, keeping the same id/DNS/Access), healing
// the re-run. The secret is written to the file, not stdout.
func (r *ExecRunner) ensureCredentials(ctx context.Context, id, credPath string) error {
// ponytail: any existing file counts as healthy; re-fetch only on absence
// Any existing file counts as healthy; re-fetch only on absence
// (the failure actually seen). A truncated/zero-byte file would still
// crash-loop — validate the JSON here if that ever shows up.
if _, err := os.Stat(credPath); err == nil {
+71 -15
View File
@@ -6,6 +6,8 @@ package config
import (
"fmt"
"net"
"net/url"
"regexp"
"strings"
@@ -25,7 +27,7 @@ type Config struct {
// AuthSources is the [[auth_source]] array-of-tables: the third-party Yggdrasil
// roots the Felis-nano hasJoined multiplexer federates over, in priority order
// (config order = priority, so array-of-tables not a map — a map would lose order
// and silently break Mojang-first). Empty = the multiplexer ships off. There is
// and silently break Mojang-first). Empty = Mojang is the only source. There is
// deliberately NO identity/trusted field here: Mojang is the single code-owned
// identity anchor (cmd/felis prepends it) and every configured source is
// namespace-rewritten, so no config can mint a source whose self-asserted UUIDs are
@@ -36,8 +38,11 @@ type Config struct {
// AuthSourceConfig is one [[auth_source]] entry: a third-party Yggdrasil root the
// Felis-nano multiplexer federates over. Tag names the source's per-source UUID
// namespace (must be unique — two sources sharing a tag would collide onto one identity);
// URL is the full hasJoined endpoint (scheme-qualified) the query string is appended to.
// namespace (must be unique — two sources sharing a tag would collide onto one identity).
// It is permanent: every player UUID of the source is hashed from it byte for byte, so
// changing it, even its case, gives all of them new UUIDs and orphans their playerdata,
// account links and bans. URL is the full hasJoined endpoint (scheme-qualified) the query
// string is appended to.
// Prefix is what a player from this source is renamed with when their name belongs to a
// Mojang player (LS_steve) — player-visible, so it is written out rather than derived from
// the tag, which cannot know that "littleskin" is meant to read LS.
@@ -253,9 +258,6 @@ func LoadNano(path string) (*Config, error) {
if err != nil {
return nil, err
}
if cfg.Server.Listen == "" {
cfg.Server.Listen = defaultListen
}
if err := cfg.validateAuthSources(); err != nil {
return nil, err
}
@@ -299,7 +301,7 @@ func (c *Config) Validate() error {
return fmt.Errorf("config: [archive] store %q is not one of tarLocal|tarS3|volumeSnapshot|longhorn", c.Archive.Store)
}
if _, ok := implementedArchiveStores[c.Archive.Store]; !ok {
return fmt.Errorf("config: [archive] store %q is recognized by §19 but not implemented in this build — only tarLocal is supported; set store = \"tarLocal\"", c.Archive.Store)
return fmt.Errorf("config: [archive] store %q is not implemented in this build — only tarLocal is supported; set store = \"tarLocal\"", c.Archive.Store)
}
switch c.K8s.EgressMode {
case "loadbalancer", "nodeport":
@@ -339,12 +341,13 @@ var authSourcePrefixRe = regexp.MustCompile(`^[A-Za-z0-9]{1,4}$`)
// validateAuthSources checks the [[auth_source]] block: each needs a namespace tag, a rename
// prefix, and a scheme-qualified hasJoined URL, and both tag and prefix must be unique. A
// blank or duplicate tag collapses two sources into one UUID namespace (cross-source
// impersonation — the exact invariant the per-source rewrite exists to hold); a duplicate
// prefix collapses two same-named players from different sources onto one in-game name
// (they stay distinct identities, but neither can be online while the other is); a
// scheme-less URL makes http.NewRequest fail so the source is silently dead (never validates
// any login). All fail fast at load, not per-login. Split out from Validate so the nano-only
// blank, duplicate or colon-bearing tag collapses two sources into one UUID namespace
// (cross-source impersonation — the exact invariant the per-source rewrite exists to
// hold); a duplicate prefix collapses two same-named players from different sources onto
// one in-game name
// (they stay distinct identities, but neither can be online while the other is); a URL
// the resolver cannot query leaves the source silently dead (never validates any login).
// All fail fast at load, not per-login. Split out from Validate so the nano-only
// LoadNano (no control-plane fields) enforces the identical rules — the impersonation guard
// has one owner, shared by full-api and nano.
func (c *Config) validateAuthSources() error {
@@ -354,6 +357,25 @@ func (c *Config) validateAuthSources() error {
if s.Tag == "" {
return fmt.Errorf("config: [[auth_source]] #%d has an empty tag; each source's tag is its per-source UUID namespace", i+1)
}
// A player's UUID is derived from tag+":"+nativeID, and the native id is whatever the
// source says it is. With a ':' allowed in tags, "guild" answering id "eu:X" hashes
// exactly like "guild:eu" answering "X", so one source could mint another's players.
// Colon-free tags make the join unambiguous. The charset is otherwise left open, because
// renaming an existing tag would move every one of its players to a new UUID.
if strings.Contains(s.Tag, ":") {
return fmt.Errorf("config: [[auth_source]] tag %q contains ':'; the tag and a player's native id are joined with ':' to derive their UUID, so a ':' in a tag would let another source mint this source's players", s.Tag)
}
// Refused for the same permanence: a stray space is invisible in the file yet is a
// different namespace, and so a different UUID for every player of the source.
if strings.TrimSpace(s.Tag) != s.Tag {
return fmt.Errorf("config: [[auth_source]] tag %q has leading or trailing whitespace; the tag is hashed into every player UUID of the source, so an invisible edit to it would give all of them new ones", s.Tag)
}
// Mojang is the built-in first source. A listed "mojang" is never it: it is asked again,
// after Mojang, on every login that reaches it, and nano's startup list then reads as if
// Mojang had been pointed at that url.
if strings.EqualFold(s.Tag, "mojang") {
return fmt.Errorf("config: [[auth_source]] tag %q is reserved: Mojang is built in as the first source and must not be listed", s.Tag)
}
if _, dup := seenTags[s.Tag]; dup {
return fmt.Errorf("config: [[auth_source]] tag %q is used twice — tags are per-source UUID namespaces and must be unique", s.Tag)
}
@@ -368,9 +390,43 @@ func (c *Config) validateAuthSources() error {
return fmt.Errorf("config: [[auth_source]] prefix %q is used twice — two sources sharing a prefix rewrite their same-named players onto the same in-game name", s.Prefix)
}
seenPrefixes[lower] = struct{}{}
if !strings.HasPrefix(s.URL, "http://") && !strings.HasPrefix(s.URL, "https://") {
return fmt.Errorf("config: [[auth_source]] %q url %q must be a scheme-qualified http(s):// hasJoined endpoint", s.Tag, s.URL)
if problem := hasJoinedURLProblem(s.URL); problem != "" {
return fmt.Errorf("config: [[auth_source]] %q url %q %s", s.Tag, s.URL, problem)
}
}
return nil
}
// hasJoinedURLProblem says why u cannot be queried as a hasJoined endpoint, or "" if it
// can. The resolver appends "?username=…&serverId=…" to it as a string, so a query or
// fragment already in it swallows those parameters, and a URL the client cannot send only
// fails one login at a time, with the source looking like it knows nobody.
func hasJoinedURLProblem(u string) string {
if strings.TrimSpace(u) != u {
return "has leading or trailing whitespace"
}
p, err := url.Parse(u)
switch {
case err != nil:
return "does not parse: " + err.Error()
case p.Scheme != "http" && p.Scheme != "https":
return "must be a scheme-qualified http(s):// hasJoined endpoint"
case p.Host == "":
return "has no host"
case strings.ContainsAny(u, "?#"):
return "must not carry a query or fragment; the username and serverId parameters are appended to it"
case p.Scheme == "http" && !plaintextHostOK(p.Hostname()):
return "sends logins in plaintext to a public host, where anyone on the path can answer as any player of this source; use https://, or http:// only for localhost or a loopback or private IP address"
}
return ""
}
// plaintextHostOK is decided on the literal host because nothing is resolved at load time,
// so a LAN root named by hostname needs its IP address or https.
func plaintextHostOK(host string) bool {
if strings.EqualFold(host, "localhost") {
return true
}
ip := net.ParseIP(host)
return ip != nil && (ip.IsLoopback() || ip.IsPrivate())
}
+120 -12
View File
@@ -233,19 +233,28 @@ url = "https://guild.example.net/sessionserver/session/minecraft/hasJoined"
// is no identity/trusted field on AuthSourceConfig, so an attempt to set one is an unknown
// key and Load rejects it loudly. A config can therefore never mint a source whose
// self-asserted UUIDs are trusted verbatim — the impersonation hole stays closed.
//
// The source is otherwise valid, so the only thing left to reject is the identity key itself:
// with a missing prefix the prefix rule would fail first and hide a loader that accepts it.
func TestLoadRejectsAuthSourceIdentityKey(t *testing.T) {
_, err := config.Load(writeTOML(t, `
const source = `
[[auth_source]]
tag = "evil"
prefix = "EV"
url = "https://evil.example.net/hasJoined"
identity = true
`
_, errFull := config.Load(writeTOML(t, `
[server]
root_domain = "mc.example.net"
[database]
url = "postgres://felis@db/felis"
[[auth_source]]
tag = "evil"
url = "https://evil.example.net/hasJoined"
identity = true
`))
if err == nil {
t.Fatal("expected error for an identity= key on [[auth_source]]")
`+source))
_, errNano := config.LoadNano(writeTOML(t, source))
for loader, err := range map[string]error{"Load": errFull, "LoadNano": errNano} {
if err == nil || !strings.Contains(err.Error(), "unknown keys") || !strings.Contains(err.Error(), "identity") {
t.Errorf("%s: err = %v, want the identity key rejected as unknown", loader, err)
}
}
}
@@ -275,6 +284,61 @@ url = "https://b.example.net/hasJoined"
}
}
// TestLoadRejectsColonInAuthSourceTag pins the separator guard. The UUID of a third-party
// player is derived from tag+":"+nativeID, and the native id is chosen by the source, so with
// "guild" and "guild:eu" both configured the "guild" root could answer id "eu:X" and receive
// the UUID of "guild:eu"'s player X. Both loaders share the check, so both are exercised.
func TestLoadRejectsColonInAuthSourceTag(t *testing.T) {
const sources = `
[[auth_source]]
tag = "guild"
prefix = "GD"
url = "https://a.example.net/hasJoined"
[[auth_source]]
tag = "guild:eu"
prefix = "GE"
url = "https://b.example.net/hasJoined"
`
_, errFull := config.Load(writeTOML(t, `
[server]
root_domain = "mc.example.net"
[database]
url = "postgres://felis@db/felis"
`+sources))
_, errNano := config.LoadNano(writeTOML(t, sources))
for loader, err := range map[string]error{"Load": errFull, "LoadNano": errNano} {
if err == nil {
t.Errorf("%s accepted a tag containing ':'", loader)
continue
}
if !strings.Contains(err.Error(), `"guild:eu"`) || !strings.Contains(err.Error(), "':'") {
t.Errorf("%s: error should name the tag and the ':' rule, got: %v", loader, err)
}
}
}
// TestLoadRejectsPaddedAuthSourceTag: whitespace around a tag cannot be seen in the file but
// is part of the namespace every player UUID of the source is hashed from.
func TestLoadRejectsPaddedAuthSourceTag(t *testing.T) {
for _, tag := range []string{"littleskin ", " littleskin", "littleskin\t"} {
_, err := config.LoadNano(writeTOML(t, "[[auth_source]]\ntag = \""+tag+"\"\nprefix = \"LS\"\nurl = \"https://a.example.net/hasJoined\"\n"))
if err == nil || !strings.Contains(err.Error(), "whitespace") {
t.Errorf("tag %q: err = %v, want a whitespace refusal", tag, err)
}
}
}
// TestLoadRejectsMojangAuthSourceTag: Mojang is prepended in code, so a listed "mojang" is a
// second, different source that only looks like a Mojang override.
func TestLoadRejectsMojangAuthSourceTag(t *testing.T) {
for _, tag := range []string{"mojang", "Mojang"} {
_, err := config.LoadNano(writeTOML(t, "[[auth_source]]\ntag = \""+tag+"\"\nprefix = \"MJ\"\nurl = \"https://sessionserver.mojang.com/session/minecraft/hasJoined\"\n"))
if err == nil || !strings.Contains(err.Error(), "built in") {
t.Errorf("tag %q: err = %v, want a refusal saying Mojang is built in", tag, err)
}
}
}
// TestLoadRejectsSchemelessAuthSourceURL pins the silently-dead-source guard: a URL with no
// http(s):// scheme makes http.NewRequest fail, so the source never validates any login yet
// felis-api boots green. Reject at load with the scheme contract spelled out. An empty tag
@@ -298,10 +362,57 @@ url = "bare.example.net/hasJoined"
}
}
// TestLoadRejectsUnqueryableAuthSourceURL covers the URL shapes that carry a scheme yet can
// never be queried: the resolver appends the query string to the URL verbatim, so each of
// these would load green and leave a source that silently validates nobody.
func TestLoadRejectsUnqueryableAuthSourceURL(t *testing.T) {
for _, u := range []string{
"ftp://a.example.net/hasJoined",
"https://",
"https://a.example.net/hasJoined?token=x",
"https://a.example.net/hasJoined?",
"https://a.example.net/hasJoined#x",
"https://a.example.net/hasJoined ",
"https://a.example.net:bad/hasJoined",
} {
_, err := config.LoadNano(writeTOML(t, "[[auth_source]]\ntag = \"a\"\nprefix = \"AA\"\nurl = \""+u+"\"\n"))
if err == nil {
t.Errorf("url %q loaded; it can never be queried", u)
}
}
if _, err := config.LoadNano(writeTOML(t, "[[auth_source]]\ntag = \"a\"\nprefix = \"AA\"\nurl = \"http://127.0.0.1:8080/hasJoined\"\n")); err != nil {
t.Errorf("a plain loopback endpoint must load: %v", err)
}
}
// A source reached over plaintext can be answered by anyone on the path, who can then log in
// as any player of that source. Only a same-host or private-network root may skip TLS, and
// that is decided on the literal host, since nothing is resolved at load time.
func TestLoadRejectsPlaintextPublicAuthSource(t *testing.T) {
load := func(u string) error {
_, err := config.LoadNano(writeTOML(t, "[[auth_source]]\ntag = \"a\"\nprefix = \"AA\"\nurl = \""+u+"\"\n"))
return err
}
for _, host := range []string{"ygg.example.net", "203.0.113.9", "ygg.lan", "172.32.0.1", "169.254.1.1", "0.0.0.0", "[2001:db8::1]"} {
u := "http://" + host + "/hasJoined"
if err := load(u); err == nil || !strings.Contains(err.Error(), "https://") {
t.Errorf("url %q: err = %v, want a refusal that asks for https://", u, err)
}
if err := load("https://" + host + "/hasJoined"); err != nil {
t.Errorf("the same host over https must load: %v", err)
}
}
for _, host := range []string{"localhost", "LOCALHOST:8080", "127.0.0.1:8080", "127.1.2.3", "[::1]:8080", "10.0.0.5", "172.16.3.4", "192.168.1.2", "[fd00::1]"} {
if err := load("http://" + host + "/hasJoined"); err != nil {
t.Errorf("a plaintext root on %s must load: %v", host, err)
}
}
}
// TestLoadNanoAcceptsMinimalConfig is the linchpin of the Felis-nano fold: a nano host has no
// Postgres and no FQDN, so LoadNano must accept a felis.toml carrying ONLY [[auth_source]] —
// the control-plane requirements (database.url, root_domain) that full Load enforces are
// deliberately skipped. It still applies the listen default and hands back the sources.
// deliberately skipped. It hands back the sources.
func TestLoadNanoAcceptsMinimalConfig(t *testing.T) {
cfg, err := config.LoadNano(writeTOML(t, `
[[auth_source]]
@@ -315,9 +426,6 @@ url = "https://littleskin.example.net/api/yggdrasil/sessionserver/session/minecr
if len(cfg.AuthSources) != 1 || cfg.AuthSources[0].Tag != "littleskin" {
t.Fatalf("auth sources = %+v, want one littleskin source", cfg.AuthSources)
}
if cfg.Server.Listen != "0.0.0.0:8080" {
t.Errorf("default listen = %q, want 0.0.0.0:8080", cfg.Server.Listen)
}
}
// TestLoadNanoStillEnforcesAuthSourceRules pins that skipping the control-plane requirements
+1 -1
View File
@@ -220,7 +220,7 @@ func list(r *os.Root, path string) Result {
// denied; a write is left alone because writing the file leaks nothing and is
// equally futile.
//
// ponytail: an exact match on one cleaned path, not a pattern. This is the whole
// An exact match on one cleaned path, not a pattern. This is the whole
// known exposure — grep FELIS_FORWARDING_SECRET across deploy/ — and if another
// image ever persists a platform secret into the mount, add its path here rather
// than inventing a matcher.
+5 -3
View File
@@ -66,9 +66,11 @@ const (
// Params parameterises the install bundle. Namespaces and the registry location
// have safe defaults; VelocityCIDRs has none — see the field comment.
type Params struct {
// ControlNamespace is where felis-api/operator/reaper run. Their SAs live here
// and the RoleBindings' subjects reference them here, even though the Roles
// they bind to live in the minecraft (and build) namespaces.
// ControlNamespace is where felis-api/operator run; their SAs live here and the
// RoleBindings' subjects reference them here, even though the Roles they bind to
// live in the minecraft (and build) namespaces. The reaper alone runs — CronJob
// and SA — in the Minecraft namespace, because a Pod can only mount a PVC and
// use a ServiceAccount from its own namespace, and its backup PVC is there.
ControlNamespace string
// MinecraftNamespace is where MinecraftServer workloads, their RCON Secrets,
// and their world PVCs live. All three identities' minecraft-scoped Roles, and
+16 -9
View File
@@ -53,7 +53,9 @@ func ControlPlaneRBAC(p Params) RBAC {
},
// Each binding lives in the Role's namespace and names the subject SA in the
// control namespace (a RoleBinding may reference an SA from another namespace;
// its roleRef must be a Role in the binding's own namespace).
// its roleRef must be a Role in the binding's own namespace). The reaper
// binding below is the one exception: its CronJob runs in the Minecraft
// namespace, so both the SA and the subject live there.
RoleBindings: []*rbacv1.RoleBinding{
bindRole(p.MinecraftNamespace, "felis-api", p.ControlNamespace, SAAPI, ComponentAPI),
bindRole(p.BuildNamespace, "felis-api-builds", p.ControlNamespace, SAAPI, ComponentAPI),
@@ -62,11 +64,14 @@ func ControlPlaneRBAC(p Params) RBAC {
}
// The destructive fourth power is conditional on its consumer (see the doc above).
if reaperEnabled(p) {
// SAReaper lives in — and its binding subject resolves in — the MINECRAFT
// namespace, because the reaper CronJob runs there (its backup PVC is there;
// a Pod can only mount a PVC and use a ServiceAccount from its own namespace).
rbac.ServiceAccounts = append(rbac.ServiceAccounts,
controlPlaneServiceAccount(p.ControlNamespace, SAReaper, ComponentReaper))
controlPlaneServiceAccount(p.MinecraftNamespace, SAReaper, ComponentReaper))
rbac.Roles = append(rbac.Roles, ReaperRole(p))
rbac.RoleBindings = append(rbac.RoleBindings,
bindRole(p.MinecraftNamespace, "felis-reaper", p.ControlNamespace, SAReaper, ComponentReaper))
bindRole(p.MinecraftNamespace, "felis-reaper", p.MinecraftNamespace, SAReaper, ComponentReaper))
}
return rbac
}
@@ -74,11 +79,13 @@ func ControlPlaneRBAC(p Params) RBAC {
// APIMinecraftRole grants felis-api exactly what it does in the minecraft
// namespace: drive MinecraftServer specs (internal/api.k8scluster — get/list/
// create/patch, never status), read RCON passwords for console writes
// (internal/api.console — secrets:get), create the restore Job
// (internal/restore — jobs:create), and stream the live console for the read
// side (internal/api.logstream — pods:list to find the server's running pod,
// then pods/log:get to follow it; spec §8 读=pods/log follow). felis-api uses a
// DIRECT client, so it needs no list/watch beyond the explicit List calls.
// (internal/api.console — secrets:get), manage the restore Job under its
// deterministic name (internal/restore — jobs:create, plus get/delete so a
// FINISHED Job whose name still blocks a retry can be replaced), and stream the
// live console for the read side (internal/api.logstream — pods:list to find
// the server's running pod, then pods/log:get to follow it; spec §8 读=pods/log
// follow). felis-api uses a DIRECT client, so it needs no list/watch beyond the
// explicit List calls.
//
// The read-side grant is deliberately minimal: pods:list + pods/log:get, NOT
// pods:get — the streamer lists pods by the server label then reads the chosen
@@ -90,7 +97,7 @@ func APIMinecraftRole(p Params) *rbacv1.Role {
return role(p.MinecraftNamespace, "felis-api", ComponentAPI, []rbacv1.PolicyRule{
rule([]string{groupFelis}, []string{"minecraftservers"}, []string{"get", "list", "create", "patch"}),
rule([]string{groupCore}, []string{"secrets"}, []string{"get"}),
rule([]string{groupBatch}, []string{"jobs"}, []string{"create"}),
rule([]string{groupBatch}, []string{"jobs"}, []string{"create", "get", "delete"}),
// Read-side console (spec §8 读=pods/log follow): list pods to find the
// server's running pod, then read its log subresource. Two separate rules so
// the verbs stay tight — list on pods, get on pods/log, and nothing else.
+4 -2
View File
@@ -73,8 +73,10 @@ func TestAPIRole_CreatesJobsInBothNamespaces(t *testing.T) {
if mc.Namespace != "minecraft" {
t.Errorf("felis-api minecraft Role namespace = %q, want minecraft", mc.Namespace)
}
if !hasRule(mc, "batch", "jobs", "create") {
t.Error("felis-api (minecraft) must have batch/jobs:create for the restore Job")
for _, v := range []string{"create", "get", "delete"} {
if !hasRule(mc, "batch", "jobs", v) {
t.Errorf("felis-api (minecraft) must have batch/jobs:%s for the restore-Job lifecycle", v)
}
}
build := roleByName(t, rbac.Roles, "felis-api-builds")
+14 -7
View File
@@ -47,11 +47,11 @@ const (
// and the service token is a credential, so writing either into a checked-in
// manifest is a hard red line. The deployment provisions both Secrets
// out-of-band before applying these workloads.
configSecretName = "felis-config"
configSecretKey = "felis.toml"
configMountPath = "/etc/felis"
configFilePath = "/etc/felis/felis.toml"
felisBinaryPath = "/usr/local/bin/felis"
configSecretName = "felis-config"
configSecretKey = "felis.toml"
configMountPath = "/etc/felis"
configFilePath = "/etc/felis/felis.toml"
felisBinaryPath = "/usr/local/bin/felis"
// Single-sourced with the operator, which injects the same Secret into the
// login system server's pod (see internal/naming).
serviceTokenSecretName = naming.ServiceTokenSecretName
@@ -472,8 +472,15 @@ func reaperCronJob(p Params) *batchv1.CronJob {
}
return &batchv1.CronJob{
TypeMeta: metav1.TypeMeta{APIVersion: "batch/v1", Kind: "CronJob"},
ObjectMeta: metav1.ObjectMeta{Name: SAReaper, Namespace: p.ControlNamespace, Labels: labels},
TypeMeta: metav1.TypeMeta{APIVersion: "batch/v1", Kind: "CronJob"},
// The CronJob lives in the MINECRAFT namespace: a Pod can only mount PVCs
// from its own namespace and the backup PVC is provisioned there alongside
// the backup Jobs. Placed under ControlNamespace it could never schedule
// (FailedScheduling: persistentvolumeclaim not found) in any stock install;
// the reaper Role/RoleBinding were already minecraft-scoped for the same
// reason, and the minecraft felis-config replica (felis setup) supplies the
// config mount.
ObjectMeta: metav1.ObjectMeta{Name: SAReaper, Namespace: p.MinecraftNamespace, Labels: labels},
Spec: batchv1.CronJobSpec{
Schedule: reaperSchedule,
ConcurrencyPolicy: batchv1.ForbidConcurrent,
+5 -2
View File
@@ -519,8 +519,11 @@ func TestReaperCronJob_Shape(t *testing.T) {
if cj.Name != SAReaper {
t.Errorf("CronJob name = %q, want %q", cj.Name, SAReaper)
}
if cj.Namespace != p.ControlNamespace {
t.Errorf("CronJob namespace = %q, want control ns %q", cj.Namespace, p.ControlNamespace)
// The CronJob must sit where its PVC lives: a Pod cannot mount a PVC across
// namespaces, and the backup PVC is provisioned in the Minecraft namespace.
// (Rendered under ControlNamespace it failed to schedule on a live cluster.)
if cj.Namespace != p.MinecraftNamespace {
t.Errorf("CronJob namespace = %q, want minecraft ns %q (its backup PVC's namespace)", cj.Namespace, p.MinecraftNamespace)
}
spec := cj.Spec
+100 -7
View File
@@ -2,17 +2,22 @@ package restore
import (
"context"
"fmt"
"time"
batchv1 "k8s.io/api/batch/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/client"
)
// K8sJobs is the production Jobs backed by a controller-runtime client (spec
// §7, §16). It creates the world-restore Job — nothing more: the restore Job is
// one-shot and self-cleaning (ttlSecondsAfterFinished), so there is no phase or
// cancel seam, and thus no config to hold (unlike build.K8sJobs, which needs the
// namespace to read and delete its Job). Every restore parameter arrives in the
// JobParams the Restorer builds from its own (defaulted) Config. The
// §7, §16). It creates the world-restore Job and, when a previous run's finished
// Job still holds the deterministic name, replaces it (see CreateRestoreJob).
// There is no other phase or cancel seam: the Job is one-shot and self-cleaning
// (ttlSecondsAfterFinished), so unlike build.K8sJobs there is no namespace to
// hold — every restore parameter arrives in the JobParams the Restorer builds
// from its own (defaulted) Config. The
// cluster-bootstrap objects (the weak felis-restore SA) are installed once by
// the deployment manifests (spec §21), not per restore, so this binding never
// creates them. It is integration-tested against a live cluster, not the
@@ -29,13 +34,87 @@ func NewK8sJobs(c client.Client) *K8sJobs {
// CreateRestoreJob renders and applies the restore Job. Its name is a
// deterministic function of the server (RestoreJobName), so a concurrent restore
// of the same server collides on Create; that collision is mapped to
// ErrAlreadyExists, which the Restorer treats as success (idempotent enqueue).
// of the same server collides on Create. The collision is answered by the state
// of the Job already holding the name:
//
// - still running (or not yet started): ErrAlreadyExists, which the Restorer
// treats as success — the idempotent coalesce.
// - finished (succeeded OR failed): the finished Job is deleted and replaced,
// so the caller's retry enqueues for real. Without this, the deterministic
// name plus the ten-minute TTL would swallow the retry — most importantly
// the retry after a FAILED restore, which must not have to wait out the TTL
// (an E2E audit found exactly that: a retry answered 202 "restoring" while
// nothing ran).
func (k *K8sJobs) CreateRestoreJob(ctx context.Context, p JobParams) error {
job, err := RestoreJob(p)
if err != nil {
return err
}
createErr := k.c.Create(ctx, job)
if createErr == nil {
return nil
}
if !apierrors.IsAlreadyExists(createErr) {
return createErr
}
var existing batchv1.Job
getErr := k.c.Get(ctx, types.NamespacedName{Namespace: job.Namespace, Name: job.Name}, &existing)
if apierrors.IsNotFound(getErr) {
// The name freed itself (TTL cleanup raced us); one retry.
return k.recreate(ctx, job)
}
if getErr != nil {
return getErr
}
if !restoreJobFinished(&existing) {
return ErrAlreadyExists
}
if deleteErr := k.c.Delete(ctx, &existing); deleteErr != nil && !apierrors.IsNotFound(deleteErr) {
return deleteErr
}
// The API server keeps the object until its job-tracking finalizer has run,
// so an immediate re-Create would collide again and swallow the retry a second
// time (found live: the E2E retry still answered 202 while nothing ran). Wait
// for the name to actually free, bounded, then replace.
if err := k.waitForNameRelease(ctx, job.Namespace, job.Name); err != nil {
return err
}
return k.recreate(ctx, job)
}
// waitForNameRelease polls until the named Job is gone or the wait budget is
// spent. The job controller releases the tracking finalizer within a second or
// two of the delete, so this normally returns on the first or second probe; the
// bound exists so a stuck finalizer surfaces as an error ("retry shortly")
// instead of another silent success.
func (k *K8sJobs) waitForNameRelease(ctx context.Context, namespace, name string) error {
const (
probeInterval = 500 * time.Millisecond
maxProbes = 20
)
for probe := 0; probe < maxProbes; probe++ {
var probeJob batchv1.Job
err := k.c.Get(ctx, types.NamespacedName{Namespace: namespace, Name: name}, &probeJob)
if apierrors.IsNotFound(err) {
return nil
}
if err != nil {
return err
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(probeInterval):
}
}
return fmt.Errorf("restore: job %s/%s is still terminating after deletion; retry shortly", namespace, name)
}
// recreate retries Create once after a finished Job released the name. A
// collision that survives means a concurrent restore re-created first, so the
// idempotent answer applies again.
func (k *K8sJobs) recreate(ctx context.Context, job *batchv1.Job) error {
if err := k.c.Create(ctx, job); err != nil {
if apierrors.IsAlreadyExists(err) {
return ErrAlreadyExists
@@ -44,3 +123,17 @@ func (k *K8sJobs) CreateRestoreJob(ctx context.Context, p JobParams) error {
}
return nil
}
// restoreJobFinished reports whether the Job has reached a terminal state. A Job
// that is merely created-but-not-started (no active pods yet, no completions)
// counts as in flight, not finished, so a duplicate enqueue during startup still
// coalesces.
func restoreJobFinished(job *batchv1.Job) bool {
if job.Status.Active > 0 {
return false
}
if job.Status.CompletionTime != nil {
return true
}
return job.Status.Succeeded > 0 || job.Status.Failed > 0
}
+112
View File
@@ -0,0 +1,112 @@
package restore
import (
"context"
"errors"
"testing"
"time"
batchv1 "k8s.io/api/batch/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/client/fake"
)
func testScheme(t *testing.T) *runtime.Scheme {
t.Helper()
scheme := runtime.NewScheme()
if err := batchv1.AddToScheme(scheme); err != nil {
t.Fatalf("scheme: %v", err)
}
return scheme
}
func testParams() JobParams {
return JobParams{
Server: "survival",
WorldPVC: "world-survival-0",
BackupPVC: "felis-backups",
Namespace: "minecraft",
Image: "felis:test",
BackupRef: "/backups/x.tar.gz",
}
}
// A finished Job still holding the deterministic name must be replaced, not
// treated as an in-flight coalesce — otherwise the retry after a failed restore
// is answered 202 while nothing runs (found by an E2E audit).
func TestCreateRestoreJobReplacesFinishedJob(t *testing.T) {
finished := &batchv1.Job{
ObjectMeta: metav1.ObjectMeta{Name: "restore-survival", Namespace: "minecraft"},
Status: batchv1.JobStatus{
Failed: 1,
Conditions: []batchv1.JobCondition{{Type: batchv1.JobFailed, Status: "True"}},
},
}
c := fake.NewClientBuilder().WithScheme(testScheme(t)).WithObjects(finished).Build()
if err := NewK8sJobs(c).CreateRestoreJob(context.Background(), testParams()); err != nil {
t.Fatalf("CreateRestoreJob: %v", err)
}
var got batchv1.Job
if err := c.Get(context.Background(), types.NamespacedName{Namespace: "minecraft", Name: "restore-survival"}, &got); err != nil {
t.Fatalf("replacement job missing: %v", err)
}
if restoreJobFinished(&got) {
t.Errorf("replacement job is already finished: %+v", got.Status)
}
}
// An in-flight Job keeps the idempotent coalesce: a duplicate enqueue during a
// running restore is absorbed, and the running Job is left untouched.
func TestCreateRestoreJobCoalescesInFlightJob(t *testing.T) {
inFlight := &batchv1.Job{
ObjectMeta: metav1.ObjectMeta{Name: "restore-survival", Namespace: "minecraft"},
Status: batchv1.JobStatus{Active: 1},
}
c := fake.NewClientBuilder().WithScheme(testScheme(t)).WithObjects(inFlight).Build()
err := NewK8sJobs(c).CreateRestoreJob(context.Background(), testParams())
if !errors.Is(err, ErrAlreadyExists) {
t.Fatalf("CreateRestoreJob = %v, want ErrAlreadyExists", err)
}
var got batchv1.Job
if err := c.Get(context.Background(), types.NamespacedName{Namespace: "minecraft", Name: "restore-survival"}, &got); err != nil {
t.Fatalf("in-flight job should stay: %v", err)
}
if got.Status.Active != 1 {
t.Errorf("in-flight job was disturbed: %+v", got.Status)
}
}
// restoreJobFinished decides whether a name collision is a genuine in-flight
// coalesce (ErrAlreadyExists) or a finished Job whose deterministic name must be
// replaced so a retry enqueues for real. Regression: a FAILED restore used to
// absorb every retry for the rest of its ten-minute TTL — the API answered 202
// "restoring" while nothing ran (found by an E2E audit against a live cluster).
func TestRestoreJobFinished(t *testing.T) {
completion := metav1.NewTime(time.Now())
cases := []struct {
name string
job batchv1.Job
want bool
}{
{"running", batchv1.Job{Status: batchv1.JobStatus{Active: 1}}, false},
{"created-not-started", batchv1.Job{}, false},
{"succeeded", batchv1.Job{Status: batchv1.JobStatus{
Succeeded: 1, CompletionTime: &completion,
Conditions: []batchv1.JobCondition{{Type: batchv1.JobComplete, Status: "True"}},
}}, true},
{"failed", batchv1.Job{Status: batchv1.JobStatus{
Failed: 1,
Conditions: []batchv1.JobCondition{{Type: batchv1.JobFailed, Status: "True"}},
}}, true},
{"failed-and-some-active", batchv1.Job{Status: batchv1.JobStatus{Active: 1, Failed: 1}}, false},
}
for _, tc := range cases {
if got := restoreJobFinished(&tc.job); got != tc.want {
t.Errorf("%s: restoreJobFinished = %v, want %v", tc.name, got, tc.want)
}
}
}
+9 -7
View File
@@ -170,9 +170,8 @@ type Restorer struct {
// serverName's world PVC. It returns once the Job is created — the extraction
// runs in the Pod — so the handler's 202 ("restoring") is honest.
//
// It is idempotent: if a restore Job for this server already exists (a restore
// is already in flight, or a just-finished one has not yet hit its TTL), the
// duplicate enqueue is treated as success rather than surfaced as an error.
// It is idempotent: a duplicate enqueue while a restore Job for this server is
// still running is treated as success rather than surfaced as an error.
//
// The coalescing key is the Job name (RestoreJobName), which depends only on the
// server, NOT on backupRef — so a second request that arrives while one is in
@@ -180,10 +179,13 @@ type Restorer struct {
// differ the second is silently dropped (the in-flight restore wins). That is
// acceptable here: restore runs only for a Stopped server (handler gate ⑥) and
// the handler always passes the latest backup, which for a stopped server does
// not change, so concurrent requests carry the same ref in practice. A caller
// that genuinely needs a different archive can re-request after the Job clears
// its TTL. This keeps the handler's 202 honest without it having to map "already
// in progress" onto a 500.
// not change, so concurrent requests carry the same ref in practice.
//
// A FINISHED Job — succeeded or failed — does not absorb the next request: its
// deterministic name is replaced so the retry enqueues for real (see
// K8sJobs.CreateRestoreJob). Distinguishing in-flight from finished is what
// keeps the handler's 202 honest in both directions — not a 500 for a genuine
// duplicate, and not a false "restoring" for a retry after a failure.
func (r *Restorer) Restore(ctx context.Context, serverName, backupRef string) error {
if err := r.Jobs.CreateRestoreJob(ctx, r.jobParams(serverName, backupRef)); err != nil {
if errors.Is(err, ErrAlreadyExists) {
+1 -1
View File
@@ -114,7 +114,7 @@ func velocityJarVersion(path string) (updates.Version, error) {
// manifestAttr returns one attribute value from a jar's META-INF/MANIFEST.MF.
//
// ponytail: this does not implement the JAR spec's 72-byte line folding (a wrapped
// This does not implement the JAR spec's 72-byte line folding (a wrapped
// value continues on the next line after a single leading space). Version values are
// far short of the wrap point, so folding cannot bite here; if this ever reads a long
// attribute, join continuation lines before splitting on ':'.
+3 -3
View File
@@ -37,10 +37,10 @@ const (
func TestVersionFromCLI(t *testing.T) {
cases := []struct {
name string
raw string
name string
raw string
wantCore [3]int
wantStr string
wantStr string
}{
{"k3s keeps +build stable", k3sVersionBanner, [3]int{1, 36, 2}, "v1.36.2+k3s1"},
{"cloudflared calver", cloudflaredVersionBanner, [3]int{2026, 6, 1}, "2026.6.1"},
+4 -4
View File
@@ -46,10 +46,10 @@ func ghFixtureServer(body string) *httptest.Server {
// while the numeric core is what comparison uses.
func TestGitHubLatestStableParsesRealTags(t *testing.T) {
cases := []struct {
name string
repo string
body string
wantMajMinPat [3]int
name string
repo string
body string
wantMajMinPat [3]int
wantRawInReport string
}{
{"cloudflared CalVer", "cloudflare/cloudflared", cloudflaredLatestFixture, [3]int{2026, 6, 1}, "2026.6.1"},
+1
View File
@@ -17,6 +17,7 @@ import (
// - the "versions" object groups the ENTIRE 3.x line under a single key "3.0.0"
// (not per-minor keys), so a parser that trusted the group key to bound the
// versions inside it would be wrong — proof the key-agnostic flatten is required.
//
// (The v2 API this replaces now returns HTTP 410.)
const velocityV3Fixture = `{
"project": {"id": "velocity", "name": "Velocity"},
+10 -10
View File
@@ -51,7 +51,7 @@ func TestWindowContains(t *testing.T) {
// The load-bearing invariants live here. Each row is a single component evaluated
// against a latest map, at a fixed `now`, asserting the Kind the plan must yield.
func TestPlanUpdatesInvariants(t *testing.T) {
now := time.Date(2026, 7, 1, 3, 30, 0, 0, time.UTC) // inside the window below
now := time.Date(2026, 7, 1, 3, 30, 0, 0, time.UTC) // inside the window below
openWin := Window{
Start: time.Date(2026, 7, 1, 3, 0, 0, 0, time.UTC),
End: time.Date(2026, 7, 1, 4, 0, 0, 0, time.UTC),
@@ -62,10 +62,10 @@ func TestPlanUpdatesInvariants(t *testing.T) {
}
cases := []struct {
name string
comp Component
latest string // "" ⇒ absent from the map (unknown latest)
want ActionKind
name string
comp Component
latest string // "" ⇒ absent from the map (unknown latest)
want ActionKind
}{
{
name: "pinned is never touched even with a newer stable upstream",
@@ -162,11 +162,11 @@ func TestPlanUpdatesPreservesOrderAndPending(t *testing.T) {
End: time.Date(2026, 7, 1, 4, 0, 0, 0, time.UTC),
}
comps := []Component{
{Name: "felis-api", Current: mustV(t, "1.4.0"), Policy: PolicyScheduled, Manageable: true, Window: win}, // apply
{Name: "k3s", Current: mustV(t, "v1.30.2+k3s1"), Policy: PolicyNotify, Manageable: true}, // notify
{Name: "cloudflared", Current: mustV(t, "2024.2.1"), Policy: PolicyScheduled, Manageable: true}, // no window ⇒ notify
{Name: "velocity", Current: mustV(t, "3.3.0"), Policy: PolicyScheduled, Manageable: false}, // off-cluster ⇒ notify
{Name: "mc-survival", Current: mustV(t, "1.20.1"), Policy: PolicyPinned}, // pinned
{Name: "felis-api", Current: mustV(t, "1.4.0"), Policy: PolicyScheduled, Manageable: true, Window: win}, // apply
{Name: "k3s", Current: mustV(t, "v1.30.2+k3s1"), Policy: PolicyNotify, Manageable: true}, // notify
{Name: "cloudflared", Current: mustV(t, "2024.2.1"), Policy: PolicyScheduled, Manageable: true}, // no window ⇒ notify
{Name: "velocity", Current: mustV(t, "3.3.0"), Policy: PolicyScheduled, Manageable: false}, // off-cluster ⇒ notify
{Name: "mc-survival", Current: mustV(t, "1.20.1"), Policy: PolicyPinned}, // pinned
}
latest := map[string]Version{
"felis-api": mustV(t, "1.5.0"),
+12 -12
View File
@@ -9,18 +9,18 @@ func TestParseTolerant(t *testing.T) {
pre string
}{
{"1.2.3", 1, 2, 3, ""},
{"v1.2.3", 1, 2, 3, ""}, // leading v
{"V1.2.3", 1, 2, 3, ""}, // leading V
{"v1.30.2+k3s1", 1, 30, 2, ""}, // k3s build suffix ignored
{"1.30.2+k3s1", 1, 30, 2, ""}, // build suffix, no v
{"2024.2.1", 2024, 2, 1, ""}, // cloudflared calendar version
{"1.2.3-rc.1", 1, 2, 3, "rc.1"}, // prerelease
{"v3.3.0-SNAPSHOT", 3, 3, 0, "SNAPSHOT"}, // velocity-style
{"1.2.3-rc.1+build.9", 1, 2, 3, "rc.1"}, // prerelease AND build
{"v0.0.0+g1a2b3c4", 0, 0, 0, ""}, // stamp of a build pinned to a ref with no tag behind it
{"v2", 2, 0, 0, ""}, // missing minor/patch fill 0
{"2.0", 2, 0, 0, ""}, // missing patch fills 0
{" v1.2.3 ", 1, 2, 3, ""}, // surrounding whitespace
{"v1.2.3", 1, 2, 3, ""}, // leading v
{"V1.2.3", 1, 2, 3, ""}, // leading V
{"v1.30.2+k3s1", 1, 30, 2, ""}, // k3s build suffix ignored
{"1.30.2+k3s1", 1, 30, 2, ""}, // build suffix, no v
{"2024.2.1", 2024, 2, 1, ""}, // cloudflared calendar version
{"1.2.3-rc.1", 1, 2, 3, "rc.1"}, // prerelease
{"v3.3.0-SNAPSHOT", 3, 3, 0, "SNAPSHOT"}, // velocity-style
{"1.2.3-rc.1+build.9", 1, 2, 3, "rc.1"}, // prerelease AND build
{"v0.0.0+g1a2b3c4", 0, 0, 0, ""}, // stamp of a build pinned to a ref with no tag behind it
{"v2", 2, 0, 0, ""}, // missing minor/patch fill 0
{"2.0", 2, 0, 0, ""}, // missing patch fills 0
{" v1.2.3 ", 1, 2, 3, ""}, // surrounding whitespace
}
for _, c := range cases {
v, err := Parse(c.in)
@@ -691,7 +691,7 @@ public final class FelisVelocityPlugin {
StringArgumentType.getString(ctx, "server"));
return Command.SINGLE_SUCCESS;
})))
// ponytail: Brigadier matches literals before arguments, so a player
// Brigadier matches literals before arguments, so a player
// actually named "accept"/"deny" cannot be invited by name. They can
// still reach the server with /felis go, and renaming the subcommands
// would break the click handlers for a case worth less than that.
@@ -868,7 +868,7 @@ public final class FelisVelocityPlugin {
NamedTextColor.YELLOW));
return;
}
// ponytail: peek-then-take is not atomic — an invite landing in that window is
// Peek-then-take is not atomic — an invite landing in that window is
// taken instead of the one just validated. "Newest wins" is already the rule the
// book enforces, so the outcome is one this player would have got anyway; make it
// a computeIfPresent if invites ever arrive fast enough for anyone to notice.
@@ -899,7 +899,7 @@ public final class FelisVelocityPlugin {
// notifyInviter closes the loop for whoever sent the invite; without it they wait on a
// prompt they can never see the answer to. Silently skipped if they left in the meantime.
//
// ponytail: ACCEPTED means the transfer was handed to the waiting queue, which is as far
// ACCEPTED means the transfer was handed to the waiting queue, which is as far
// as this can see synchronously — a wake that fails later is reported to the guest only.
private void notifyInviter(InviteBook.Invite invite, String who, Answer answer) {
proxy.getPlayer(invite.from()).ifPresent(p -> {
@@ -56,7 +56,7 @@ final class InviteBook {
* cooldownRemaining is how long the sender must still wait, in millis, or 0 when they
* may send now.
*
* <p>ponytail: one global stamp per sender, so inviting Alex also holds off inviting
* <p>One global stamp per sender, so inviting Alex also holds off inviting
* Steve. That is the shape that actually stops the spam — a per-(sender, invitee) key
* would let one sender paper every player on the proxy at once, which is the thing
* being rate-limited. Key it per pair only if a real group of players complains.