Commit Graph
267 Commits
Author SHA1 Message Date
flyemoji ff81295aa9 fix(nano): report failing sources instead of treating them as a no
A source that timed out, answered 5xx or 429, redirected, or sent a 200
without a usable profile was skipped exactly like one that answered 204.
With nobody else validating, the login got a 204 and Velocity told the
player their account is offline-mode. Nothing was logged, so a dead or
mistyped source URL, or an http:// root that now redirects to https since
redirects stopped being followed, failed every one of its players with
no trace.

Each such failure now logs the source tag and the cause; for a 3xx it
names the Location to configure instead. When no source validates and at
least one failed, the answer is 503, which Velocity reports as the auth
servers being down and logs with the status. A source answering 204 is
still a plain no, and a validating source still wins regardless of
failures before it.

The new subtest puts a 503 source, a redirecting source and an
unreachable one each behind a Mojang that answers 204, and expects 503.
Against the previous handler every case returns 204.
2026-09-22 13:18:33 +09:00
flyemoji a0f54df2a6 fix(bootstrap): fail the nano install when the unit does not stay up
install_nano_service printed "enabled and started" straight after
systemctl restart, which returns as soon as the process is forked. An
upgrade that keeps an old felis.toml the new binary rejects (an
[[auth_source]] without a prefix, say) left the unit crash-looping in
auto-restart while the installer reported success, and every login
through the proxy failed.

The install now waits two seconds and asks systemctl is-active. A unit
that exited is in "activating (auto-restart)", which is-active does not
count as active; on real systemd a unit whose process exits 1 under
Restart=on-failure reads activating/auto-restart and is-active returns
non-zero, while a running one reads active/running and returns 0. On
failure the install prints the unit's last 20 journal lines and stops.
This also surfaces a nano unit locked out of an existing 0700 /etc/felis.

The harness runs the extracted function with systemctl stubbed both
ways. Without the check, the dead-unit cases fail.
2026-09-22 13:05:42 +09:00
flyemoji 26f685be0e fix(bootstrap): create the nano config dir world-searchable
felis-nano runs as a systemd DynamicUser, so it can read
/etc/felis/felis.toml only if others may search /etc/felis.
write_nano_config made the directory with a bare mkdir -p, which takes
its mode from root's umask. On a host hardened to umask 027 that is
0750: nano exits on "permission denied", the unit restarts every five
seconds, and no login gets through.

A missing directory is now created 0755 explicitly. An existing one
keeps its mode, because the full install sets it to 0700 to protect its
secrets and widening that from the nano path would expose them. A nano
unit locked out that way is left for the install to report.

The harness runs the extracted function under umask 027 and checks both
cases. Reverting to the bare mkdir fails the first; an unconditional
chmod 0755 fails the second. The mode checks skip on filesystems that
ignore chmod, such as Git Bash on NTFS.
2026-09-22 13:04:34 +09:00
flyemoji 1dd62a9bdc fix(nano): cap upstream response headers at 16 KiB
The hasJoined and name-lookup clients limited the body to 64 KiB but
left headers at the transport default of 1 MiB. A configured root could
answer with a megabyte of headers and stall the body, holding a few MiB
of heap per in-flight login for the full five seconds; enough parallel
logins take down the host, and every source's logins with it.

Both clients now share a transport with MaxResponseHeaderBytes set to
16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends
338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over
the cap fails the request and the resolver moves on to the next one.

The new subtest puts a source with 64 KiB of headers and a valid profile
ahead of an honest one and expects the honest player. Without the cap
the padded source wins.
2026-09-22 13:02:48 +09:00
flyemoji 2180e77cf5 chore: drop tool-name markers from source comments
Seventeen comments opened with a tag naming the tool that wrote them.
The tag goes and each comment keeps its reasoning, now starting as a
plain sentence. None of the reasoning changes.

The AGENTS.md note in .gitignore drops the story of how the file got
into the tree and keeps the one fact a reader needs: its advice to run
go fmt is destructive on this CRLF working tree.

Comments only; no code, build or test changes.
2026-09-22 12:57:19 +09:00
flyemoji 4e98ae6e56 fix(config): reject auth-source tags that contain a colon
A third-party player's canonical UUID is UUIDv3 over tag+":"+nativeID,
and the native id is whatever the source answers. Tags were only
checked for being non-empty and unique, so both "guild" and "guild:eu"
could be configured. The "guild" root could then answer hasJoined with
id "eu:X" and receive exactly the UUID of "guild:eu"'s player X, along
with their playerdata, permissions and account links. Real native ids
are 32 hex digits, so only the shorter tag's source can do this, and
only when the operator has configured such a pair; when they have, it
is a full impersonation.

Reject a ':' in a tag at load. With colon-free tags the join is
unambiguous: two different (tag, id) pairs can no longer produce the
same input, since equal inputs force equal tags and duplicate tags are
already refused. The tag is deliberately not narrowed any further.
It is a permanent UUID namespace, and forcing an operator to rename a
tag that has no colon would move every one of its players to a new
UUID. The hash input and the native id are left exactly as they were,
so no existing player's UUID changes.

Load and LoadNano share validateAuthSources; the new test runs both
against the guild / guild:eu pair and fails on the previous config.go.
2026-09-22 12:53:35 +09:00
flyemoji 1976fca809 fix(nano): always relay properties as an array
sessionProfile tagged properties with omitempty, so an upstream answer
of "properties": [] (or null, or no key at all) reached Velocity with
no properties key. A Yggdrasil root may legitimately answer that way for
a player without a skin. Velocity 3.5.1's GameProfile deserializer
passes the missing key on as null and ImmutableList.copyOf throws, so
that player hangs at login with nothing logged, even though the same
answer sent straight to Velocity is accepted. Mojang always sends
textures, which is why the premium path and the hardware runs never hit
it.

Drop omitempty and replace a nil slice with an empty one before the
response is written. Removing omitempty alone is not enough: a nil
slice marshals as null, which Velocity rejects the same way.

The new subtest feeds the relay [], null and a missing key and expects
"properties":[] every time. The previous handler fails all three.
2026-09-22 12:52:04 +09:00
flyemoji 28d3638952 fix(nano): stop following redirects from upstream Yggdrasil roots
authHTTPClient kept net/http's default redirect policy, so a configured
third-party root that answered hasJoined with a 3xx made this host fetch
whatever URL it named, up to ten hops. That is a blind SSRF into
anything the host can reach, and it includes the multiplexer's own
listener: a root that redirects back to /session/minecraft/hasJoined
re-enters the handler, which queries Mojang and every source again and
gets redirected again, until the outer 5s client timeout fires. With a
50ms Mojang stub, one login produced 86 nested handler calls and 86
Mojang requests from this host's egress IP. The loopback default does
not help, because the redirect target is resolved from this host.

Return the 3xx as the response instead. resolveHasJoined already skips
any non-200 answer and closes its body, so a redirecting source is now
treated like one that is down, and the next source gets its turn. The
same probe now makes one handler call and one Mojang request.
Neither Mojang's nor LittleSkin's hasJoined redirects.

The new subtest puts a redirecting root ahead of an honest one and
checks that the redirect target is never contacted and the honest
source's player is returned. The pre-fix handler fails it.
2026-09-22 12:50:56 +09:00
flyemoji 07bafebf0d fix(bootstrap): keep the operator's auth sources across re-runs
write_felis_toml regenerates felis.host.toml and felis.pod.toml with a
wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin
block in that heredoc. Re-running the installer, which is also what
`felis setup` does, threw away any edit to the list: a root the operator
added stopped admitting logins, and a root they removed came back. The
generated comment invited exactly that edit.

Carry the tables forward the way [smtp] already is: read every
[[auth_source]] table from the existing felis.host.toml (falling back to
felis.pod.toml) and emit the LittleSkin default only when there is no
earlier file at all. An earlier file with no tables stays empty, because
that is a Mojang-only server rather than a missing value; felis-api now
treats an empty list that way.

The file header and the comment above the list now say what survives a
re-run, and point at felis.host.toml, which is what the next run reads.

bootstrap_test.sh extracts the new function from bootstrap.sh and checks
the fresh-install default, an operator's own table carried without the
default or the following section, an empty list staying empty, and the
indented form the setup TUI writes. It passes under dash with gawk and
with mawk; forcing the function to always return the default fails five
of the new cases.
2026-09-22 12:49:25 +09:00
flyemoji 8fe255e38f fix(api): relay Mojang logins when no auth source is configured
felis api wired the hasJoined multiplexer only when felis.toml had at
least one [[auth_source]]. With none, the source list stayed nil and
every hasJoined answer was a 204. That was harmless while nothing
pointed at the route, but the installer now starts Velocity with
-Dmojang.sessionserver aimed at felis-api unconditionally, and the
generated felis.toml tells the operator to delete the LittleSkin block
for a Mojang-only server. Doing exactly that turned every login away,
premium accounts included, and felis-api logged nothing about it.

Always build the list through authSourcesFromConfig, which prepends
Mojang in code, so an empty config is a Mojang-only relay. felis nano
already behaves this way with the same file.

The new test pins authSourcesFromConfig itself: Mojang first, the only
Identity source, and still present when nothing is configured. Marking
a configured source Identity makes it fail. The call site in cmdAPI is
now a single unconditional assignment and has no unit test of its own.
2026-09-22 12:46:00 +09:00
flyemoji 800a9042a1 test: use placeholder domains in setup and system-server tests
Three tests carried the maintainer's production root domain, a personal
mailbox and the public IP of a live demo host as fixture values. None of
them needs the value to be real: the re-domain test only needs two
different roots, and the setup flow only needs a well-formed address.

Swap them for the placeholders the rest of the suite already uses
(mc.example.net, [email protected]), and move the "before" root in the
re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737
documentation range, so the stale install the test models still has an
IP-derived hostname, which is the case the refresh exists for.
2026-09-22 12:43:54 +09:00
flyemoji d9246ddae6 feat(bootstrap): verify Paper and Velocity jars against Fill's digest 2026-08-04 17:36:47 +09:00
flyemoji 3f2b28d0ec fix(bootstrap): ship deploy/paper in the embedded game-stack tar 2026-08-04 16:34:09 +09:00
flyemoji 587f183191 Merge pull request #19 from MliroLirrorsIngenuity/chore/issue-sweep
chore: work the tracker items that need no cluster
2026-07-29 01:52:15 +09:00
flyemoji 503240db7b ci: stop running the whole suite twice on every pull-request push
The header of this file argues that release.yml must not be repeated here,
because a private repository is billed twice for one answer. The push trigger
it shipped with then did exactly that: `branches: ['**']` plus `pull_request`
means a branch with an open PR runs everything once for refs/heads/<branch>
and once for refs/pull/N/merge. The concurrency group is keyed on github.ref,
which differs between the two, so neither cancels the other. Visible on this
branch's own checks: go 3m14s and go 3m3s, shell 7s and 7s, panel 25s and 24s.

Limiting the push trigger to main keeps both gates that matter -- a PR is
still checked before merge, main is still checked after -- and drops only the
duplicate. The one case that loses coverage is a branch pushed with no PR
open, where nothing has asked for the answer yet.

Verified by parsing the result with the repository's own sigs.k8s.io/yaml:
triggers are {"pull_request":null,"push":{"branches":["main"]}} and the three
jobs go/panel/shell are intact. Worth recording for the next person who
parses a workflow: YAML 1.1 reads the bare key `on` as the boolean true, so
it arrives as the string "true" after the YAML-to-JSON conversion, and a
struct tag of `json:"on"` silently matches nothing. GitHub's own parser does
not have this problem; a local check of the triggers does.

Refs #7
2026-07-28 18:20:40 +09:00
flyemoji 584d31fc49 feat(bootstrap): refuse to install an unverified Velocity fork jar
FELIS_VELOCITY_FORK_JAR replaces the proxy every player connects through, and the
only thing checked about it was that the path pointed at a readable file. A
truncated copy, a stale build left at the same path, or the two-patch jar where
the three-patch one was meant all installed silently.

It now requires FELIS_VELOCITY_FORK_JAR_SHA256 and refuses on a mismatch, hashing
stdin rather than the path for the reason install_via_plugins already documents:
sha256sum escapes its output line for a filename carrying a backslash or newline,
and the leading "\" that adds fails every comparison. The absent-digest refusal
prints the jar's actual hash, so the first run after a deliberate rebuild is one
copy-paste rather than an investigation.

The comparison ignores case and internal spaces. The fork is built on a developer
machine, which is usually Windows, and nothing there prints a digest the way
sha256sum does: Get-FileHash returns uppercase and certutil has shipped the bytes
space-separated. Comparing raw would refuse two of the three spellings of the
correct answer and word the refusal as tampering.

No digest is hardcoded, which is the half of the request this does not deliver.
The fork is built from Felis-Legacy and has never been reproduced on a second
machine, so a constant here would pin one machine's output rather than the fork.
The comment that previously asserted the build "is not byte-reproducible" is gone
too -- it was stated more confidently than the evidence supports. The fork jars on
disk carry Gradle's constant 1980-02-01 entry timestamps, so the usual reason a
jar differs between builds is already absent; that is not proof it reproduces, and
neither claim should sit in the script unmeasured.

This is deliberately not a supply-chain signature and the comment says so: an
operator who can write the jar can write the digest. What it buys is that a path
stops being an identity, and that every later re-run re-checks the same build.

Scope: the fork jar only. The else branch still curls stock Velocity from PaperMC
with no verification at all, and that is the branch a default install takes. The
digest is already in hand there -- Fill v3 returns checksums.sha256 and its
download URL is content-addressed on that same value -- and papermc_latest_jar
discards it. Left alone rather than widened into this change.

deploy/bootstrap_test.sh covers the gate's two refusals, its happy path, and the
two Windows digest spellings. Each case extracts the block under test out of
bootstrap.sh with awk and runs it with die/log stubbed, rather than transcribing
it -- a transcribed copy passes forever after someone edits the original. The
extraction is length-bounded: awk runs an unmatched end pattern to EOF, which
would quietly feed the rest of bootstrap.sh to the shell under test. bootstrap.sh
itself cannot run here; it wants root, a package manager and k3s.

A `shell` CI job runs that plus a syntax check over every tracked script. The
syntax step dispatches on each file's shebang instead of running `sh -n` across
the board. The blanket form looks fine and is a false green: on a developer
machine `sh` is usually bash and accepts everything, while the runner's `sh` is
dash. Verified against the real thing rather than an approximation -- inside
ubuntu:24.04, where /bin/sh is /usr/bin/dash, the dispatching loop passes all six
scripts and the blanket loop dies at bootstrap.sh:191 on the first of its 14
arrays.

Refs: Felis-Legacy #19
2026-07-28 18:01:39 +09:00
flyemoji 4f5014d033 feat(bootstrap): make the legacy-forwarding backend list overridable
Which backends receive their forwarded identity through the handshake address --
rather than proxy-wide modern forwarding -- was the literal string "legacy18",
assigned inside write_velocity_service. Standing up a second protocol-47 backend
therefore meant editing this script, on every host, and remembering to.

It is now FELIS_LEGACY_FORWARDING_SERVERS, defaulting to legacy18, declared beside
FELIS_NANO_LISTEN and documented in the Tunables block like every other knob. The
default is unchanged, so an existing install re-runs to the same systemd unit it
already has.

This is deliberately only half of what the list should eventually do. It is a JVM
system property, read once when Velocity starts, so it is fixed for the life of
the proxy process and a change still needs a restart -- an environment variable is
as far as a startup property can be pushed. Having the list follow the
MinecraftServer CRs is a larger change than it looks: the forwarding decision is
made by the fork's patch to Velocity core, not by the Felis plugin, so core would
have to read state the plugin owns and refreshes. The plugin already maintains a
dynamic backend registry, which is where that state would come from, but the
bridge from core to it does not exist. The comment at the assignment now says so
instead of leaving "the upgrade path is to have the operator render this list from
the MinecraftServer CRs" as though it were a small step.

The -D is now double-quoted in ExecStart. The fork trims each element -- it parses
the property as `split(",")` into a Set, mapped through String::trim with empties
filtered -- so it accepts "legacy18, legacy112", but systemd splits ExecStart on
whitespace before java sees it. Unquoted, that spelling handed java a stray
"legacy112" argument and the unit failed to start; documenting the knob as
comma-separated without quoting it would have shipped that as a footgun.

`bash -n` passes; the default resolves to legacy18, an override to the value given,
and a value containing a space renders inside a single quoted ExecStart item.
2026-07-28 17:27:02 +09:00
flyemoji 4e5a809dad docs(nano): drop the promise of a felis setup --nano that should not exist
nano.go's package comment told the reader that `felis setup --nano` installs the
multiplexer as a service. No such flag has ever existed -- `felis setup` defines
only -config and -dev -- so anyone following the comment gets "flag provided but
not defined: -nano" and exit 2.

Adding the flag was the obvious reading, and it is the wrong one. setup does not
install anything selectively: it re-images the host by running the full bootstrap
TUI, and no install-mode parameter is threaded anywhere -- nothing in Go reads or
writes FELIS_INSTALL_MODE, which is a shell variable bootstrap.sh consumes on its
own. So `--nano` could only mean one of two things. Re-run the installer in nano
mode, which is what pointing at the installer already does. Or convert a
provisioned full host into a nano one, which means tearing down k3s, Postgres and
the proxy -- an uninstall, not a flag.

setup's --dev already settled this shape once. It looked like a channel selector,
silently installed release, and the fix was to refuse it and name the installer
rather than pretend to choose. The same answer applies here, so the comment now
names the real entry point -- the installer's `[2] Felis-nano` prompt, or
FELIS_INSTALL_MODE=nano -- and records why there is no flag, so the next reader
does not reopen it.

No refusing --nano flag is added: --dev exists because it used to be a silent
no-op that people passed, and nothing has ever accepted --nano, so flag's own
"not defined" error is already the correct and clearer failure.
2026-07-28 17:27:02 +09:00
flyemoji afdbfac7a8 docs: index the deferred integration seams and correct two stale markers
INTEGRATION-ONLY and KNOWN-LIMITATION are grep-able, but the grep answers the
wrong question. Thirty-four Go sites share the two markers and they carry four
different meanings: "declared, nothing implements it" reads exactly like
"implemented, only its I/O is unreachable from here", and neither reads
differently from a limitation that was accepted on purpose and is not coming
back. docs/deferred-seams.md sorts them, following the bucketed shape
internal/updater/doc.go already uses for its own package rather than starting a
second convention.

Sorting them turned up two markers that had outlived the condition they describe.

config.go called the modpack upload transport a deferred integration after both
backends had shipped -- LocalContextStore and S3ContextStore, selected in
cmd/felis by the shape of user_uploads_context, with the uploads PVC mounted and
the felis-uploads-s3 Secret rendered. What is still deferred is the far end:
Kaniko reading that context from inside the build Pod.

updater/doc.go listed the `felis update` CLI and the off-cluster Velocity jar read
under REMAINING INTEGRATION. Both exist -- cmd/felis/update.go, and
gatherer_host.go, which answers Velocity from the installed jar's manifest and
felis-api from the running binary's build stamp. The two nil seams that bullet
also names are real, but they belong to the in-cluster gatherer only, so the
bullet now says which caller has what and which is still empty.

The index also records the collision that makes a naive grep misleading:
docs/troubleshooting.md uses [INTEGRATION-ONLY] for something else, defined in its
own opening at :19 -- the symptom is produced by the kubelet, kaniko or a live
handshake, so it cannot be reproduced from the repository. Those twelve marks say
where a failure comes from, not that something is unbuilt, and are excluded.

Both code changes are comments. Every file:line the index cites was checked
against the line it points at.
2026-07-28 17:27:01 +09:00
flyemoji 23792d6251 fix(crd): remove spec.storage.retainOnDelete rather than leave it inert
The field validated, shipped in the CRD, and reached no controller. The world
PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate,
StatefulSet deletion does not cascade to template PVCs, and no finalizer exists
anywhere in the operator. So setting it true described what already happened,
and setting it false did nothing at all. False is the worse half: it reads as a
request to delete a world, and was silently ignored.

This departs from spec v4.1 §5, which asks for
"删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half
was ever built. Restoring that line means adding a finalizer whose other listed
duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs,
so the only work it would newly do is delete worlds, on a path that does not
pass the reaper's verified-backup check. The reaper is the one thing in the
system allowed to destroy a world and it earns that by proving a backup first.
A second door without that check is not an improvement.

The spec is a frozen versioned document, so it is left alone and the departure
is recorded in troubleshooting.md §13, beside the behaviour it explains. §12
loses its inert row and its opening sentence, which existed to introduce this
one field: every field in that table is now read by a controller.

Deployed installs need nothing. A CR still carrying retainOnDelete keeps
working, because a v1 CRD prunes unknown keys on the next write and the
behaviour the field claimed to control was never conditional.

go build, go vet and go test ./... pass on Linux with zero failures; the CRD
still parses and storage keeps size and storageClassName.
2026-07-28 17:27:01 +09:00
flyemoji 417769407f ci: run the checks on push and pull request
release.yml was the only workflow and it fires on v* tags, so `go vet` and
`go test` first met a change once that change was already on the release path,
where the only remedy is another tag. The panel suite ran nowhere at all: a
release goes through the Dockerfile and the Dockerfile runs `npm run build`,
never `npm test`. 111 assertions across 8 files existed and nothing outside a
developer's checkout ever executed them.

Both jobs are green as of this commit, checked before writing it rather than
after: go vet and go test ./... (24 packages, 0 failures, on Linux), npm test
(8 files, 111 tests) and npm run typecheck. A gate that lands red is a gate
everyone learns to ignore.

The panel's Node version is read out of the Dockerfile instead of repeated
here. `FROM node:<major>` is the only place the tree declares it -- no .nvmrc,
no engines field -- so a copy in this file would keep testing 22 the first time
the image moved. That is the class of drift this workflow exists to catch, not
to introduce. The step fails loudly if the FROM line stops matching.

Tags are excluded from the push trigger. A v* push already runs release.yml,
which repeats the Go job, and this is a private repository billed for both.
2026-07-28 17:27:00 +09:00
flyemoji 82a1275fcf docs(plugins): say which plugin jars an install actually produces
The module table listed felis-fabric, felis-forge and felis-neoforge next to the
two jars a finished install really has, with nothing distinguishing them. Neither
deploy/bootstrap.sh nor the embed set in bootstrap_asset.go builds a loader mod,
so someone reading the table expected three jars that are not there after setup
and had no way to tell from this file. The mods do build -- the wrapper commands
under Building work -- they are just never installed for you, which is what the
new column says.

Two further disagreements with the code, in the same table:

  limbo/ was missing entirely. It is embedded, built by bootstrap.sh and running
  on the login gate, so the one module the table omitted was a shipped one. It is
  also the only module that reaches the account-link endpoint without a command:
  it mints the code on join for anyone unlinked and holds them until they redeem
  it, so the opening claim that every module except the lobby ships /link named
  the wrong exception.

  velocity was listed as felis-velocity-0.2.0.jar; plugins/velocity/build.gradle:6
  says 0.1.0, as does every other module. Nothing breaks on this because
  bootstrap.sh globs felis-velocity-*.jar and installs it under a fixed name, but
  the version in the table was not a version anything produces.

The Gradle table and the JDK note now carry limbo's Java-21 toolchain, which it
needs for the same reason paper does and for a different cause: LOOHP/Limbo
releases are class-file major 65, so the compiler JDK must be able to read them.
It still emits release 17 bytecode.
2026-07-28 14:33:37 +09:00
flyemoji 3af5cc360c docs(troubleshooting): correct four fields the runbook documents as inert
Sections 11 and 12 told the operator that idle auto-stop, both startup
budgets, and the player tally are read by nobody. All four are read, and
§1 repeated the same claim in its strongest form: "the operator has no
start timeout ... loops forever".

  spec.idle.autoStopEnabled       reconciler.go:175
  spec.idle.emptySecondsBeforeStop  reconciler.go:175
  spec.startup.timeoutSeconds     reconciler.go:479, called at :126
  spec.startup.readinessTimeoutSeconds  reconciler.go:490, called at :157
  status.players.online           reconciler.go:413 (markRunningReady)

The repository already contained the disproof:
TestReconcileRunning_StartupTimeoutConvertsToFailed and
TestReconcileRunning_ReadinessTimeoutConvertsToFailed both assert the
escalation §1 said does not exist. The test §1 cited,
TestReconcileRunning_RconProbeFailureStaysStarting, only asserts that a
single failed probe does not flap the phase; that was read as "forever".

§11 was the costly one, because it misdiagnosed a configuration problem
as a missing feature. Idle auto-stop and the player tally both hang off
spec.rcon.enabled -- the tally is a by-product of the RCON readiness
probe (prober.go:63 runs `list`), and reconciler.go:175 carries the RCON
condition explicitly so a never-sampled zero cannot stop a server full of
people. Following the old text, an operator whose RCON was never enabled
would conclude the feature was unwritten and stop. The section now opens
with the jsonpath that reads spec.rcon.enabled.

Also separated the prober's fixed 5s dial timeout (prober.go:45) from
spec.startup.readinessTimeoutSeconds, which §1c conflated: the former
bounds one probe, the latter is a deadline for the whole start measured
from status.startRequestedAt.

spec.storage.retainOnDelete is the one field still genuinely inert, so
the [INERT] legend and §12 stay -- §12 now records the condition each
read field depends on instead of claiming none of them are read.
2026-07-28 13:49:07 +09:00
flyemoji 71e1664c36 build: pin line endings to LF so the embedded bootstrap.sh ships without CRs
The repository had no .gitattributes. With core.autocrlf=true a Windows checkout
handed deploy/bootstrap.sh 2374 CRs, and bootstrap_asset.go embeds that file from
the working tree verbatim, so a dev-built felis piped a CRLF script into `bash -s`
on the target host. CI builds on Linux, which is why released binaries were clean
and only local builds carried it.

eol=lf is global rather than scoped to *.sh because go:embed reaches further than
the installer: deploy/*/Dockerfile, deploy/*/entrypoint.sh, plugins/*/src, the
migrations and internal/panel/static are all compiled in and read on Linux. *.bat
is the one exception, for the gradle wrappers' Windows launchers.

Renormalizing the index touched exactly one tracked file, cmd/felis/version.go,
and only its line endings: `git diff --cached --ignore-cr-at-eol` reports nothing
outside .gitattributes itself.

TestBootstrapPinsViaBlockConnectionsOff used to strip \r\n before asserting, with a
comment stating that the repository pinned no eol attribute. That is no longer true,
and the stripping hid the regression this commit prevents. It now asserts the absence
of CRs, so losing the attribute reports itself as line endings rather than as a
missing serverside-blockconnections pin.

Closes #5
2026-07-28 12:58:46 +09:00
flyemoji 5d4f3063a9 feat(operator): make any Paper image joinable behind the forwarding proxy
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed
handshake does not degrade -- it rejects every login the proxy forwards. Until now the
only backends that could verify it were the two images Felis builds itself
(deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own
entrypoints. An arbitrary Paper image a user brings does not, so it passed admission,
started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the
lobby image as a base for a user's own world (0018_recommended_images.sql), which was
never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining
player away, the exact opposite of a server you mean to stay on.

The fix configures forwarding from OUTSIDE the image instead of requiring it inside.
The operator now injects a root `felis init-forwarding` initContainer into every user
server; it writes the proxies.velocity block into config/paper-global.yml and forces
online-mode=false in server.properties on the /data PVC before the main container
starts. The image needs no forwarding logic of its own, so the joinable set stops being
"images that self-configure forwarding" and becomes every Paper-family image the
platform runs.

buildStatefulSet gates the injection on the ABSENCE of the system-role label: the
Felis-built system servers already consume the secret in their entrypoints and the
login gate is a limbo, not Paper. It is also gated on a non-empty felis image name --
the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it
skips the injection rather than failing, because a cluster whose proxy is not in modern
mode has nothing to configure.

The init runs as root deliberately. The world volume's ownership comes from the storage
provisioner and the main container runs as whatever UID its image declares, so root is
the only UID that can reliably write these files; it then chmods them 0666/0777 so that
non-root main container can rewrite them on boot. The privilege is bounded -- the init
exits before the server container starts and the server container keeps its own UID.
The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the
init ever stops running as root.

The writer merges rather than overwrites, both because Paper expands paper-global.yml to
its full default tree on first boot and because the panel file editor may edit either
file between boots. It sets proxies.velocity.* and the single online-mode key and leaves
every other setting alone. It is a no-op on an empty secret, for the same reason the env
var is optional: a proxy that is not in modern mode provisions no Secret, and wedging
every server's init on a missing optional value would be worse than the status quo.

felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and
0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu
plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the
online-player list and permission commands work out of the box. 0018's row is left in
place -- an admin who kept it can keep it; this only adds the better default beside it.

Three fixes ride along, each of which the 1.8 path hit in practice.

bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only
builds its block-connection provider when Via's lowest supported protocol is below 1.13;
under modern forwarding the Velocity injector reports 393, so init() returns early,
blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences
it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining.
Every call site is behind isServersideBlockConnections(), so switching it off skips all
of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected.
ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with
this one key suffices -- Config#loadConfig parses the bundled default as the base map and
merges the on-disk file over it, so every other option stays current across version
bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this
worked: init() gates on the protocol version too, and that half fails on its own, so the
line is missing either way. The config value is the only evidence, which is what the test
asserts.

The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend
sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it
down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13
and has nowhere to go. Only the handshake address field survives Via, so the Felis fork
forwards the named servers BungeeCord-style while every other backend keeps modern+secret
untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer
CRs is the upgrade path.

deploy/lobby's set_prop escapes the value before substituting it. The RCON password is
operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare
`sed s|...|...|` and silently kills the key -- taking the console, the online-player list
and permission commands with it. deploy/paper was written with the escaping, so the lobby
gets the same rather than leaving the sibling caller broken.

Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where
TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping.
The new tests cover the initContainer's image, root UID, world mount and secret env; the
merge preserving unrelated config trees; the properties upsert including the commented-key
case; and the bootstrap script both writing the Via key and still calling the function
that writes it.

Not verified: the initContainer has never run in a real cluster, and the felis-paper
image is code-only here as the other game-stack images are -- no Go CI builds them.

The ViaVersion pin is the one piece with live evidence, and that evidence is what it was
written from. Before it, a client was cut within a second of "logged in with entity id"
on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into
Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on
2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same
player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login
from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes.

Neither session says which client version it was. The proxy never logged a protocol
number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16
players and below" warning that fired for that player on the lobby leg, and no lower --
Via floors every handshake to the proxy's 393, so anything from 47 up is admissible.
Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain
runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the
NPE proves is that the pin was load-bearing, not who was holding the mouse.

That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one
now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8
behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the
fault above; off, holds. No automated test in THIS repository exercises the leg.
2026-07-28 09:13:58 +09:00
flyemoji 0798f903b0 feat(bootstrap): allow the Felis-Legacy Velocity fork to be installed as the proxy
Stock Velocity will not offer the login-plugin-message exchange below 1.13, so a
1.8 client reaching a modern-forwarding backend today is a side effect of Via
replacing the channel initialisers before that check runs. It works, and nobody
designed it. FL-008's fork registers the two login packets from 1.7.2 and drops
the handshake gate, which makes the same outcome deliberate — and its gateonly
control shows the registry half is the load-bearing one.

FELIS_VELOCITY_FORK_JAR points at such a build; unset, the default, nothing
changes and the stock 3.5.1 download runs as before. It stays opt-in because the
fork is unmeasured where it counts: FL-008's probe runs offline-mode against a
stub backend, while this jar would carry every real Mojang session on the server.

No digest is pinned for it. The gradle build is not byte-reproducible across
machines, so a hash here would assert a provenance that does not exist; the jar
is trusted because that probe certified a build, and the path is checked for
readability before anything is replaced.
2026-07-23 03:31:08 +09:00
flyemoji c9f3aa68f6 fix(velocity): stop the lobby tile failing for players already in the lobby
Asking for the server you are standing on is a no-op, but it went through the
whole wake-and-queue path and ended in a Connect to the current server, which
Velocity answers ALREADY_CONNECTED. The player saw a failure for a move that
was never real.

The guard belongs in authorizeAndWait rather than at each entry point: the
menu tile, /felis go and an accepted invite all funnel through it. The lobby
is where it shows up most, since the tile for the lobby itself sits in front
of every player already standing in it.
2026-07-23 00:00:59 +09:00
flyemoji 8480387c50 fix(velocity): pin ViaRewind 4.1.3, the release that matches ViaVersion 5.11.0
The previous commit staged ViaRewind 4.1.2 alongside ViaVersion and
ViaBackwards 5.11.0. The staging mechanism was right; the pairing was not.
On startup the proxy logs

  ERROR [viaversion]: Error during loading of Protocol1_9To1_8
  java.lang.IllegalArgumentException: Invalid version: 1

and Protocol1_9To1_8 is the one protocol every 1.8 client needs. Without it
a 1.8 player clears the handshake gate we just opened and then lands on a
translation layer that never initialised.

The three versions are a set, not three independent pins. ViaRewind 4.1.3's
release notes state it adds compatibility with ViaVersion and ViaBackwards
5.11.0; 4.1.2 predates that. Measured rather than assumed: a two-arm run of
the same proxy image, same velocity.toml, same ViaVersion/ViaBackwards jars,
differing only in the ViaRewind jar, logs the error three times on 4.1.2 and
zero times on 4.1.3.

The full FL-007 forwarding matrix was re-run against 4.1.3 rather than
inferred from the two-arm result — a different jar is a different
configuration under test. It closes 20/20, including the three
force-key-authentication assertions, with the proxy log confirming
"Loaded plugin viarewind 4.1.3" and no Protocol1_9To1_8 error.

Still unproven, and deliberately not claimed here: a real 1.8.9 client
authenticated against Mojang. The harness has no Mojang account, so
online-mode = true remains the one untested axis.
2026-07-22 18:10:53 +09:00
flyemoji f532684249 feat(velocity): let 1.8.x players through the modern-forwarding proxy
Felis pins player-info-forwarding-mode = "modern", and Velocity's
HandshakeSessionHandler#handleLogin refuses anything below 1.13 outright: it reads
handshake.getProtocolVersion() and disconnects with
velocity.error.modern-forwarding-needs-new-client before the backend is ever
contacted. A player pinned to 1.8.9 never reaches the login gate, never sees the
onboarding link, and gets an error string that tells them to upgrade their client.

install_via_plugins now stages ViaVersion, ViaBackwards and ViaRewind into
/opt/felis/velocity/plugins alongside felis-velocity.jar. Nothing else moves: same
forwarding mode, same secret, no backend patched and no backend downgraded. The 1.13
floor turns out to be a property of the unassisted proxy pipeline rather than of the
forwarding protocol, so lifting it costs three jars and no source change.

This was measured, not assumed. Felis-Legacy's FL-007 probe stands a protocol-47
client in front of a stock Paper 1.21.11 backend behind a modern-forwarding proxy and
watches it join. The proof is the join itself rather than the log line: that backend
runs velocity.enabled with a shared secret, and Paper in that state rejects any login
not carrying forwarding data signed with a matching HMAC. The control cell without
Via is rejected before the backend is contacted, so Via is the only difference. A
second cell re-runs the same join with force-key-authentication = true, the way Felis
sets it, because a result that only holds under a config Felis does not run is not a
result about Felis; the 1.19+ signed chat key a protocol-47 client cannot produce is
never demanded, and it cannot be, since the pre-1.19 wire format has no player-key
field to decode.

The jars are pinned by sha256 and not by a moving tag. They sit in front of every
packet on the proxy and they are the exact bytes FL-007 measured; "latest" would
quietly make this an unmeasured configuration. The digests come from the GitHub
releases FL-006 locked, which is deliberate — Hangar's VELOCITY/download endpoint
serves different bytes for the same version numbers, so an installer that only
checked for HTTP 200 would ship artifacts nothing has tested. A version bump means a
digest bump here.

Two limits are worth writing down. Only protocol 47 was measured; the rest of Via's
documented 1.7-1.12 range is inference from that one point. And the probe runs
online-mode = false because it has no Mojang account, so Felis's online-mode = true
is untested — the untested part is the Mojang auth handshake specifically, which puts
the residual in Via's own login handling rather than in forwarding or in the
hasJoined multiplexer, whose request is protocol-independent.

Checks: a fresh install lands all three jars at the pinned digests and sizes with no
temp files left behind; a re-run downloads nothing; a tampered jar is restored to the
pinned bytes; and a deliberately wrong digest aborts without installing anything.
Existing installs pick this up by re-running bootstrap, which re-enters
install_velocity because bootstrap.done is written but never read as an early exit.
2026-07-22 17:03:01 +09:00
flyemoji 0e69ade1b3 docs(license): state AGPL-3.0 in the READMEs, matching the LICENSE file
LICENSE became AGPL-3.0-only in 037eb24, but both READMEs still told readers
the project was MIT — the one place a user actually looks before deciding what
they may do with it. The two "license notes" underneath were MIT-shaped as
well: attribution and a disclaimer, with no mention of copyleft at all.

They now say what AGPL actually requires, including section 13, which is the
clause that matters most here: Felis is a hosting platform reached over a
network, so a modified deployment owes its source to the people using it even
if no binary is ever distributed. Leaving that unsaid was the part that could
mislead someone into a violation they never intended.
2026-07-22 15:01:25 +09:00
flyemoji 37ab87f07c docs(mail): say plainly what a green SMTP self-test does not prove
The Ping comment claimed the self-test could not produce a false negative and
left the impression it therefore proved deliverability. It does not, and the
distinction is the whole trap: a relay that gates sender identity at
end-of-DATA gates it on the way OUT. Fastmail answers 250 for
[email protected] addressed to the account's own mailbox and 551 5.7.1 "Not
authorised to send from this header address" for that same From addressed to
anyone else — and only the account's exact authorized identity passes the
second one, another local-part on the same domain is refused too.

So a green Ping means connect, TLS, AUTH and message shape are good, and
nothing more; the operator still has to have authorized From as a sending
identity with their provider, and the first real OTP is what proves they did.
Addressing the self-test somewhere external would not fix it either — the only
mailbox an operator can reliably check is usually inside the same account — so
the honest move is to scope the claim rather than buy false confidence with a
bigger probe.
2026-07-22 14:40:58 +09:00
flyemoji a8c4077202 fix(api): log the panic value and stack behind the opaque 500
withRecover turned a panicking handler into a 500 envelope and threw the panic
away. The client is meant to get an opaque "internal error" — that part is
right — but nothing was written server-side, so a recovered panic was an
untraceable 500: an operator holding "internal error" has no message, no
stack and no request to grep for, and diagnosis degrades into guessing against
a live install. That is what it cost during the email-OTP report.

The panic value, a stack and the method+path are now logged first, keyed by
the same request_id writeError already stamps on unmapped errors, so the
client envelope and the server log can be joined. The test pins all three
markers plus the unchanged 500/"panic" response, because a silent recover
looks exactly like a working one from the outside.
2026-07-22 14:40:42 +09:00
flyemoji c4c964578e fix(setup): stop the forced-onboarding gate trapping players who have no email
setup_required is what the SPA polls to decide whether the onboarding wall is
still owed, and it disagreed with the middleware that actually enforces the
wall. requireOnboarded lifts on a verified email OR an enrolled passkey;
setup_required answered `u.Email == "" || !hasPasskey`. A console-tier player
joins through the bind-code door with no email at all — by design, there is no
SMTP at that point — so the email term never clears and the SPA keeps them on
the setup screen forever, even after they enroll the passkey that already
unlocked the API for them.

The predicate now lives in one place (setupRequired) and both endpoints call
it, so the next edit to the unlock condition cannot drift them apart again.
Keying it on EmailVerified rather than email presence is the deliberate part:
presence is exactly the term that trapped the no-email player, and it was also
wrong on its own terms — an unverified address is not an authentication
factor, so it was never what the lockdown could safely lift on.

Also lands the regression test for the mechanism behind the live claim-403
report: /me/servers answers 200 for a bind-onboarded player (which is why the
dashboard renders the 认领 button at all) while claim, wake and status all
answer 403 with code "setup_required" — i.e. the refusal comes from
requireOnboarded before the handler, not from isOwnerOrAdmin inside it, which
would have said "forbidden". Enrolling a passkey and changing nothing else
lifts all three, which isolates the gate as the sole cause. The backend authz
is correct; the button that leads a locked-down player into a 403 is the
frontend's to hide.
2026-07-22 14:40:28 +09:00
flyemoji 60b0ec97ac feat(invite): let a player bring a friend to the server they're on
/invite <player> posts a chat card to the invitee with a green [Accept] and a
red [Deny] button, and Accept walks them to the server the inviter is standing
on. It is a UX wrapper over `/felis go` and nothing more: the accept runs the
same doGo path on the ACCEPTING player's own verified uuid, so a stored invite
carries a server name and never an identity to act as, and the prompt needs no
unguessable token.

Why it can be this simple: an invite can only name the server its sender is
currently on, so the target is running by construction, and a running felis
server already admits any linked player through <name>.<root-domain> on the
link check alone (WaitingRouter.onServerPreConnect) — no wake, no
autostartPolicy consultation. Hence enqueueFromInvite: a READY backend is
joined directly, and only the not-ready case falls through to the policy-gated
wake path unchanged. Routing an accept through wakeAndWaitLinked would have
asked the API to wake a server that needs no waking, and autostartPolicy
defaults to ownerOnly, so the API would answer 403 and the green button would
do nothing for exactly the people you would invite.

It is not consequence-free, and the inviter is told so at send time rather
than in a comment only we read. Landing on a felis server records the player
in its allowlist (onServerConnected -> join-event -> RecordJoin); on an
autostartPolicy=allowlist server that row is what lets them come back and
START the thing later. The same row they would earn by walking in unaided —
the invite shortened the walk, it did not widen the door — but it outlives the
invite, so accessNotice says so. The wording follows the policy: only under
allowlist does it claim they will be able to start the server themselves,
because under the ownerOnly default (and the empty string the API reports for
an unset field) that row grants no waking and the claim would be a lie.

InviteBook also holds a 30s per-sender cooldown, because the one capability
/invite genuinely adds is "make a chat card appear on any online player", and
unrated that is a way to follow someone around their own chat log. It is
charged in put() rather than at the top of the command, so an invite refused
for an offline name or a player already on the server costs the sender
nothing, and the gate sits after every other validation for the same reason.
The stamp is global per sender on purpose: a per-(sender, invitee) key would
wave through one player papering the whole proxy, which is the thing being
limited.

The card lives in InviteCard as a pure function so the buttons — the whole
point of the feature — can be asserted without a live proxy, and the button
clicks are pinned to the server the card named, so a stale card cannot answer
a newer invite (checked with peek before the invite is spent, so refusing a
superseded card leaves the live one answerable).

Verified: production gradle 8.14 + JDK 21 build; jar carries the plugin classes
and no test classes; InviteBookTest (36 checks) and InviteCardTest (48 checks);
and a live Velocity 3.5.1 that loads the jar, registers /invite <player> and
answers /invite accept <server>, with a malformed subcommand as a negative
control.
2026-07-22 14:32:51 +09:00
flyemoji a0064adfdd fix(setup): stop the login gate sending players to the old console after a re-domain
Changing root_domain updated felis.toml and the panel, but the login gate kept
pointing players at the hostname it was created with. FELIS_ROOT_DOMAIN and
FELIS_PANEL_HOSTNAME are baked into the login MinecraftServer at provisioning time,
FelisLimboPlugin reads them to build the link an unauthenticated player is told to
open, and ensureSystemServers is create-if-absent — so nothing in the install ever
rewrote them. On the demo host the CR still carried
console.159.223.32.51.nip.io hours after the domain had moved to
mc.flyemoji.network: every joining player was handed a link that bypasses the
tunnel, hits the node directly and trips a certificate warning, on the one screen
someone with no account is guaranteed to see.

setup now converges these values on an existing system service instead of skipping
it. Create-if-absent stays the rule for everything else, and the comment on it is
still true — an operator's edits to a system service must survive a re-run. These
three names are the exception because they are not the operator's to own: they are a
copy of config that is wrong the moment config changes, and there is no other writer
who could notice.

The convergence is deliberately narrow. Only a name already present with a different
value is rewritten, so env the operator added by hand is untouched and the rest of
the spec — image, memory, storage — is not read at all. A derived name that is
absent from the live object is left absent rather than added back: a deliberate
removal and drift look identical from here, and re-adding it would mean fighting the
operator on every run. The outcome string reports the refresh so a setup run does not
silently rewrite the front door.

This closes one surface of a re-domain, not the whole of it. The write-once panel
certificate at deploy/bootstrap.sh keeps its old SANs, and so do the Velocity config
and the forwarding material; the warning in bootstrap.sh that says so is still
accurate. What changes is that the surface players actually walk through now catches
up when setup is re-run.

Checks: a login gate built with the old domain converges onto the new one and says
so; an env var the operator added and a hand-raised javaMemory both survive that
same run; and an install whose config already matches reports no refresh, so a
routine setup does not read like a re-domain. The middle one is the one worth having
— converging config must not turn into a licence to clobber the edits
create-if-absent exists to protect.
2026-07-21 00:47:04 +09:00
flyemoji 80a29ba653 feat(lobby): ship LuckPerms in the lobby image so the panel's permission controls work
The panel has a full permission surface — internal/api/handlers_access.go issues
`lp user <player> permission set/unset` and `lp user <player> parent add/remove`
over RCON, and projects the result back at
GET /api/v1/servers/{name}/access/luckperms/{player} — but nothing in this tree
ever installed LuckPerms. The lobby image copied felis-paper.jar into the plugin
directory and stopped there, so every grant the panel sent reached a server that
answered "Unknown command". Confirmed on the demo host: /data/plugins held only
FelisPaper/, bStats/, felis-paper.jar and spark/.

This is the other half of the RCON change. That one gave the control plane a
channel to send commands on; this one puts something at the far end that
understands them. Neither is useful alone.

The jar is resolved at build time rather than pinned in the Dockerfile, the same
way PAPER_JAR_URL already is: metadata.luckperms.net publishes the current build
for every platform, and asking upstream keeps this tree from going stale on every
LuckPerms release. Unlike Paper it is not version-matched to MC_VERSION — LuckPerms
ships one Bukkit build covering the whole supported Minecraft range, so there is no
per-version endpoint to ask. The resolver's pattern pins the /bukkit/loader/ path
segment deliberately: the metadata endpoint hands back the fabric, forge, velocity
and bukkit-legacy URLs in the same payload, and a looser match would happily return
a jar Paper cannot load, or the legacy build that targets Minecraft 1.8-1.12.

A missing LUCKPERMS_JAR_URL fails the build. That is a harsher default than the
RCON password, which only warns, and the difference is where the failure surfaces:
a lobby without RCON degrades visibly at once, whereas a lobby without LuckPerms
starts perfectly, runs perfectly, and only reveals itself when an owner tries to
grant somebody a permission. Build time is the cheap place to notice.

The entrypoint refreshes the jar from the image seed on every boot exactly as it
does for paper.jar and felis-paper.jar, so the executable artifact tracks the image
while LuckPerms' H2 database and config under plugins/LuckPerms/ stay on the PVC.
That split is the point: every grant ever issued lives in that directory, so the
refresh must never become a wipe.

Check: the three files that have to agree about LuckPerms — bootstrap.sh resolving
and passing the build-arg, the Dockerfile requiring that arg name and writing a
fixed path, the entrypoint copying from that same path — are pinned against each
other. Nothing compiles them together, and a typo in the path is invisible until a
lobby boots and `set -e` turns the failed cp into a crashloop on the hub every
authenticated player is transferred to. The test reads all three back out of the
embedded FS rather than off disk, since that is what the TUI install path ships.

Deployed installs are NOT fixed by this commit, for the same reason the RCON change
was not: the felis-lobby image has to be rebuilt and re-imported, and the pods
recreated, before the jar exists on the volume.
2026-07-21 00:39:46 +09:00
flyemoji 694e3cb800 feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.

The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.

Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.

RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.

Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.

That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.

Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.

Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.

Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
2026-07-21 00:12:37 +09:00
flyemoji b3989fa4af fix(mail): prove SMTP deliverability before saving, and stop losing the relay
A live install passed the SMTP setup screen and then failed every one-time
code with a bare `internal error`. Four separate defects had to line up for
that, and each is fixed here.

The relay was configured with `from = noreply@<domain-A>` on an account
authenticated as `<user>@<domain-B>`. Providers that validate sender identity
— Fastmail among them — answer MAIL FROM with an unconditional 250 and only
refuse at end-of-DATA. Ping stopped at NOOP, so it never saw the refusal: the
wizard reported success, wrote the config, rolled felis-api, and every OTP
afterwards died at w.Close().

Ping now runs the same transaction a real code takes — connect, (STARTTLS,)
AUTH, MAIL FROM, RCPT TO, DATA — delivering one self-test message to the From
address, and SendOTP and Ping share deliver() so the check cannot drift from
the thing it checks. The self-test recipient cannot cause a false negative:
an authenticated submission relay accepts RCPT for any destination by
definition, while the sender identity it does validate is exactly what we
want tested. The setup screen now says a message will be sent, names the
address it went to, and warns that From must be an address the account is
allowed to send as.

A relay refusal also answered 500 `internal`, which reads as a broken panel
and sends the operator hunting through handler code instead of their [smtp]
block. It is now 502 `mail_undeliverable`, mapped inside deliverOTP so all
four doors that mail a code (onboarding, email login, op-login, migrate
step-up) answer alike. The relay's own text stays out of the response — it
can name the SMTP account, and these routes are reachable by any signed-in
player — and goes to the log instead.

writeError logged nothing when it collapsed an unmapped error to 500, so an
operator holding an `internal error` had nothing to grep for and diagnosis
degraded into guessing against a live install. It now logs the method, path,
wrapped chain and the same request_id the caller is shown.

Finally, write_felis_toml regenerated the config wholesale and never emitted
[smtp], so re-running the installer — the documented way to update felis-api —
silently erased a working relay and reverted OTP delivery to the no-Mailer
path, logging codes instead of sending them. It now carries the block forward,
cached on first read because the host toml is clobbered before the pod toml is
written. Same defect family as the root_domain loss fixed in ecbeb20: a
generated file holding a hand-set value with no carry-forward.

Tests cover the case a MAIL FROM probe cannot see: a fake relay that answers
250 to MAIL FROM and 550 at end-of-DATA must fail both Ping and SendOTP, and
the 502 must carry a distinct machine code without leaking the relay's text.
2026-07-21 00:11:40 +09:00
flyemoji 32be3e17b7 docs(readme): give an install command that works against a private repo
The documented one-liner fetches bootstrap.sh from raw.githubusercontent.com
unauthenticated, which 404s for as long as this repository stays private -- so
the single command the README exists to provide did not work for anyone.

The authenticated form goes through the contents API with the raw media type,
matching what github_api already does, and hands the token to curl over stdin
via --config rather than -H. argv is world-readable through /proc, and a token
on the command line would leak to any local user during the install; bootstrap
avoids that in its own fetches for the same reason and the README should not
teach the opposite.

sudo -E, because the installer needs that same token to resolve and download the
release. Without it sudo drops the variable and the run fails later, at the
release lookup, for a reason the operator has no way to connect to this command.

The public one-liner stays first: it is what this becomes once the repository is
public, and the note is scoped to the current state.

Also records that re-running the installer is how felis-api moves to a newer
release, that it now keeps the installed root domain, and that it does not keep
the channel.
2026-07-20 21:20:09 +09:00
flyemoji ecbeb20761 fix(bootstrap): reuse the installed root domain instead of re-deriving it
detect_node_ip recomputed FELIS_ROOT_DOMAIN from scratch on every run and fell
back to <node-ip>.nip.io. Nothing read the domain back out of the felis.toml an
earlier run wrote, so it survived only as long as the operator kept passing the
same environment.

That made re-running the installer destructive on any install with a real
domain, and re-running it is not optional: it is the only way to move felis-api
to a newer release, which is what `felis update` points operators at. A bare
re-run rewrote root_domain, panel_hostname and admin_hostname to nip.io names
while ensure_panel_tls_cert returned early on the certificate it had already
written, leaving the console serving a cert for hostnames it no longer answered
to -- with no re-domain flow to recover through.

Precedence is now explicit FELIS_ROOT_DOMAIN, then the domain the last run
persisted, then the nip.io default. First installs are unaffected. Deliberate
re-domains still work, because there is no other route to one, but they now warn
that the write-once certificate is not reissued and that the proxy and login
config carry the old name too.

Secrets were never exposed to this: load_or_make_secrets has always sourced
secrets.env before generating anything. The domain was the one piece of install
identity with no read-back.

The channel is deliberately left alone. FELIS_VERSION_BOOTSTRAP is not persisted
either, but defaulting a re-run to the release channel installs a working build
rather than breaking one, so cmd/felis/update.go states that instead. Its
warning about the domain went with the bug and would now be false.

Verified against the shipped function text: the ladder holds for a fresh host,
a re-run with and without the variable set, a re-domain, and a felis.toml whose
root_domain is missing or empty. Reverting the one line reproduces the nip.io
overwrite.
2026-07-20 21:20:08 +09:00
flyemoji f7815629bf fix(update): say that setup cannot move felis-api to a newer release
`felis update` offers `sudo felis setup` for every planner-backed target and
closed with a trailer calling setup idempotent. That is true for velocity --
install_velocity re-resolves the newest build of the pinned minor on each run --
and misleading for felis-api, which both --panel and --plugins resolve to.

setup hands bootstrap the binary it is itself running and takes the
bootstrap_from_tui arm, which skips the release lookup. The run rebuilds the
image and rolls the deployment off that SAME binary: it reports success and
leaves the version exactly where it was. An operator following this guidance to
apply a felis-api update would watch it appear to work and then see the same
version reported again.

Only the bootstrap installer moves felis-api, and naming it is where this gets
dangerous, so the warning ships with it. The installer is not an updater. Every
run re-derives FELIS_ROOT_DOMAIN through detect_node_ip and defaults it to
<node-ip>.nip.io; nothing reads the domain back out of the felis.toml an earlier
run wrote. A bare re-run -- which is exactly what README documents, with no
environment at all -- rewrites root-domain, panel-hostname and admin-hostname to
nip.io names, while ensure_panel_tls_cert returns early on the certificate it
already wrote and keeps serving the old hostnames. The console then fails to
match its own certificate, and there is no re-domain flow to recover with.
Persisted secrets are not at risk: load_or_make_secrets sources secrets.env
before it generates anything.

felis update stays report-only, so no command changed; only the claim about what
the offered one accomplishes, and the conditions under which the alternative is
safe to run.

Tested four ways, because the scoping and the warning are both the point:
--panel carries the caveat AND names FELIS_ROOT_DOMAIN, --velocity keeps the
ordinary trailer without either, and --mc, which offers no command at all, gets
neither.
2026-07-20 21:20:04 +09:00
flyemoji 8b25114fb4 fix(update): say that setup cannot move felis-api to a newer release
`felis update` offers `sudo felis setup` for every planner-backed target and
closed with a trailer calling setup idempotent. That is true for velocity --
install_velocity re-resolves the newest build of the pinned minor on each run --
and misleading for felis-api, which both --panel and --plugins resolve to.

setup hands bootstrap the binary it is itself running and takes the
bootstrap_from_tui arm, which skips the release lookup. The run rebuilds the
image and rolls the deployment off that SAME binary: it reports success and
leaves the version exactly where it was. An operator following this guidance to
apply a felis-api update would watch it appear to work and then see the same
version reported again.

The report now says so, scoped to runs that actually offered a felis-api
target, and points at the bootstrap installer -- the path that resolves and
downloads a release. felis update stays report-only, so no command changed;
only the claim about what the offered one accomplishes.

Tested three ways, because the scoping is the whole point: --panel carries the
caveat, --velocity keeps the ordinary trailer without it, and --mc, which
offers no command at all, gets neither.
2026-07-20 19:53:55 +09:00
flyemoji 8675cda001 fix(setup): refuse --dev rather than silently installing the release channel
`felis setup --dev` promised "install the dev channel (main HEAD)" and
installed release: the flag only exported FELIS_CHANNEL, a variable nothing in
the tree reads. deploy/bootstrap.sh reads FELIS_VERSION_BOOTSTRAP.

Renaming the variable would have been a worse bug than the dead one, because it
would look wired. setup runs bootstrap with FELIS_BOOTSTRAP_FROM_TUI=1, and on
that arm every reader of FELIS_VERSION_BOOTSTRAP is unreachable: the channel
case and its validation live in resolve_install_ref, which the TUI path skips
outright, and use_release_binary is only consulted by the elif that
`if bootstrap_from_tui` already short-circuited. setup re-images the host from
the felis binary it is itself running; there is no channel to pick.

So the flag refuses, exits 2 and names FELIS_VERSION_BOOTSTRAP=dev on the
installer, which is the mechanism that does work. Refusing beats defaulting:
the operator asked for dev, and release is the one answer they did not want.
The refusal precedes the root check, or an unprivileged operator gets told
about sudo instead of about the channel.

channelName had no other caller and goes with it. Nothing else referenced
--dev -- no doc, no script, no test -- so this removes a promise the tree only
ever made to itself.
2026-07-20 19:53:32 +09:00
flyemoji e5ea51c0db fix(bootstrap): wrap the downloaded binary in the same base CI ships
build_image_from_binary built on distroless/base-debian12 while the repo
Dockerfile's final stage uses distroless/static-debian12, so the image an
install runs did not match the image CI publishes.

Every binary that can reach HOST_BIN traces back to the Dockerfile's
CGO_ENABLED=0 build -- the downloaded CI asset, the binary the TUI is already
running, and the one build_image_from_source docker-cp's out of the image it
just built. None link glibc, so base-debian12 bought nothing and only widened
the runtime surface.

This mattered little while build_image_from_binary was the rare fallback.
659c8e5 made the release channel download a binary and wrap it here, which
makes this the image most installs actually run.
2026-07-20 19:53:04 +09:00
flyemoji 646d514a65 test(updates): pin the ordering of a dev build's own version stamp
659c8e5 chose "+g<sha>" over "-g<sha>" for the dev channel's stamp on the
grounds that "+" is SemVer build metadata, ignored for precedence, so a build
some commits past v1.2.3 still reads as v1.2.3 rather than as something older.
That claim is load-bearing and was asserted only in prose: if metadata ever
counted for ordering, every dev install would report an upgrade available onto
a release it already contains.

Parse does handle it -- build metadata is stripped before the prerelease tail,
so the "-earlyAccess+g<sha>" order a real build stamps splits correctly too --
but nothing tested it. The existing coverage compares "+k3s1" against "+k3s2",
which is metadata on BOTH sides; the case that matters here is metadata on one
side only, against the bare tag.

Four pairs in the compare table, plus "v0.0.0+g<sha>" in the parse table for a
build pinned to a ref with no tag behind it. Verified by mutation: disabling the
"+" split in Parse fails these cases specifically.

They sit at the end of the table rather than beside the other build-metadata
pairs. The entries are wide and carry no trailing comment, and inserting them
mid-table splits the contiguous comment block, which makes gofmt rewrite the
alignment of seven untouched lines.
2026-07-20 19:05:19 +09:00
flyemoji d9ef5e4ffd chore(deps): tidy the module files
Plain `go mod tidy` output, no hand edits, so the files match what the tool
produces from the current import graph:

  - github.com/minio/minio-go/v7 becomes direct. internal/submit/s3store.go has
    imported it since 598f3d3, which added the S3 upload backend without tidying.
  - golang.org/x/crypto becomes indirect. Nothing imports it directly any more —
    the passwordless migration removed the last caller.
  - github.com/clipperhouse/stringish drops out of the graph entirely.
  - go.sum loses 26 stale lines.

No version changes and no behaviour change; vet and the full test suite pass
unchanged.
2026-07-20 18:53:25 +09:00
flyemoji c4300cb005 feat(updater): authenticate GitHub polling and track the real release repo
felis-api's coord was the placeholder "felis/felis", which resolves against
nothing on real GitHub. It is now MliroLirrorsIngenuity/Felis — the same slug
deploy/bootstrap.sh clones from — so update reporting for the control plane
itself is live rather than parked.

That repo is private today, so the github source gained an optional token, read
from FELIS_GITHUB_TOKEN: the variable bootstrap already needs, so an operator
sets one value once. It comes from the environment and is never compiled in. A
constant would be committed to the very repository it protects, ship inside
every felis binary where strings(1) recovers it, reach every node the image is
imported onto, and need a rebuild and a redeploy to rotate.

Empty stays the correct posture for the other tracked components — k3s and
cloudflared are public — and an empty token sends no Authorization header at
all rather than an empty one.

GitHub answers 404, not 401 or 403, for a private repo the caller cannot see,
so "no token" and "no stable release published yet" arrive as the same status.
On an unauthenticated 404 the error now names both causes and the variable that
fixes the actionable one. With a token already set that hint would be wrong, so
it is suppressed.

Tests pin both halves: the Bearer header is sent only when the token is set,
and the diagnostic names the variable only when it is not.

doc.go's CAVEATS bullet still described the coord as a placeholder and the
component as "dark at runtime". Both were true only until this change; it now
records the real condition, which is that the component resolves like the others
but needs a credential while the repo is private.
2026-07-20 18:53:12 +09:00
flyemoji 659c8e5e9f feat(bootstrap): install the published release build instead of compiling on the host
deploy/bootstrap.sh now resolves the newest published GitHub release, downloads
the binary CI built for that tag, and builds a thin image around it. Compiling
on the target host becomes the fallback and the opt-in, not the default.

The panel is not a separate artifact. The Dockerfile copies panel/dist into
internal/panel/static before the go build, so the control plane — panel and
backend — ships as ONE file. The release channel therefore downloads exactly
one asset, felis-linux-<arch>, and needs no registry, no Go toolchain and no
checkout on the host.

The Minecraft game stack (limbo, lobby, the Velocity plugin) is still always
built locally. game_stack_source now keys on HAVE_PREBUILT_BINARY — the same
flag build_image uses — so on any prebuilt path it unpacks the tar embedded in
that binary instead of trusting a checkout an earlier install left behind.
Trusting the checkout would build the plugin from an old commit against a
freshly downloaded control plane: a silent version skew across the plugin/API
boundary.

Channels:

    (default)                      newest published release, downloaded
    FELIS_VERSION_BOOTSTRAP=dev    clone main and compile
    FELIS_REF=<ref>                pins the tree, forces the source path

The download is best-effort. A tag whose assets are not uploaded yet, an
architecture with no published asset, or an asset that fails validation each
warn and fall back to compiling THE SAME TAG from source — never a different
commit.

The ref is resolved right after install_base, the first point curl exists and
well before docker and k3s, so a missing FELIS_GITHUB_TOKEN or an unpublished
release costs the operator seconds instead of a k3s install they then have to
unwind. It is skipped on exactly the paths that never consume the result: the
TUI, which rebuilds the binary it is already running, and FELIS_SKIP_FETCH,
which builds whatever is staged. Resolving anyway would set FELIS_VERSION to
the newest tag and stamp a staged tree as that release.

The asset is staged next to HOST_BIN rather than in TMPDIR. Validation EXECUTES
it, and /tmp is noexec on CIS-hardened images, where the exec dies 126, the
check reads it as a bad asset, and every such host silently falls back to the
full on-host compile this path exists to avoid. It also keeps a private-repo
artifact out of a world-readable 1777 directory.

git_auth, which supplies the token to git for a private-repo clone, passes an EMPTY
credential.helper before the inline one. credential.helper is multi-valued: a bare
`-c credential.helper=...` APPENDS to whatever the host has configured rather than
replacing it, and an empty value is git's documented list reset. Without it, on a host
with a persistent helper (Git for Windows ships `manager` at SYSTEM scope) two things
go wrong. Git runs `credential approve` automatically after a successful clone and
feeds every helper in the list, so a `store` helper writes the PAT to
~/.git-credentials in cleartext — the token outlives the install, in a file bootstrap
never created and never cleans up. And because the inline helper is LAST, a
pre-existing helper answers `fill` first, so a stale cached credential can win and the
clone authenticates as the wrong account — surfacing as exactly the 404-on-private-repo
the surrounding code works hard to explain. Reproduced both against a real clone, and
confirmed the reset closes both.

internal/panel parses the new stamp. The dev channel now emits "<tag>+g<sha>", which
matched neither describeSuffix ("-N-g<sha>") nor releaseTag, so a dev build fell through
to the default case and the version badge rendered the entire stamp as the release with
no commit. A devSuffix case handles it; the git-describe case stays for hand-rolled
`-ldflags "-X main.version=$(git describe)"` builds. Table test covers both forms plus
the release, dirty and unstamped cases.

CRD application no longer branches on the install path: it is always
`felis bootstrap-assets crd`. That output is byte-identical to deploy/crd/ —
bootstrap_asset.go embeds that very file — and needs no checkout, so one source
replaces a branch whose two arms had to be kept in agreement by hand.

Dockerfile gains a FELIS_VERSION build arg wired into -X main.version, declared
after `go mod download` so a version bump does not invalidate that layer. Both
build stages are pinned to $BUILDPLATFORM so a multi-platform buildx run never
emulates them: the panel's output is architecture-independent and the Go stage
cross-compiles via TARGETARCH. The final stage stays on the target platform and
is COPY-only, which BuildKit performs without QEMU.

.github/workflows/release.yml publishes on a vX.Y.Z tag: vet, tests, then one
buildx run producing both architectures through the repo Dockerfile. Not a bare
`go build` — internal/panel/static holds a tracked placeholder index.html so the
//go:embed compiles without node, which means a direct build succeeds and
quietly ships a release whose panel is that placeholder.

The stamp is asserted end to end, because it fails silently: an unstamped binary
reports "dev", which the updater refuses to compare, disabling update reporting
for every install built from that release. The arm64 artifact is checked by ELF
machine type rather than by running it — runners have binfmt registered, so
executing an amd64 binary misnamed arm64 would succeed.

Prerelease tags are flagged explicitly. The trigger glob is v*, gh does not read
semver out of a tag name, and an RC published as a full release becomes
/releases/latest — the single endpoint the default channel installs from and
`felis update` polls.

No SHA256SUMS. A checksum fetched over the same TLS session, with the same
credential, from the same host as the binary adds no trust root; signing is the
real answer and is a separate decision.

Not verified: the download -> validate -> image -> k3s path has never run on a
host against a real published release, because no tag exists yet. The shell
logic around it is verified out of tree; the network and exec behaviour is not.
2026-07-20 18:53:01 +09:00
flyemoji d146f1ccd7 fix(bootstrap): keep the install alive on a host with only felis-api
restart_existing_control_plane ended in an and-list per deployment:

    [ "$had_api" = "1" ] && kube ... rollout restart deployment/felis-api
    [ "$had_operator" = "1" ] && kube ... rollout restart deployment/felis-operator

As the LAST command of a function, an and-list whose test is false returns 1,
and that becomes the function's exit status. The call site is bare, so under
`set -Eeuo pipefail` the installer dies there — after the bundle has been
applied and before the rollout wait, leaving a half-finished upgrade and no
message naming the cause.

It fires on any host carrying one control-plane deployment but not the other:
felis-api present without felis-operator restarts the api, then exits 1 on the
second test. Both present, or neither, happened to work, which is why it
survived.

Rewritten as explicit `if` statements, which return 0 when the test is false.
Verified out of tree against all four had_api/had_operator combinations.
2026-07-20 18:30:19 +09:00
flyemoji fe4c92c1c5 feat(files): add the server file editor
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.

felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.

What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.

Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.

The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.

Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:

  * A write accepts arbitrary bytes at any path, so an owner can place a
    loadable plugin jar. This is deliberate — it is what a hosting panel's
    file manager does, scoped to a server the caller already owns and
    already drives through /command — but it is the one owner-tier route
    that lands executable code in a backend pod, since images are
    admin-only and modpack submissions need an admin verdict.
  * config/paper-global.yml is refused on read. felis-lobby's entrypoint
    writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
    identical on every backend, so reading it from a server you own would
    hand you the handshake key for everyone else's. It is the only path in
    the mount that is not the caller's own data, and therefore the only
    denial. The comparison is on the cleaned path, or ./config/... would
    walk straight through it.

Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.

The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
2026-07-20 14:33:29 +09:00