chore: work the tracker items that need no cluster #19

Merged
FLYEMOJ1 merged 10 commits from chore/issue-sweep into main 2026-07-29 01:52:15 +09:00
10 Commits
Author SHA1 Message Date
flyemoji 503240db7b ci: stop running the whole suite twice on every pull-request push
The header of this file argues that release.yml must not be repeated here,
because a private repository is billed twice for one answer. The push trigger
it shipped with then did exactly that: `branches: ['**']` plus `pull_request`
means a branch with an open PR runs everything once for refs/heads/<branch>
and once for refs/pull/N/merge. The concurrency group is keyed on github.ref,
which differs between the two, so neither cancels the other. Visible on this
branch's own checks: go 3m14s and go 3m3s, shell 7s and 7s, panel 25s and 24s.

Limiting the push trigger to main keeps both gates that matter -- a PR is
still checked before merge, main is still checked after -- and drops only the
duplicate. The one case that loses coverage is a branch pushed with no PR
open, where nothing has asked for the answer yet.

Verified by parsing the result with the repository's own sigs.k8s.io/yaml:
triggers are {"pull_request":null,"push":{"branches":["main"]}} and the three
jobs go/panel/shell are intact. Worth recording for the next person who
parses a workflow: YAML 1.1 reads the bare key `on` as the boolean true, so
it arrives as the string "true" after the YAML-to-JSON conversion, and a
struct tag of `json:"on"` silently matches nothing. GitHub's own parser does
not have this problem; a local check of the triggers does.

Refs #7
2026-07-28 18:20:40 +09:00
flyemoji 584d31fc49 feat(bootstrap): refuse to install an unverified Velocity fork jar
FELIS_VELOCITY_FORK_JAR replaces the proxy every player connects through, and the
only thing checked about it was that the path pointed at a readable file. A
truncated copy, a stale build left at the same path, or the two-patch jar where
the three-patch one was meant all installed silently.

It now requires FELIS_VELOCITY_FORK_JAR_SHA256 and refuses on a mismatch, hashing
stdin rather than the path for the reason install_via_plugins already documents:
sha256sum escapes its output line for a filename carrying a backslash or newline,
and the leading "\" that adds fails every comparison. The absent-digest refusal
prints the jar's actual hash, so the first run after a deliberate rebuild is one
copy-paste rather than an investigation.

The comparison ignores case and internal spaces. The fork is built on a developer
machine, which is usually Windows, and nothing there prints a digest the way
sha256sum does: Get-FileHash returns uppercase and certutil has shipped the bytes
space-separated. Comparing raw would refuse two of the three spellings of the
correct answer and word the refusal as tampering.

No digest is hardcoded, which is the half of the request this does not deliver.
The fork is built from Felis-Legacy and has never been reproduced on a second
machine, so a constant here would pin one machine's output rather than the fork.
The comment that previously asserted the build "is not byte-reproducible" is gone
too -- it was stated more confidently than the evidence supports. The fork jars on
disk carry Gradle's constant 1980-02-01 entry timestamps, so the usual reason a
jar differs between builds is already absent; that is not proof it reproduces, and
neither claim should sit in the script unmeasured.

This is deliberately not a supply-chain signature and the comment says so: an
operator who can write the jar can write the digest. What it buys is that a path
stops being an identity, and that every later re-run re-checks the same build.

Scope: the fork jar only. The else branch still curls stock Velocity from PaperMC
with no verification at all, and that is the branch a default install takes. The
digest is already in hand there -- Fill v3 returns checksums.sha256 and its
download URL is content-addressed on that same value -- and papermc_latest_jar
discards it. Left alone rather than widened into this change.

deploy/bootstrap_test.sh covers the gate's two refusals, its happy path, and the
two Windows digest spellings. Each case extracts the block under test out of
bootstrap.sh with awk and runs it with die/log stubbed, rather than transcribing
it -- a transcribed copy passes forever after someone edits the original. The
extraction is length-bounded: awk runs an unmatched end pattern to EOF, which
would quietly feed the rest of bootstrap.sh to the shell under test. bootstrap.sh
itself cannot run here; it wants root, a package manager and k3s.

A `shell` CI job runs that plus a syntax check over every tracked script. The
syntax step dispatches on each file's shebang instead of running `sh -n` across
the board. The blanket form looks fine and is a false green: on a developer
machine `sh` is usually bash and accepts everything, while the runner's `sh` is
dash. Verified against the real thing rather than an approximation -- inside
ubuntu:24.04, where /bin/sh is /usr/bin/dash, the dispatching loop passes all six
scripts and the blanket loop dies at bootstrap.sh:191 on the first of its 14
arrays.

Refs: Felis-Legacy #19
2026-07-28 18:01:39 +09:00
flyemoji 4f5014d033 feat(bootstrap): make the legacy-forwarding backend list overridable
Which backends receive their forwarded identity through the handshake address --
rather than proxy-wide modern forwarding -- was the literal string "legacy18",
assigned inside write_velocity_service. Standing up a second protocol-47 backend
therefore meant editing this script, on every host, and remembering to.

It is now FELIS_LEGACY_FORWARDING_SERVERS, defaulting to legacy18, declared beside
FELIS_NANO_LISTEN and documented in the Tunables block like every other knob. The
default is unchanged, so an existing install re-runs to the same systemd unit it
already has.

This is deliberately only half of what the list should eventually do. It is a JVM
system property, read once when Velocity starts, so it is fixed for the life of
the proxy process and a change still needs a restart -- an environment variable is
as far as a startup property can be pushed. Having the list follow the
MinecraftServer CRs is a larger change than it looks: the forwarding decision is
made by the fork's patch to Velocity core, not by the Felis plugin, so core would
have to read state the plugin owns and refreshes. The plugin already maintains a
dynamic backend registry, which is where that state would come from, but the
bridge from core to it does not exist. The comment at the assignment now says so
instead of leaving "the upgrade path is to have the operator render this list from
the MinecraftServer CRs" as though it were a small step.

The -D is now double-quoted in ExecStart. The fork trims each element -- it parses
the property as `split(",")` into a Set, mapped through String::trim with empties
filtered -- so it accepts "legacy18, legacy112", but systemd splits ExecStart on
whitespace before java sees it. Unquoted, that spelling handed java a stray
"legacy112" argument and the unit failed to start; documenting the knob as
comma-separated without quoting it would have shipped that as a footgun.

`bash -n` passes; the default resolves to legacy18, an override to the value given,
and a value containing a space renders inside a single quoted ExecStart item.
2026-07-28 17:27:02 +09:00
flyemoji 4e5a809dad docs(nano): drop the promise of a felis setup --nano that should not exist
nano.go's package comment told the reader that `felis setup --nano` installs the
multiplexer as a service. No such flag has ever existed -- `felis setup` defines
only -config and -dev -- so anyone following the comment gets "flag provided but
not defined: -nano" and exit 2.

Adding the flag was the obvious reading, and it is the wrong one. setup does not
install anything selectively: it re-images the host by running the full bootstrap
TUI, and no install-mode parameter is threaded anywhere -- nothing in Go reads or
writes FELIS_INSTALL_MODE, which is a shell variable bootstrap.sh consumes on its
own. So `--nano` could only mean one of two things. Re-run the installer in nano
mode, which is what pointing at the installer already does. Or convert a
provisioned full host into a nano one, which means tearing down k3s, Postgres and
the proxy -- an uninstall, not a flag.

setup's --dev already settled this shape once. It looked like a channel selector,
silently installed release, and the fix was to refuse it and name the installer
rather than pretend to choose. The same answer applies here, so the comment now
names the real entry point -- the installer's `[2] Felis-nano` prompt, or
FELIS_INSTALL_MODE=nano -- and records why there is no flag, so the next reader
does not reopen it.

No refusing --nano flag is added: --dev exists because it used to be a silent
no-op that people passed, and nothing has ever accepted --nano, so flag's own
"not defined" error is already the correct and clearer failure.
2026-07-28 17:27:02 +09:00
flyemoji afdbfac7a8 docs: index the deferred integration seams and correct two stale markers
INTEGRATION-ONLY and KNOWN-LIMITATION are grep-able, but the grep answers the
wrong question. Thirty-four Go sites share the two markers and they carry four
different meanings: "declared, nothing implements it" reads exactly like
"implemented, only its I/O is unreachable from here", and neither reads
differently from a limitation that was accepted on purpose and is not coming
back. docs/deferred-seams.md sorts them, following the bucketed shape
internal/updater/doc.go already uses for its own package rather than starting a
second convention.

Sorting them turned up two markers that had outlived the condition they describe.

config.go called the modpack upload transport a deferred integration after both
backends had shipped -- LocalContextStore and S3ContextStore, selected in
cmd/felis by the shape of user_uploads_context, with the uploads PVC mounted and
the felis-uploads-s3 Secret rendered. What is still deferred is the far end:
Kaniko reading that context from inside the build Pod.

updater/doc.go listed the `felis update` CLI and the off-cluster Velocity jar read
under REMAINING INTEGRATION. Both exist -- cmd/felis/update.go, and
gatherer_host.go, which answers Velocity from the installed jar's manifest and
felis-api from the running binary's build stamp. The two nil seams that bullet
also names are real, but they belong to the in-cluster gatherer only, so the
bullet now says which caller has what and which is still empty.

The index also records the collision that makes a naive grep misleading:
docs/troubleshooting.md uses [INTEGRATION-ONLY] for something else, defined in its
own opening at :19 -- the symptom is produced by the kubelet, kaniko or a live
handshake, so it cannot be reproduced from the repository. Those twelve marks say
where a failure comes from, not that something is unbuilt, and are excluded.

Both code changes are comments. Every file:line the index cites was checked
against the line it points at.
2026-07-28 17:27:01 +09:00
flyemoji 23792d6251 fix(crd): remove spec.storage.retainOnDelete rather than leave it inert
The field validated, shipped in the CRD, and reached no controller. The world
PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate,
StatefulSet deletion does not cascade to template PVCs, and no finalizer exists
anywhere in the operator. So setting it true described what already happened,
and setting it false did nothing at all. False is the worse half: it reads as a
request to delete a world, and was silently ignored.

This departs from spec v4.1 §5, which asks for
"删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half
was ever built. Restoring that line means adding a finalizer whose other listed
duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs,
so the only work it would newly do is delete worlds, on a path that does not
pass the reaper's verified-backup check. The reaper is the one thing in the
system allowed to destroy a world and it earns that by proving a backup first.
A second door without that check is not an improvement.

The spec is a frozen versioned document, so it is left alone and the departure
is recorded in troubleshooting.md §13, beside the behaviour it explains. §12
loses its inert row and its opening sentence, which existed to introduce this
one field: every field in that table is now read by a controller.

Deployed installs need nothing. A CR still carrying retainOnDelete keeps
working, because a v1 CRD prunes unknown keys on the next write and the
behaviour the field claimed to control was never conditional.

go build, go vet and go test ./... pass on Linux with zero failures; the CRD
still parses and storage keeps size and storageClassName.
2026-07-28 17:27:01 +09:00
flyemoji 417769407f ci: run the checks on push and pull request
release.yml was the only workflow and it fires on v* tags, so `go vet` and
`go test` first met a change once that change was already on the release path,
where the only remedy is another tag. The panel suite ran nowhere at all: a
release goes through the Dockerfile and the Dockerfile runs `npm run build`,
never `npm test`. 111 assertions across 8 files existed and nothing outside a
developer's checkout ever executed them.

Both jobs are green as of this commit, checked before writing it rather than
after: go vet and go test ./... (24 packages, 0 failures, on Linux), npm test
(8 files, 111 tests) and npm run typecheck. A gate that lands red is a gate
everyone learns to ignore.

The panel's Node version is read out of the Dockerfile instead of repeated
here. `FROM node:<major>` is the only place the tree declares it -- no .nvmrc,
no engines field -- so a copy in this file would keep testing 22 the first time
the image moved. That is the class of drift this workflow exists to catch, not
to introduce. The step fails loudly if the FROM line stops matching.

Tags are excluded from the push trigger. A v* push already runs release.yml,
which repeats the Go job, and this is a private repository billed for both.
2026-07-28 17:27:00 +09:00
flyemoji 82a1275fcf docs(plugins): say which plugin jars an install actually produces
The module table listed felis-fabric, felis-forge and felis-neoforge next to the
two jars a finished install really has, with nothing distinguishing them. Neither
deploy/bootstrap.sh nor the embed set in bootstrap_asset.go builds a loader mod,
so someone reading the table expected three jars that are not there after setup
and had no way to tell from this file. The mods do build -- the wrapper commands
under Building work -- they are just never installed for you, which is what the
new column says.

Two further disagreements with the code, in the same table:

  limbo/ was missing entirely. It is embedded, built by bootstrap.sh and running
  on the login gate, so the one module the table omitted was a shipped one. It is
  also the only module that reaches the account-link endpoint without a command:
  it mints the code on join for anyone unlinked and holds them until they redeem
  it, so the opening claim that every module except the lobby ships /link named
  the wrong exception.

  velocity was listed as felis-velocity-0.2.0.jar; plugins/velocity/build.gradle:6
  says 0.1.0, as does every other module. Nothing breaks on this because
  bootstrap.sh globs felis-velocity-*.jar and installs it under a fixed name, but
  the version in the table was not a version anything produces.

The Gradle table and the JDK note now carry limbo's Java-21 toolchain, which it
needs for the same reason paper does and for a different cause: LOOHP/Limbo
releases are class-file major 65, so the compiler JDK must be able to read them.
It still emits release 17 bytecode.
2026-07-28 14:33:37 +09:00
flyemoji 3af5cc360c docs(troubleshooting): correct four fields the runbook documents as inert
Sections 11 and 12 told the operator that idle auto-stop, both startup
budgets, and the player tally are read by nobody. All four are read, and
§1 repeated the same claim in its strongest form: "the operator has no
start timeout ... loops forever".

  spec.idle.autoStopEnabled       reconciler.go:175
  spec.idle.emptySecondsBeforeStop  reconciler.go:175
  spec.startup.timeoutSeconds     reconciler.go:479, called at :126
  spec.startup.readinessTimeoutSeconds  reconciler.go:490, called at :157
  status.players.online           reconciler.go:413 (markRunningReady)

The repository already contained the disproof:
TestReconcileRunning_StartupTimeoutConvertsToFailed and
TestReconcileRunning_ReadinessTimeoutConvertsToFailed both assert the
escalation §1 said does not exist. The test §1 cited,
TestReconcileRunning_RconProbeFailureStaysStarting, only asserts that a
single failed probe does not flap the phase; that was read as "forever".

§11 was the costly one, because it misdiagnosed a configuration problem
as a missing feature. Idle auto-stop and the player tally both hang off
spec.rcon.enabled -- the tally is a by-product of the RCON readiness
probe (prober.go:63 runs `list`), and reconciler.go:175 carries the RCON
condition explicitly so a never-sampled zero cannot stop a server full of
people. Following the old text, an operator whose RCON was never enabled
would conclude the feature was unwritten and stop. The section now opens
with the jsonpath that reads spec.rcon.enabled.

Also separated the prober's fixed 5s dial timeout (prober.go:45) from
spec.startup.readinessTimeoutSeconds, which §1c conflated: the former
bounds one probe, the latter is a deadline for the whole start measured
from status.startRequestedAt.

spec.storage.retainOnDelete is the one field still genuinely inert, so
the [INERT] legend and §12 stay -- §12 now records the condition each
read field depends on instead of claiming none of them are read.
2026-07-28 13:49:07 +09:00
flyemoji 71e1664c36 build: pin line endings to LF so the embedded bootstrap.sh ships without CRs
The repository had no .gitattributes. With core.autocrlf=true a Windows checkout
handed deploy/bootstrap.sh 2374 CRs, and bootstrap_asset.go embeds that file from
the working tree verbatim, so a dev-built felis piped a CRLF script into `bash -s`
on the target host. CI builds on Linux, which is why released binaries were clean
and only local builds carried it.

eol=lf is global rather than scoped to *.sh because go:embed reaches further than
the installer: deploy/*/Dockerfile, deploy/*/entrypoint.sh, plugins/*/src, the
migrations and internal/panel/static are all compiled in and read on Linux. *.bat
is the one exception, for the gradle wrappers' Windows launchers.

Renormalizing the index touched exactly one tracked file, cmd/felis/version.go,
and only its line endings: `git diff --cached --ignore-cr-at-eol` reports nothing
outside .gitattributes itself.

TestBootstrapPinsViaBlockConnectionsOff used to strip \r\n before asserting, with a
comment stating that the repository pinned no eol attribute. That is no longer true,
and the stripping hid the regression this commit prevents. It now asserts the absence
of CRs, so losing the attribute reports itself as line endings rather than as a
missing serverside-blockconnections pin.

Closes #5
2026-07-28 12:58:46 +09:00