Commit Graph
8 Commits
Author SHA1 Message Date
Lemon-miaow 9931f64fff feat(release): CI 一次构建各架构二进制、镜像包与 velocity 插件 2026-09-26 14:22:46 +08:00
Lemon-miaow d9453a6488 build(plugins): 插件构建统一钉到 gradle 9.8 镜像与带 sha256 的 wrapper,paper-api/Limbo 对齐锁文件并启用依赖校验 2026-09-25 23:45:33 +08:00
Lemon-miaow 82545548e7 feat(bootstrap): game stack 按 lock 文件固定构建并校验 sha256,基础镜像按 digest 固定,JRE 固定补丁版本 2026-09-25 00:24:02 +08:00
flyemoji d9246ddae6 feat(bootstrap): verify Paper and Velocity jars against Fill's digest 2026-08-04 17:36:47 +09:00
flyemoji 3f2b28d0ec fix(bootstrap): ship deploy/paper in the embedded game-stack tar 2026-08-04 16:34:09 +09:00
flyemoji 71e1664c36 build: pin line endings to LF so the embedded bootstrap.sh ships without CRs
The repository had no .gitattributes. With core.autocrlf=true a Windows checkout
handed deploy/bootstrap.sh 2374 CRs, and bootstrap_asset.go embeds that file from
the working tree verbatim, so a dev-built felis piped a CRLF script into `bash -s`
on the target host. CI builds on Linux, which is why released binaries were clean
and only local builds carried it.

eol=lf is global rather than scoped to *.sh because go:embed reaches further than
the installer: deploy/*/Dockerfile, deploy/*/entrypoint.sh, plugins/*/src, the
migrations and internal/panel/static are all compiled in and read on Linux. *.bat
is the one exception, for the gradle wrappers' Windows launchers.

Renormalizing the index touched exactly one tracked file, cmd/felis/version.go,
and only its line endings: `git diff --cached --ignore-cr-at-eol` reports nothing
outside .gitattributes itself.

TestBootstrapPinsViaBlockConnectionsOff used to strip \r\n before asserting, with a
comment stating that the repository pinned no eol attribute. That is no longer true,
and the stripping hid the regression this commit prevents. It now asserts the absence
of CRs, so losing the attribute reports itself as line endings rather than as a
missing serverside-blockconnections pin.

Closes #5
2026-07-28 12:58:46 +09:00
flyemoji 5d4f3063a9 feat(operator): make any Paper image joinable behind the forwarding proxy
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed
handshake does not degrade -- it rejects every login the proxy forwards. Until now the
only backends that could verify it were the two images Felis builds itself
(deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own
entrypoints. An arbitrary Paper image a user brings does not, so it passed admission,
started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the
lobby image as a base for a user's own world (0018_recommended_images.sql), which was
never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining
player away, the exact opposite of a server you mean to stay on.

The fix configures forwarding from OUTSIDE the image instead of requiring it inside.
The operator now injects a root `felis init-forwarding` initContainer into every user
server; it writes the proxies.velocity block into config/paper-global.yml and forces
online-mode=false in server.properties on the /data PVC before the main container
starts. The image needs no forwarding logic of its own, so the joinable set stops being
"images that self-configure forwarding" and becomes every Paper-family image the
platform runs.

buildStatefulSet gates the injection on the ABSENCE of the system-role label: the
Felis-built system servers already consume the secret in their entrypoints and the
login gate is a limbo, not Paper. It is also gated on a non-empty felis image name --
the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it
skips the injection rather than failing, because a cluster whose proxy is not in modern
mode has nothing to configure.

The init runs as root deliberately. The world volume's ownership comes from the storage
provisioner and the main container runs as whatever UID its image declares, so root is
the only UID that can reliably write these files; it then chmods them 0666/0777 so that
non-root main container can rewrite them on boot. The privilege is bounded -- the init
exits before the server container starts and the server container keeps its own UID.
The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the
init ever stops running as root.

The writer merges rather than overwrites, both because Paper expands paper-global.yml to
its full default tree on first boot and because the panel file editor may edit either
file between boots. It sets proxies.velocity.* and the single online-mode key and leaves
every other setting alone. It is a no-op on an empty secret, for the same reason the env
var is optional: a proxy that is not in modern mode provisions no Secret, and wedging
every server's init on a missing optional value would be worse than the status quo.

felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and
0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu
plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the
online-player list and permission commands work out of the box. 0018's row is left in
place -- an admin who kept it can keep it; this only adds the better default beside it.

Three fixes ride along, each of which the 1.8 path hit in practice.

bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only
builds its block-connection provider when Via's lowest supported protocol is below 1.13;
under modern forwarding the Velocity injector reports 393, so init() returns early,
blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences
it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining.
Every call site is behind isServersideBlockConnections(), so switching it off skips all
of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected.
ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with
this one key suffices -- Config#loadConfig parses the bundled default as the base map and
merges the on-disk file over it, so every other option stays current across version
bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this
worked: init() gates on the protocol version too, and that half fails on its own, so the
line is missing either way. The config value is the only evidence, which is what the test
asserts.

The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend
sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it
down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13
and has nowhere to go. Only the handshake address field survives Via, so the Felis fork
forwards the named servers BungeeCord-style while every other backend keeps modern+secret
untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer
CRs is the upgrade path.

deploy/lobby's set_prop escapes the value before substituting it. The RCON password is
operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare
`sed s|...|...|` and silently kills the key -- taking the console, the online-player list
and permission commands with it. deploy/paper was written with the escaping, so the lobby
gets the same rather than leaving the sibling caller broken.

Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where
TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping.
The new tests cover the initContainer's image, root UID, world mount and secret env; the
merge preserving unrelated config trees; the properties upsert including the commented-key
case; and the bootstrap script both writing the Via key and still calling the function
that writes it.

Not verified: the initContainer has never run in a real cluster, and the felis-paper
image is code-only here as the other game-stack images are -- no Go CI builds them.

The ViaVersion pin is the one piece with live evidence, and that evidence is what it was
written from. Before it, a client was cut within a second of "logged in with entity id"
on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into
Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on
2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same
player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login
from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes.

Neither session says which client version it was. The proxy never logged a protocol
number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16
players and below" warning that fired for that player on the lobby leg, and no lower --
Via floors every handshake to the proxy's 393, so anything from 47 up is admissible.
Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain
runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the
NPE proves is that the pin was load-bearing, not who was holding the mouse.

That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one
now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8
behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the
fault above; off, holds. No automated test in THIS repository exercises the leg.
2026-07-28 09:13:58 +09:00
flyemoji 80a29ba653 feat(lobby): ship LuckPerms in the lobby image so the panel's permission controls work
The panel has a full permission surface — internal/api/handlers_access.go issues
`lp user <player> permission set/unset` and `lp user <player> parent add/remove`
over RCON, and projects the result back at
GET /api/v1/servers/{name}/access/luckperms/{player} — but nothing in this tree
ever installed LuckPerms. The lobby image copied felis-paper.jar into the plugin
directory and stopped there, so every grant the panel sent reached a server that
answered "Unknown command". Confirmed on the demo host: /data/plugins held only
FelisPaper/, bStats/, felis-paper.jar and spark/.

This is the other half of the RCON change. That one gave the control plane a
channel to send commands on; this one puts something at the far end that
understands them. Neither is useful alone.

The jar is resolved at build time rather than pinned in the Dockerfile, the same
way PAPER_JAR_URL already is: metadata.luckperms.net publishes the current build
for every platform, and asking upstream keeps this tree from going stale on every
LuckPerms release. Unlike Paper it is not version-matched to MC_VERSION — LuckPerms
ships one Bukkit build covering the whole supported Minecraft range, so there is no
per-version endpoint to ask. The resolver's pattern pins the /bukkit/loader/ path
segment deliberately: the metadata endpoint hands back the fabric, forge, velocity
and bukkit-legacy URLs in the same payload, and a looser match would happily return
a jar Paper cannot load, or the legacy build that targets Minecraft 1.8-1.12.

A missing LUCKPERMS_JAR_URL fails the build. That is a harsher default than the
RCON password, which only warns, and the difference is where the failure surfaces:
a lobby without RCON degrades visibly at once, whereas a lobby without LuckPerms
starts perfectly, runs perfectly, and only reveals itself when an owner tries to
grant somebody a permission. Build time is the cheap place to notice.

The entrypoint refreshes the jar from the image seed on every boot exactly as it
does for paper.jar and felis-paper.jar, so the executable artifact tracks the image
while LuckPerms' H2 database and config under plugins/LuckPerms/ stay on the PVC.
That split is the point: every grant ever issued lives in that directory, so the
refresh must never become a wipe.

Check: the three files that have to agree about LuckPerms — bootstrap.sh resolving
and passing the build-arg, the Dockerfile requiring that arg name and writing a
fixed path, the entrypoint copying from that same path — are pinned against each
other. Nothing compiles them together, and a typo in the path is invisible until a
lobby boots and `set -e` turns the failed cp into a crashloop on the hub every
authenticated player is transferred to. The test reads all three back out of the
embedded FS rather than off disk, since that is what the TUI install path ships.

Deployed installs are NOT fixed by this commit, for the same reason the RCON change
was not: the felis-lobby image has to be rebuilt and re-imported, and the pods
recreated, before the jar exists on the volume.
2026-07-21 00:39:46 +09:00