6a4c30b3f1b34bbc9b793c551e306bd77c23c37c
95
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0faec2b02a |
fix(bootstrap): open the nano port to the proxy alone
For a non-loopback bind, configure_nano_firewall opened the nano port in firewalld to every source, while the summary told the operator to restrict it to the proxy. hasJoined takes no token, so on a public host that port is an auth relay anyone can point a proxy at, spending this host's Mojang egress until Mojang rate-limits it and the operator's own players stop getting in. A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets one rich rule that admits the port from that source only, ipv4 or ipv6 by the address given. Without it, no port is opened and the summary prints the rule to add. A re-run closes the port an earlier installer opened to every source. A rule for a previous FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand. Hosts without firewalld are handled as before. The value goes into the rule text, so it is checked up front for an address with one prefix length and nothing else. firewalld's own parser accepts both rule forms and refuses an ipv6 address under the ipv4 family. The harness covers the rule for each family, the closed-by-default case, the re-run cleanup, the loopback case and the CIDR check. |
||
|
|
fa7b54f5ab |
fix(bootstrap): refuse an unbracketed ipv6 nano listen address
validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081
passed. Go refuses that form ("too many colons in address") and needs
[::1]:8081, so the unit crash-looped on every start. A host part that
contains a colon must now be in brackets.
With that, the bare ::1 pattern in nano_listen_is_loopback can no
longer match an address that gets this far, so it goes. [::1] stays.
The harness adds ::1:8081 to the refused addresses, and [::]:8081 and
:8081, both of which Go binds, to the accepted ones.
|
||
|
|
928a1fdfff |
docs(bootstrap): credit velocity, not authlib, with the hasjoined call
Two installer comments still said authlib makes the hasJoined request and sends no token. Velocity reads -Dmojang.sessionserver and sends the request itself. Comment text only. |
||
|
|
cf65ffdae5 |
fix(bootstrap): open up a nano-only config dir an older run left 0750
write_nano_config creates a missing /etc/felis as 0755, but it left an existing one alone. On a nano-only host an older installer made that directory with a bare mkdir -p, so under a root umask of 027 it is 0750. The DynamicUser unit cannot search it, so felis-nano cannot read its config, and a re-run stops at the service check instead of repairing the directory. An existing directory is now set to 0755 unless it holds the full install's secrets.env or bootstrap.done. The full install locks the directory to 0700 and writes secrets.env right after, so its directory keeps that mode, and install_nano_service still reports the lockout rather than this widening it. The mode cases run only where chmod works; on a filesystem that ignores it the harness skips them. |
||
|
|
2458ee1722 |
docs(bootstrap): say auth_source tags are permanent and order is trust
Both config templates the installer writes, the nano felis.toml and the comment above [[auth_source]] in the generated felis tomls, now state two things an operator editing the list needs to know. A tag is hashed verbatim into every player UUID of its source, with no case folding, so renaming it gives all of those players new UUIDs and orphans their data, links and bans. The list is scanned in order and the first source that validates wins, so order is trust, and a compromised root has to be removed, not moved down. Comment text only. |
||
|
|
6794e66c4d |
fix(bootstrap): fail a tokenless private clone instead of prompting
A source build against a private repository with no FELIS_GITHUB_TOKEN, or a wrong one, made git ask for a username on /dev/tty, and a piped install sat there waiting. git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git fails at once with "terminal prompts disabled". Both fetch_source failures name FELIS_GITHUB_TOKEN in their message: the fresh clone, and the fetch into an existing checkout, which had no message of its own before. |
||
|
|
34f73ba19f |
fix(bootstrap): detect a missing terminal by opening /dev/tty
prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which never fires on Linux: /dev/tty is mode 0666 whether or not the process has a controlling terminal, and only opening it fails. Without a terminal the menu was printed, the read failed with "No such device or address", and the default was taken by accident rather than by the documented path. The guard now opens /dev/tty in a subshell and takes the "no terminal for a prompt" path when that fails. |
||
|
|
0758b9c5d7 |
docs(bootstrap): pass tunables on the sudo line, not by export
The header said to export tunables before running, but its own `curl ... | sudo bash` entrypoint resets the environment, so an exported FELIS_INSTALL_MODE or FELIS_NANO_LISTEN never reached the installer. The header now shows the two forms that do arrive: the variable named on the sudo line, or export followed by sudo -E. The nano summary's hint for a proxy on another machine now prints a sudo line that can be pasted as is, instead of "re-run with FELIS_NANO_LISTEN=...". Comment and log text only. |
||
|
|
02c079c893 |
fix(bootstrap): verify the go toolchain tarball against a pinned digest
install_go_toolchain downloaded the tarball to a fixed /tmp name and unpacked it into /usr/local as root, with no digest check. Another local user could plant that file first, and nothing would notice a tampered download. The tarball is now staged in a mktemp -d directory that the exit cleanup removes, and its sha256 must match before the old toolchain is touched, so a refusal leaves the host as it was. The default 1.26.4 carries pinned amd64 and arm64 digests next to its version; they are the ones https://go.dev/dl/?mode=json&include=all publishes. Any other FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in the header, because no pin can cover a version chosen at run time. Where and which version gets installed is unchanged. |
||
|
|
3b0fc7a3e0 |
fix(bootstrap): print the address nano binds in the install summary
summary_nano printed the node's primary IP for every non-loopback bind and 127.0.0.1 for every loopback one. A bind to a second private address, or to [::1], handed the operator a hasJoined URL that nothing listens on. The host is now the part of FELIS_NANO_LISTEN before the last ':'. The node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]), which names no address a proxy could dial. The loopback and public-bind notes are unchanged. |
||
|
|
404d1172a6 |
fix(bootstrap): refuse a nano listen address without a usable port
FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in the firewall while nano bound nothing, a bare 127.0.0.1 printed http://127.0.0.1:127.0.0.1/... in the summary, and the unit crash-looped either way. validate_settings now requires a ':' and a decimal port of 1-65535 after the last one. It runs after resolve_nano_listen, so an address read back from an existing unit is checked too, and the default always passes. [::1]:8081 and 0.0.0.0:8081 are accepted. |
||
|
|
3918a4b11a |
fix(bootstrap): install only the full control plane under felis setup
prompt_install_mode also runs inside felis setup. Setup then goes on to the Owner and edge setup, which need the control plane, so choosing nano there always ended in a setup error. Under felis setup the mode is now full before any prompt or default is considered, and an explicit FELIS_INSTALL_MODE=nano stops with a message pointing at deploy/bootstrap.sh. That leaves the felis setup branch of acquire_nano_binary unreachable, so it goes. install_embedded_binary stays, since the full install still uses it. |
||
|
|
515c4a6496 |
fix(bootstrap): keep a nano host's listen address and mode on re-run
Re-running the installer is how a nano host updates. That re-run reset FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost its endpoint and every login through it failed. It also offered the full control plane as the default, which on a nano host means k3s and Postgres nobody asked for. The listen address is now settled by resolve_nano_listen, the first step of main, so the later checks see the result. The operator's value wins, then the -listen argument of the installed felis-nano unit, then loopback. The install mode defaults to nano, at the prompt and without a terminal, when the felis-nano unit exists and the full install's bootstrap.done marker does not. Only the full install writes that marker. The harness reads back the unit it wrote earlier, and checks the mode default on a nano-only host, a host with the full install, and a fresh host. |
||
|
|
17b4396460 |
fix(bootstrap): carry auth_source tables with spaced or quoted headers
A re-run copies the operator's [[auth_source]] tables from the existing felis toml into the new one. The awk program that finds them matched only the literal header [[auth_source]], so a table written as [[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid TOML, was taken for some other section and dropped from the config. Each section header now decides afresh whether it opens an auth_source table, through one regex that allows inner whitespace and a single- or double-quoted key. The single quote is spelled \047, which gawk and mawk both honour inside a bracket expression. The harness carries each spelling and checks that the table still stops at the next section. |
||
|
|
a0f54df2a6 |
fix(bootstrap): fail the nano install when the unit does not stay up
install_nano_service printed "enabled and started" straight after systemctl restart, which returns as soon as the process is forked. An upgrade that keeps an old felis.toml the new binary rejects (an [[auth_source]] without a prefix, say) left the unit crash-looping in auto-restart while the installer reported success, and every login through the proxy failed. The install now waits two seconds and asks systemctl is-active. A unit that exited is in "activating (auto-restart)", which is-active does not count as active; on real systemd a unit whose process exits 1 under Restart=on-failure reads activating/auto-restart and is-active returns non-zero, while a running one reads active/running and returns 0. On failure the install prints the unit's last 20 journal lines and stops. This also surfaces a nano unit locked out of an existing 0700 /etc/felis. The harness runs the extracted function with systemctl stubbed both ways. Without the check, the dead-unit cases fail. |
||
|
|
26f685be0e |
fix(bootstrap): create the nano config dir world-searchable
felis-nano runs as a systemd DynamicUser, so it can read /etc/felis/felis.toml only if others may search /etc/felis. write_nano_config made the directory with a bare mkdir -p, which takes its mode from root's umask. On a host hardened to umask 027 that is 0750: nano exits on "permission denied", the unit restarts every five seconds, and no login gets through. A missing directory is now created 0755 explicitly. An existing one keeps its mode, because the full install sets it to 0700 to protect its secrets and widening that from the nano path would expose them. A nano unit locked out that way is left for the install to report. The harness runs the extracted function under umask 027 and checks both cases. Reverting to the bare mkdir fails the first; an unconditional chmod 0755 fails the second. The mode checks skip on filesystems that ignore chmod, such as Git Bash on NTFS. |
||
|
|
07bafebf0d |
fix(bootstrap): keep the operator's auth sources across re-runs
write_felis_toml regenerates felis.host.toml and felis.pod.toml with a wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin block in that heredoc. Re-running the installer, which is also what `felis setup` does, threw away any edit to the list: a root the operator added stopped admitting logins, and a root they removed came back. The generated comment invited exactly that edit. Carry the tables forward the way [smtp] already is: read every [[auth_source]] table from the existing felis.host.toml (falling back to felis.pod.toml) and emit the LittleSkin default only when there is no earlier file at all. An earlier file with no tables stays empty, because that is a Mojang-only server rather than a missing value; felis-api now treats an empty list that way. The file header and the comment above the list now say what survives a re-run, and point at felis.host.toml, which is what the next run reads. bootstrap_test.sh extracts the new function from bootstrap.sh and checks the fresh-install default, an operator's own table carried without the default or the following section, an empty list staying empty, and the indented form the setup TUI writes. It passes under dash with gawk and with mawk; forcing the function to always return the default fails five of the new cases. |
||
|
|
d9246ddae6 | feat(bootstrap): verify Paper and Velocity jars against Fill's digest | ||
|
|
584d31fc49 |
feat(bootstrap): refuse to install an unverified Velocity fork jar
FELIS_VELOCITY_FORK_JAR replaces the proxy every player connects through, and the only thing checked about it was that the path pointed at a readable file. A truncated copy, a stale build left at the same path, or the two-patch jar where the three-patch one was meant all installed silently. It now requires FELIS_VELOCITY_FORK_JAR_SHA256 and refuses on a mismatch, hashing stdin rather than the path for the reason install_via_plugins already documents: sha256sum escapes its output line for a filename carrying a backslash or newline, and the leading "\" that adds fails every comparison. The absent-digest refusal prints the jar's actual hash, so the first run after a deliberate rebuild is one copy-paste rather than an investigation. The comparison ignores case and internal spaces. The fork is built on a developer machine, which is usually Windows, and nothing there prints a digest the way sha256sum does: Get-FileHash returns uppercase and certutil has shipped the bytes space-separated. Comparing raw would refuse two of the three spellings of the correct answer and word the refusal as tampering. No digest is hardcoded, which is the half of the request this does not deliver. The fork is built from Felis-Legacy and has never been reproduced on a second machine, so a constant here would pin one machine's output rather than the fork. The comment that previously asserted the build "is not byte-reproducible" is gone too -- it was stated more confidently than the evidence supports. The fork jars on disk carry Gradle's constant 1980-02-01 entry timestamps, so the usual reason a jar differs between builds is already absent; that is not proof it reproduces, and neither claim should sit in the script unmeasured. This is deliberately not a supply-chain signature and the comment says so: an operator who can write the jar can write the digest. What it buys is that a path stops being an identity, and that every later re-run re-checks the same build. Scope: the fork jar only. The else branch still curls stock Velocity from PaperMC with no verification at all, and that is the branch a default install takes. The digest is already in hand there -- Fill v3 returns checksums.sha256 and its download URL is content-addressed on that same value -- and papermc_latest_jar discards it. Left alone rather than widened into this change. deploy/bootstrap_test.sh covers the gate's two refusals, its happy path, and the two Windows digest spellings. Each case extracts the block under test out of bootstrap.sh with awk and runs it with die/log stubbed, rather than transcribing it -- a transcribed copy passes forever after someone edits the original. The extraction is length-bounded: awk runs an unmatched end pattern to EOF, which would quietly feed the rest of bootstrap.sh to the shell under test. bootstrap.sh itself cannot run here; it wants root, a package manager and k3s. A `shell` CI job runs that plus a syntax check over every tracked script. The syntax step dispatches on each file's shebang instead of running `sh -n` across the board. The blanket form looks fine and is a false green: on a developer machine `sh` is usually bash and accepts everything, while the runner's `sh` is dash. Verified against the real thing rather than an approximation -- inside ubuntu:24.04, where /bin/sh is /usr/bin/dash, the dispatching loop passes all six scripts and the blanket loop dies at bootstrap.sh:191 on the first of its 14 arrays. Refs: Felis-Legacy #19 |
||
|
|
4f5014d033 |
feat(bootstrap): make the legacy-forwarding backend list overridable
Which backends receive their forwarded identity through the handshake address --
rather than proxy-wide modern forwarding -- was the literal string "legacy18",
assigned inside write_velocity_service. Standing up a second protocol-47 backend
therefore meant editing this script, on every host, and remembering to.
It is now FELIS_LEGACY_FORWARDING_SERVERS, defaulting to legacy18, declared beside
FELIS_NANO_LISTEN and documented in the Tunables block like every other knob. The
default is unchanged, so an existing install re-runs to the same systemd unit it
already has.
This is deliberately only half of what the list should eventually do. It is a JVM
system property, read once when Velocity starts, so it is fixed for the life of
the proxy process and a change still needs a restart -- an environment variable is
as far as a startup property can be pushed. Having the list follow the
MinecraftServer CRs is a larger change than it looks: the forwarding decision is
made by the fork's patch to Velocity core, not by the Felis plugin, so core would
have to read state the plugin owns and refreshes. The plugin already maintains a
dynamic backend registry, which is where that state would come from, but the
bridge from core to it does not exist. The comment at the assignment now says so
instead of leaving "the upgrade path is to have the operator render this list from
the MinecraftServer CRs" as though it were a small step.
The -D is now double-quoted in ExecStart. The fork trims each element -- it parses
the property as `split(",")` into a Set, mapped through String::trim with empties
filtered -- so it accepts "legacy18, legacy112", but systemd splits ExecStart on
whitespace before java sees it. Unquoted, that spelling handed java a stray
"legacy112" argument and the unit failed to start; documenting the knob as
comma-separated without quoting it would have shipped that as a footgun.
`bash -n` passes; the default resolves to legacy18, an override to the value given,
and a value containing a space renders inside a single quoted ExecStart item.
|
||
|
|
5d4f3063a9 |
feat(operator): make any Paper image joinable behind the forwarding proxy
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed handshake does not degrade -- it rejects every login the proxy forwards. Until now the only backends that could verify it were the two images Felis builds itself (deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own entrypoints. An arbitrary Paper image a user brings does not, so it passed admission, started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the lobby image as a base for a user's own world (0018_recommended_images.sql), which was never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining player away, the exact opposite of a server you mean to stay on. The fix configures forwarding from OUTSIDE the image instead of requiring it inside. The operator now injects a root `felis init-forwarding` initContainer into every user server; it writes the proxies.velocity block into config/paper-global.yml and forces online-mode=false in server.properties on the /data PVC before the main container starts. The image needs no forwarding logic of its own, so the joinable set stops being "images that self-configure forwarding" and becomes every Paper-family image the platform runs. buildStatefulSet gates the injection on the ABSENCE of the system-role label: the Felis-built system servers already consume the secret in their entrypoints and the login gate is a limbo, not Paper. It is also gated on a non-empty felis image name -- the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it skips the injection rather than failing, because a cluster whose proxy is not in modern mode has nothing to configure. The init runs as root deliberately. The world volume's ownership comes from the storage provisioner and the main container runs as whatever UID its image declares, so root is the only UID that can reliably write these files; it then chmods them 0666/0777 so that non-root main container can rewrite them on boot. The privilege is bounded -- the init exits before the server container starts and the server container keeps its own UID. The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the init ever stops running as root. The writer merges rather than overwrites, both because Paper expands paper-global.yml to its full default tree on first boot and because the panel file editor may edit either file between boots. It sets proxies.velocity.* and the single online-mode key and leaves every other setting alone. It is a no-op on an empty secret, for the same reason the env var is optional: a proxy that is not in modern mode provisions no Secret, and wedging every server's init on a missing optional value would be worse than the status quo. felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and 0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the online-player list and permission commands work out of the box. 0018's row is left in place -- an admin who kept it can keep it; this only adds the better default beside it. Three fixes ride along, each of which the 1.8 path hit in practice. bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only builds its block-connection provider when Via's lowest supported protocol is below 1.13; under modern forwarding the Velocity injector reports 393, so init() returns early, blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining. Every call site is behind isServersideBlockConnections(), so switching it off skips all of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected. ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with this one key suffices -- Config#loadConfig parses the bundled default as the base map and merges the on-disk file over it, so every other option stays current across version bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this worked: init() gates on the protocol version too, and that half fails on its own, so the line is missing either way. The config value is the only evidence, which is what the test asserts. The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13 and has nowhere to go. Only the handshake address field survives Via, so the Felis fork forwards the named servers BungeeCord-style while every other backend keeps modern+secret untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer CRs is the upgrade path. deploy/lobby's set_prop escapes the value before substituting it. The RCON password is operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare `sed s|...|...|` and silently kills the key -- taking the console, the online-player list and permission commands with it. deploy/paper was written with the escaping, so the lobby gets the same rather than leaving the sibling caller broken. Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping. The new tests cover the initContainer's image, root UID, world mount and secret env; the merge preserving unrelated config trees; the properties upsert including the commented-key case; and the bootstrap script both writing the Via key and still calling the function that writes it. Not verified: the initContainer has never run in a real cluster, and the felis-paper image is code-only here as the other game-stack images are -- no Go CI builds them. The ViaVersion pin is the one piece with live evidence, and that evidence is what it was written from. Before it, a client was cut within a second of "logged in with entity id" on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on 2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes. Neither session says which client version it was. The proxy never logged a protocol number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16 players and below" warning that fired for that player on the lobby leg, and no lower -- Via floors every handshake to the proxy's 393, so anything from 47 up is admissible. Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the NPE proves is that the pin was load-bearing, not who was holding the mouse. That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8 behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the fault above; off, holds. No automated test in THIS repository exercises the leg. |
||
|
|
0798f903b0 |
feat(bootstrap): allow the Felis-Legacy Velocity fork to be installed as the proxy
Stock Velocity will not offer the login-plugin-message exchange below 1.13, so a 1.8 client reaching a modern-forwarding backend today is a side effect of Via replacing the channel initialisers before that check runs. It works, and nobody designed it. FL-008's fork registers the two login packets from 1.7.2 and drops the handshake gate, which makes the same outcome deliberate — and its gateonly control shows the registry half is the load-bearing one. FELIS_VELOCITY_FORK_JAR points at such a build; unset, the default, nothing changes and the stock 3.5.1 download runs as before. It stays opt-in because the fork is unmeasured where it counts: FL-008's probe runs offline-mode against a stub backend, while this jar would carry every real Mojang session on the server. No digest is pinned for it. The gradle build is not byte-reproducible across machines, so a hash here would assert a provenance that does not exist; the jar is trusted because that probe certified a build, and the path is checked for readability before anything is replaced. |
||
|
|
8480387c50 |
fix(velocity): pin ViaRewind 4.1.3, the release that matches ViaVersion 5.11.0
The previous commit staged ViaRewind 4.1.2 alongside ViaVersion and ViaBackwards 5.11.0. The staging mechanism was right; the pairing was not. On startup the proxy logs ERROR [viaversion]: Error during loading of Protocol1_9To1_8 java.lang.IllegalArgumentException: Invalid version: 1 and Protocol1_9To1_8 is the one protocol every 1.8 client needs. Without it a 1.8 player clears the handshake gate we just opened and then lands on a translation layer that never initialised. The three versions are a set, not three independent pins. ViaRewind 4.1.3's release notes state it adds compatibility with ViaVersion and ViaBackwards 5.11.0; 4.1.2 predates that. Measured rather than assumed: a two-arm run of the same proxy image, same velocity.toml, same ViaVersion/ViaBackwards jars, differing only in the ViaRewind jar, logs the error three times on 4.1.2 and zero times on 4.1.3. The full FL-007 forwarding matrix was re-run against 4.1.3 rather than inferred from the two-arm result — a different jar is a different configuration under test. It closes 20/20, including the three force-key-authentication assertions, with the proxy log confirming "Loaded plugin viarewind 4.1.3" and no Protocol1_9To1_8 error. Still unproven, and deliberately not claimed here: a real 1.8.9 client authenticated against Mojang. The harness has no Mojang account, so online-mode = true remains the one untested axis. |
||
|
|
f532684249 |
feat(velocity): let 1.8.x players through the modern-forwarding proxy
Felis pins player-info-forwarding-mode = "modern", and Velocity's HandshakeSessionHandler#handleLogin refuses anything below 1.13 outright: it reads handshake.getProtocolVersion() and disconnects with velocity.error.modern-forwarding-needs-new-client before the backend is ever contacted. A player pinned to 1.8.9 never reaches the login gate, never sees the onboarding link, and gets an error string that tells them to upgrade their client. install_via_plugins now stages ViaVersion, ViaBackwards and ViaRewind into /opt/felis/velocity/plugins alongside felis-velocity.jar. Nothing else moves: same forwarding mode, same secret, no backend patched and no backend downgraded. The 1.13 floor turns out to be a property of the unassisted proxy pipeline rather than of the forwarding protocol, so lifting it costs three jars and no source change. This was measured, not assumed. Felis-Legacy's FL-007 probe stands a protocol-47 client in front of a stock Paper 1.21.11 backend behind a modern-forwarding proxy and watches it join. The proof is the join itself rather than the log line: that backend runs velocity.enabled with a shared secret, and Paper in that state rejects any login not carrying forwarding data signed with a matching HMAC. The control cell without Via is rejected before the backend is contacted, so Via is the only difference. A second cell re-runs the same join with force-key-authentication = true, the way Felis sets it, because a result that only holds under a config Felis does not run is not a result about Felis; the 1.19+ signed chat key a protocol-47 client cannot produce is never demanded, and it cannot be, since the pre-1.19 wire format has no player-key field to decode. The jars are pinned by sha256 and not by a moving tag. They sit in front of every packet on the proxy and they are the exact bytes FL-007 measured; "latest" would quietly make this an unmeasured configuration. The digests come from the GitHub releases FL-006 locked, which is deliberate — Hangar's VELOCITY/download endpoint serves different bytes for the same version numbers, so an installer that only checked for HTTP 200 would ship artifacts nothing has tested. A version bump means a digest bump here. Two limits are worth writing down. Only protocol 47 was measured; the rest of Via's documented 1.7-1.12 range is inference from that one point. And the probe runs online-mode = false because it has no Mojang account, so Felis's online-mode = true is untested — the untested part is the Mojang auth handshake specifically, which puts the residual in Via's own login handling rather than in forwarding or in the hasJoined multiplexer, whose request is protocol-independent. Checks: a fresh install lands all three jars at the pinned digests and sizes with no temp files left behind; a re-run downloads nothing; a tampered jar is restored to the pinned bytes; and a deliberately wrong digest aborts without installing anything. Existing installs pick this up by re-running bootstrap, which re-enters install_velocity because bootstrap.done is written but never read as an early exit. |
||
|
|
80a29ba653 |
feat(lobby): ship LuckPerms in the lobby image so the panel's permission controls work
The panel has a full permission surface — internal/api/handlers_access.go issues
`lp user <player> permission set/unset` and `lp user <player> parent add/remove`
over RCON, and projects the result back at
GET /api/v1/servers/{name}/access/luckperms/{player} — but nothing in this tree
ever installed LuckPerms. The lobby image copied felis-paper.jar into the plugin
directory and stopped there, so every grant the panel sent reached a server that
answered "Unknown command". Confirmed on the demo host: /data/plugins held only
FelisPaper/, bStats/, felis-paper.jar and spark/.
This is the other half of the RCON change. That one gave the control plane a
channel to send commands on; this one puts something at the far end that
understands them. Neither is useful alone.
The jar is resolved at build time rather than pinned in the Dockerfile, the same
way PAPER_JAR_URL already is: metadata.luckperms.net publishes the current build
for every platform, and asking upstream keeps this tree from going stale on every
LuckPerms release. Unlike Paper it is not version-matched to MC_VERSION — LuckPerms
ships one Bukkit build covering the whole supported Minecraft range, so there is no
per-version endpoint to ask. The resolver's pattern pins the /bukkit/loader/ path
segment deliberately: the metadata endpoint hands back the fabric, forge, velocity
and bukkit-legacy URLs in the same payload, and a looser match would happily return
a jar Paper cannot load, or the legacy build that targets Minecraft 1.8-1.12.
A missing LUCKPERMS_JAR_URL fails the build. That is a harsher default than the
RCON password, which only warns, and the difference is where the failure surfaces:
a lobby without RCON degrades visibly at once, whereas a lobby without LuckPerms
starts perfectly, runs perfectly, and only reveals itself when an owner tries to
grant somebody a permission. Build time is the cheap place to notice.
The entrypoint refreshes the jar from the image seed on every boot exactly as it
does for paper.jar and felis-paper.jar, so the executable artifact tracks the image
while LuckPerms' H2 database and config under plugins/LuckPerms/ stay on the PVC.
That split is the point: every grant ever issued lives in that directory, so the
refresh must never become a wipe.
Check: the three files that have to agree about LuckPerms — bootstrap.sh resolving
and passing the build-arg, the Dockerfile requiring that arg name and writing a
fixed path, the entrypoint copying from that same path — are pinned against each
other. Nothing compiles them together, and a typo in the path is invisible until a
lobby boots and `set -e` turns the failed cp into a crashloop on the hub every
authenticated player is transferred to. The test reads all three back out of the
embedded FS rather than off disk, since that is what the TUI install path ships.
Deployed installs are NOT fixed by this commit, for the same reason the RCON change
was not: the felis-lobby image has to be rebuilt and re-imported, and the pods
recreated, before the jar exists on the volume.
|
||
|
|
b3989fa4af |
fix(mail): prove SMTP deliverability before saving, and stop losing the relay
A live install passed the SMTP setup screen and then failed every one-time
code with a bare `internal error`. Four separate defects had to line up for
that, and each is fixed here.
The relay was configured with `from = noreply@<domain-A>` on an account
authenticated as `<user>@<domain-B>`. Providers that validate sender identity
— Fastmail among them — answer MAIL FROM with an unconditional 250 and only
refuse at end-of-DATA. Ping stopped at NOOP, so it never saw the refusal: the
wizard reported success, wrote the config, rolled felis-api, and every OTP
afterwards died at w.Close().
Ping now runs the same transaction a real code takes — connect, (STARTTLS,)
AUTH, MAIL FROM, RCPT TO, DATA — delivering one self-test message to the From
address, and SendOTP and Ping share deliver() so the check cannot drift from
the thing it checks. The self-test recipient cannot cause a false negative:
an authenticated submission relay accepts RCPT for any destination by
definition, while the sender identity it does validate is exactly what we
want tested. The setup screen now says a message will be sent, names the
address it went to, and warns that From must be an address the account is
allowed to send as.
A relay refusal also answered 500 `internal`, which reads as a broken panel
and sends the operator hunting through handler code instead of their [smtp]
block. It is now 502 `mail_undeliverable`, mapped inside deliverOTP so all
four doors that mail a code (onboarding, email login, op-login, migrate
step-up) answer alike. The relay's own text stays out of the response — it
can name the SMTP account, and these routes are reachable by any signed-in
player — and goes to the log instead.
writeError logged nothing when it collapsed an unmapped error to 500, so an
operator holding an `internal error` had nothing to grep for and diagnosis
degraded into guessing against a live install. It now logs the method, path,
wrapped chain and the same request_id the caller is shown.
Finally, write_felis_toml regenerated the config wholesale and never emitted
[smtp], so re-running the installer — the documented way to update felis-api —
silently erased a working relay and reverted OTP delivery to the no-Mailer
path, logging codes instead of sending them. It now carries the block forward,
cached on first read because the host toml is clobbered before the pod toml is
written. Same defect family as the root_domain loss fixed in
|
||
|
|
ecbeb20761 |
fix(bootstrap): reuse the installed root domain instead of re-deriving it
detect_node_ip recomputed FELIS_ROOT_DOMAIN from scratch on every run and fell back to <node-ip>.nip.io. Nothing read the domain back out of the felis.toml an earlier run wrote, so it survived only as long as the operator kept passing the same environment. That made re-running the installer destructive on any install with a real domain, and re-running it is not optional: it is the only way to move felis-api to a newer release, which is what `felis update` points operators at. A bare re-run rewrote root_domain, panel_hostname and admin_hostname to nip.io names while ensure_panel_tls_cert returned early on the certificate it had already written, leaving the console serving a cert for hostnames it no longer answered to -- with no re-domain flow to recover through. Precedence is now explicit FELIS_ROOT_DOMAIN, then the domain the last run persisted, then the nip.io default. First installs are unaffected. Deliberate re-domains still work, because there is no other route to one, but they now warn that the write-once certificate is not reissued and that the proxy and login config carry the old name too. Secrets were never exposed to this: load_or_make_secrets has always sourced secrets.env before generating anything. The domain was the one piece of install identity with no read-back. The channel is deliberately left alone. FELIS_VERSION_BOOTSTRAP is not persisted either, but defaulting a re-run to the release channel installs a working build rather than breaking one, so cmd/felis/update.go states that instead. Its warning about the domain went with the bug and would now be false. Verified against the shipped function text: the ladder holds for a fresh host, a re-run with and without the variable set, a re-domain, and a felis.toml whose root_domain is missing or empty. Reverting the one line reproduces the nip.io overwrite. |
||
|
|
e5ea51c0db |
fix(bootstrap): wrap the downloaded binary in the same base CI ships
build_image_from_binary built on distroless/base-debian12 while the repo
Dockerfile's final stage uses distroless/static-debian12, so the image an
install runs did not match the image CI publishes.
Every binary that can reach HOST_BIN traces back to the Dockerfile's
CGO_ENABLED=0 build -- the downloaded CI asset, the binary the TUI is already
running, and the one build_image_from_source docker-cp's out of the image it
just built. None link glibc, so base-debian12 bought nothing and only widened
the runtime surface.
This mattered little while build_image_from_binary was the rare fallback.
|
||
|
|
659c8e5e9f |
feat(bootstrap): install the published release build instead of compiling on the host
deploy/bootstrap.sh now resolves the newest published GitHub release, downloads
the binary CI built for that tag, and builds a thin image around it. Compiling
on the target host becomes the fallback and the opt-in, not the default.
The panel is not a separate artifact. The Dockerfile copies panel/dist into
internal/panel/static before the go build, so the control plane — panel and
backend — ships as ONE file. The release channel therefore downloads exactly
one asset, felis-linux-<arch>, and needs no registry, no Go toolchain and no
checkout on the host.
The Minecraft game stack (limbo, lobby, the Velocity plugin) is still always
built locally. game_stack_source now keys on HAVE_PREBUILT_BINARY — the same
flag build_image uses — so on any prebuilt path it unpacks the tar embedded in
that binary instead of trusting a checkout an earlier install left behind.
Trusting the checkout would build the plugin from an old commit against a
freshly downloaded control plane: a silent version skew across the plugin/API
boundary.
Channels:
(default) newest published release, downloaded
FELIS_VERSION_BOOTSTRAP=dev clone main and compile
FELIS_REF=<ref> pins the tree, forces the source path
The download is best-effort. A tag whose assets are not uploaded yet, an
architecture with no published asset, or an asset that fails validation each
warn and fall back to compiling THE SAME TAG from source — never a different
commit.
The ref is resolved right after install_base, the first point curl exists and
well before docker and k3s, so a missing FELIS_GITHUB_TOKEN or an unpublished
release costs the operator seconds instead of a k3s install they then have to
unwind. It is skipped on exactly the paths that never consume the result: the
TUI, which rebuilds the binary it is already running, and FELIS_SKIP_FETCH,
which builds whatever is staged. Resolving anyway would set FELIS_VERSION to
the newest tag and stamp a staged tree as that release.
The asset is staged next to HOST_BIN rather than in TMPDIR. Validation EXECUTES
it, and /tmp is noexec on CIS-hardened images, where the exec dies 126, the
check reads it as a bad asset, and every such host silently falls back to the
full on-host compile this path exists to avoid. It also keeps a private-repo
artifact out of a world-readable 1777 directory.
git_auth, which supplies the token to git for a private-repo clone, passes an EMPTY
credential.helper before the inline one. credential.helper is multi-valued: a bare
`-c credential.helper=...` APPENDS to whatever the host has configured rather than
replacing it, and an empty value is git's documented list reset. Without it, on a host
with a persistent helper (Git for Windows ships `manager` at SYSTEM scope) two things
go wrong. Git runs `credential approve` automatically after a successful clone and
feeds every helper in the list, so a `store` helper writes the PAT to
~/.git-credentials in cleartext — the token outlives the install, in a file bootstrap
never created and never cleans up. And because the inline helper is LAST, a
pre-existing helper answers `fill` first, so a stale cached credential can win and the
clone authenticates as the wrong account — surfacing as exactly the 404-on-private-repo
the surrounding code works hard to explain. Reproduced both against a real clone, and
confirmed the reset closes both.
internal/panel parses the new stamp. The dev channel now emits "<tag>+g<sha>", which
matched neither describeSuffix ("-N-g<sha>") nor releaseTag, so a dev build fell through
to the default case and the version badge rendered the entire stamp as the release with
no commit. A devSuffix case handles it; the git-describe case stays for hand-rolled
`-ldflags "-X main.version=$(git describe)"` builds. Table test covers both forms plus
the release, dirty and unstamped cases.
CRD application no longer branches on the install path: it is always
`felis bootstrap-assets crd`. That output is byte-identical to deploy/crd/ —
bootstrap_asset.go embeds that very file — and needs no checkout, so one source
replaces a branch whose two arms had to be kept in agreement by hand.
Dockerfile gains a FELIS_VERSION build arg wired into -X main.version, declared
after `go mod download` so a version bump does not invalidate that layer. Both
build stages are pinned to $BUILDPLATFORM so a multi-platform buildx run never
emulates them: the panel's output is architecture-independent and the Go stage
cross-compiles via TARGETARCH. The final stage stays on the target platform and
is COPY-only, which BuildKit performs without QEMU.
.github/workflows/release.yml publishes on a vX.Y.Z tag: vet, tests, then one
buildx run producing both architectures through the repo Dockerfile. Not a bare
`go build` — internal/panel/static holds a tracked placeholder index.html so the
//go:embed compiles without node, which means a direct build succeeds and
quietly ships a release whose panel is that placeholder.
The stamp is asserted end to end, because it fails silently: an unstamped binary
reports "dev", which the updater refuses to compare, disabling update reporting
for every install built from that release. The arm64 artifact is checked by ELF
machine type rather than by running it — runners have binfmt registered, so
executing an amd64 binary misnamed arm64 would succeed.
Prerelease tags are flagged explicitly. The trigger glob is v*, gh does not read
semver out of a tag name, and an RC published as a full release becomes
/releases/latest — the single endpoint the default channel installs from and
`felis update` polls.
No SHA256SUMS. A checksum fetched over the same TLS session, with the same
credential, from the same host as the binary adds no trust root; signing is the
real answer and is a separate decision.
Not verified: the download -> validate -> image -> k3s path has never run on a
host against a real published release, because no tag exists yet. The shell
logic around it is verified out of tree; the network and exec behaviour is not.
|
||
|
|
d146f1ccd7 |
fix(bootstrap): keep the install alive on a host with only felis-api
restart_existing_control_plane ended in an and-list per deployment:
[ "$had_api" = "1" ] && kube ... rollout restart deployment/felis-api
[ "$had_operator" = "1" ] && kube ... rollout restart deployment/felis-operator
As the LAST command of a function, an and-list whose test is false returns 1,
and that becomes the function's exit status. The call site is bare, so under
`set -Eeuo pipefail` the installer dies there — after the bundle has been
applied and before the rollout wait, leaving a half-finished upgrade and no
message naming the cause.
It fires on any host carrying one control-plane deployment but not the other:
felis-api present without felis-operator restarts the api, then exits 1 on the
second test. Both present, or neither, happened to work, which is why it
survived.
Rewritten as explicit `if` statements, which return 0 when the test is false.
Verified out of tree against all four had_api/had_operator combinations.
|
||
|
|
7860152f57 |
feat(auth)!: go fully passwordless and fix cross-check review findings
Remove password authentication everywhere; the only session doors are passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and op-login vouching. Remediates the 33-finding cross-check review across backend, CLI, panel, plugins, and docs. Backend/CLI: - Drop password routes and fields from account/user/onboard/auth handlers; align tests (new account subtests, naming reserves "console", op-login/onboard/qr-login test updates). - Add migrations 0016_op_login.sql and 0017_drop_password.sql. - Thread panel/admin hostnames from hostcfg through api.go, setup_panel.go, tui_root.go and tui_preflight.go instead of hardcoding; bootstrap.sh writes panel-hostname/admin-hostname into felis.toml. - Reword breakglass and TUI copy for passwordless flows. Panel: - Delete the ChangePassword page and all password UI; align login/auth/api/types with the passwordless contract; add the migration and op-login approval flows. - i18n: convert ImageBuildPage durations/status badges and ServerLuckPerms strings to translation keys; drop 72 orphan keys per locale; unify the title as "Felis - Console". Plugins (all six rebuilt): - Velocity waiting router returns 503 at_capacity during wake; MOTD/control-channel copy and config comments. - Paper zh menu title; Limbo bind-code TTL 600s with panel_url preference; unified /link lines in fabric/forge/neoforge; shared link-client javadoc contract fixes. Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and plugins/README.md aligned with the implementation. BREAKING CHANGE: migration 0017 irreversibly drops users.password_hash and users.must_change_password; password login cannot be restored after migrating. |
||
|
|
c96b36a41f |
fix(bootstrap): retry transient Fill failures when resolving build jars
A single HTTP 502 from fill.papermc.io aborted the entire bootstrap. Build resolution used a one-shot curl, so one gateway blip from an upstream that flaps was indistinguishable from a permanent failure, and the run died before Docker, k3s or any game server was provisioned. Pass --retry 5 --retry-delay 2 to the build-resolution fetches. 502/503/504 are already in curl's built-in transient set, so the tool had solved this; the flags were simply never passed. papermc_latest_jar is shared by the Paper and the Velocity resolve, so hardening it once covers both callers. The LOOHP/Limbo CI metadata fetch feeds the same step and gets the same treatment. Deliberately no --retry-connrefused. It only adds ECONNREFUSED to a set that already covers this incident, and it needs curl 7.52.0 while the yum (el7) path the script supports ships 7.29.0, where an unrecognised long option is a parse error rather than a warning: # centos:7, curl 7.29.0 $ curl -fsSL --retry 5 --retry-delay 2 --retry-connrefused https://example.com curl: option --retry-connrefused: is unknown Under set -Eeuo pipefail that exits 2 and trips the || die, so both hardened fetches would hard-fail on a host where they used to work, each naming a cause that is not the real one. A comment above papermc_latest_jar records this so the flag does not come back. The failure message was actively misleading. "no Paper build for Minecraft 26.2 (the login gate speaks only that protocol)" reads as "that Minecraft version is unsupported", sending the reader after a version-pinning problem that does not exist: Paper 26.2 build 60 resolved fine minutes later. Say what is actually known instead, that the build likely exists and Fill is flapping. Verified against a local always-502 server: curl now issues 6 requests (1 initial + 5 retries) over 10.1s before giving up, where it previously issued 1 and died. Verified on centos:7 that this flag set is accepted, and against the live Fill v3 API that the resolve still returns a jar URL. Known and deliberately unchanged: no fetch sets --max-time, so an upstream that accepts a connection and never answers still blocks forever. --retry does not cover that, as it fires only once a request completes with a failure. The remaining single-shot downloads (cloudflared, the Docker GPG key and repo list, k3s, the Temurin JRE, the Velocity jar, the Go toolchain) keep their existing no-retry shape rather than widen this diff on a script that is about to provision a live host. |
||
|
|
60732a6283 |
feat(operator): gate op.console to staff and land owner setup there
The operator console (op.console.<root>) requires internal permission verification on top of Zero-Trust: a passkey is not access. requireExternal now refuses any non-admin principal arriving on the admin host, before any handler, so op.console is staff-only at the door rather than per-route — including on the passwordless demo face where Cloudflare Access is not in front. The gate is inert on the player console (console.<root>). Owner first-run setup is staff onboarding, so `felis setup` mints the one-time setup URL on op.console.<root>/setup (was console.<root>). The passkey verifier lists both console and op.console in RPOrigins so the one-time binding asserts on either face under the shared console.<root> RP-ID. Session admin-access now includes role=owner, not only admin: the owner is a superset of admin, so excluding it left IsOwner() unreachable through a passwordless session. No path assigns role=owner yet — this is forward consistency. The bootstrap summary now names console.<root> the player panel and op.console.<root> the operator console where the Owner runs setup, fixing text that told operators not to run setup there. Tests: op.console door gate (non-admin refused, player console unaffected, admin passes) and owner session admin-access; the setup-bind default-host test follows the move to op.console. |
||
|
|
26b62e91cd |
feat(bootstrap): federate LittleSkin by default and let Velocity own its runtime dir
Point Velocity's authlib (mojang.sessionserver) at the felis-api hasJoined multiplexer so a full install federates Mojang plus the configured [[auth_source]] set out of the box, not just the standalone `felis nano`, and ship a default LittleSkin auth_source in the generated felis.toml (delete the block for a Mojang-only server). Velocity now owns velocity.toml and its working tree — it migrates the config version and extracts localizations on start — while the jar and forwarding secret stay root-owned read-only and ReadWritePaths widens to VELOCITY_DIR; the felis-api-internal ClusterIP lookup is factored into a felis_internal_ip helper shared by the link config and the sessionserver override. Re-include plugins/velocity in the docker context because the felis binary now embeds all four plugin trees for `felis bootstrap-assets game-stack`. |
||
|
|
be0c4c41f4 | feat(bootstrap): install authenticated game stack | ||
|
|
fd062882ed |
feat(nano): give a Mojang player's name back to them, by prefixing the squatter
A premium player and a third-party player sharing a username could not both be online. Whichever logged in second was kicked with "You are already connected to this proxy!" -- even though the UUID rewrite had already made them two distinct players on the backend. Velocity's player registry is keyed on the NAME (lowercased), not the UUID, so two identities holding one name are one player as far as the proxy is concerned, and the reclaim invariant the rewrite buys is invisible to it. The fix needs no plugin and no state, because Velocity honours the name in the hasJoined RESPONSE rather than pinning the one the client sent at login-start -- established by a real login, not by reading the source. So the multiplexer hands back a different name and the collision is simply gone. A third-party player whose name belongs to a Mojang account now joins as PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only on an actual collision, decided by asking api.mojang.com whether the name is registered. The name's owner is never the one renamed, which is 正版优先 falling out for free -- the identity source is never rewritten, so there is no policy to encode and no 30-day hold to track. The premium-name answer is cached asymmetrically, because the two directions have very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and is trusted for a day; "free" can stop being true the moment someone buys that name, and a stale "free" leaves a squatter holding a name its real owner has just bought, so it is trusted for ten minutes. A lookup that fails with nothing cached fails CLOSED -- assume premium, rename the third-party player: a Mojang outage must not become an opportunity to hold someone else's name, and being wrong that way costs a cosmetic prefix while being wrong the other way bounces the name's owner off the proxy. The lookup gets its own 2s client rather than sharing the 5s auth client, since it is a SECOND Mojang round-trip on a login that already spent one. prefix is a required, unique, 1-4 character config field rather than something derived from the tag, because it is player-visible and no derivation can know that "littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their same-named players onto one name, so uniqueness is enforced case-insensitively -- the proxy folds case, and LS/ls would collide there while reading as distinct here. Also close a pre-existing hole on the path this touches: a third-party source's profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put "§4admin", an empty string, or 200 characters straight into the proxy's player list. The name is now checked against the Minecraft username charset and a bad one is a 204, the same way a bad UUID already was. Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches: premium FLYEMOJ1 -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged LittleSkin FLYEMOJ1 -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668 both online at once, zero "already connected" rejections LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3 The last line is the one that matters: an ordinary third-party player collides with nobody and keeps their name, while the rewrite that keeps identities apart still ran. Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other half of it -- the rename moved the player's display name and their playerdata came along untouched, because every server-side key is the UUID and the UUID does not depend on the name. Known ceiling, left alone deliberately: two players of one source whose names agree on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a prefixed name may itself happen to be a premium name. Both cost an "already connected" bounce, not an identity -- the UUID rewrite does not depend on the name at all. BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or digits, unique across sources). An existing nano felis.toml without it fails to load with an error naming the field, rather than silently keeping the collision. |
||
|
|
b323975ddb |
fix(nano): -Dmojang.sessionserver takes the full hasJoined URL, not the base
|
||
|
|
d417efc8bf |
fix(nano): make the installer work on EL10 and stop serving hasJoined to the world
Deploying `felis nano` to a real Rocky Linux 10 host surfaced four defects that no local check could see. Fixed together because they all sit on the same path from `curl|bash` to a running felis-nano.service. * docker killed the nano install on EL10. `acquire_nano_binary` pulled in docker purely to build the binary; on Rocky 10.2 the docker-ce el10 rpms install but dockerd refuses to start, so the install died at `systemctl enable --now docker`. nano needs one static binary, not an image, so the docker dependency is gone: fetch_source -> install_go_toolchain (pinned FELIS_GO_VERSION, default 1.26.4, amd64/arm64) -> build_nano_binary. * the built binary could not be exec'd by systemd (203/EXEC). The Go linker renames its output out of $TMPDIR, and a same-filesystem rename carries the source SELinux label, so `go build -o /usr/local/bin/felis` produced a binary labelled user_tmp_t rather than bin_t. root is unconfined and could run it by hand, which is what made this look fine, but the DynamicUser service could not. build_nano_binary now stages the output and installs it as a fresh file so the policy type transition labels it bin_t, with restorecon as a belt. * re-running the installer did not converge. `systemctl enable --now` is a no-op on an already-active unit, so a rebuilt binary was installed while the old process kept running. Now enable + restart. * the Velocity wiring comment in cmd/felis/nano.go was wrong. authlib appends /session/minecraft/hasJoined itself, so -Dmojang.sessionserver takes the base URL only, as the installer has always printed. Also bind to loopback by default. hasJoined is unauthenticated by protocol -- authlib speaks the vanilla sessionserver dialect and sends no token -- so a public bind is an open auth relay: anyone can point their own proxy at it and spend this host's egress IP on Mojang until Mojang rate-limits it and the operator's own players stop getting in. It is not an identity bypass (a caller still needs a serverId hash bound to their own server key, which the upstream Yggdrasil validates), but it is someone else's traffic on your address. FELIS_NANO_LISTEN and the -listen flag now default to 127.0.0.1:8081, which a same-host Velocity reaches unchanged; serving an off-host proxy is an explicit opt-in. configure_nano_firewall no longer opens a port for a loopback bind, and summary_nano prints the real bind address plus the relay warning. Verified on the target host: installs with no docker present, service active, binary labelled bin_t, `ss` shows LISTEN 127.0.0.1:8081, an external request is unreachable, and an in-host request returns 204 with the login logged. BREAKING CHANGE: felis nano defaults to 127.0.0.1:8081 instead of 0.0.0.0:8081. A Velocity proxy on another machine must now set FELIS_NANO_LISTEN (or -listen) to a reachable address, and should allow that port only from the proxy's IP. |
||
|
|
87aa9ea625 |
feat(deploy): bootstrap can install Felis-nano only, chosen at the start prompt
bootstrap.sh now asks up front whether to install the full Felis control
plane or only Felis-nano, and grows a parallel install path for the
nano-only case.
- prompt_install_mode() runs right after OS detection and reads /dev/tty
(so it works under `curl ... | sudo bash`) offering [1] Felis / [2]
Felis-nano, default full. FELIS_INSTALL_MODE=full|nano skips the prompt
for non-interactive runs; no tty falls back to full.
- main_nano() installs only what nano needs: the felis binary (reusing
the embedded-binary / docker-build acquisition), a template felis.toml
carrying a commented [[auth_source]] example (Mojang-only until edited),
a felis-nano.service unit running `felis nano -config ... -listen ...`
under DynamicUser hardening, and a firewalld port-open for the listen
port. None of the k3s / Postgres / migrate / bundle steps run.
- write_nano_config is idempotent (leaves any existing config untouched)
and its template is valid as-is. summary_nano prints the hasJoined
endpoint and the Velocity -Dmojang.sessionserver flag, offering the
127.0.0.1 form when the proxy is on the same host.
Verified on WSL: `bash -n` clean; the emitted template loads via
config.LoadNano and the exact systemd ExecStart command serves 204 on a
miss ("Mojang + 0 third-party source(s)"); a duplicate-tag config still
exits non-zero citing "unique". Not exercised: a full main_nano run,
systemd activation of the unit, and shellcheck (unavailable in this env).
|
||
|
|
e5f1682898 | refactor(deploy)!: TUI | ||
|
|
318a724f98 | feat(deploy): add pacman support for Arch Linux | ||
|
|
deaa2f8c85 | feat(deploy): add zypper support for openSUSE/SLES | ||
|
|
9c46632929 |
feat(cli): add felis setup first-run console with reclaim protection and cfsetup idempotency
- Add `felis setup` TUI for initial Owner provisioning and optional Cloudflare edge - Refactor breakGlass to share console TUI model (runConsoleTUI) with setup mode - Session auth respects configured [auth].admin_hostname; fallback to op.console.<root> - Protect linked Yggdrasil admins from Mojang-priority reclaim (spec §B3) - cfsetup: idempotent Access app/policy creation, better 401/403 errors, GET + lookup - Bootstrap: auto-install cloudflared, symlink /etc/felis/felis.toml - Add sequence diagrams for ping-to-join, claim, and link flows |
||
|
|
94a3b7b5e8 | fix(deploy): harden bootstrap for RHEL-family Linux | ||
|
|
58fa4b0af8 |
feat(deploy): add one-line bootstrap installer and container image
bootstrap.sh auto-detects the host package manager (apt/dnf) and installs whatever is missing: Docker, k3s, and PostgreSQL. It builds and imports the felis image, opens pg_hba to the pod CIDR, runs migrations, and applies the rendered control-plane bundle, leaving Web disabled pending 'felis setup'. The Dockerfile builds the distroless felis image; deploy/crd holds the MinecraftServer CRD. |