The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.
Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).
Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.
push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.
persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.
Found while landing the registry-hosting work, which depends on those same
keys surviving.
- persisted_registry_block carries the operator-owned [registry] keys and the
[registry.s3] subtable forward, same first-readable-file rule as
persisted_smtp_block; url/build_namespace stay installer-owned (they must
match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
idempotence (a second re-run writes a byte-identical file).
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:
- Every image the installer builds is now a registry ref
(registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
containerd under that exact name (first boot needs no registry round-trip)
and mirrored into the registry after deploy_bundle (push_image_to_registry:
push endpoint 127.0.0.1:5000, and only the path after the host matters to the
registry — a push there lands where kubelet's mirrored pull looks). A ref
outside the registry is warned about, not silently unmirrored.
- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
registry Deployment binds (node containerd cannot dial the Service VIP — live
drill: "Empty reply"). k3s regenerates containerd config only at agent start,
so a CONTENT change restarts k3s and an identical file (every re-run)
restarts nothing.
- import_registry_image caches registry:2 into containerd so the registry
Deployment can start on a box that cannot reach Docker Hub.
- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
'felis-paper:demo') at the registry refs — a user server created from those
rows must not strand when GC collects the bare tag. Only recommended rows
still holding the old seed are touched; enabled is preserved; a pre-existing
target row wins over a duplicate.
bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:
- submit: derived context refs become the internal-face URL
/api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
image's new fetch-context entrypoint) that streams the blob with the
namespace-local service-token Secret — never mounted into Kaniko — and
extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
build namespace gets the token Secret through the existing replica mechanism
(bootstrap.sh + felis setup); the build egress lock opens exactly the control
namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
escapes/symlinks/non-gzip).
Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.
Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
Three faces of one gap, all on the supported install path:
- Backup/restore answered 503 out of the box: nothing ever rendered the
archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
'felis manifests' names it by default (--backup-pvc= is the explicit
no-store shape); bootstrap passes it through so the generated felis.toml
[archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
nothing deleted).
- Even when enabled it could not find a world on a stock install:
resolveWorldDir now also resolves the exact local-path directory
<pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
glob, so a stale deleted PV's bytes can't be archived in place of the
current world). Reaper Role gains persistentvolumeclaims:get (weaker
than the delete it already held).
README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
The source build of the nano binary installed Go at /usr/local/go and
replaced whatever version was already there. On a host that also
builds other things, the operator's own toolchain was removed and
swapped for Felis's pinned version without a word.
GOROOT_DIR is now /opt/felis/go, next to the source, the Velocity
install and the JRE Felis already keeps under /opt/felis, and
install_go_toolchain creates the parent before unpacking. A host where
an earlier run put Go at /usr/local/go downloads it once more on the
next re-run and keeps the old tree untouched; removing it is the
operator's call. The harness now requires the toolchain directory to
be under /opt/felis.
For a non-loopback bind, configure_nano_firewall opened the nano port
in firewalld to every source, while the summary told the operator to
restrict it to the proxy. hasJoined takes no token, so on a public
host that port is an auth relay anyone can point a proxy at, spending
this host's Mojang egress until Mojang rate-limits it and the
operator's own players stop getting in.
A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets
one rich rule that admits the port from that source only, ipv4 or
ipv6 by the address given. Without it, no port is opened and the
summary prints the rule to add. A re-run closes the port an earlier
installer opened to every source. A rule for a previous
FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand.
Hosts without firewalld are handled as before.
The value goes into the rule text, so it is checked up front for an
address with one prefix length and nothing else. firewalld's own
parser accepts both rule forms and refuses an ipv6 address under the
ipv4 family. The harness covers the rule for each family, the
closed-by-default case, the re-run cleanup, the loopback case and the
CIDR check.
validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081
passed. Go refuses that form ("too many colons in address") and needs
[::1]:8081, so the unit crash-looped on every start. A host part that
contains a colon must now be in brackets.
With that, the bare ::1 pattern in nano_listen_is_loopback can no
longer match an address that gets this far, so it goes. [::1] stays.
The harness adds ::1:8081 to the refused addresses, and [::]:8081 and
:8081, both of which Go binds, to the accepted ones.
Two installer comments still said authlib makes the hasJoined request
and sends no token. Velocity reads -Dmojang.sessionserver and sends
the request itself. Comment text only.
write_nano_config creates a missing /etc/felis as 0755, but it left
an existing one alone. On a nano-only host an older installer made
that directory with a bare mkdir -p, so under a root umask of 027 it
is 0750. The DynamicUser unit cannot search it, so felis-nano cannot
read its config, and a re-run stops at the service check instead of
repairing the directory.
An existing directory is now set to 0755 unless it holds the full
install's secrets.env or bootstrap.done. The full install locks the
directory to 0700 and writes secrets.env right after, so its directory
keeps that mode, and install_nano_service still reports the lockout
rather than this widening it. The mode cases run only where chmod
works; on a filesystem that ignores it the harness skips them.
Both config templates the installer writes, the nano felis.toml and
the comment above [[auth_source]] in the generated felis tomls, now
state two things an operator editing the list needs to know.
A tag is hashed verbatim into every player UUID of its source, with no
case folding, so renaming it gives all of those players new UUIDs and
orphans their data, links and bans. The list is scanned in order and
the first source that validates wins, so order is trust, and a
compromised root has to be removed, not moved down. Comment text only.
A source build against a private repository with no FELIS_GITHUB_TOKEN,
or a wrong one, made git ask for a username on /dev/tty, and a piped
install sat there waiting.
git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git
fails at once with "terminal prompts disabled". Both fetch_source
failures name FELIS_GITHUB_TOKEN in their message: the fresh clone,
and the fetch into an existing checkout, which had no message of its
own before.
prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which
never fires on Linux: /dev/tty is mode 0666 whether or not the process
has a controlling terminal, and only opening it fails. Without a
terminal the menu was printed, the read failed with "No such device or
address", and the default was taken by accident rather than by the
documented path.
The guard now opens /dev/tty in a subshell and takes the "no terminal
for a prompt" path when that fails.
The header said to export tunables before running, but its own
`curl ... | sudo bash` entrypoint resets the environment, so an
exported FELIS_INSTALL_MODE or FELIS_NANO_LISTEN never reached the
installer. The header now shows the two forms that do arrive: the
variable named on the sudo line, or export followed by sudo -E.
The nano summary's hint for a proxy on another machine now prints a
sudo line that can be pasted as is, instead of "re-run with
FELIS_NANO_LISTEN=...". Comment and log text only.
install_go_toolchain downloaded the tarball to a fixed /tmp name and
unpacked it into /usr/local as root, with no digest check. Another
local user could plant that file first, and nothing would notice a
tampered download.
The tarball is now staged in a mktemp -d directory that the exit
cleanup removes, and its sha256 must match before the old toolchain is
touched, so a refusal leaves the host as it was. The default 1.26.4
carries pinned amd64 and arm64 digests next to its version; they are
the ones https://go.dev/dl/?mode=json&include=all publishes. Any other
FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in
the header, because no pin can cover a version chosen at run time.
Where and which version gets installed is unchanged.
summary_nano printed the node's primary IP for every non-loopback bind
and 127.0.0.1 for every loopback one. A bind to a second private
address, or to [::1], handed the operator a hasJoined URL that nothing
listens on.
The host is now the part of FELIS_NANO_LISTEN before the last ':'. The
node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]),
which names no address a proxy could dial. The loopback and
public-bind notes are unchanged.
FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in
the firewall while nano bound nothing, a bare 127.0.0.1 printed
http://127.0.0.1:127.0.0.1/... in the summary, and the unit
crash-looped either way.
validate_settings now requires a ':' and a decimal port of 1-65535
after the last one. It runs after resolve_nano_listen, so an address
read back from an existing unit is checked too, and the default always
passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
prompt_install_mode also runs inside felis setup. Setup then goes on
to the Owner and edge setup, which need the control plane, so choosing
nano there always ended in a setup error.
Under felis setup the mode is now full before any prompt or default is
considered, and an explicit FELIS_INSTALL_MODE=nano stops with a
message pointing at deploy/bootstrap.sh. That leaves the felis setup
branch of acquire_nano_binary unreachable, so it goes.
install_embedded_binary stays, since the full install still uses it.
Re-running the installer is how a nano host updates. That re-run reset
FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost
its endpoint and every login through it failed. It also offered the
full control plane as the default, which on a nano host means k3s and
Postgres nobody asked for.
The listen address is now settled by resolve_nano_listen, the first
step of main, so the later checks see the result. The operator's value
wins, then the -listen argument of the installed felis-nano unit, then
loopback. The install mode defaults to nano, at the prompt and without
a terminal, when the felis-nano unit exists and the full install's
bootstrap.done marker does not. Only the full install writes that
marker.
The harness reads back the unit it wrote earlier, and checks the mode
default on a nano-only host, a host with the full install, and a fresh
host.
A re-run copies the operator's [[auth_source]] tables from the existing
felis toml into the new one. The awk program that finds them matched
only the literal header [[auth_source]], so a table written as
[[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid
TOML, was taken for some other section and dropped from the config.
Each section header now decides afresh whether it opens an auth_source
table, through one regex that allows inner whitespace and a single- or
double-quoted key. The single quote is spelled \047, which gawk and
mawk both honour inside a bracket expression. The harness carries each
spelling and checks that the table still stops at the next section.
install_nano_service printed "enabled and started" straight after
systemctl restart, which returns as soon as the process is forked. An
upgrade that keeps an old felis.toml the new binary rejects (an
[[auth_source]] without a prefix, say) left the unit crash-looping in
auto-restart while the installer reported success, and every login
through the proxy failed.
The install now waits two seconds and asks systemctl is-active. A unit
that exited is in "activating (auto-restart)", which is-active does not
count as active; on real systemd a unit whose process exits 1 under
Restart=on-failure reads activating/auto-restart and is-active returns
non-zero, while a running one reads active/running and returns 0. On
failure the install prints the unit's last 20 journal lines and stops.
This also surfaces a nano unit locked out of an existing 0700 /etc/felis.
The harness runs the extracted function with systemctl stubbed both
ways. Without the check, the dead-unit cases fail.