FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in
the firewall while nano bound nothing, a bare 127.0.0.1 printed
http://127.0.0.1:127.0.0.1/... in the summary, and the unit
crash-looped either way.
validate_settings now requires a ':' and a decimal port of 1-65535
after the last one. It runs after resolve_nano_listen, so an address
read back from an existing unit is checked too, and the default always
passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
prompt_install_mode also runs inside felis setup. Setup then goes on
to the Owner and edge setup, which need the control plane, so choosing
nano there always ended in a setup error.
Under felis setup the mode is now full before any prompt or default is
considered, and an explicit FELIS_INSTALL_MODE=nano stops with a
message pointing at deploy/bootstrap.sh. That leaves the felis setup
branch of acquire_nano_binary unreachable, so it goes.
install_embedded_binary stays, since the full install still uses it.
Re-running the installer is how a nano host updates. That re-run reset
FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost
its endpoint and every login through it failed. It also offered the
full control plane as the default, which on a nano host means k3s and
Postgres nobody asked for.
The listen address is now settled by resolve_nano_listen, the first
step of main, so the later checks see the result. The operator's value
wins, then the -listen argument of the installed felis-nano unit, then
loopback. The install mode defaults to nano, at the prompt and without
a terminal, when the felis-nano unit exists and the full install's
bootstrap.done marker does not. Only the full install writes that
marker.
The harness reads back the unit it wrote earlier, and checks the mode
default on a nano-only host, a host with the full install, and a fresh
host.
A re-run copies the operator's [[auth_source]] tables from the existing
felis toml into the new one. The awk program that finds them matched
only the literal header [[auth_source]], so a table written as
[[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid
TOML, was taken for some other section and dropped from the config.
Each section header now decides afresh whether it opens an auth_source
table, through one regex that allows inner whitespace and a single- or
double-quoted key. The single quote is spelled \047, which gawk and
mawk both honour inside a bracket expression. The harness carries each
spelling and checks that the table still stops at the next section.
install_nano_service printed "enabled and started" straight after
systemctl restart, which returns as soon as the process is forked. An
upgrade that keeps an old felis.toml the new binary rejects (an
[[auth_source]] without a prefix, say) left the unit crash-looping in
auto-restart while the installer reported success, and every login
through the proxy failed.
The install now waits two seconds and asks systemctl is-active. A unit
that exited is in "activating (auto-restart)", which is-active does not
count as active; on real systemd a unit whose process exits 1 under
Restart=on-failure reads activating/auto-restart and is-active returns
non-zero, while a running one reads active/running and returns 0. On
failure the install prints the unit's last 20 journal lines and stops.
This also surfaces a nano unit locked out of an existing 0700 /etc/felis.
The harness runs the extracted function with systemctl stubbed both
ways. Without the check, the dead-unit cases fail.
felis-nano runs as a systemd DynamicUser, so it can read
/etc/felis/felis.toml only if others may search /etc/felis.
write_nano_config made the directory with a bare mkdir -p, which takes
its mode from root's umask. On a host hardened to umask 027 that is
0750: nano exits on "permission denied", the unit restarts every five
seconds, and no login gets through.
A missing directory is now created 0755 explicitly. An existing one
keeps its mode, because the full install sets it to 0700 to protect its
secrets and widening that from the nano path would expose them. A nano
unit locked out that way is left for the install to report.
The harness runs the extracted function under umask 027 and checks both
cases. Reverting to the bare mkdir fails the first; an unconditional
chmod 0755 fails the second. The mode checks skip on filesystems that
ignore chmod, such as Git Bash on NTFS.
write_felis_toml regenerates felis.host.toml and felis.pod.toml with a
wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin
block in that heredoc. Re-running the installer, which is also what
`felis setup` does, threw away any edit to the list: a root the operator
added stopped admitting logins, and a root they removed came back. The
generated comment invited exactly that edit.
Carry the tables forward the way [smtp] already is: read every
[[auth_source]] table from the existing felis.host.toml (falling back to
felis.pod.toml) and emit the LittleSkin default only when there is no
earlier file at all. An earlier file with no tables stays empty, because
that is a Mojang-only server rather than a missing value; felis-api now
treats an empty list that way.
The file header and the comment above the list now say what survives a
re-run, and point at felis.host.toml, which is what the next run reads.
bootstrap_test.sh extracts the new function from bootstrap.sh and checks
the fresh-install default, an operator's own table carried without the
default or the following section, an empty list staying empty, and the
indented form the setup TUI writes. It passes under dash with gawk and
with mawk; forcing the function to always return the default fails five
of the new cases.
FELIS_VELOCITY_FORK_JAR replaces the proxy every player connects through, and the
only thing checked about it was that the path pointed at a readable file. A
truncated copy, a stale build left at the same path, or the two-patch jar where
the three-patch one was meant all installed silently.
It now requires FELIS_VELOCITY_FORK_JAR_SHA256 and refuses on a mismatch, hashing
stdin rather than the path for the reason install_via_plugins already documents:
sha256sum escapes its output line for a filename carrying a backslash or newline,
and the leading "\" that adds fails every comparison. The absent-digest refusal
prints the jar's actual hash, so the first run after a deliberate rebuild is one
copy-paste rather than an investigation.
The comparison ignores case and internal spaces. The fork is built on a developer
machine, which is usually Windows, and nothing there prints a digest the way
sha256sum does: Get-FileHash returns uppercase and certutil has shipped the bytes
space-separated. Comparing raw would refuse two of the three spellings of the
correct answer and word the refusal as tampering.
No digest is hardcoded, which is the half of the request this does not deliver.
The fork is built from Felis-Legacy and has never been reproduced on a second
machine, so a constant here would pin one machine's output rather than the fork.
The comment that previously asserted the build "is not byte-reproducible" is gone
too -- it was stated more confidently than the evidence supports. The fork jars on
disk carry Gradle's constant 1980-02-01 entry timestamps, so the usual reason a
jar differs between builds is already absent; that is not proof it reproduces, and
neither claim should sit in the script unmeasured.
This is deliberately not a supply-chain signature and the comment says so: an
operator who can write the jar can write the digest. What it buys is that a path
stops being an identity, and that every later re-run re-checks the same build.
Scope: the fork jar only. The else branch still curls stock Velocity from PaperMC
with no verification at all, and that is the branch a default install takes. The
digest is already in hand there -- Fill v3 returns checksums.sha256 and its
download URL is content-addressed on that same value -- and papermc_latest_jar
discards it. Left alone rather than widened into this change.
deploy/bootstrap_test.sh covers the gate's two refusals, its happy path, and the
two Windows digest spellings. Each case extracts the block under test out of
bootstrap.sh with awk and runs it with die/log stubbed, rather than transcribing
it -- a transcribed copy passes forever after someone edits the original. The
extraction is length-bounded: awk runs an unmatched end pattern to EOF, which
would quietly feed the rest of bootstrap.sh to the shell under test. bootstrap.sh
itself cannot run here; it wants root, a package manager and k3s.
A `shell` CI job runs that plus a syntax check over every tracked script. The
syntax step dispatches on each file's shebang instead of running `sh -n` across
the board. The blanket form looks fine and is a false green: on a developer
machine `sh` is usually bash and accepts everything, while the runner's `sh` is
dash. Verified against the real thing rather than an approximation -- inside
ubuntu:24.04, where /bin/sh is /usr/bin/dash, the dispatching loop passes all six
scripts and the blanket loop dies at bootstrap.sh:191 on the first of its 14
arrays.
Refs: Felis-Legacy #19