build_image_from_binary built on distroless/base-debian12 while the repo
Dockerfile's final stage uses distroless/static-debian12, so the image an
install runs did not match the image CI publishes.
Every binary that can reach HOST_BIN traces back to the Dockerfile's
CGO_ENABLED=0 build -- the downloaded CI asset, the binary the TUI is already
running, and the one build_image_from_source docker-cp's out of the image it
just built. None link glibc, so base-debian12 bought nothing and only widened
the runtime surface.
This mattered little while build_image_from_binary was the rare fallback.
659c8e5 made the release channel download a binary and wrap it here, which
makes this the image most installs actually run.
deploy/bootstrap.sh now resolves the newest published GitHub release, downloads
the binary CI built for that tag, and builds a thin image around it. Compiling
on the target host becomes the fallback and the opt-in, not the default.
The panel is not a separate artifact. The Dockerfile copies panel/dist into
internal/panel/static before the go build, so the control plane — panel and
backend — ships as ONE file. The release channel therefore downloads exactly
one asset, felis-linux-<arch>, and needs no registry, no Go toolchain and no
checkout on the host.
The Minecraft game stack (limbo, lobby, the Velocity plugin) is still always
built locally. game_stack_source now keys on HAVE_PREBUILT_BINARY — the same
flag build_image uses — so on any prebuilt path it unpacks the tar embedded in
that binary instead of trusting a checkout an earlier install left behind.
Trusting the checkout would build the plugin from an old commit against a
freshly downloaded control plane: a silent version skew across the plugin/API
boundary.
Channels:
(default) newest published release, downloaded
FELIS_VERSION_BOOTSTRAP=dev clone main and compile
FELIS_REF=<ref> pins the tree, forces the source path
The download is best-effort. A tag whose assets are not uploaded yet, an
architecture with no published asset, or an asset that fails validation each
warn and fall back to compiling THE SAME TAG from source — never a different
commit.
The ref is resolved right after install_base, the first point curl exists and
well before docker and k3s, so a missing FELIS_GITHUB_TOKEN or an unpublished
release costs the operator seconds instead of a k3s install they then have to
unwind. It is skipped on exactly the paths that never consume the result: the
TUI, which rebuilds the binary it is already running, and FELIS_SKIP_FETCH,
which builds whatever is staged. Resolving anyway would set FELIS_VERSION to
the newest tag and stamp a staged tree as that release.
The asset is staged next to HOST_BIN rather than in TMPDIR. Validation EXECUTES
it, and /tmp is noexec on CIS-hardened images, where the exec dies 126, the
check reads it as a bad asset, and every such host silently falls back to the
full on-host compile this path exists to avoid. It also keeps a private-repo
artifact out of a world-readable 1777 directory.
git_auth, which supplies the token to git for a private-repo clone, passes an EMPTY
credential.helper before the inline one. credential.helper is multi-valued: a bare
`-c credential.helper=...` APPENDS to whatever the host has configured rather than
replacing it, and an empty value is git's documented list reset. Without it, on a host
with a persistent helper (Git for Windows ships `manager` at SYSTEM scope) two things
go wrong. Git runs `credential approve` automatically after a successful clone and
feeds every helper in the list, so a `store` helper writes the PAT to
~/.git-credentials in cleartext — the token outlives the install, in a file bootstrap
never created and never cleans up. And because the inline helper is LAST, a
pre-existing helper answers `fill` first, so a stale cached credential can win and the
clone authenticates as the wrong account — surfacing as exactly the 404-on-private-repo
the surrounding code works hard to explain. Reproduced both against a real clone, and
confirmed the reset closes both.
internal/panel parses the new stamp. The dev channel now emits "<tag>+g<sha>", which
matched neither describeSuffix ("-N-g<sha>") nor releaseTag, so a dev build fell through
to the default case and the version badge rendered the entire stamp as the release with
no commit. A devSuffix case handles it; the git-describe case stays for hand-rolled
`-ldflags "-X main.version=$(git describe)"` builds. Table test covers both forms plus
the release, dirty and unstamped cases.
CRD application no longer branches on the install path: it is always
`felis bootstrap-assets crd`. That output is byte-identical to deploy/crd/ —
bootstrap_asset.go embeds that very file — and needs no checkout, so one source
replaces a branch whose two arms had to be kept in agreement by hand.
Dockerfile gains a FELIS_VERSION build arg wired into -X main.version, declared
after `go mod download` so a version bump does not invalidate that layer. Both
build stages are pinned to $BUILDPLATFORM so a multi-platform buildx run never
emulates them: the panel's output is architecture-independent and the Go stage
cross-compiles via TARGETARCH. The final stage stays on the target platform and
is COPY-only, which BuildKit performs without QEMU.
.github/workflows/release.yml publishes on a vX.Y.Z tag: vet, tests, then one
buildx run producing both architectures through the repo Dockerfile. Not a bare
`go build` — internal/panel/static holds a tracked placeholder index.html so the
//go:embed compiles without node, which means a direct build succeeds and
quietly ships a release whose panel is that placeholder.
The stamp is asserted end to end, because it fails silently: an unstamped binary
reports "dev", which the updater refuses to compare, disabling update reporting
for every install built from that release. The arm64 artifact is checked by ELF
machine type rather than by running it — runners have binfmt registered, so
executing an amd64 binary misnamed arm64 would succeed.
Prerelease tags are flagged explicitly. The trigger glob is v*, gh does not read
semver out of a tag name, and an RC published as a full release becomes
/releases/latest — the single endpoint the default channel installs from and
`felis update` polls.
No SHA256SUMS. A checksum fetched over the same TLS session, with the same
credential, from the same host as the binary adds no trust root; signing is the
real answer and is a separate decision.
Not verified: the download -> validate -> image -> k3s path has never run on a
host against a real published release, because no tag exists yet. The shell
logic around it is verified out of tree; the network and exec behaviour is not.
restart_existing_control_plane ended in an and-list per deployment:
[ "$had_api" = "1" ] && kube ... rollout restart deployment/felis-api
[ "$had_operator" = "1" ] && kube ... rollout restart deployment/felis-operator
As the LAST command of a function, an and-list whose test is false returns 1,
and that becomes the function's exit status. The call site is bare, so under
`set -Eeuo pipefail` the installer dies there — after the bundle has been
applied and before the rollout wait, leaving a half-finished upgrade and no
message naming the cause.
It fires on any host carrying one control-plane deployment but not the other:
felis-api present without felis-operator restarts the api, then exits 1 on the
second test. Both present, or neither, happened to work, which is why it
survived.
Rewritten as explicit `if` statements, which return 0 when the test is false.
Verified out of tree against all four had_api/had_operator combinations.
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.
Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
handlers; align tests (new account subtests, naming reserves
"console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
setup_panel.go, tui_root.go and tui_preflight.go instead of
hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.
Panel:
- Delete the ChangePassword page and all password UI; align
login/auth/api/types with the passwordless contract; add the
migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
ServerLuckPerms strings to translation keys; drop 72 orphan keys
per locale; unify the title as "Felis - Console".
Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
preference; unified /link lines in fabric/forge/neoforge; shared
link-client javadoc contract fixes.
Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.
BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
A single HTTP 502 from fill.papermc.io aborted the entire bootstrap. Build
resolution used a one-shot curl, so one gateway blip from an upstream that
flaps was indistinguishable from a permanent failure, and the run died before
Docker, k3s or any game server was provisioned.
Pass --retry 5 --retry-delay 2 to the build-resolution fetches. 502/503/504
are already in curl's built-in transient set, so the tool had solved this; the
flags were simply never passed. papermc_latest_jar is shared by the Paper and
the Velocity resolve, so hardening it once covers both callers. The
LOOHP/Limbo CI metadata fetch feeds the same step and gets the same treatment.
Deliberately no --retry-connrefused. It only adds ECONNREFUSED to a set that
already covers this incident, and it needs curl 7.52.0 while the yum (el7)
path the script supports ships 7.29.0, where an unrecognised long option is a
parse error rather than a warning:
# centos:7, curl 7.29.0
$ curl -fsSL --retry 5 --retry-delay 2 --retry-connrefused https://example.com
curl: option --retry-connrefused: is unknown
Under set -Eeuo pipefail that exits 2 and trips the || die, so both hardened
fetches would hard-fail on a host where they used to work, each naming a cause
that is not the real one. A comment above papermc_latest_jar records this so
the flag does not come back.
The failure message was actively misleading. "no Paper build for Minecraft
26.2 (the login gate speaks only that protocol)" reads as "that Minecraft
version is unsupported", sending the reader after a version-pinning problem
that does not exist: Paper 26.2 build 60 resolved fine minutes later. Say what
is actually known instead, that the build likely exists and Fill is flapping.
Verified against a local always-502 server: curl now issues 6 requests
(1 initial + 5 retries) over 10.1s before giving up, where it previously
issued 1 and died. Verified on centos:7 that this flag set is accepted, and
against the live Fill v3 API that the resolve still returns a jar URL.
Known and deliberately unchanged: no fetch sets --max-time, so an upstream
that accepts a connection and never answers still blocks forever. --retry does
not cover that, as it fires only once a request completes with a failure. The
remaining single-shot downloads (cloudflared, the Docker GPG key and repo
list, k3s, the Temurin JRE, the Velocity jar, the Go toolchain) keep their
existing no-retry shape rather than widen this diff on a script that is about
to provision a live host.
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).
Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.
Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.
The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.
Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
Point Velocity's authlib (mojang.sessionserver) at the felis-api hasJoined multiplexer so
a full install federates Mojang plus the configured [[auth_source]] set out of the box,
not just the standalone `felis nano`, and ship a default LittleSkin auth_source in the
generated felis.toml (delete the block for a Mojang-only server). Velocity now owns
velocity.toml and its working tree — it migrates the config version and extracts
localizations on start — while the jar and forwarding secret stay root-owned read-only and
ReadWritePaths widens to VELOCITY_DIR; the felis-api-internal ClusterIP lookup is factored
into a felis_internal_ip helper shared by the link config and the sessionserver override.
Re-include plugins/velocity in the docker context because the felis binary now embeds all
four plugin trees for `felis bootstrap-assets game-stack`.
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.
The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.
A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.
The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.
prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.
Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.
Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:
premium FLYEMOJ1 -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
LittleSkin FLYEMOJ1 -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
both online at once, zero "already connected" rejections
LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3
The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.
Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.
BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
d417efc got this backwards, in both the code comment and the installer summary.
It claimed authlib appends /session/minecraft/hasJoined itself, so the property
should be given the base URL only. Velocity does not work that way, and a real
login says so: pointed at http://127.0.0.1:8081, a Mojang login arrives at nano
as
GET /?username=FLYEMOJ1&serverId=-23ae0b50...
with no path at all. Velocity appends the query string to the property verbatim
and issues the request itself; authlib is not in the loop. nano has no route on
/, so it answers 404 and Velocity kicks the player with authservers_down.
Velocity's own default for the property is the full URL,
https://sessionserver.mojang.com/session/minecraft/hasJoined, which is the same
thing said another way.
With the full endpoint URL the same account logs straight in, so both the
nano.go header and summary_nano now print
-Dmojang.sessionserver=http://127.0.0.1:8081/session/minecraft/hasJoined
and note that the flag belongs between `java` and `-jar`.
Verified against Velocity 3.5.1 + Paper 26.2 on the deploy host: a Mojang login
reaches the backend with its real Mojang UUID unchanged, and a LittleSkin login
under the same username reaches it as UUIDv3(felisAuthNS, "littleskin:"+id) --
two different players on the backend, which is the point.
Deploying `felis nano` to a real Rocky Linux 10 host surfaced four defects that
no local check could see. Fixed together because they all sit on the same path
from `curl|bash` to a running felis-nano.service.
* docker killed the nano install on EL10. `acquire_nano_binary` pulled in
docker purely to build the binary; on Rocky 10.2 the docker-ce el10 rpms
install but dockerd refuses to start, so the install died at
`systemctl enable --now docker`. nano needs one static binary, not an image,
so the docker dependency is gone: fetch_source -> install_go_toolchain
(pinned FELIS_GO_VERSION, default 1.26.4, amd64/arm64) -> build_nano_binary.
* the built binary could not be exec'd by systemd (203/EXEC). The Go linker
renames its output out of $TMPDIR, and a same-filesystem rename carries the
source SELinux label, so `go build -o /usr/local/bin/felis` produced a binary
labelled user_tmp_t rather than bin_t. root is unconfined and could run it by
hand, which is what made this look fine, but the DynamicUser service could
not. build_nano_binary now stages the output and installs it as a fresh file
so the policy type transition labels it bin_t, with restorecon as a belt.
* re-running the installer did not converge. `systemctl enable --now` is a
no-op on an already-active unit, so a rebuilt binary was installed while the
old process kept running. Now enable + restart.
* the Velocity wiring comment in cmd/felis/nano.go was wrong. authlib appends
/session/minecraft/hasJoined itself, so -Dmojang.sessionserver takes the base
URL only, as the installer has always printed.
Also bind to loopback by default. hasJoined is unauthenticated by protocol --
authlib speaks the vanilla sessionserver dialect and sends no token -- so a
public bind is an open auth relay: anyone can point their own proxy at it and
spend this host's egress IP on Mojang until Mojang rate-limits it and the
operator's own players stop getting in. It is not an identity bypass (a caller
still needs a serverId hash bound to their own server key, which the upstream
Yggdrasil validates), but it is someone else's traffic on your address.
FELIS_NANO_LISTEN and the -listen flag now default to 127.0.0.1:8081, which a
same-host Velocity reaches unchanged; serving an off-host proxy is an explicit
opt-in. configure_nano_firewall no longer opens a port for a loopback bind, and
summary_nano prints the real bind address plus the relay warning.
Verified on the target host: installs with no docker present, service active,
binary labelled bin_t, `ss` shows LISTEN 127.0.0.1:8081, an external request is
unreachable, and an in-host request returns 204 with the login logged.
BREAKING CHANGE: felis nano defaults to 127.0.0.1:8081 instead of 0.0.0.0:8081.
A Velocity proxy on another machine must now set FELIS_NANO_LISTEN (or -listen)
to a reachable address, and should allow that port only from the proxy's IP.
bootstrap.sh now asks up front whether to install the full Felis control
plane or only Felis-nano, and grows a parallel install path for the
nano-only case.
- prompt_install_mode() runs right after OS detection and reads /dev/tty
(so it works under `curl ... | sudo bash`) offering [1] Felis / [2]
Felis-nano, default full. FELIS_INSTALL_MODE=full|nano skips the prompt
for non-interactive runs; no tty falls back to full.
- main_nano() installs only what nano needs: the felis binary (reusing
the embedded-binary / docker-build acquisition), a template felis.toml
carrying a commented [[auth_source]] example (Mojang-only until edited),
a felis-nano.service unit running `felis nano -config ... -listen ...`
under DynamicUser hardening, and a firewalld port-open for the listen
port. None of the k3s / Postgres / migrate / bundle steps run.
- write_nano_config is idempotent (leaves any existing config untouched)
and its template is valid as-is. summary_nano prints the hasJoined
endpoint and the Velocity -Dmojang.sessionserver flag, offering the
127.0.0.1 form when the proxy is on the same host.
Verified on WSL: `bash -n` clean; the emitted template loads via
config.LoadNano and the exact systemd ExecStart command serves 204 on a
miss ("Mojang + 0 third-party source(s)"); a duplicate-tag config still
exits non-zero citing "unique". Not exercised: a full main_nano run,
systemd activation of the unit, and shellcheck (unavailable in this env).
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.
Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).
Manifest-level fix; the live packet path is pending real-cluster verification.
demo-up.sh collapses bootstrap -> build+import the limbo/lobby images -> wire [velocity] login_image/lobby_image into felis.host.toml -> felis setup into a single command, ending in the interactive Owner-creation TUI (the only step it cannot automate). Prefers prebuilt tars under deploy/images, else builds on the host, auto-resolving the LOOHP/Limbo CI jar and the latest stable Paper jar (all overridable by env); SKIP_BOOTSTRAP/SKIP_SETUP toggles for reruns.
The lobby image had never been built and two defects blocked it: the felis image .dockerignore excluded plugins/* and only re-included limbo/shared, so the lobby Dockerfile's COPY plugins/paper landed empty; and the plugin stage used eclipse-temurin:21-jdk, which ships no gradle (and the tree vendors no wrapper), failing with 'gradle: not found'. Re-include plugins/paper and build the paper plugin on gradle:8.14-jdk21, matching the limbo image. Verified: both images build and boot (limbo /healthz 200 on 25565; lobby reaches 'Done' with felis-paper enabled).
deploy/limbo assembles LOOHP/Limbo from its loose CI artifacts plus the felis-limbo plugin (and the shared link core), with an entrypoint that pins server-port to the operator's GamePort (25565) on every start. deploy/lobby carries the Paper + felis-paper hub image. .dockerignore re-includes plugins/limbo and plugins/shared so the plugin image build sees them.
StartupSpec.HealthHTTPPort/Path switch pod readiness from plain-TCP to an HTTP GET for RCON-less loaders (LOOHP/Limbo) that report 'started' only after the first tick. User servers now default FallbackServer to the login gate, never the lobby, so a stopped/starting backend keeps authentication in front of a fresh connection.
Wire the fourth mandated §23 metric to a real producer. The histogram
spans two reconcile passes, so anchor and observation must persist in
status:
- Add status.startRequestedAt, set once on the first Starting reconcile
of a start attempt and cleared on Stopped so the next start re-anchors.
- Observe felis_start_duration_seconds exactly when readiness is first
reached (ReadySignalAt - StartRequestedAt), guarded so a server that
reaches ready without a Starting pass records nothing.
- Mirror the field into the deepcopy and the structural CRD schema so the
apiserver does not prune it on patchStatus round-trips.
- Promote prometheus/client_golang and client_model to direct deps now
that the operator and its tests import them.
Tests drive a step clock through Starting -> Running asserting the exact
observed duration, and through Running -> Stopped asserting the metric is
observed once and the anchor clears.
bootstrap.sh auto-detects the host package manager (apt/dnf) and installs whatever is missing: Docker, k3s, and PostgreSQL. It builds and imports the felis image, opens pg_hba to the pod CIDR, runs migrations, and applies the rendered control-plane bundle, leaving Web disabled pending 'felis setup'. The Dockerfile builds the distroless felis image; deploy/crd holds the MinecraftServer CRD.