Commit Graph
51 Commits
Author SHA1 Message Date
flyemoji 5d4f3063a9 feat(operator): make any Paper image joinable behind the forwarding proxy
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed
handshake does not degrade -- it rejects every login the proxy forwards. Until now the
only backends that could verify it were the two images Felis builds itself
(deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own
entrypoints. An arbitrary Paper image a user brings does not, so it passed admission,
started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the
lobby image as a base for a user's own world (0018_recommended_images.sql), which was
never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining
player away, the exact opposite of a server you mean to stay on.

The fix configures forwarding from OUTSIDE the image instead of requiring it inside.
The operator now injects a root `felis init-forwarding` initContainer into every user
server; it writes the proxies.velocity block into config/paper-global.yml and forces
online-mode=false in server.properties on the /data PVC before the main container
starts. The image needs no forwarding logic of its own, so the joinable set stops being
"images that self-configure forwarding" and becomes every Paper-family image the
platform runs.

buildStatefulSet gates the injection on the ABSENCE of the system-role label: the
Felis-built system servers already consume the secret in their entrypoints and the
login gate is a limbo, not Paper. It is also gated on a non-empty felis image name --
the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it
skips the injection rather than failing, because a cluster whose proxy is not in modern
mode has nothing to configure.

The init runs as root deliberately. The world volume's ownership comes from the storage
provisioner and the main container runs as whatever UID its image declares, so root is
the only UID that can reliably write these files; it then chmods them 0666/0777 so that
non-root main container can rewrite them on boot. The privilege is bounded -- the init
exits before the server container starts and the server container keeps its own UID.
The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the
init ever stops running as root.

The writer merges rather than overwrites, both because Paper expands paper-global.yml to
its full default tree on first boot and because the panel file editor may edit either
file between boots. It sets proxies.velocity.* and the single online-mode key and leaves
every other setting alone. It is a no-op on an empty secret, for the same reason the env
var is optional: a proxy that is not in modern mode provisions no Secret, and wedging
every server's init on a missing optional value would be worse than the status quo.

felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and
0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu
plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the
online-player list and permission commands work out of the box. 0018's row is left in
place -- an admin who kept it can keep it; this only adds the better default beside it.

Three fixes ride along, each of which the 1.8 path hit in practice.

bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only
builds its block-connection provider when Via's lowest supported protocol is below 1.13;
under modern forwarding the Velocity injector reports 393, so init() returns early,
blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences
it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining.
Every call site is behind isServersideBlockConnections(), so switching it off skips all
of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected.
ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with
this one key suffices -- Config#loadConfig parses the bundled default as the base map and
merges the on-disk file over it, so every other option stays current across version
bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this
worked: init() gates on the protocol version too, and that half fails on its own, so the
line is missing either way. The config value is the only evidence, which is what the test
asserts.

The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend
sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it
down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13
and has nowhere to go. Only the handshake address field survives Via, so the Felis fork
forwards the named servers BungeeCord-style while every other backend keeps modern+secret
untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer
CRs is the upgrade path.

deploy/lobby's set_prop escapes the value before substituting it. The RCON password is
operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare
`sed s|...|...|` and silently kills the key -- taking the console, the online-player list
and permission commands with it. deploy/paper was written with the escaping, so the lobby
gets the same rather than leaving the sibling caller broken.

Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where
TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping.
The new tests cover the initContainer's image, root UID, world mount and secret env; the
merge preserving unrelated config trees; the properties upsert including the commented-key
case; and the bootstrap script both writing the Via key and still calling the function
that writes it.

Not verified: the initContainer has never run in a real cluster, and the felis-paper
image is code-only here as the other game-stack images are -- no Go CI builds them.

The ViaVersion pin is the one piece with live evidence, and that evidence is what it was
written from. Before it, a client was cut within a second of "logged in with entity id"
on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into
Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on
2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same
player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login
from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes.

Neither session says which client version it was. The proxy never logged a protocol
number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16
players and below" warning that fired for that player on the lobby leg, and no lower --
Via floors every handshake to the proxy's 393, so anything from 47 up is admissible.
Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain
runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the
NPE proves is that the pin was load-bearing, not who was holding the mouse.

That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one
now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8
behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the
fault above; off, holds. No automated test in THIS repository exercises the leg.
2026-07-28 09:13:58 +09:00
flyemoji a0064adfdd fix(setup): stop the login gate sending players to the old console after a re-domain
Changing root_domain updated felis.toml and the panel, but the login gate kept
pointing players at the hostname it was created with. FELIS_ROOT_DOMAIN and
FELIS_PANEL_HOSTNAME are baked into the login MinecraftServer at provisioning time,
FelisLimboPlugin reads them to build the link an unauthenticated player is told to
open, and ensureSystemServers is create-if-absent — so nothing in the install ever
rewrote them. On the demo host the CR still carried
console.159.223.32.51.nip.io hours after the domain had moved to
mc.flyemoji.network: every joining player was handed a link that bypasses the
tunnel, hits the node directly and trips a certificate warning, on the one screen
someone with no account is guaranteed to see.

setup now converges these values on an existing system service instead of skipping
it. Create-if-absent stays the rule for everything else, and the comment on it is
still true — an operator's edits to a system service must survive a re-run. These
three names are the exception because they are not the operator's to own: they are a
copy of config that is wrong the moment config changes, and there is no other writer
who could notice.

The convergence is deliberately narrow. Only a name already present with a different
value is rewritten, so env the operator added by hand is untouched and the rest of
the spec — image, memory, storage — is not read at all. A derived name that is
absent from the live object is left absent rather than added back: a deliberate
removal and drift look identical from here, and re-adding it would mean fighting the
operator on every run. The outcome string reports the refresh so a setup run does not
silently rewrite the front door.

This closes one surface of a re-domain, not the whole of it. The write-once panel
certificate at deploy/bootstrap.sh keeps its old SANs, and so do the Velocity config
and the forwarding material; the warning in bootstrap.sh that says so is still
accurate. What changes is that the surface players actually walk through now catches
up when setup is re-run.

Checks: a login gate built with the old domain converges onto the new one and says
so; an env var the operator added and a hand-raised javaMemory both survive that
same run; and an install whose config already matches reports no refresh, so a
routine setup does not read like a re-domain. The middle one is the one worth having
— converging config must not turn into a licence to clobber the edits
create-if-absent exists to protect.
2026-07-21 00:47:04 +09:00
flyemoji 694e3cb800 feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.

The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.

Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.

RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.

Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.

That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.

Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.

Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.

Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
2026-07-21 00:12:37 +09:00
flyemoji b3989fa4af fix(mail): prove SMTP deliverability before saving, and stop losing the relay
A live install passed the SMTP setup screen and then failed every one-time
code with a bare `internal error`. Four separate defects had to line up for
that, and each is fixed here.

The relay was configured with `from = noreply@<domain-A>` on an account
authenticated as `<user>@<domain-B>`. Providers that validate sender identity
— Fastmail among them — answer MAIL FROM with an unconditional 250 and only
refuse at end-of-DATA. Ping stopped at NOOP, so it never saw the refusal: the
wizard reported success, wrote the config, rolled felis-api, and every OTP
afterwards died at w.Close().

Ping now runs the same transaction a real code takes — connect, (STARTTLS,)
AUTH, MAIL FROM, RCPT TO, DATA — delivering one self-test message to the From
address, and SendOTP and Ping share deliver() so the check cannot drift from
the thing it checks. The self-test recipient cannot cause a false negative:
an authenticated submission relay accepts RCPT for any destination by
definition, while the sender identity it does validate is exactly what we
want tested. The setup screen now says a message will be sent, names the
address it went to, and warns that From must be an address the account is
allowed to send as.

A relay refusal also answered 500 `internal`, which reads as a broken panel
and sends the operator hunting through handler code instead of their [smtp]
block. It is now 502 `mail_undeliverable`, mapped inside deliverOTP so all
four doors that mail a code (onboarding, email login, op-login, migrate
step-up) answer alike. The relay's own text stays out of the response — it
can name the SMTP account, and these routes are reachable by any signed-in
player — and goes to the log instead.

writeError logged nothing when it collapsed an unmapped error to 500, so an
operator holding an `internal error` had nothing to grep for and diagnosis
degraded into guessing against a live install. It now logs the method, path,
wrapped chain and the same request_id the caller is shown.

Finally, write_felis_toml regenerated the config wholesale and never emitted
[smtp], so re-running the installer — the documented way to update felis-api —
silently erased a working relay and reverted OTP delivery to the no-Mailer
path, logging codes instead of sending them. It now carries the block forward,
cached on first read because the host toml is clobbered before the pod toml is
written. Same defect family as the root_domain loss fixed in ecbeb20: a
generated file holding a hand-set value with no carry-forward.

Tests cover the case a MAIL FROM probe cannot see: a fake relay that answers
250 to MAIL FROM and 550 at end-of-DATA must fail both Ping and SendOTP, and
the 502 must carry a distinct machine code without leaking the relay's text.
2026-07-21 00:11:40 +09:00
flyemoji ecbeb20761 fix(bootstrap): reuse the installed root domain instead of re-deriving it
detect_node_ip recomputed FELIS_ROOT_DOMAIN from scratch on every run and fell
back to <node-ip>.nip.io. Nothing read the domain back out of the felis.toml an
earlier run wrote, so it survived only as long as the operator kept passing the
same environment.

That made re-running the installer destructive on any install with a real
domain, and re-running it is not optional: it is the only way to move felis-api
to a newer release, which is what `felis update` points operators at. A bare
re-run rewrote root_domain, panel_hostname and admin_hostname to nip.io names
while ensure_panel_tls_cert returned early on the certificate it had already
written, leaving the console serving a cert for hostnames it no longer answered
to -- with no re-domain flow to recover through.

Precedence is now explicit FELIS_ROOT_DOMAIN, then the domain the last run
persisted, then the nip.io default. First installs are unaffected. Deliberate
re-domains still work, because there is no other route to one, but they now warn
that the write-once certificate is not reissued and that the proxy and login
config carry the old name too.

Secrets were never exposed to this: load_or_make_secrets has always sourced
secrets.env before generating anything. The domain was the one piece of install
identity with no read-back.

The channel is deliberately left alone. FELIS_VERSION_BOOTSTRAP is not persisted
either, but defaulting a re-run to the release channel installs a working build
rather than breaking one, so cmd/felis/update.go states that instead. Its
warning about the domain went with the bug and would now be false.

Verified against the shipped function text: the ladder holds for a fresh host,
a re-run with and without the variable set, a re-domain, and a felis.toml whose
root_domain is missing or empty. Reverting the one line reproduces the nip.io
overwrite.
2026-07-20 21:20:08 +09:00
flyemoji f7815629bf fix(update): say that setup cannot move felis-api to a newer release
`felis update` offers `sudo felis setup` for every planner-backed target and
closed with a trailer calling setup idempotent. That is true for velocity --
install_velocity re-resolves the newest build of the pinned minor on each run --
and misleading for felis-api, which both --panel and --plugins resolve to.

setup hands bootstrap the binary it is itself running and takes the
bootstrap_from_tui arm, which skips the release lookup. The run rebuilds the
image and rolls the deployment off that SAME binary: it reports success and
leaves the version exactly where it was. An operator following this guidance to
apply a felis-api update would watch it appear to work and then see the same
version reported again.

Only the bootstrap installer moves felis-api, and naming it is where this gets
dangerous, so the warning ships with it. The installer is not an updater. Every
run re-derives FELIS_ROOT_DOMAIN through detect_node_ip and defaults it to
<node-ip>.nip.io; nothing reads the domain back out of the felis.toml an earlier
run wrote. A bare re-run -- which is exactly what README documents, with no
environment at all -- rewrites root-domain, panel-hostname and admin-hostname to
nip.io names, while ensure_panel_tls_cert returns early on the certificate it
already wrote and keeps serving the old hostnames. The console then fails to
match its own certificate, and there is no re-domain flow to recover with.
Persisted secrets are not at risk: load_or_make_secrets sources secrets.env
before it generates anything.

felis update stays report-only, so no command changed; only the claim about what
the offered one accomplishes, and the conditions under which the alternative is
safe to run.

Tested four ways, because the scoping and the warning are both the point:
--panel carries the caveat AND names FELIS_ROOT_DOMAIN, --velocity keeps the
ordinary trailer without either, and --mc, which offers no command at all, gets
neither.
2026-07-20 21:20:04 +09:00
flyemoji 8b25114fb4 fix(update): say that setup cannot move felis-api to a newer release
`felis update` offers `sudo felis setup` for every planner-backed target and
closed with a trailer calling setup idempotent. That is true for velocity --
install_velocity re-resolves the newest build of the pinned minor on each run --
and misleading for felis-api, which both --panel and --plugins resolve to.

setup hands bootstrap the binary it is itself running and takes the
bootstrap_from_tui arm, which skips the release lookup. The run rebuilds the
image and rolls the deployment off that SAME binary: it reports success and
leaves the version exactly where it was. An operator following this guidance to
apply a felis-api update would watch it appear to work and then see the same
version reported again.

The report now says so, scoped to runs that actually offered a felis-api
target, and points at the bootstrap installer -- the path that resolves and
downloads a release. felis update stays report-only, so no command changed;
only the claim about what the offered one accomplishes.

Tested three ways, because the scoping is the whole point: --panel carries the
caveat, --velocity keeps the ordinary trailer without it, and --mc, which
offers no command at all, gets neither.
2026-07-20 19:53:55 +09:00
flyemoji 8675cda001 fix(setup): refuse --dev rather than silently installing the release channel
`felis setup --dev` promised "install the dev channel (main HEAD)" and
installed release: the flag only exported FELIS_CHANNEL, a variable nothing in
the tree reads. deploy/bootstrap.sh reads FELIS_VERSION_BOOTSTRAP.

Renaming the variable would have been a worse bug than the dead one, because it
would look wired. setup runs bootstrap with FELIS_BOOTSTRAP_FROM_TUI=1, and on
that arm every reader of FELIS_VERSION_BOOTSTRAP is unreachable: the channel
case and its validation live in resolve_install_ref, which the TUI path skips
outright, and use_release_binary is only consulted by the elif that
`if bootstrap_from_tui` already short-circuited. setup re-images the host from
the felis binary it is itself running; there is no channel to pick.

So the flag refuses, exits 2 and names FELIS_VERSION_BOOTSTRAP=dev on the
installer, which is the mechanism that does work. Refusing beats defaulting:
the operator asked for dev, and release is the one answer they did not want.
The refusal precedes the root check, or an unprivileged operator gets told
about sudo instead of about the channel.

channelName had no other caller and goes with it. Nothing else referenced
--dev -- no doc, no script, no test -- so this removes a promise the tree only
ever made to itself.
2026-07-20 19:53:32 +09:00
flyemoji 659c8e5e9f feat(bootstrap): install the published release build instead of compiling on the host
deploy/bootstrap.sh now resolves the newest published GitHub release, downloads
the binary CI built for that tag, and builds a thin image around it. Compiling
on the target host becomes the fallback and the opt-in, not the default.

The panel is not a separate artifact. The Dockerfile copies panel/dist into
internal/panel/static before the go build, so the control plane — panel and
backend — ships as ONE file. The release channel therefore downloads exactly
one asset, felis-linux-<arch>, and needs no registry, no Go toolchain and no
checkout on the host.

The Minecraft game stack (limbo, lobby, the Velocity plugin) is still always
built locally. game_stack_source now keys on HAVE_PREBUILT_BINARY — the same
flag build_image uses — so on any prebuilt path it unpacks the tar embedded in
that binary instead of trusting a checkout an earlier install left behind.
Trusting the checkout would build the plugin from an old commit against a
freshly downloaded control plane: a silent version skew across the plugin/API
boundary.

Channels:

    (default)                      newest published release, downloaded
    FELIS_VERSION_BOOTSTRAP=dev    clone main and compile
    FELIS_REF=<ref>                pins the tree, forces the source path

The download is best-effort. A tag whose assets are not uploaded yet, an
architecture with no published asset, or an asset that fails validation each
warn and fall back to compiling THE SAME TAG from source — never a different
commit.

The ref is resolved right after install_base, the first point curl exists and
well before docker and k3s, so a missing FELIS_GITHUB_TOKEN or an unpublished
release costs the operator seconds instead of a k3s install they then have to
unwind. It is skipped on exactly the paths that never consume the result: the
TUI, which rebuilds the binary it is already running, and FELIS_SKIP_FETCH,
which builds whatever is staged. Resolving anyway would set FELIS_VERSION to
the newest tag and stamp a staged tree as that release.

The asset is staged next to HOST_BIN rather than in TMPDIR. Validation EXECUTES
it, and /tmp is noexec on CIS-hardened images, where the exec dies 126, the
check reads it as a bad asset, and every such host silently falls back to the
full on-host compile this path exists to avoid. It also keeps a private-repo
artifact out of a world-readable 1777 directory.

git_auth, which supplies the token to git for a private-repo clone, passes an EMPTY
credential.helper before the inline one. credential.helper is multi-valued: a bare
`-c credential.helper=...` APPENDS to whatever the host has configured rather than
replacing it, and an empty value is git's documented list reset. Without it, on a host
with a persistent helper (Git for Windows ships `manager` at SYSTEM scope) two things
go wrong. Git runs `credential approve` automatically after a successful clone and
feeds every helper in the list, so a `store` helper writes the PAT to
~/.git-credentials in cleartext — the token outlives the install, in a file bootstrap
never created and never cleans up. And because the inline helper is LAST, a
pre-existing helper answers `fill` first, so a stale cached credential can win and the
clone authenticates as the wrong account — surfacing as exactly the 404-on-private-repo
the surrounding code works hard to explain. Reproduced both against a real clone, and
confirmed the reset closes both.

internal/panel parses the new stamp. The dev channel now emits "<tag>+g<sha>", which
matched neither describeSuffix ("-N-g<sha>") nor releaseTag, so a dev build fell through
to the default case and the version badge rendered the entire stamp as the release with
no commit. A devSuffix case handles it; the git-describe case stays for hand-rolled
`-ldflags "-X main.version=$(git describe)"` builds. Table test covers both forms plus
the release, dirty and unstamped cases.

CRD application no longer branches on the install path: it is always
`felis bootstrap-assets crd`. That output is byte-identical to deploy/crd/ —
bootstrap_asset.go embeds that very file — and needs no checkout, so one source
replaces a branch whose two arms had to be kept in agreement by hand.

Dockerfile gains a FELIS_VERSION build arg wired into -X main.version, declared
after `go mod download` so a version bump does not invalidate that layer. Both
build stages are pinned to $BUILDPLATFORM so a multi-platform buildx run never
emulates them: the panel's output is architecture-independent and the Go stage
cross-compiles via TARGETARCH. The final stage stays on the target platform and
is COPY-only, which BuildKit performs without QEMU.

.github/workflows/release.yml publishes on a vX.Y.Z tag: vet, tests, then one
buildx run producing both architectures through the repo Dockerfile. Not a bare
`go build` — internal/panel/static holds a tracked placeholder index.html so the
//go:embed compiles without node, which means a direct build succeeds and
quietly ships a release whose panel is that placeholder.

The stamp is asserted end to end, because it fails silently: an unstamped binary
reports "dev", which the updater refuses to compare, disabling update reporting
for every install built from that release. The arm64 artifact is checked by ELF
machine type rather than by running it — runners have binfmt registered, so
executing an amd64 binary misnamed arm64 would succeed.

Prerelease tags are flagged explicitly. The trigger glob is v*, gh does not read
semver out of a tag name, and an RC published as a full release becomes
/releases/latest — the single endpoint the default channel installs from and
`felis update` polls.

No SHA256SUMS. A checksum fetched over the same TLS session, with the same
credential, from the same host as the binary adds no trust root; signing is the
real answer and is a separate decision.

Not verified: the download -> validate -> image -> k3s path has never run on a
host against a real published release, because no tag exists yet. The shell
logic around it is verified out of tree; the network and exec behaviour is not.
2026-07-20 18:53:01 +09:00
flyemoji fe4c92c1c5 feat(files): add the server file editor
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.

felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.

What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.

Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.

The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.

Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:

  * A write accepts arbitrary bytes at any path, so an owner can place a
    loadable plugin jar. This is deliberate — it is what a hosting panel's
    file manager does, scoped to a server the caller already owns and
    already drives through /command — but it is the one owner-tier route
    that lands executable code in a backend pod, since images are
    admin-only and modpack submissions need an admin verdict.
  * config/paper-global.yml is refused on read. felis-lobby's entrypoint
    writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
    identical on every backend, so reading it from a server you own would
    hand you the handshake key for everyone else's. It is the only path in
    the mount that is not the caller's own data, and therefore the only
    denial. The comparison is on the cleaned path, or ./config/... would
    walk straight through it.

Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.

The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
2026-07-20 14:33:29 +09:00
flyemoji 05cb8f6320 feat(cli): report component updates and make the router a data table
Add `felis update`, which reports which platform components have newer
versions available, and route `felis version`, which shipped implemented but
unreachable.

That bug is why the subcommand router is now a map rather than a switch.
cmdVersion existed with nothing dispatching to it and no usage line, so
`felis version` fell through to "unknown command" and no test noticed — a
switch offers no way to enumerate what it routes, so the usage text and the
router could not be compared. As data, they can: a test now walks the
Commands: block and the table in both directions, failing an entry added to
one without the other. bootstrap-assets stays deliberately undocumented and
is listed as such, which makes its absence a decision rather than an
oversight.

The host gatherer answers the two seams NewSysGatherer leaves nil, for the
one caller that can satisfy them without a cluster client. felis-api is
answered from the running binary's own build stamp rather than the
Deployment's image tag: deploy/bootstrap.sh builds the image from the same
checkout it installs /usr/local/bin/felis from and stamps both with one git
describe, so it is the same artifact, and it is the identity `felis version`
reports. Reading the Deployment answers a slightly different question — what
is rolled out — and stays the right seam for the in-cluster path.

Velocity is read from the jar's own META-INF/MANIFEST.MF
Implementation-Version, which is what the proxy reports about itself at
runtime, because bootstrap installs the jar under a fixed name with no
version in it. The filename extractor remains only as a fallback for a
hand-placed velocity-3.5.1.jar.

An unstamped local build reports "dev" and is refused with an actionable
message rather than being treated as 0.0.0, which would make every release
upstream look like an upgrade. The panel and the plugin jars have no version
of their own on purpose: they are embedded in or built alongside the felis
binary, so the felis version is theirs.
2026-07-20 14:32:58 +09:00
flyemoji f0b79e9edd feat(mail): deliver email one-time codes over SMTP and add the setup email screen
Felis never actually sent mail: OTP codes for onboarding, email login and
op-login were only written to the felis-api log behind a "demo has no SMTP"
limitation, and the Settings/SMTP flow those comments promised was never
built. Combined with the bootstrap Owner's address being recorded unverified
(87279a1), op-login start always took the anti-enumeration neutral branch and
minted a fake request_id, so the in-game approve inevitably answered "No
pending operator sign-in with that code".

Give the codes a real delivery path, configured in felis.toml rather than a
web settings page so config keeps a single source of truth:

- config: new [smtp] table (host, port defaulting to 587, from, username,
  password_ref). Validation requires a plausible from address and a sane
  port; the password itself never enters the config file.
- internal/mail (new): stdlib net/smtp mailer implementing the api.OTPMailer
  seam. Port 465 dials implicit TLS, other ports upgrade via STARTTLS when
  advertised; AUTH only when a username is configured (PlainAuth itself
  refuses plaintext, so the password cannot leak to a TLS-less relay).
  Ping() proves reachability and credentials without sending mail. The
  message shape (CRLF, Q-encoded bilingual subject) is pinned by test.
- platform: felis-smtp Secret constants and an optional FELIS_SMTP_PASSWORD
  env var on the felis-api Deployment, mirroring felis-uploads-s3.
- cmd/felis api: construct the real mailer when [smtp] is configured; keep
  the log fallback otherwise and say so at startup. Warn when a username is
  set but the credentials env is empty.
- setup TUI: "e" on the summary/status screen opens the email form (host,
  port, from, optional auth). Apply order: Ping preflight, [smtp] into both
  host and pod config files, felis-smtp Secret piped to kubectl via stdin,
  config Secret, felis-api rollout. A failed preflight leaves the install
  untouched. SMTP is deliberately not a wizard rail step: first-run stays
  mail-less by design, and the passkey minted at onboarding is the pre-SMTP
  owner credential.

Also make PGRepo.UserByEmail match case-insensitively (lower(email) =
lower($1)), honoring the interface contract and the users_verified_email_
unique partial index; the fake repo already matched with EqualFold.

Existing installs need the felis-api Deployment manifest re-applied (e.g. a
bootstrap re-run) before the new env var exists; a rollout restart alone
cannot add it.
2026-07-20 10:24:52 +09:00
flyemoji 7860152f57 feat(auth)!: go fully passwordless and fix cross-check review findings
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.

Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
  handlers; align tests (new account subtests, naming reserves
  "console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
  setup_panel.go, tui_root.go and tui_preflight.go instead of
  hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
  into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.

Panel:
- Delete the ChangePassword page and all password UI; align
  login/auth/api/types with the passwordless contract; add the
  migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
  ServerLuckPerms strings to translation keys; drop 72 orphan keys
  per locale; unify the title as "Felis - Console".

Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
  MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
  preference; unified /link lines in fabric/forge/neoforge; shared
  link-client javadoc contract fixes.

Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.

BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
2026-07-20 04:47:32 +09:00
flyemoji 60732a6283 feat(operator): gate op.console to staff and land owner setup there
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).

Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.

Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.

The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.

Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
2026-07-16 18:03:30 +09:00
flyemoji 9ea35304e3 fix(setup): source the console host from the panel hostname, not op.console
The owner setup URL and the limbo login link were built from the admin host
(op.console.<root>, with an op.console.localhost fallback) and a hardcoded
console.<root>, so an operator who set a custom panel_hostname got an unreachable setup
link and a wrong login target. Thread the resolved panel host (defaultPanelHostname)
through performSetupMCBind, the MC-bind TUI, and the login system-server env
(new FELIS_PANEL_HOSTNAME); the limbo plugin prefers it and keeps console.<root> only as
the fallback for an older operator whose env predates it. This also matters for security:
the only wired WebAuthn verifier is scoped to the panel host, so passkey enrollment must
land on the panel face, never op.console.

While here, the limbo login handler checks link status before minting a bind code: an
already-linked player is sent straight to the lobby instead of being shown a useless code.
2026-07-16 13:27:02 +09:00
flyemoji be0c4c41f4 feat(bootstrap): install authenticated game stack 2026-07-14 02:58:21 +09:00
flyemoji dab8fc214b feat(setup): bind owner through login gate 2026-07-14 02:57:15 +09:00
flyemoji 5dc8eb92a8 feat(operator): secure system server workloads 2026-07-14 02:56:19 +09:00
flyemoji fd062882ed feat(nano): give a Mojang player's name back to them, by prefixing the squatter
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.

The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.

A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.

The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.

prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.

Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.

Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:

  premium FLYEMOJ1     -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
  LittleSkin FLYEMOJ1  -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
  both online at once, zero "already connected" rejections
  LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3

The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.

Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.

BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
2026-07-13 13:00:54 +09:00
flyemoji b323975ddb fix(nano): -Dmojang.sessionserver takes the full hasJoined URL, not the base
d417efc got this backwards, in both the code comment and the installer summary.
It claimed authlib appends /session/minecraft/hasJoined itself, so the property
should be given the base URL only. Velocity does not work that way, and a real
login says so: pointed at http://127.0.0.1:8081, a Mojang login arrives at nano
as

    GET /?username=FLYEMOJ1&serverId=-23ae0b50...

with no path at all. Velocity appends the query string to the property verbatim
and issues the request itself; authlib is not in the loop. nano has no route on
/, so it answers 404 and Velocity kicks the player with authservers_down.
Velocity's own default for the property is the full URL,
https://sessionserver.mojang.com/session/minecraft/hasJoined, which is the same
thing said another way.

With the full endpoint URL the same account logs straight in, so both the
nano.go header and summary_nano now print

    -Dmojang.sessionserver=http://127.0.0.1:8081/session/minecraft/hasJoined

and note that the flag belongs between `java` and `-jar`.

Verified against Velocity 3.5.1 + Paper 26.2 on the deploy host: a Mojang login
reaches the backend with its real Mojang UUID unchanged, and a LittleSkin login
under the same username reaches it as UUIDv3(felisAuthNS, "littleskin:"+id) --
two different players on the backend, which is the point.
2026-07-13 11:05:45 +09:00
flyemoji d417efc8bf fix(nano): make the installer work on EL10 and stop serving hasJoined to the world
Deploying `felis nano` to a real Rocky Linux 10 host surfaced four defects that
no local check could see. Fixed together because they all sit on the same path
from `curl|bash` to a running felis-nano.service.

* docker killed the nano install on EL10. `acquire_nano_binary` pulled in
  docker purely to build the binary; on Rocky 10.2 the docker-ce el10 rpms
  install but dockerd refuses to start, so the install died at
  `systemctl enable --now docker`. nano needs one static binary, not an image,
  so the docker dependency is gone: fetch_source -> install_go_toolchain
  (pinned FELIS_GO_VERSION, default 1.26.4, amd64/arm64) -> build_nano_binary.

* the built binary could not be exec'd by systemd (203/EXEC). The Go linker
  renames its output out of $TMPDIR, and a same-filesystem rename carries the
  source SELinux label, so `go build -o /usr/local/bin/felis` produced a binary
  labelled user_tmp_t rather than bin_t. root is unconfined and could run it by
  hand, which is what made this look fine, but the DynamicUser service could
  not. build_nano_binary now stages the output and installs it as a fresh file
  so the policy type transition labels it bin_t, with restorecon as a belt.

* re-running the installer did not converge. `systemctl enable --now` is a
  no-op on an already-active unit, so a rebuilt binary was installed while the
  old process kept running. Now enable + restart.

* the Velocity wiring comment in cmd/felis/nano.go was wrong. authlib appends
  /session/minecraft/hasJoined itself, so -Dmojang.sessionserver takes the base
  URL only, as the installer has always printed.

Also bind to loopback by default. hasJoined is unauthenticated by protocol --
authlib speaks the vanilla sessionserver dialect and sends no token -- so a
public bind is an open auth relay: anyone can point their own proxy at it and
spend this host's egress IP on Mojang until Mojang rate-limits it and the
operator's own players stop getting in. It is not an identity bypass (a caller
still needs a serverId hash bound to their own server key, which the upstream
Yggdrasil validates), but it is someone else's traffic on your address.
FELIS_NANO_LISTEN and the -listen flag now default to 127.0.0.1:8081, which a
same-host Velocity reaches unchanged; serving an off-host proxy is an explicit
opt-in. configure_nano_firewall no longer opens a port for a loopback bind, and
summary_nano prints the real bind address plus the relay warning.

Verified on the target host: installs with no docker present, service active,
binary labelled bin_t, `ss` shows LISTEN 127.0.0.1:8081, an external request is
unreachable, and an in-host request returns 204 with the login logged.

BREAKING CHANGE: felis nano defaults to 127.0.0.1:8081 instead of 0.0.0.0:8081.
A Velocity proxy on another machine must now set FELIS_NANO_LISTEN (or -listen)
to a reachable address, and should allow that port only from the proxy's IP.
2026-07-13 10:30:18 +09:00
flyemoji d177428fb0 feat(felis): add felis nano — Yggdrasil hasJoined multiplexer without a control plane
`felis nano` serves the vanilla sessionserver protocol
(GET /session/minecraft/hasJoined) as a federating multiplexer over
Mojang plus any number of third-party Yggdrasil roots, with no k3s,
Postgres, or panel — a MultiLogin-style auth front-end delivered as a
subcommand of the single felis binary rather than a separate build.

- config.LoadNano reads only [[auth_source]] blocks; it skips the
  database.url / root_domain / archive requirements the full server
  needs. Zero sources is valid (Mojang-only).
- Mojang is prepended in code (Identity:true), never from config, so it
  is always the sole identity root. Third-party profiles are rewritten
  to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID).
- validateAuthSources rejects unknown keys, duplicate tags, and
  scheme-less URLs — a malformed nano config fails loud at load.
- Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend);
  a rejected login is a 204, matching the vanilla sessionserver.
- nano.go binds the -listen flag and ignores [server] listen in config.

Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime
smoke against the template config returns 204 on a miss and logs
"Mojang + 0 third-party source(s)"; a duplicate-tag config exits
non-zero citing "unique".
2026-07-13 02:39:41 +09:00
flyemoji ecea20ee7c feat(nano): configure hasJoined auth sources via [[auth_source]], Mojang-anchored
Step 2 of Felis-nano: a [[auth_source]] array-of-tables (tag + full hasJoined
url, config order = priority) supplies the multiplexer's third-party Yggdrasil
roots; cmd/felis prepends Mojang as the sole code-owned identity anchor and
wires them into API.AuthSources. With no sources configured the endpoint stays
inert (204s), unchanged from step 1.

The config deliberately has no identity/trusted field: Mojang is the only source
whose self-asserted UUIDs are trusted verbatim, so no misconfiguration can
reopen the impersonation hole the per-source UUID rewrite closes. An identity=
key is an unknown key and Load rejects it. Validate adds two fail-fast guards:
unique tags (namespace collision) and a scheme-qualified url (else the source is
silently dead, never validating any login).
2026-07-12 02:27:02 +09:00
flyemoji fc748d3462 feat(breakglass): add "back up a world now" console peer (§B4 Sync)
Adds a break-glass console operation that snapshots a stopped world by
calling the felis-api internal face while the API is alive, rather than
rendering the backup Job locally: the Job needs felis-api deployment
coordinates the console does not hold.

The peer resolves the felis-api-internal ClusterIP Service + service
token from the control namespace, POSTs the internal backup endpoint
with the operator os_user for audit attribution, and maps 409/503/404
to friendly outcome cards. Core decision logic lives in backupnow.go
(unit-tested against a fake client + httptest); tui_backupnow.go is the
untested bubbletea glue mirroring tui_halt.go.
2026-07-07 19:32:13 +09:00
flyemoji 7a7c0d53ab feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
2026-07-07 10:04:30 +09:00
flyemoji c2ee21ae05 feat(breakglass): add halt-a-server op to the recovery console (§B4)
Add a root-gated "Halt a running server" operation to the break-glass
console. The operator picks a server from the live fleet and the console
flips that MinecraftServer CRD's spec.desiredState to Stopped via a
spec-only merge patch, disjoint from the operator's status writes, so it
cannot race or clobber reconciliation. It is the panel-independent
emergency stop for when the box still has root plus a kubeconfig.

System servers (login/lobby) are allowed but flagged: a system tag in the
picker and an explicit WARNING in the post-exit summary, since halting
login takes the shared auth front door down with no fallback. Audit is
best-effort so a halt still works with the audit sink down. Already-stopped
is a distinct no-op. The core (halt.go) is unit-tested against a real
controller-runtime fake client that applies the patch.
2026-07-06 23:50:09 +09:00
flyemoji 0c1cc598c1 feat(auth): migrate console login to passwordless
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.

- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
  login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
  methods an email can use. The single sanctioned existence oracle; methods
  are computed with no role branch, so staff and player accounts in the same
  credential state return byte-identical bodies (staffness invisible by
  construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
  /auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
  (migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
  tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.

Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
2026-07-04 21:47:12 +09:00
Lemon-miaow 3347cc05d5 feat(panel): implement user management administration panel with sessions and minecraft link support 2026-07-04 04:08:14 +08:00
Lemon-miaow 598f3d31f4 feat(submit): local + S3 backends for modpack upload contexts, installer-selectable 2026-07-02 23:38:46 +08:00
flyemoji f554d525d4 feat(cli): provision login/lobby system servers with login env and token replica
setup builds the always-on, reaper-exempt login/lobby MinecraftServers (create-if-absent), bakes the login limbo's non-secret config (internal API URL, root domain, lobby name) into spec.env, and replicates the felis-service-token Secret from the control namespace into the minecraft namespace so the operator's namespace-local secretKeyRef on the login pod resolves.
2026-07-02 19:38:37 +09:00
flyemoji 3c1d64749f fix(api): cap concurrent SSE streams per principal
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.

Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.

This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
2026-07-01 22:04:19 +09:00
flyemoji c6c0772a7a fix(api): set read/idle timeouts on the felis-api listeners
The three felis-api http.Servers (internal, external, https) were built with
only Addr and Handler, leaving ReadHeaderTimeout, IdleTimeout, and ReadTimeout
at zero. A zero ReadHeaderTimeout is a Slowloris hole — a client trickling
header bytes pins a connection indefinitely — and a zero IdleTimeout lets
kept-alive connections accumulate (gosec G112).

Route all three listeners through a newAPIServer factory that sets a 10s
ReadHeaderTimeout and a 120s IdleTimeout. WriteTimeout and ReadTimeout are
left unset on purpose: the external and https faces stream Server-Sent Events
(console / build logs) for the lifetime of a client attachment, and a
WriteTimeout would sever a healthy long-lived stream. Slowloris is closed by
ReadHeaderTimeout, which bounds only the header phase.
2026-07-01 21:12:03 +09:00
flyemoji 7a51c1d9c3 fix(api): bound concurrent login bcrypt to shed CPU-pin floods
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.

Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.

The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
2026-07-01 21:01:28 +09:00
flyemoji e058a64a9b feat(edge): close the panel NodePort to the public after the tunnel is up
After the Cloudflare tunnel connector is installed and the origin has rolled
out, applyCloudflareEdge now fences the panel NodePort so the origin is
reachable only over loopback -- the hop the host-side connector uses -- and
never from a public interface. This closes the Access-bypass hole where a direct
https://<node-ip>:<nodeport>/ with the right Host header reached the origin
behind Cloudflare Access.

The fence is an nftables table hooked at prerouting priority -300 (raw), before
kube-proxy's NodePort DNAT (dstnat, -100), so it catches the packet on its
original destination port; a filter/INPUT rule would miss the DNAT'd, then
FORWARDed NodePort packet. Loopback is accepted first, so the connector origin
hop is untouched; the inet family fences a public IPv6 NodePort too.

It is gated on the connector actually serving (verifyConnectorServing polls
`cloudflared tunnel info`): fencing a dead tunnel would sever the only web path
to a still-up origin. If serving cannot be confirmed the port is left open (its
pre-tunnel state) and the failure is surfaced loudly. unfenceOriginNodePort is
the on-host break-glass reversal. The nft/cloudflared calls are INTEGRATION-ONLY;
the ruleset shape and the conn-count gate are pure and unit-tested.

KNOWN-LIMITATION: targets nftables; firewalld-native coordination is not yet
handled (a firewalld reload can flush the standalone table).
2026-07-01 15:39:05 +09:00
flyemoji fce0fceac4 feat(passkey): wire enrollment verifier into felis-api
Construct the go-webauthn verifier at the composition root and attach
it to the API when auth.panel_hostname is configured (RP id = panel
hostname, origin = https://<panel hostname>, display name Felis). When
the hostname is unset or the verifier fails to build it stays nil and
the passkey ceremony routes report 503, matching the existing
nil-when-unconfigured subsystem pattern. An admin passkey, if ever
added, is a separate relying party on the admin host and is
intentionally not wired here.
2026-07-01 14:36:28 +09:00
flyemoji eb5875a699 feat(felis): add Operator break-glass op behind an operation menu
When a staff account already exists, the break-glass console now opens on a
thin top-level menu (menuModel) where account operations are peers rather than
tails of one wizard: provision/reset the Owner, or add an Operator. A fresh
machine with no Owner skips the menu and goes straight to Owner bootstrap, since
minting an Operator first would create a staff account the login gate rejects.

The Operator path reuses ownerModel via a bgOperation discriminator. It is
insert-only (performAddOperator -> InsertOperator), wraps a duplicate username as
api.ErrConflict and routes back to the provision form for a retry rather than
tearing down, and deliberately never flips the global local_auth toggle the way
the Owner thread does. The post-exit summary and audit trail distinguish the two
outcomes (isOperator); only the Owner provision claims local-password login was
enabled.

Tests cover the operator-model defaults, path selection (insert vs upsert and
the local-auth gate), conflict-retry versus generic teardown, isOperator
propagation, and the root menu routing for both fresh and admin-present
machines.
2026-06-30 15:40:24 +09:00
flyemoji 79eae7f669 feat(metrics): publish felis_servers_total from a fleet snapshot
A per-object reconcile cannot maintain felis_servers_total (spec §23): it
sees one server per call, so it could never Set a correct fleet-wide gauge
and inc/dec on transitions would drift on any missed event. Add a snapshot
producer instead.

metrics.SyncServerGauge Resets the GaugeVec then Sets one child per state,
so a state that drains to zero reports 0 rather than a stale last value.
operator.GaugeSyncer is a manager.Runnable that periodically Lists the
fleet and republishes from it, defaulting an unset desiredState to Stopped.

SyncOnce is exercised end-to-end against a fake client (List, default,
republish); the ticker loop in Start is the only untested I/O edge.
2026-06-30 12:52:53 +09:00
flyemoji 75642d90cf feat(metrics): add named felis_* Prometheus collectors
Introduce internal/metrics exposing the four metric families spec §23
mandates at minimum: felis_servers_total (gauge by desired state),
felis_start_duration_seconds (histogram with Minecraft cold-start
buckets), felis_image_build_failures_total and
felis_reaper_worlds_deleted_total (counters). Collectors are
package-level vars so any subsystem records without an import cycle;
Register wires them into a prometheus.Registerer and is idempotent.

Wire registration into the operator against controller-runtime's global
Registry, so /metrics on the manager's existing metrics endpoint carries
the felis_* families. Instrument the reaper to increment
felis_reaper_worlds_deleted_total in lockstep with Summary.WorldsReaped,
at the one point a world's PVC has actually been deleted.
2026-06-30 12:38:55 +09:00
Lemon-miaow 346ec68e51 refactor(deploy): improved cloudflare walkthrough 2026-06-30 04:20:36 +08:00
flyemoji a94b0015c4 feat(deploy): add break-glass Operator account provisioning
Add the insert-only Operator-creation path to the break-glass console
(felis breakGlass). An Operator is an additional staff admin: role=admin
with must_change_password=true, identical in shape to the Owner, since
Felis has no separate operator DB role (migration 0003).

Unlike the Owner upsert, provisioning is insert-only -- a username already
taken returns ErrConflict (ON CONFLICT DO NOTHING + zero RowsAffected)
rather than silently resetting a live account, so adding an Operator can
never clobber the Owner's or another Operator's credential. A typed
password is used as-is; an empty one is replaced with a generated
one-time credential returned for display. Operator-add does not touch
local_auth_enabled -- that global gate belongs to the Owner thread alone.
Accountability is recorded best-effort under a break_glass.operator_create
audit action, written only after a successful provision.

The TUI menu router that reaches this path is deferred; this lands the
fully unit-testable logic layer (provisionOperator, performAddOperator,
auditAddOperator) with the PGRepo insert kept integration-only.
2026-06-30 04:38:40 +09:00
Lemon-miaow 28c3eee361 refactor(deploy): improved TUI walkthrough 2026-06-30 02:25:18 +08:00
Lemon-miaow e5f1682898 refactor(deploy)!: TUI 2026-06-29 20:17:17 +08:00
flyemoji 5450c268f4 chore: normalize line endings and apply formatting
- Convert CRLF to LF across Go, panel, and plugin files
- Add Cloudflare API token template URL to breakGlass TUI edge intro
- Verify API token in cfsetup before creating tunnel, DNS, or Access app
2026-06-28 16:43:10 +09:00
flyemoji 9c46632929 feat(cli): add felis setup first-run console with reclaim protection and cfsetup idempotency
- Add `felis setup` TUI for initial Owner provisioning and optional Cloudflare edge
- Refactor breakGlass to share console TUI model (runConsoleTUI) with setup mode
- Session auth respects configured [auth].admin_hostname; fallback to op.console.<root>
- Protect linked Yggdrasil admins from Mojang-priority reclaim (spec §B3)
- cfsetup: idempotent Access app/policy creation, better 401/403 errors, GET + lookup
- Bootstrap: auto-install cloudflared, symlink /etc/felis/felis.toml
- Add sequence diagrams for ping-to-join, claim, and link flows
2026-06-28 16:41:37 +09:00
Lemon-miaow f5d00f389e feat(cli): implement felis apply command for direct CRD creation 2026-06-28 02:20:43 +08:00
flyemoji ba13839341 feat(breakglass): optional Cloudflare Tunnel + Access setup in the TUI
Add an optional edge-setup flow to the `felis breakGlass` sudo TUI,
reachable as an independent peer of Owner provisioning through a new
top-level menu (so reaching it never forces an Owner password reset).

The flow drives the operator's own Cloudflare consent (interactive
`cloudflared tunnel login`, suspending the alt-screen, plus an API
token) and then calls cfsetup to stand up a Tunnel routing the
configured admin and panel hosts and a fail-closed Access application.

It stays gated shut unless an admin hostname is configured and the
operator is logged in (edgeReady), and refuses empty or bare-domain
credentials before any side effect. On success the TUI surfaces the
issued Access aud and an explicit ACTION REQUIRED note; it never edits
felis.toml. The live cloudflared and Cloudflare API calls are
integration-only and exercised against a real account.
2026-06-27 15:03:48 +09:00
flyemoji 2d0bbb0c37 feat(cli): attribute break-glass recovery to the SysAdmin who runs it
Root is machine authority, not a human identity, so `felis breakGlass`
now also records WHICH SysAdmin broke the glass. Even under
`sudo felis breakGlass` an account and password are entered in the TUI;
the root gate is necessary but no longer sufficient for accountability.

The console resolves one of three modes up front and audits the
difference:

- bootstrap (no staff account exists yet): the typed credential mints
  the first Owner; the act is attributed to the OS user ($SUDO_USER,
  else root) and recorded verified:false.
- recovery (an admin already exists): the operator authenticates as an
  existing admin via bcrypt; the verified identity is the accountable
  actor and the row is recorded verified:true.
- root override (the typed credential did not verify): a deliberate
  OVERRIDE token proceeds under local-root authority, attributed to the
  OS user and recorded verified:false. Break-glass never refuses -
  recovering when no admin password can be produced is its whole job.

Attribution is best-effort, not proof (whoever runs this is root and can
edit Postgres directly); the audit row is honest about which it is.

- internal/api: AuditEntry gains an optional jsonb Payload (nil maps to
  SQL NULL, so existing callers are unaffected); PGRepo.Audit writes it
  and a new PGRepo.AdminExists drives the bootstrap-vs-recovery switch.
- the accountability row is written the instant the credential changes,
  before local auth is enabled, so a failed toggle write can never leave
  a reset credential with no "who did it" record.
- local_auth_enabled is now one exported api.LocalAuthEnabledKey shared
  by the break-glass writer and the per-request reader, replacing two
  drifting copies of the literal.
- break-glass password entry reuses the panel's 8-72-byte rule so a
  credential set here is never later rejected by web change-password.

Covered by Go unit tests over a fake owner store: auth match/non-match,
the three audit modes and their payloads, that a dead audit sink does
not fail the recovery, that the audit precedes the toggle write, and a
headless drive of the TUI state machine asserting no credential reaches
provisioning without a verified admin or an explicit OVERRIDE.
2026-06-27 11:48:41 +09:00
flyemoji e108a3709a feat(cli): break-glass emergency console TUI
Add `felis breakGlass`, a root-only interactive TUI that provisions or
resets the Owner account directly against Postgres and enables local
password login. It is the local-root recovery path that bypasses web
Zero Trust by design - used to bootstrap the first Owner credential and
to recover when the web login is unreachable.

- Bare `felis` prints CLI usage only; breakGlass is the sole subcommand
  that enters a TUI rather than running as a CLI.
- Refuses to run unless euid is 0 (try: sudo felis breakGlass); on
  non-Unix platforms the euid check also refuses.
- Generates a one-time Owner password, sets must_change_password, and
  prints a durable summary (username, one-time password, op.console
  login URL derived from the configured root domain) after the
  alt-screen TUI is torn down.

Covered by Go unit tests over a fake owner store.
2026-06-27 04:23:21 +09:00
flyemoji af14f02f38 feat(api): local-password authentication backend
Add username+password login for Owner/Operator staff accounts on
op.console, the primary web login when Zero Trust is not in front of the
API. Three handlers form the whole surface: login mints a server-side
session cookie, logout revokes it idempotently, and change-password
re-verifies the current password before rotating the hash and clearing
must_change_password.

- Session cookies are HttpOnly+Secure+SameSite=Lax, host-only, stored
  server-side as a SHA-256 hash with a 12h TTL.
- Login is anti-enumeration: every failure runs a uniform bcrypt compare
  against a dummy hash and returns the same vague error.
- Credential-bearing writes require Content-Type: application/json,
  returning 415 otherwise, to close the cross-site form-POST forgery
  vector as a belt to the SameSite cookie.
- Local auth fails closed: login is rejected unless local_auth_enabled
  is set, so a Zero-Trust-only deployment never accepts a local password.
- Extend the users table with a nullable password_hash and
  must_change_password; staff are role=admin rows with a hash, players
  are role=user rows with hash NULL.
- /me now reports must_change_password so the panel can force a
  first-login change.

Covered by Go unit tests (handlers, content-type guard, anti-enumeration,
forced-change lockdown) and the OpenAPI route-parity gate.
2026-06-27 04:22:45 +09:00
flyemoji 7d913737af fix(migrate): honor -config flag placed after the up verb
Go's flag.Parse stops at the first non-flag token, so a -config given as 'felis migrate up -config path' was silently dropped and the default path used instead. Pull the up verb off the front, then parse the remaining flags so the configured path is honored.
2026-06-27 00:40:38 +09:00