580032056ea89aa205be4df45d91ebf2d59cbf14
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f79e5ebb5e |
feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:
- submit: derived context refs become the internal-face URL
/api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
image's new fetch-context entrypoint) that streams the blob with the
namespace-local service-token Secret — never mounted into Kaniko — and
extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
build namespace gets the token Secret through the existing replica mechanism
(bootstrap.sh + felis setup); the build egress lock opens exactly the control
namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
escapes/symlinks/non-gzip).
Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
|
||
|
|
5d4f3063a9 |
feat(operator): make any Paper image joinable behind the forwarding proxy
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed handshake does not degrade -- it rejects every login the proxy forwards. Until now the only backends that could verify it were the two images Felis builds itself (deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own entrypoints. An arbitrary Paper image a user brings does not, so it passed admission, started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the lobby image as a base for a user's own world (0018_recommended_images.sql), which was never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining player away, the exact opposite of a server you mean to stay on. The fix configures forwarding from OUTSIDE the image instead of requiring it inside. The operator now injects a root `felis init-forwarding` initContainer into every user server; it writes the proxies.velocity block into config/paper-global.yml and forces online-mode=false in server.properties on the /data PVC before the main container starts. The image needs no forwarding logic of its own, so the joinable set stops being "images that self-configure forwarding" and becomes every Paper-family image the platform runs. buildStatefulSet gates the injection on the ABSENCE of the system-role label: the Felis-built system servers already consume the secret in their entrypoints and the login gate is a limbo, not Paper. It is also gated on a non-empty felis image name -- the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it skips the injection rather than failing, because a cluster whose proxy is not in modern mode has nothing to configure. The init runs as root deliberately. The world volume's ownership comes from the storage provisioner and the main container runs as whatever UID its image declares, so root is the only UID that can reliably write these files; it then chmods them 0666/0777 so that non-root main container can rewrite them on boot. The privilege is bounded -- the init exits before the server container starts and the server container keeps its own UID. The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the init ever stops running as root. The writer merges rather than overwrites, both because Paper expands paper-global.yml to its full default tree on first boot and because the panel file editor may edit either file between boots. It sets proxies.velocity.* and the single online-mode key and leaves every other setting alone. It is a no-op on an empty secret, for the same reason the env var is optional: a proxy that is not in modern mode provisions no Secret, and wedging every server's init on a missing optional value would be worse than the status quo. felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and 0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the online-player list and permission commands work out of the box. 0018's row is left in place -- an admin who kept it can keep it; this only adds the better default beside it. Three fixes ride along, each of which the 1.8 path hit in practice. bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only builds its block-connection provider when Via's lowest supported protocol is below 1.13; under modern forwarding the Velocity injector reports 393, so init() returns early, blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining. Every call site is behind isServersideBlockConnections(), so switching it off skips all of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected. ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with this one key suffices -- Config#loadConfig parses the bundled default as the base map and merges the on-disk file over it, so every other option stays current across version bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this worked: init() gates on the protocol version too, and that half fails on its own, so the line is missing either way. The config value is the only evidence, which is what the test asserts. The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13 and has nowhere to go. Only the handshake address field survives Via, so the Felis fork forwards the named servers BungeeCord-style while every other backend keeps modern+secret untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer CRs is the upgrade path. deploy/lobby's set_prop escapes the value before substituting it. The RCON password is operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare `sed s|...|...|` and silently kills the key -- taking the console, the online-player list and permission commands with it. deploy/paper was written with the escaping, so the lobby gets the same rather than leaving the sibling caller broken. Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping. The new tests cover the initContainer's image, root UID, world mount and secret env; the merge preserving unrelated config trees; the properties upsert including the commented-key case; and the bootstrap script both writing the Via key and still calling the function that writes it. Not verified: the initContainer has never run in a real cluster, and the felis-paper image is code-only here as the other game-stack images are -- no Go CI builds them. The ViaVersion pin is the one piece with live evidence, and that evidence is what it was written from. Before it, a client was cut within a second of "logged in with entity id" on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on 2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes. Neither session says which client version it was. The proxy never logged a protocol number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16 players and below" warning that fired for that player on the lobby leg, and no lower -- Via floors every handshake to the proxy's 393, so anything from 47 up is admissible. Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the NPE proves is that the pin was load-bearing, not who was holding the mouse. That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8 behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the fault above; off, holds. No automated test in THIS repository exercises the leg. |
||
|
|
fe4c92c1c5 |
feat(files): add the server file editor
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.
felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.
What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.
Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.
The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.
Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:
* A write accepts arbitrary bytes at any path, so an owner can place a
loadable plugin jar. This is deliberate — it is what a hosting panel's
file manager does, scoped to a server the caller already owns and
already drives through /command — but it is the one owner-tier route
that lands executable code in a backend pod, since images are
admin-only and modpack submissions need an admin verdict.
* config/paper-global.yml is refused on read. felis-lobby's entrypoint
writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
identical on every backend, so reading it from a server you own would
hand you the handshake key for everyone else's. It is the only path in
the mount that is not the caller's own data, and therefore the only
denial. The comparison is on the cleaned path, or ./config/... would
walk straight through it.
Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.
The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
|
||
|
|
05cb8f6320 |
feat(cli): report component updates and make the router a data table
Add `felis update`, which reports which platform components have newer versions available, and route `felis version`, which shipped implemented but unreachable. That bug is why the subcommand router is now a map rather than a switch. cmdVersion existed with nothing dispatching to it and no usage line, so `felis version` fell through to "unknown command" and no test noticed — a switch offers no way to enumerate what it routes, so the usage text and the router could not be compared. As data, they can: a test now walks the Commands: block and the table in both directions, failing an entry added to one without the other. bootstrap-assets stays deliberately undocumented and is listed as such, which makes its absence a decision rather than an oversight. The host gatherer answers the two seams NewSysGatherer leaves nil, for the one caller that can satisfy them without a cluster client. felis-api is answered from the running binary's own build stamp rather than the Deployment's image tag: deploy/bootstrap.sh builds the image from the same checkout it installs /usr/local/bin/felis from and stamps both with one git describe, so it is the same artifact, and it is the identity `felis version` reports. Reading the Deployment answers a slightly different question — what is rolled out — and stays the right seam for the in-cluster path. Velocity is read from the jar's own META-INF/MANIFEST.MF Implementation-Version, which is what the proxy reports about itself at runtime, because bootstrap installs the jar under a fixed name with no version in it. The filename extractor remains only as a fallback for a hand-placed velocity-3.5.1.jar. An unstamped local build reports "dev" and is refused with an actionable message rather than being treated as 0.0.0, which would make every release upstream look like an upgrade. The panel and the plugin jars have no version of their own on purpose: they are embedded in or built alongside the felis binary, so the felis version is theirs. |
||
|
|
d177428fb0 |
feat(felis): add felis nano — Yggdrasil hasJoined multiplexer without a control plane
`felis nano` serves the vanilla sessionserver protocol (GET /session/minecraft/hasJoined) as a federating multiplexer over Mojang plus any number of third-party Yggdrasil roots, with no k3s, Postgres, or panel — a MultiLogin-style auth front-end delivered as a subcommand of the single felis binary rather than a separate build. - config.LoadNano reads only [[auth_source]] blocks; it skips the database.url / root_domain / archive requirements the full server needs. Zero sources is valid (Mojang-only). - Mojang is prepended in code (Identity:true), never from config, so it is always the sole identity root. Third-party profiles are rewritten to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID). - validateAuthSources rejects unknown keys, duplicate tags, and scheme-less URLs — a malformed nano config fails loud at load. - Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend); a rejected login is a 204, matching the vanilla sessionserver. - nano.go binds the -listen flag and ignores [server] listen in config. Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime smoke against the template config returns 204 on a miss and logs "Mojang + 0 third-party source(s)"; a duplicate-tag config exits non-zero citing "unique". |
||
|
|
7a7c0d53ab |
feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.
felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).
Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
|
||
|
|
e5f1682898 | refactor(deploy)!: TUI | ||
|
|
9c46632929 |
feat(cli): add felis setup first-run console with reclaim protection and cfsetup idempotency
- Add `felis setup` TUI for initial Owner provisioning and optional Cloudflare edge - Refactor breakGlass to share console TUI model (runConsoleTUI) with setup mode - Session auth respects configured [auth].admin_hostname; fallback to op.console.<root> - Protect linked Yggdrasil admins from Mojang-priority reclaim (spec §B3) - cfsetup: idempotent Access app/policy creation, better 401/403 errors, GET + lookup - Bootstrap: auto-install cloudflared, symlink /etc/felis/felis.toml - Add sequence diagrams for ping-to-join, claim, and link flows |
||
|
|
f5d00f389e | feat(cli): implement felis apply command for direct CRD creation | ||
|
|
e108a3709a |
feat(cli): break-glass emergency console TUI
Add `felis breakGlass`, a root-only interactive TUI that provisions or resets the Owner account directly against Postgres and enables local password login. It is the local-root recovery path that bypasses web Zero Trust by design - used to bootstrap the first Owner credential and to recover when the web login is unreachable. - Bare `felis` prints CLI usage only; breakGlass is the sole subcommand that enters a TUI rather than running as a CLI. - Refuses to run unless euid is 0 (try: sudo felis breakGlass); on non-Unix platforms the euid check also refuses. - Generates a one-time Owner password, sets must_change_password, and prints a durable summary (username, one-time password, op.console login URL derived from the configured root domain) after the alt-screen TUI is torn down. Covered by Go unit tests over a fake owner store. |
||
|
|
47fcd90f75 |
feat(platform): add node orchestration and the felis entrypoint
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them. |