The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:
- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
Pod never scheduled on the platform's own starter node: FailedScheduling /
Insufficient memory, Pending forever. Requests are now a small floor
(250m / 512Mi, never above a configured cap) while the limits stay the
safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
the source owner; a 65532-owned context (the distroless felis image uid)
fails that under the pod's dropped capabilities ('copying dockerfile:
chown /kaniko/Dockerfile: operation not permitted'). The fetch container
now extracts as root — the uid Kaniko already runs as — so the copy
succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
renders --db-repository, and docs/troubleshooting.md §8e now carries the
verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
internal registry; --insecure already covers its plain HTTP).
Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:
- submit: derived context refs become the internal-face URL
/api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
image's new fetch-context entrypoint) that streams the blob with the
namespace-local service-token Secret — never mounted into Kaniko — and
extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
build namespace gets the token Secret through the existing replica mechanism
(bootstrap.sh + felis setup); the build egress lock opens exactly the control
namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
escapes/symlinks/non-gzip).
Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
build_mem_limit overrides; empty keeps the compiled-in defaults. An
air-gapped or mirrored install has no route to gcr.io/aquasec (the
build egress policy allows only DNS + registry + package mirrors), so
builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
impossible (PVCs cannot cross namespaces) and that the s3 lane also
lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
13b rewritten (verified eviction refusal, 5m pressure-transition,
image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
volume (worlds + config + plugins + cache) and a restore rolls all of
it back — OpenAPI/README wording updated to match (same-tag images are
still watched for regressions by the openapi parity gate).
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
felis api wired the hasJoined multiplexer only when felis.toml had at
least one [[auth_source]]. With none, the source list stayed nil and
every hasJoined answer was a 204. That was harmless while nothing
pointed at the route, but the installer now starts Velocity with
-Dmojang.sessionserver aimed at felis-api unconditionally, and the
generated felis.toml tells the operator to delete the LittleSkin block
for a Mojang-only server. Doing exactly that turned every login away,
premium accounts included, and felis-api logged nothing about it.
Always build the list through authSourcesFromConfig, which prepends
Mojang in code, so an empty config is a Mojang-only relay. felis nano
already behaves this way with the same file.
The new test pins authSourcesFromConfig itself: Mojang first, the only
Identity source, and still present when nothing is configured. Marking
a configured source Identity makes it fail. The call site in cmdAPI is
now a single unconditional assignment and has no unit test of its own.
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.
felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.
What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.
Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.
The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.
Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:
* A write accepts arbitrary bytes at any path, so an owner can place a
loadable plugin jar. This is deliberate — it is what a hosting panel's
file manager does, scoped to a server the caller already owns and
already drives through /command — but it is the one owner-tier route
that lands executable code in a backend pod, since images are
admin-only and modpack submissions need an admin verdict.
* config/paper-global.yml is refused on read. felis-lobby's entrypoint
writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
identical on every backend, so reading it from a server you own would
hand you the handshake key for everyone else's. It is the only path in
the mount that is not the caller's own data, and therefore the only
denial. The comparison is on the cleaned path, or ./config/... would
walk straight through it.
Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.
The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
Felis never actually sent mail: OTP codes for onboarding, email login and
op-login were only written to the felis-api log behind a "demo has no SMTP"
limitation, and the Settings/SMTP flow those comments promised was never
built. Combined with the bootstrap Owner's address being recorded unverified
(87279a1), op-login start always took the anti-enumeration neutral branch and
minted a fake request_id, so the in-game approve inevitably answered "No
pending operator sign-in with that code".
Give the codes a real delivery path, configured in felis.toml rather than a
web settings page so config keeps a single source of truth:
- config: new [smtp] table (host, port defaulting to 587, from, username,
password_ref). Validation requires a plausible from address and a sane
port; the password itself never enters the config file.
- internal/mail (new): stdlib net/smtp mailer implementing the api.OTPMailer
seam. Port 465 dials implicit TLS, other ports upgrade via STARTTLS when
advertised; AUTH only when a username is configured (PlainAuth itself
refuses plaintext, so the password cannot leak to a TLS-less relay).
Ping() proves reachability and credentials without sending mail. The
message shape (CRLF, Q-encoded bilingual subject) is pinned by test.
- platform: felis-smtp Secret constants and an optional FELIS_SMTP_PASSWORD
env var on the felis-api Deployment, mirroring felis-uploads-s3.
- cmd/felis api: construct the real mailer when [smtp] is configured; keep
the log fallback otherwise and say so at startup. Warn when a username is
set but the credentials env is empty.
- setup TUI: "e" on the summary/status screen opens the email form (host,
port, from, optional auth). Apply order: Ping preflight, [smtp] into both
host and pod config files, felis-smtp Secret piped to kubectl via stdin,
config Secret, felis-api rollout. A failed preflight leaves the install
untouched. SMTP is deliberately not a wizard rail step: first-run stays
mail-less by design, and the passkey minted at onboarding is the pre-SMTP
owner credential.
Also make PGRepo.UserByEmail match case-insensitively (lower(email) =
lower($1)), honoring the interface contract and the users_verified_email_
unique partial index; the fake repo already matched with EqualFold.
Existing installs need the felis-api Deployment manifest re-applied (e.g. a
bootstrap re-run) before the new env var exists; a rollout restart alone
cannot add it.
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.
Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
handlers; align tests (new account subtests, naming reserves
"console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
setup_panel.go, tui_root.go and tui_preflight.go instead of
hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.
Panel:
- Delete the ChangePassword page and all password UI; align
login/auth/api/types with the passwordless contract; add the
migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
ServerLuckPerms strings to translation keys; drop 72 orphan keys
per locale; unify the title as "Felis - Console".
Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
preference; unified /link lines in fabric/forge/neoforge; shared
link-client javadoc contract fixes.
Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.
BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).
Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.
Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.
The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.
Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.
The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.
A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.
The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.
prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.
Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.
Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:
premium FLYEMOJ1 -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
LittleSkin FLYEMOJ1 -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
both online at once, zero "already connected" rejections
LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3
The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.
Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.
BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
`felis nano` serves the vanilla sessionserver protocol
(GET /session/minecraft/hasJoined) as a federating multiplexer over
Mojang plus any number of third-party Yggdrasil roots, with no k3s,
Postgres, or panel — a MultiLogin-style auth front-end delivered as a
subcommand of the single felis binary rather than a separate build.
- config.LoadNano reads only [[auth_source]] blocks; it skips the
database.url / root_domain / archive requirements the full server
needs. Zero sources is valid (Mojang-only).
- Mojang is prepended in code (Identity:true), never from config, so it
is always the sole identity root. Third-party profiles are rewritten
to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID).
- validateAuthSources rejects unknown keys, duplicate tags, and
scheme-less URLs — a malformed nano config fails loud at load.
- Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend);
a rejected login is a 204, matching the vanilla sessionserver.
- nano.go binds the -listen flag and ignores [server] listen in config.
Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime
smoke against the template config returns 204 on a miss and logs
"Mojang + 0 third-party source(s)"; a duplicate-tag config exits
non-zero citing "unique".
Step 2 of Felis-nano: a [[auth_source]] array-of-tables (tag + full hasJoined
url, config order = priority) supplies the multiplexer's third-party Yggdrasil
roots; cmd/felis prepends Mojang as the sole code-owned identity anchor and
wires them into API.AuthSources. With no sources configured the endpoint stays
inert (204s), unchanged from step 1.
The config deliberately has no identity/trusted field: Mojang is the only source
whose self-asserted UUIDs are trusted verbatim, so no misconfiguration can
reopen the impersonation hole the per-source UUID rewrite closes. An identity=
key is an unknown key and Load rejects it. Validate adds two fail-fast guards:
unique tags (namespace collision) and a scheme-qualified url (else the source is
silently dead, never validating any login).
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.
felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).
Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.
- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
methods an email can use. The single sanctioned existence oracle; methods
are computed with no role branch, so staff and player accounts in the same
credential state return byte-identical bodies (staffness invisible by
construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
/auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
(migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.
Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.
Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.
This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
The three felis-api http.Servers (internal, external, https) were built with
only Addr and Handler, leaving ReadHeaderTimeout, IdleTimeout, and ReadTimeout
at zero. A zero ReadHeaderTimeout is a Slowloris hole — a client trickling
header bytes pins a connection indefinitely — and a zero IdleTimeout lets
kept-alive connections accumulate (gosec G112).
Route all three listeners through a newAPIServer factory that sets a 10s
ReadHeaderTimeout and a 120s IdleTimeout. WriteTimeout and ReadTimeout are
left unset on purpose: the external and https faces stream Server-Sent Events
(console / build logs) for the lifetime of a client attachment, and a
WriteTimeout would sever a healthy long-lived stream. Slowloris is closed by
ReadHeaderTimeout, which bounds only the header phase.
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.
Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.
The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
Construct the go-webauthn verifier at the composition root and attach
it to the API when auth.panel_hostname is configured (RP id = panel
hostname, origin = https://<panel hostname>, display name Felis). When
the hostname is unset or the verifier fails to build it stays nil and
the passkey ceremony routes report 503, matching the existing
nil-when-unconfigured subsystem pattern. An admin passkey, if ever
added, is a separate relying party on the admin host and is
intentionally not wired here.
Add username+password login for Owner/Operator staff accounts on
op.console, the primary web login when Zero Trust is not in front of the
API. Three handlers form the whole surface: login mints a server-side
session cookie, logout revokes it idempotently, and change-password
re-verifies the current password before rotating the hash and clearing
must_change_password.
- Session cookies are HttpOnly+Secure+SameSite=Lax, host-only, stored
server-side as a SHA-256 hash with a 12h TTL.
- Login is anti-enumeration: every failure runs a uniform bcrypt compare
against a dummy hash and returns the same vague error.
- Credential-bearing writes require Content-Type: application/json,
returning 415 otherwise, to close the cross-site form-POST forgery
vector as a belt to the SameSite cookie.
- Local auth fails closed: login is rejected unless local_auth_enabled
is set, so a Zero-Trust-only deployment never accepts a local password.
- Extend the users table with a nullable password_hash and
must_change_password; staff are role=admin rows with a hash, players
are role=user rows with hash NULL.
- /me now reports must_change_password so the panel can force a
first-login change.
Covered by Go unit tests (handlers, content-type guard, anti-enumeration,
forced-change lockdown) and the OpenAPI route-parity gate.
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them.