The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:
- Every image the installer builds is now a registry ref
(registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
containerd under that exact name (first boot needs no registry round-trip)
and mirrored into the registry after deploy_bundle (push_image_to_registry:
push endpoint 127.0.0.1:5000, and only the path after the host matters to the
registry — a push there lands where kubelet's mirrored pull looks). A ref
outside the registry is warned about, not silently unmirrored.
- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
registry Deployment binds (node containerd cannot dial the Service VIP — live
drill: "Empty reply"). k3s regenerates containerd config only at agent start,
so a CONTENT change restarts k3s and an identical file (every re-run)
restarts nothing.
- import_registry_image caches registry:2 into containerd so the registry
Deployment can start on a box that cannot reach Docker Hub.
- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
'felis-paper:demo') at the registry refs — a user server created from those
rows must not strand when GC collects the bare tag. Only recommended rows
still holding the old seed are touched; enabled is preserved; a pre-existing
target row wins over a duplicate.
bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
The design has claimed since migration 0010 that at most one account can hold
a PROVEN email address, with ErrEmailTaken as the 409 a second verifier sees.
Neither half ever shipped: no migration created users_verified_email_unique,
and VerifyEmailOTP had no guard at all — the sentinel was defined but never
returned, so two accounts could both verify one address. The damage is not
cosmetic: the pre-session login resolves accounts BY verified email, so the
duplicate decided which identity a mailed sign-in code belonged to.
- Migration 0020 creates the partial unique index (lower(email) WHERE
email_verified) the comments have been citing — the database-level backstop.
- VerifyEmailOTP now refuses the take-over with ErrEmailTaken BEFORE consuming
the code (the address, not the code, is the problem), charges no attempt,
and maps a lost cross-user race (unique violation) to the same answer.
- The verify handler answers 409 email_taken instead of a generic 500.
Covered by the pgint suite (sequential double-verify refused with the code
still live, a direct duplicate write still loses to the index, the refused
account can still prove its own address) and a hermetic 409 case.
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed
handshake does not degrade -- it rejects every login the proxy forwards. Until now the
only backends that could verify it were the two images Felis builds itself
(deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own
entrypoints. An arbitrary Paper image a user brings does not, so it passed admission,
started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the
lobby image as a base for a user's own world (0018_recommended_images.sql), which was
never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining
player away, the exact opposite of a server you mean to stay on.
The fix configures forwarding from OUTSIDE the image instead of requiring it inside.
The operator now injects a root `felis init-forwarding` initContainer into every user
server; it writes the proxies.velocity block into config/paper-global.yml and forces
online-mode=false in server.properties on the /data PVC before the main container
starts. The image needs no forwarding logic of its own, so the joinable set stops being
"images that self-configure forwarding" and becomes every Paper-family image the
platform runs.
buildStatefulSet gates the injection on the ABSENCE of the system-role label: the
Felis-built system servers already consume the secret in their entrypoints and the
login gate is a limbo, not Paper. It is also gated on a non-empty felis image name --
the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it
skips the injection rather than failing, because a cluster whose proxy is not in modern
mode has nothing to configure.
The init runs as root deliberately. The world volume's ownership comes from the storage
provisioner and the main container runs as whatever UID its image declares, so root is
the only UID that can reliably write these files; it then chmods them 0666/0777 so that
non-root main container can rewrite them on boot. The privilege is bounded -- the init
exits before the server container starts and the server container keeps its own UID.
The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the
init ever stops running as root.
The writer merges rather than overwrites, both because Paper expands paper-global.yml to
its full default tree on first boot and because the panel file editor may edit either
file between boots. It sets proxies.velocity.* and the single online-mode key and leaves
every other setting alone. It is a no-op on an empty secret, for the same reason the env
var is optional: a proxy that is not in modern mode provisions no Secret, and wedging
every server's init on a missing optional value would be worse than the status quo.
felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and
0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu
plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the
online-player list and permission commands work out of the box. 0018's row is left in
place -- an admin who kept it can keep it; this only adds the better default beside it.
Three fixes ride along, each of which the 1.8 path hit in practice.
bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only
builds its block-connection provider when Via's lowest supported protocol is below 1.13;
under modern forwarding the Velocity injector reports 393, so init() returns early,
blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences
it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining.
Every call site is behind isServersideBlockConnections(), so switching it off skips all
of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected.
ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with
this one key suffices -- Config#loadConfig parses the bundled default as the base map and
merges the on-disk file over it, so every other option stays current across version
bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this
worked: init() gates on the protocol version too, and that half fails on its own, so the
line is missing either way. The config value is the only evidence, which is what the test
asserts.
The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend
sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it
down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13
and has nowhere to go. Only the handshake address field survives Via, so the Felis fork
forwards the named servers BungeeCord-style while every other backend keeps modern+secret
untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer
CRs is the upgrade path.
deploy/lobby's set_prop escapes the value before substituting it. The RCON password is
operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare
`sed s|...|...|` and silently kills the key -- taking the console, the online-player list
and permission commands with it. deploy/paper was written with the escaping, so the lobby
gets the same rather than leaving the sibling caller broken.
Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where
TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping.
The new tests cover the initContainer's image, root UID, world mount and secret env; the
merge preserving unrelated config trees; the properties upsert including the commented-key
case; and the bootstrap script both writing the Via key and still calling the function
that writes it.
Not verified: the initContainer has never run in a real cluster, and the felis-paper
image is code-only here as the other game-stack images are -- no Go CI builds them.
The ViaVersion pin is the one piece with live evidence, and that evidence is what it was
written from. Before it, a client was cut within a second of "logged in with entity id"
on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into
Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on
2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same
player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login
from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes.
Neither session says which client version it was. The proxy never logged a protocol
number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16
players and below" warning that fired for that player on the lobby leg, and no lower --
Via floors every handshake to the proxy's 393, so anything from 47 up is admissible.
Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain
runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the
NPE proves is that the pin was load-bearing, not who was holding the mouse.
That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one
now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8
behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the
fault above; off, holds. No automated test in THIS repository exercises the leg.
The create-server form has no way to tell a user which of the whitelisted
images is a sensible starting point. Add 'recommended' as a third
image_whitelist.source alongside 'built' and 'external', and seed it with
the one image that has earned it.
The marker is presentation only. ImageAdmitted still turns solely on
enabled, so a recommended row is admitted by exactly the rule that governs
every other row and carries no extra privilege; a test pins both halves,
because the failure modes are silent and opposite — make admission
source-aware and the curated images vanish from the form, or let curation
bypass the disable switch and an admin who pulled an image finds it still
creatable.
Only one image is seeded, and the restraint is the point. Velocity runs
proxy-wide modern forwarding, so a backend that cannot verify the signed
handshake rejects every login the proxy sends it. The operator injects
FELIS_FORWARDING_SECRET into every backend but cannot make an image consume
it. An arbitrary public Minecraft image therefore passes admission, builds,
schedules, reports Ready — and then refuses every join, with nothing in the
server's status explaining why. Exactly two images read that variable,
deploy/limbo and deploy/lobby; limbo is the login gate and is nonsense as a
base for a user's server, which leaves lobby. The list grows when Felis
ships another forwarding-aware image, not before.
AdmitBuiltImage now preserves a 'recommended' source through its ON CONFLICT
path. Rebuilding a curated tag is the expected way to patch it, and that
rebuild arrives through this exact path, so a blind SET source = 'built'
would demote the curation on the first rebuild with nothing in the request
saying so. AddExternalImage deliberately does not preserve it: an admin
POSTing the ref is an explicit, named re-admission, and the 201 body reports
the Image it constructed without re-reading the row, so a sticky source
there would report a value the database does not hold.
The migration is idempotent via ON CONFLICT DO NOTHING, so an admin who
disabled or re-pointed the row does not have that decision undone on the
next apply.
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.
Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
handlers; align tests (new account subtests, naming reserves
"console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
setup_panel.go, tui_root.go and tui_preflight.go instead of
hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.
Panel:
- Delete the ChangePassword page and all password UI; align
login/auth/api/types with the passwordless contract; add the
migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
ServerLuckPerms strings to translation keys; drop 72 orphan keys
per locale; unify the title as "Felis - Console".
Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
preference; unified /link lines in fabric/forge/neoforge; shared
link-client javadoc contract fixes.
Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.
BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.
- 0015 migration: account_migrations state machine (initiated -> confirmed ->
code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem
Quota is grandfathered at redeem: no per-target quota re-check when servers move.
A resource-cache migration (0013_resource_cache.sql) was merged onto main concurrently and also claimed version 0013. LoadMigrations rejects any duplicate migration version, so the app would refuse to boot with both files present.
Renumber the discoverable-login migration to 0014. The two migrations touch disjoint objects (0013 ALTERs servers to add cached_* columns; this one CREATEs webauthn_discoverable_challenges), so their relative order does not matter, and the table name is unchanged -- no Go reference moves.
Renaming a just-published migration is safe here because neither version has been applied to a persistent database yet: there is no schema_migrations row for version 13 to reconcile. This is a pre-application renumber, not a history rewrite of an already-applied migration.
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.
Routes (both Public):
POST /api/v1/auth/passkey/login/discoverable/begin
POST /api/v1/auth/passkey/login/discoverable/finish
Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.
The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.
Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
Add per-user resource quota enforcement across all four dimensions:
max_servers, max_cpu_milli, max_memory_mb, and max_storage_gb.
- Migration 0013: add cached_cpu_milli, cached_memory_mb, cached_storage_mb
columns to servers table for pure-SQL per-owner aggregation
- SeedServer now writes resource cache alongside server row
- QuotaCheck replaces QuotaAvailable at claim time, checking all four caps
against the owning user's cumulative usage
- handlePatchServer checks owner's quota before allowing memory/resource
changes on owned servers; unowned servers skip the gate
- handleInternalClaim mirrors the full quota check
- UpdateServerResources keeps the cache in sync after spec mutations
- Reaper zeros resource cache on ReleaseWorld so released resources
are not counted against a former owner
- quantityToMilli/quantityToMB helpers convert K8s quantities to
quota-comparable integers
19 test packages pass.
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.
- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
methods an email can use. The single sanctioned existence oracle; methods
are computed with no role branch, so staff and player accounts in the same
credential state return byte-identical bodies (staffness invisible by
construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
/auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
(migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.
Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
Enrollment set no AuthenticatorSelection, so user verification defaulted to preferred (not enforced), and the UV/backup flags the ceremony reported were discarded. Set UserVerification=required so a bound passkey always proves possession AND user (a UV-incapable device falls back to email-OTP), and capture user_verified/backup_eligible/backup_state through VerifiedCredential -> PasskeyCredential -> webauthn_credentials (migration 0009) so a future login path can enforce UV per credential. Adds a negative test proving a presence-only authenticator is rejected, and asserts the roundtrip records UV=true.
webauthn_credentials.user_id and webauthn_challenges.user_id referenced users(id) with the default ON DELETE NO ACTION, so a future user-delete would either fail or leave orphaned auth material. Recreate both FKs ON DELETE CASCADE: a bound passkey and a pending challenge are ephemeral and must not outlive the account. Scoped to the passkey tables only, not blanket, so retention-bearing child data (world_backups) is not swept away with an account.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14). Adds the data
layer an already-authenticated principal needs to bind and manage passkeys:
- migration 0007: webauthn_credentials (one bound passkey per row, public
attestation material only) and webauthn_challenges (server-stashed ceremony
state between begin and finish, single-use via consumed_at). Both rows are
bound to a known user_id; there is no usernameless login lookup, since the
assertion/login path is a deferred slice.
- PasskeyCredential type and five Repo methods (create/consume challenge,
create/list/delete credential) with the PG semantics the handlers rely on:
supersede-prior-live on begin, expiry-before-consume single-use on finish,
credential_id UNIQUE -> ErrConflict, owner-scoped delete -> ErrNotFound.
- ErrPasskeyChallengeInvalid sentinel for a missing/expired/consumed ceremony.
When the configured third-party Yggdrasil and the official Mojang service
issue the same username under different UUIDs, the non-genuine squatter is
displaced in favour of the real Mojang owner (正版优先). This adds the
Go-verifiable data layer of that flow on the internal (velocity) face.
- migration 0006: username_blacklist (barred squatter UUIDs) and
player_data_holds (the displaced account's 30-day data stash), both keyed
by mc_uuid so the genuine Mojang player — identical username, different
UUID — is never caught by the bar.
- POST /api/v1/internal/player/reclaim bars the squatter UUID and stashes
its data in one transaction (all-or-nothing). It is idempotent on a
retried callback and returns the hold's effective expiry — the first
reclaim's window, never a fresh now()+30d — so the rejected player is told
the truth about how long their data is kept.
- GET /api/v1/internal/player/blacklist/{mc_uuid} is the login-gate check
velocity calls to reject a barred squatter before admitting them.
Scope: velocity collision-routing, the limbo prompt, the authlib
dual-backend and the data-inherit flow are code-only (Java plus a QR-bound
device session a row cannot express) and are not part of this slice. Unit
tests cover the handlers and the in-memory repo contract; the Postgres SQL
path is exercised by integration only.
Capture which Yggdrasil authenticated an in-game UUID when a link code is
minted (spec §10 dual-Yggdrasil) and copy it onto the durable account_links
row at verify. The value originates in-game — the web verify side never sees
the authentication — so it threads through account_link_codes, mirroring how
mc_uuid (not user_id) lives on a code.
- migration 0005: add link_auth_source enum + auth_source column on both
account_link_codes and account_links; DEFAULT 'mojang' backfills existing
rows and sets the Mojang-priority default for a mint that omits the field
- mint validates an explicit auth_source (unknown value -> 400); verify
surfaces it in the 200 body and refreshes it on idempotent re-verify
Forced web onboarding proves a player controls an email before it is
bound to their account. POST /api/v1/account/email/start mints a random
6-digit code, mails it (or logs it server-side when no Mailer is wired —
the demo has no SMTP), and POST /api/v1/account/email/verify redeems it,
flipping users.email_verified in the same transaction that consumes the
code.
Brute force is bounded two ways: a 10-minute TTL and a 5-attempt cap,
both enforced in the repo so the fake and Postgres agree. Only the
sha-256 of the code is stored; the digits live only in the email. Both
routes are app-tier external — verifying your own email is scoped to the
principal, never names another user.
Add username+password login for Owner/Operator staff accounts on
op.console, the primary web login when Zero Trust is not in front of the
API. Three handlers form the whole surface: login mints a server-side
session cookie, logout revokes it idempotently, and change-password
re-verifies the current password before rotating the hash and clearing
must_change_password.
- Session cookies are HttpOnly+Secure+SameSite=Lax, host-only, stored
server-side as a SHA-256 hash with a 12h TTL.
- Login is anti-enumeration: every failure runs a uniform bcrypt compare
against a dummy hash and returns the same vague error.
- Credential-bearing writes require Content-Type: application/json,
returning 415 otherwise, to close the cross-site form-POST forgery
vector as a belt to the SameSite cookie.
- Local auth fails closed: login is rejected unless local_auth_enabled
is set, so a Zero-Trust-only deployment never accepts a local password.
- Extend the users table with a nullable password_hash and
must_change_password; staff are role=admin rows with a hash, players
are role=user rows with hash NULL.
- /me now reports must_change_password so the panel can force a
first-login change.
Covered by Go unit tests (handlers, content-type guard, anti-enumeration,
forced-change lockdown) and the OpenAPI route-parity gate.
Foundational libraries: deterministic resource naming, the RCON client, the Postgres store with embedded SQL migrations, configuration loading, and container image-build helpers.