A single HTTP 502 from fill.papermc.io aborted the entire bootstrap. Build
resolution used a one-shot curl, so one gateway blip from an upstream that
flaps was indistinguishable from a permanent failure, and the run died before
Docker, k3s or any game server was provisioned.
Pass --retry 5 --retry-delay 2 to the build-resolution fetches. 502/503/504
are already in curl's built-in transient set, so the tool had solved this; the
flags were simply never passed. papermc_latest_jar is shared by the Paper and
the Velocity resolve, so hardening it once covers both callers. The
LOOHP/Limbo CI metadata fetch feeds the same step and gets the same treatment.
Deliberately no --retry-connrefused. It only adds ECONNREFUSED to a set that
already covers this incident, and it needs curl 7.52.0 while the yum (el7)
path the script supports ships 7.29.0, where an unrecognised long option is a
parse error rather than a warning:
# centos:7, curl 7.29.0
$ curl -fsSL --retry 5 --retry-delay 2 --retry-connrefused https://example.com
curl: option --retry-connrefused: is unknown
Under set -Eeuo pipefail that exits 2 and trips the || die, so both hardened
fetches would hard-fail on a host where they used to work, each naming a cause
that is not the real one. A comment above papermc_latest_jar records this so
the flag does not come back.
The failure message was actively misleading. "no Paper build for Minecraft
26.2 (the login gate speaks only that protocol)" reads as "that Minecraft
version is unsupported", sending the reader after a version-pinning problem
that does not exist: Paper 26.2 build 60 resolved fine minutes later. Say what
is actually known instead, that the build likely exists and Fill is flapping.
Verified against a local always-502 server: curl now issues 6 requests
(1 initial + 5 retries) over 10.1s before giving up, where it previously
issued 1 and died. Verified on centos:7 that this flag set is accepted, and
against the live Fill v3 API that the resolve still returns a jar URL.
Known and deliberately unchanged: no fetch sets --max-time, so an upstream
that accepts a connection and never answers still blocks forever. --retry does
not cover that, as it fires only once a request completes with a failure. The
remaining single-shot downloads (cloudflared, the Docker GPG key and repo
list, k3s, the Temurin JRE, the Velocity jar, the Go toolchain) keep their
existing no-retry shape rather than widen this diff on a script that is about
to provision a live host.
At bootstrap there is no SMTP, so the old /setup flow was unreachable: it
requested an emailed OTP that could never arrive. Setup now records the
Owner's email address unverified (no OTP round-trip) and requires a passkey,
deferring SMTP configuration to a later Settings page. Setup completes on
email-recorded + passkey-enrolled, and the lockdown lifts on the passkey, not
on email_verified: a passkey is the Owner's only pre-SMTP login credential
(email-OTP login refuses admin accounts).
The record-email endpoint (POST /account/email) now clears email_verified in
the same write. Only VerifyEmailOTP, which proves control of the address, may
set that flag; recording a fresh unproven address must never leave a stale
email_verified=true asserting a proof the user never gave. The change strictly
tightens the invariant, so no existing reader breaks.
Remove the dead ErrEmailTaken path and its documented 409: no migration puts a
unique index on users.email and the codebase does not enforce email
uniqueness, so the unique-violation branch was unreachable and the 409 an
impossible response.
The /setup route (Setup.tsx, setEmail helper, setup i18n copy) is rewritten to
match: record-email, mandatory passkey, no skip-for-now. The SMTP settings
page and post-setup configure-SMTP nudge are deferred.
cloudflared writes the tunnel credentials JSON only at `tunnel create`. An
idempotent re-run against a tunnel that already exists — or a reset +
re-bootstrap where the old box's ~/.cloudflared was wiped but the
Cloudflare-side tunnel survived — finds no local credentials file, and the
connector crash-loops with "Tunnel credentials file doesn't exist". A tunnel
that never comes up leaves op.console unreachable, so the one-time setup link
minted just before it ages out (30-min TTL) unredeemed.
CreateTunnel now resolves the tunnel id on both paths (fresh create and
already-exists) and routes through ensureCredentials, which re-fetches the
token with `cloudflared tunnel token --cred-file` (authenticating via cert.pem,
preserving the same id / DNS / Access) when the file is absent. The secret is
written to the file, not stdout, and the file is chmod 0600 so it is not left
world-readable next to cert.pem.
The self-heal is unconditional on re-bootstrap: Setup gates on Pre.check()
(cert.pem present) before CreateTunnel, so the token re-fetch always has its
cert.pem authority.
The /setup route redeems the one-time token from `felis setup`, then walks
the Owner through email-OTP verification and passkey enrollment before
handing off to the console. It sits outside RequireAuth — the visitor
arrives without a session and the redeem is what mints one — and is
reload-safe: a spent token resumes from the surviving session via
/auth/setup/status.
Adds the Setup page and its /setup route, the setup API client methods
(redeem/status), and the en-US/zh-CN onboarding strings.
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).
Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.
Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.
The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.
Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
Point Velocity's authlib (mojang.sessionserver) at the felis-api hasJoined multiplexer so
a full install federates Mojang plus the configured [[auth_source]] set out of the box,
not just the standalone `felis nano`, and ship a default LittleSkin auth_source in the
generated felis.toml (delete the block for a Mojang-only server). Velocity now owns
velocity.toml and its working tree — it migrates the config version and extracts
localizations on start — while the jar and forwarding secret stay root-owned read-only and
ReadWritePaths widens to VELOCITY_DIR; the felis-api-internal ClusterIP lookup is factored
into a felis_internal_ip helper shared by the link config and the sessionserver override.
Re-include plugins/velocity in the docker context because the felis binary now embeds all
four plugin trees for `felis bootstrap-assets game-stack`.
The owner setup URL and the limbo login link were built from the admin host
(op.console.<root>, with an op.console.localhost fallback) and a hardcoded
console.<root>, so an operator who set a custom panel_hostname got an unreachable setup
link and a wrong login target. Thread the resolved panel host (defaultPanelHostname)
through performSetupMCBind, the MC-bind TUI, and the login system-server env
(new FELIS_PANEL_HOSTNAME); the limbo plugin prefers it and keeps console.<root> only as
the fallback for an older operator whose env predates it. This also matters for security:
the only wired WebAuthn verifier is scoped to the panel host, so passkey enrollment must
land on the panel face, never op.console.
While here, the limbo login handler checks link status before minting a bind code: an
already-linked player is sent straight to the lobby instead of being shown a useless code.
go-webauthn marshals CredentialCreation/CredentialAssertion as {"publicKey": {...}},
but the panel's register (Account.tsx) and username-first login (Login.tsx) read the
options flat (options.challenge, options.user.id), so base64urlToBytes(undefined) threw
"Cannot read properties of undefined (reading 'replace')" and neither ceremony could
start. Strip the envelope in the register-begin and username-login-begin handlers via a
small unwrapPublicKey helper; discoverable login keeps the envelope because it reads
options.publicKey.* plus a top-level options.login_id. The begin tests now feed a wrapped
body and assert the handlers return it flat, so they genuinely exercise the unwrap.
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.
The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.
A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.
The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.
prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.
Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.
Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:
premium FLYEMOJ1 -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
LittleSkin FLYEMOJ1 -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
both online at once, zero "already connected" rejections
LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3
The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.
Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.
BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
d417efc got this backwards, in both the code comment and the installer summary.
It claimed authlib appends /session/minecraft/hasJoined itself, so the property
should be given the base URL only. Velocity does not work that way, and a real
login says so: pointed at http://127.0.0.1:8081, a Mojang login arrives at nano
as
GET /?username=FLYEMOJ1&serverId=-23ae0b50...
with no path at all. Velocity appends the query string to the property verbatim
and issues the request itself; authlib is not in the loop. nano has no route on
/, so it answers 404 and Velocity kicks the player with authservers_down.
Velocity's own default for the property is the full URL,
https://sessionserver.mojang.com/session/minecraft/hasJoined, which is the same
thing said another way.
With the full endpoint URL the same account logs straight in, so both the
nano.go header and summary_nano now print
-Dmojang.sessionserver=http://127.0.0.1:8081/session/minecraft/hasJoined
and note that the flag belongs between `java` and `-jar`.
Verified against Velocity 3.5.1 + Paper 26.2 on the deploy host: a Mojang login
reaches the backend with its real Mojang UUID unchanged, and a LittleSkin login
under the same username reaches it as UUIDv3(felisAuthNS, "littleskin:"+id) --
two different players on the backend, which is the point.
Deploying `felis nano` to a real Rocky Linux 10 host surfaced four defects that
no local check could see. Fixed together because they all sit on the same path
from `curl|bash` to a running felis-nano.service.
* docker killed the nano install on EL10. `acquire_nano_binary` pulled in
docker purely to build the binary; on Rocky 10.2 the docker-ce el10 rpms
install but dockerd refuses to start, so the install died at
`systemctl enable --now docker`. nano needs one static binary, not an image,
so the docker dependency is gone: fetch_source -> install_go_toolchain
(pinned FELIS_GO_VERSION, default 1.26.4, amd64/arm64) -> build_nano_binary.
* the built binary could not be exec'd by systemd (203/EXEC). The Go linker
renames its output out of $TMPDIR, and a same-filesystem rename carries the
source SELinux label, so `go build -o /usr/local/bin/felis` produced a binary
labelled user_tmp_t rather than bin_t. root is unconfined and could run it by
hand, which is what made this look fine, but the DynamicUser service could
not. build_nano_binary now stages the output and installs it as a fresh file
so the policy type transition labels it bin_t, with restorecon as a belt.
* re-running the installer did not converge. `systemctl enable --now` is a
no-op on an already-active unit, so a rebuilt binary was installed while the
old process kept running. Now enable + restart.
* the Velocity wiring comment in cmd/felis/nano.go was wrong. authlib appends
/session/minecraft/hasJoined itself, so -Dmojang.sessionserver takes the base
URL only, as the installer has always printed.
Also bind to loopback by default. hasJoined is unauthenticated by protocol --
authlib speaks the vanilla sessionserver dialect and sends no token -- so a
public bind is an open auth relay: anyone can point their own proxy at it and
spend this host's egress IP on Mojang until Mojang rate-limits it and the
operator's own players stop getting in. It is not an identity bypass (a caller
still needs a serverId hash bound to their own server key, which the upstream
Yggdrasil validates), but it is someone else's traffic on your address.
FELIS_NANO_LISTEN and the -listen flag now default to 127.0.0.1:8081, which a
same-host Velocity reaches unchanged; serving an off-host proxy is an explicit
opt-in. configure_nano_firewall no longer opens a port for a loopback bind, and
summary_nano prints the real bind address plus the relay warning.
Verified on the target host: installs with no docker present, service active,
binary labelled bin_t, `ss` shows LISTEN 127.0.0.1:8081, an external request is
unreachable, and an in-host request returns 204 with the login logged.
BREAKING CHANGE: felis nano defaults to 127.0.0.1:8081 instead of 0.0.0.0:8081.
A Velocity proxy on another machine must now set FELIS_NANO_LISTEN (or -listen)
to a reachable address, and should allow that port only from the proxy's IP.
bootstrap.sh now asks up front whether to install the full Felis control
plane or only Felis-nano, and grows a parallel install path for the
nano-only case.
- prompt_install_mode() runs right after OS detection and reads /dev/tty
(so it works under `curl ... | sudo bash`) offering [1] Felis / [2]
Felis-nano, default full. FELIS_INSTALL_MODE=full|nano skips the prompt
for non-interactive runs; no tty falls back to full.
- main_nano() installs only what nano needs: the felis binary (reusing
the embedded-binary / docker-build acquisition), a template felis.toml
carrying a commented [[auth_source]] example (Mojang-only until edited),
a felis-nano.service unit running `felis nano -config ... -listen ...`
under DynamicUser hardening, and a firewalld port-open for the listen
port. None of the k3s / Postgres / migrate / bundle steps run.
- write_nano_config is idempotent (leaves any existing config untouched)
and its template is valid as-is. summary_nano prints the hasJoined
endpoint and the Velocity -Dmojang.sessionserver flag, offering the
127.0.0.1 form when the proxy is on the same host.
Verified on WSL: `bash -n` clean; the emitted template loads via
config.LoadNano and the exact systemd ExecStart command serves 204 on a
miss ("Mojang + 0 third-party source(s)"); a duplicate-tag config still
exits non-zero citing "unique". Not exercised: a full main_nano run,
systemd activation of the unit, and shellcheck (unavailable in this env).
`felis nano` serves the vanilla sessionserver protocol
(GET /session/minecraft/hasJoined) as a federating multiplexer over
Mojang plus any number of third-party Yggdrasil roots, with no k3s,
Postgres, or panel — a MultiLogin-style auth front-end delivered as a
subcommand of the single felis binary rather than a separate build.
- config.LoadNano reads only [[auth_source]] blocks; it skips the
database.url / root_domain / archive requirements the full server
needs. Zero sources is valid (Mojang-only).
- Mojang is prepended in code (Identity:true), never from config, so it
is always the sole identity root. Third-party profiles are rewritten
to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID).
- validateAuthSources rejects unknown keys, duplicate tags, and
scheme-less URLs — a malformed nano config fails loud at load.
- Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend);
a rejected login is a 204, matching the vanilla sessionserver.
- nano.go binds the -listen flag and ignores [server] listen in config.
Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime
smoke against the template config returns 204 on a miss and logs
"Mojang + 0 third-party source(s)"; a duplicate-tag config exits
non-zero citing "unique".
Step 2 of Felis-nano: a [[auth_source]] array-of-tables (tag + full hasJoined
url, config order = priority) supplies the multiplexer's third-party Yggdrasil
roots; cmd/felis prepends Mojang as the sole code-owned identity anchor and
wires them into API.AuthSources. With no sources configured the endpoint stays
inert (204s), unchanged from step 1.
The config deliberately has no identity/trusted field: Mojang is the only source
whose self-asserted UUIDs are trusted verbatim, so no misconfiguration can
reopen the impersonation hole the per-source UUID rewrite closes. An identity=
key is an unknown key and Load rejects it. Validate adds two fail-fast guards:
unique tags (namespace collision) and a scheme-qualified url (else the source is
silently dead, never validating any login).
Round-2 mutation audit of the backup/restore data-safety surface (7 fail-open gates pinned, 1 coverage gap found and closed by the owner-gate test in 85b8a92). Reverts the Pending section and indexes c67a4d3 + 85b8a92 into the committed ledger.
handleRestoreBackup's owner-or-admin gate was not pinned by any test: the former-owner gate backstopped every non-owner case the suite exercised, so a broken owner gate would not redden. Add the mirror of the former-owner test — a released former owner (still the backup's former_owner, no longer the current owner) must get 403 — the sole subtest that fails when the owner gate is disabled. Found by the round-2 backup/restore mutation audit; production code unchanged.
The operational break-glass ops (Owner provision, OP create, halt, Sync backup) are built and oracle-verified. The fourth §B4 line item — the tarS3 archive backend — is recorded as a deliberate deferral, not a silent gap: config.Load fail-closes store=tarS3 (frozen by TestLoadRejectsUnimplementedArchiveStore), tarLocal is the tested baseline every backup/restore path uses today, and offsite/cross-cluster DR is opt-in future work. Supersedes the "S3 remains open" note in the phase-2b peer doc.
Break each load-bearing safety gate the change ledger names as "unit-tested" and
confirm the specific test turns red — passing proves GREEN, not that the test would
catch a regression. All 18 fail-open crown-jewel gates across the subsystems (auto-update
pin/no-downgrade/no-prerelease/window, passkey clone-refuse, modpack CAS + Trivy scan-gate,
cfsetup fail-closed + NodePort fence conn-count, SSE cap, OTP atomic reserve, idle stop,
startup/readiness timeout, /readyz deps, naming reservation, service-token login-pod-only)
are mutation-proven to pin behaviour. Every documented "unit-tested" claim is reconciled
one-for-one; the fence conn-count gate, previously mis-classified as integration-only, is
corrected and verified. No code changed — read-and-verify only.
A completeness re-check of the ledger backfill (in-scope pre-fad48ff
feat/fix/refactor commits vs SHAs actually cited in detail docs, excluding
the auto-generated ledger table) surfaced one backend functional commit with
no detail-doc home: 346ec68 refactor(deploy) — an 868-insertion rework of the
cloudflare-edge TUI walkthrough plus a tested cfsetup integration-runner path.
The "improved cloudflare walkthrough" subject undersold a behaviour change, so
it is folded into the existing cloudflare-tunnel-access-edge detail doc and its
INDEX row rather than left orphaned. Backend detail-doc coverage of the
pre-ledger history is now complete (0 backend orphans).
Retroactively author 15 grouped detail docs covering the backend
functional (feat/fix) commits made before the change ledger was
established (fad48ff), closing the ledger's detail-doc axis for the
pre-convention history. Each doc groups a feature's constituent commits,
lists their SHAs with subjects, and carries a backfill note stating it
was reconstructed from git history on 2026-07-07 and not independently
re-verified (current tree green at 9911b8c).
Add a Detail docs section to INDEX.md linking every detail doc (the 6
existing + 15 backfill) to the commit(s) it covers, so a doc is findable
from the index without a column on the auto-generated ledger table. Catch
the table up with the missing 9911b8c row.
Scope: backend (Go/Java/K8s) only, per the ledger's stated convention
that frontend/panel commits are the collaborator's UI work; non-functional
commits (docs/style/chore/refactor) keep their table row without a
dedicated detail doc.
Adds a break-glass console operation that snapshots a stopped world by
calling the felis-api internal face while the API is alive, rather than
rendering the backup Job locally: the Job needs felis-api deployment
coordinates the console does not hold.
The peer resolves the felis-api-internal ClusterIP Service + service
token from the control namespace, POSTs the internal backup endpoint
with the operator os_user for audit attribution, and maps 409/503/404
to friendly outcome cards. Core decision logic lives in backupnow.go
(unit-tested against a fake client + httptest); tui_backupnow.go is the
untested bubbletea glue mirroring tui_halt.go.
Detail doc + ledger row for 2ba9948: the separate felis-api-internal ClusterIP
Service that gives the login pod (and the break-glass console) a routable 8081.
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.
Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).
Manifest-level fix; the live packet path is pending real-cluster verification.
Detail doc + ledger row for f2fc57c: the internal (service-token) twin of the
on-demand world backup endpoint that the break-glass console peer will call.
Add POST /api/v1/internal/servers/{name}/backup so the on-node break-glass
console can snapshot a stopped world while felis-api is alive. It goes through
the API (not direct-to-CRD like halt) because rendering the backup Job needs
deployment coordinates (FELIS_IMAGE, FELIS_BACKUP_PVC) only felis-api holds.
Service-token auth (no Principal); the middleware IS the authorization, since
the operator already has root on the node. Refactor the RWO stopped-gate,
optional-Backuper 503, async hand-off and audit+202 into a shared enqueueBackup
tail so the external (owner/admin) and internal (break-glass) faces cannot
diverge on the security-critical stopped-gate. The internal audit is attributed
to break-glass/internal so a console-initiated backup is distinguishable from an
owner self-service one.
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.
felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).
Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
Add docs/changes/ — a durable, in-repo map of every functional change and
the commit that records it, independent of git log. INDEX.md carries the
convention (each functional change gets a dated detail doc plus a ledger
row) and the full oldest-first ledger, regenerable losslessly from git.
Seed detail docs for the two changes just landed: the break-glass halt op
(c2ee21a) and the /felis migrate command (c1aa38b).
Add the in-game /felis migrate command that a player runs to open an
account migration, the entry point for handing their owned servers to
another account (§B3 inherit, scenario A). The command posts the player's
Mojang-verified UUID to the existing handleMigrateStart backend, which puts
the account into migrate mode; the player then finishes on the web console
(prove identity, name the receiving account, redeem a one-time code).
Mirrors the existing /felis claim path: requires a real player past login
limbo, acts on the caller's account rather than the current server, expects
201 Created affirming started=true (a 201 without it is a contract breach,
not a refusal), and maps the handler refusals (404 not_linked, 409
account_retired) to player-facing guidance. On success it points the player
at https://console.<root_domain>, derived from config, never hardcoded.
Compile-verified against velocity-api:3.3.0-SNAPSHOT via the podman gradle
toolchain. Closes the code-only gap named in handlers_account_migrate.go.
Add a root-gated "Halt a running server" operation to the break-glass
console. The operator picks a server from the live fleet and the console
flips that MinecraftServer CRD's spec.desiredState to Stopped via a
spec-only merge patch, disjoint from the operator's status writes, so it
cannot race or clobber reconciliation. It is the panel-independent
emergency stop for when the box still has root plus a kubeconfig.
System servers (login/lobby) are allowed but flagged: a system tag in the
picker and an explicit WARNING in the post-exit summary, since halting
login takes the shared auth front door down with no fallback. Audit is
best-effort so a halt still works with the audit sink down. Already-stopped
is a distinct no-op. The core (halt.go) is unit-tested against a real
controller-runtime fake client that applies the patch.
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.
- 0015 migration: account_migrations state machine (initiated -> confirmed ->
code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem
Quota is grandfathered at redeem: no per-target quota re-check when servers move.
Both login doors (username-first and discoverable) now run a shared applyAssertionCounter after a verified assertion. A signature-counter regression — go-webauthn's CloneWarning, the possible-cloned-authenticator signal — is refused fail-closed with the same opaque passkey_login_invalid envelope any other finish failure returns (no clone oracle to a prober) and audited distinctly as auth.passkey_clone_rejected under the resolved account. A clean assertion advances the stored sign_count to the asserted value and stamps last_used_at, before any session is minted.
Counter-less/synced authenticators report 0 and never warn, so they pass through and simply re-stamp 0; the check gates only counter-keeping hardware authenticators, where a rollback is the meaningful signal. Email-OTP and username-first passkey remain fallbacks, so a rejected clone is never bricked.
Adds Repo.AdvanceCredentialSignCount (pgrepo UPDATE by credential_id) and surfaces CloneWarning from the internal/passkey adapter's FinishLogin/FinishDiscoverableLogin. Proven by real-crypto adapter tests (a counter regression still verifies but flags CloneWarning), handler tests (advance-and-stamp on success, fail-closed on clone), and a symmetric test on each door so both call sites of the shared helper are covered.
A resource-cache migration (0013_resource_cache.sql) was merged onto main concurrently and also claimed version 0013. LoadMigrations rejects any duplicate migration version, so the app would refuse to boot with both files present.
Renumber the discoverable-login migration to 0014. The two migrations touch disjoint objects (0013 ALTERs servers to add cached_* columns; this one CREATEs webauthn_discoverable_challenges), so their relative order does not matter, and the table name is unchanged -- no Go reference moves.
Renaming a just-published migration is safe here because neither version has been applied to a persistent database yet: there is no schema_migrations row for version 13 to reconcile. This is a pre-application renumber, not a history rewrite of an already-applied migration.
Anchor the username-collision reclaim's UUID-keyed, proxy-detected design to
the multi-Yggdrasil reference: CaaMoe/MultiLogin v6 binds identity as
serviceId+online-UUID via "identity cards" that decouple the in-game name from
online identity — keyed by UUID, never by name. Note that §B3's Mojang-priority
reclaim goes beyond the common "protect the first-bound name" behavior by
evicting a squatter once the genuine Mojang owner appears and stashing the
squatter's data for the code-only inherit path.
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.
Routes (both Public):
POST /api/v1/auth/passkey/login/discoverable/begin
POST /api/v1/auth/passkey/login/discoverable/finish
Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.
The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.
Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
Give the Runner a way to read each component's CURRENT version so it can be
compared against the release sources already wired. Three pure extractors turn
raw system text into an updates.Version, each fail-closed:
- versionFromCLI — a `<tool> --version` banner (k3s, cloudflared)
- versionFromImageRef — a container image tag (felis-api)
- versionFromJarName — a proxy jar filename (velocity)
sysGatherer routes each Topology component to the right extractor over an
injected seam; every path is exercised with a fake runner, mirroring how the
release sources are proven against httptest.
The load-bearing case is k3s: its Git tag "v1.36.2+k3s1" parses stable, but a
registry cannot store '+', so the same build ships as image tag "v1.36.2-k3s1",
which parses as a prerelease unless repaired. versionFromImageRef normalizes
"-k3sN"/"-rke2rN" back to "+", so an image read and a CLI read agree instead of
the image masquerading as a prerelease and being barred from comparison.
Honest runtime state after this slice — a green suite is not "the updater runs
against real infra": only the CLI seam (execRunner) is wired, so of the four
tracked components just cloudflared is live end to end (gatherable AND
Scheduled/appliable). k3s is CLI-gatherable but Notify-only. felis-api and
velocity are NOT yet runtime-gatherable: their producing seams — a k8s read of
the control-plane Deployment image, and an off-cluster jar inspection — are left
nil, so both surface an explicit "gather seam not wired" error rather than a
wrong version. felis-api self-update is therefore not functional yet.
Remaining integration (tracked in doc.go): the two producing seams, the concrete
Notifier (SMTP + in-game), the Applier (image bump, cloudflared swap), the
`felis update` CLI + CronJob entry point, and the runtime append of the Pinned
Minecraft fleet.
Give RoutingSource its second upstream so every non-pinned component now
resolves a real latest-stable: Velocity via PaperMC (already wired), and
felis-api, k3s and cloudflared via the GitHub REST API.
github.go queries /repos/{repo}/releases/latest (one request, rate-limit
friendly) and fails closed: a transport error, a non-200 status (404 = no
stable release), an undecodable body, a draft/prerelease flag, or an
unparseable / prerelease-parsing tag all return an error, never a zero
version. It sends the User-Agent GitHub requires (a UA-less request is
403'd) and tolerates the two live tag styles -- cloudflared's CalVer
"2026.6.1" and k3s's v-prefixed, build-tagged "v1.36.2+k3s1" -- while
String() keeps the raw tag for the report.
source.go routes sourceGitHub to it and drops the errGitHubNotWired stub;
velocity still routes to PaperMC.
Tests: github_test.go covers both tag styles, the User-Agent gate, and
fail-closed on 404 / prerelease-flag / unparseable tag, with fixtures
captured from api.github.com on 2026-07-05. runner_test.go now drives
PaperMC and GitHub through dual httptest servers end to end with no source
degrading to an error.
doc.go re-tiers the verification boundary: both release sources are now
built and live-grounded; the VersionGatherer's version-extraction core is
the next verifiable slice (logic over an exec seam, not pure I/O); the
genuine I/O remainder is the Notifier, Applier and felis update CLI/CronJob.
felis-api's coord is still a placeholder slug, so that component is dark at
runtime until a real repository is configured.
An out-of-band curl of the live Fill v3 endpoint contradicted two claims the
previous commit shipped and surfaced a mis-tiering:
- User-Agent is NOT enforced: fill.papermc.io/v3/projects/velocity returned
HTTP 200 to a bare curl UA. The comments claimed a generic UA "is refused"
and the API "REQUIRES" a contact UA. Reword to what is true — PaperMC's usage
policy asks for a descriptive UA and may block generic ones, but sending it is
etiquette/defensive here, not a gate Felis depends on.
- The test fixture's shape was invented, not captured: the real "versions"
object groups the entire 3.x line under a single key "3.0.0", not the
per-minor keys the fixture used. Replace it with the real body (keys and
version strings as returned). The key-agnostic parser already produced the
right answer, and an independent max-stable check confirms 3.4.0.
- Re-tier doc.go: the GitHub Releases source is verifiable-here (the same
httptest-testable shape as PaperMC), not integration remainder. It is why
3 of 4 components report "latest unknown" today and is the next verifiable
slice — the release-source work is only ~half done until it exists.
No production logic changed. WSL oracle: build + vet clean, internal/updater
10/10, full tree go test RC=0 (19 ok, 0 fail).
internal/updates is a pure, fakes-tested decision core with no production caller,
so nothing could produce its "版本号状态" report. Add internal/updater as that caller:
- topology: the fixed platform components and their user-set policies (felis-api
and cloudflared Scheduled+manageable; k3s Notify, high-blast-radius single node;
velocity Notify, off-cluster and unmanageable). Minecraft is pinned by ABSENCE,
never force-tracked here, appended from the live fleet at runtime.
- PaperMC Fill v3 release source: the v2 API (api.papermc.io) was retired
2026-07-01 and returns HTTP 410, so this targets fill.papermc.io/v3, sends the
required non-generic User-Agent, and returns the newest STABLE version, filtering
the -SNAPSHOT/rc prereleases the plan would otherwise suppress. Its test fixture
is captured from the live v3 response shape (2026-07-04).
- RoutingSource: the single ReleaseSource updates.Run requires, dispatching
velocity to PaperMC and returning errGitHubNotWired for the GitHub-backed
components so they degrade to "latest unknown" honestly, never a fabricated one.
- Runner: gather current versions (seam) -> assemble Components -> updates.Run ->
Report; report-only when notifier and applier are nil.
Verification boundary: the parse/plan/compose logic is unit-tested (httptest +
fakes, fixture grounded in the live v3 shape). Live network/TLS/User-Agent
enforcement, the GitHub Releases source, the concrete version gatherer, the
notifier and applier, and the felis update CLI/CronJob remain integration work,
enumerated in doc.go.
Added by the tooling in commit 0c1cc59, not authored guidance. Untrack and gitignore it: the file claims precedence over CLAUDE.md and tells agents to run go fmt, which rewrites the CRLF working tree. The file stays on disk (git rm --cached) so local tooling keeps it, but it is no longer tracked or committed.
The passwordless migration left ResetMailer (SendPasswordReset) and its API field with zero callers and no wiring; the web console authenticates via email-OTP and passkey only. Remove both, plus the now-orphaned context import that the interface was the last user of in handlers_users.go.
Reconcile the DeleteAllPasskeyCredentialsForUser docs in repo.go and pgrepo.go: they claimed there was no production caller, but 2f22027 wired the owner-tier DELETE /users/{id}/passkeys. Both now note that a complete authenticator remediation pairs the unbind with a session revoke (unbinding alone leaves the live hijacked session; revoking alone leaves a re-enrollable credential), and the OpenAPI operation carries the same guidance in a new description. Reword the stale local-password test-fake header, since the passwordless fakes carry no must_change_password field.
No behavior change. gofmt, build, and the full test tree are green; OpenAPI parity and passkey-unbind tests pass; a grep confirms ResetMailer/SendPasswordReset are gone from the Go tree.
Add DELETE /api/v1/users/{id}/passkeys (owner-only) to unbind every passkey a
target account holds — the authenticator remediation that stops a passkey planted
or retained via a transiently-hijacked session from surviving as a standing login
foothold. It wires the previously-uncalled DeleteAllPasskeyCredentialsForUser and
is deliberately not a lockout: the account re-enters via the email-OTP door
(players) or op-login's in-game approval (staff), then re-enrolls. Documented in
the OpenAPI, so the served/documented parity gate covers it.
Remove RevokeUserSessionsExcept: a change-password-era orphan with no callers
since the passwordless migration. Its keep-one ("log out my other devices")
semantics is inherently self-service, and no such slice is on the roadmap; the
admin remediation path already uses RevokeAllUserSessions.
The passwordless migration (b330d77) removed the password-login route, leaving
concurrencyLimiter — its bcrypt concurrency cap — with no caller, and scattered
stale "local-password" / "change-password" references through the surviving auth
code's comments.
- Remove the dead concurrencyLimiter (type + newConcurrencyLimiter + acquire):
no caller, no struct field, no test. Reword the one streamLimiter doc that
contrasted against it.
- Realign comments in repo.go, pgrepo.go, session.go, util.go to the passwordless
reality: staff lookups feed email-OTP / passkey / setup redeem, not a password
compare; RevokeUserSessionsExcept and DeleteAllPasskeyCredentialsForUser are
retained (uncalled) for the P5 account-remediation path (#78); "local sessions"
no longer implies a password.
Comments and dead code only; no behavior change. Full WSL test tree green.
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.
- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
methods an email can use. The single sanctioned existence oracle; methods
are computed with no role branch, so staff and player accounts in the same
credential state return byte-identical bodies (staffness invisible by
construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
/auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
(migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.
Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
demo-up.sh collapses bootstrap -> build+import the limbo/lobby images -> wire [velocity] login_image/lobby_image into felis.host.toml -> felis setup into a single command, ending in the interactive Owner-creation TUI (the only step it cannot automate). Prefers prebuilt tars under deploy/images, else builds on the host, auto-resolving the LOOHP/Limbo CI jar and the latest stable Paper jar (all overridable by env); SKIP_BOOTSTRAP/SKIP_SETUP toggles for reruns.
The lobby image had never been built and two defects blocked it: the felis image .dockerignore excluded plugins/* and only re-included limbo/shared, so the lobby Dockerfile's COPY plugins/paper landed empty; and the plugin stage used eclipse-temurin:21-jdk, which ships no gradle (and the tree vendors no wrapper), failing with 'gradle: not found'. Re-include plugins/paper and build the paper plugin on gradle:8.14-jdk21, matching the limbo image. Verified: both images build and boot (limbo /healthz 200 on 25565; lobby reaches 'Done' with felis-paper enabled).
deploy/limbo assembles LOOHP/Limbo from its loose CI artifacts plus the felis-limbo plugin (and the shared link core), with an entrypoint that pins server-port to the operator's GamePort (25565) on every start. deploy/lobby carries the Paper + felis-paper hub image. .dockerignore re-includes plugins/limbo and plugins/shared so the plugin image build sees them.
The login limbo now performs the onboarding inside Limbo: on join it checks the collision blacklist, mints a bind code, opens a book linking the player to console.<root_domain> (guiding them to the system browser), polls link-status, and transfers to the lobby via BungeeCord Connect — fail-closed on blacklist, mint/transport error, or window elapse. FelisApiClient gains linkStatus/isBlacklisted on the existing internal transport.
WebAuthn is unusable inside the WeChat/QQ in-app WebViews, so a document navigation carrying those UAs is served a bilingual 'open in your system browser' interstitial instead of the passkey-centric SPA. API/config/health/asset requests pass through, and an ack cookie (ua_ack) lets a determined user or false-positive continue. Backend-only; the SPA is untouched.
setup builds the always-on, reaper-exempt login/lobby MinecraftServers (create-if-absent), bakes the login limbo's non-secret config (internal API URL, root domain, lobby name) into spec.env, and replicates the felis-service-token Secret from the control namespace into the minecraft namespace so the operator's namespace-local secretKeyRef on the login pod resolves.
InternalAPIBaseURL builds the felis-api internal-face DNS from SAAPI and the internal port for cross-namespace callers (the login limbo). The service-token Secret name/key now reference the shared naming constants so the Deployment wiring and the operator's login-pod injection cannot drift.
buildStatefulSet gates readiness on an HTTP probe when HealthHTTPPort is set (exposing it as a named container port). buildEnv injects FELIS_SERVICE_TOKEN into the login server only — keyed off the reserved name so it can never leak into a user pod — sourced from a Secret via secretKeyRef, never inlined into the CRD.
StartupSpec.HealthHTTPPort/Path switch pod readiness from plain-TCP to an HTTP GET for RCON-less loaders (LOOHP/Limbo) that report 'started' only after the first tick. User servers now default FallbackServer to the login gate, never the lobby, so a stopped/starting backend keeps authentication in front of a fresh connection.
SystemLoginServer/SystemLobbyServer plus ValidateSystemServerName (format rule without the reservation check) let the platform provision the reserved login/lobby names users can never claim. ServiceTokenSecretName/Key are the one source of truth for the internal-API credential Secret, shared by the platform renderer and the operator's login-pod injection.
setup provisions the always-on login/lobby system services only when these image refs are set; empty means skip-and-say-so (the same fail-loud stance manifests takes), since no official LOOHP/Limbo image exists and a deployment must build its own.
handleChangePassword revoked other sessions but never cleared webauthn_credentials, and enrollment needs no step-up. A passkey planted through a transiently-hijacked session needs no password, so it survived the reset + session-revoke as a standing login foothold. Add DeleteAllPasskeyCredentialsForUser and call it in the change-password remediation so every passkey is unbound alongside the session revoke. Removing zero rows is a successful no-op. Email-OTP remains the fallback factor, so this never locks anyone out; the user re-enrolls a passkey afterward if they want one.
Enrollment set no AuthenticatorSelection, so user verification defaulted to preferred (not enforced), and the UV/backup flags the ceremony reported were discarded. Set UserVerification=required so a bound passkey always proves possession AND user (a UV-incapable device falls back to email-OTP), and capture user_verified/backup_eligible/backup_state through VerifiedCredential -> PasskeyCredential -> webauthn_credentials (migration 0009) so a future login path can enforce UV per credential. Adds a negative test proving a presence-only authenticator is rejected, and asserts the roundtrip records UV=true.
webauthn_credentials.user_id and webauthn_challenges.user_id referenced users(id) with the default ON DELETE NO ACTION, so a future user-delete would either fail or leave orphaned auth material. Recreate both FKs ON DELETE CASCADE: a bound passkey and a pending challenge are ephemeral and must not outlive the account. Scoped to the passkey tables only, not blanket, so retention-bearing child data (world_backups) is not swept away with an account.
The supersede DELETE in CreatePasskeyChallenge filtered consumed_at IS NULL, so it only reaped the prior LIVE challenge; the row that each finish stamps consumed_at on was left behind. A begin->finish loop therefore accumulated one dead row per cycle, unbounded. Drop the consumed_at clause so a fresh begin reaps ALL prior rows for (user, purpose), bounding the table at one row per (user, purpose) with zero net growth per cycle. Deleting an already-consumed row is safe: it has been redeemed and nothing reads it. The fake mirrors the widened supersede.
handlePasskeyRegisterFinish logged an empty target for account.passkey.registered, while the delete half logs the credential id. An operator auditing the log could see that a passkey was bound but not which one. Pass cred.ID as the audit target so bind and unbind are symmetric, and tighten the enrollment test to assert both halves name the credential id.
The MyServers query lists both a user's own servers and unclaimed (owner_id IS NULL) servers, but computed owned as s.owner_id = $1. For an ownerless row that comparison is SQL NULL, which fails to scan into the Go bool and 500s the whole listing. Wrap it in COALESCE(..., false) so an ownerless row reports owned=false while still surfacing as claimable.
The per-write deadline that severs a stalled SSE reader was never cleared on
return. Server.WriteTimeout is deliberately unset -- a WriteTimeout would sever
a healthy long-lived stream -- and with it unset net/http never resets the
connection write deadline between keep-alive requests. So the deadline the last
writeChunk left set leaks onto the next request that reuses the pooled
connection and fails its first write for no reason. Clear it to the zero value
on return via a deferred rc.SetWriteDeadline; best-effort, a no-op on writers
without deadline support.
Also record honestly at the header flush that the connect-time stall stays
bounded only by the per-principal stream cap, not severed by this guard -- only
the mid-stream stall is closed. Adds a test pinning the clear (fails closed:
neutering the deferred clear leaves a +writeTimeout deadline set on return).
QuotaAvailable and ClaimServer run as two separate statements, so the
count read is not serialized against a concurrent claim's UPDATE: two
claims by one user for two different ownerless servers can both pass the
gate and both succeed, leaving the user one server over quota. It is low
severity — quota over-provisioning under a deliberate burst, not an
authorization, ownership, or isolation break, since each server is still
claimed atomically via UPDATE ... WHERE owner_id IS NULL.
Closing it requires Postgres transaction semantics (advisory-xact-lock on
the user, or SERIALIZABLE with retry) folding the gate into a single repo
method — verifiable only against a real Postgres, not the hermetic
fakeRepo suite. Documented at QuotaAvailable with back-references from the
two claim gates (handleClaim and the internal UUID claim) rather than
patched blind.
relayLogStream copied a pod-log follow to the client with a plain
flusher.Flush per event. On a client that stays connected but stops
reading (its TCP receive window shut), net/http buffers the small
"data:" line and only touches the socket at Flush, which then blocks
forever inside the write. The select's <-ctx.Done() branch is never
reached, because r.Context() cancels on an actual disconnect, not on a
stall, so the relay goroutine and its upstream apiserver follow leak for
the life of the process.
Route every event's write+flush through http.ResponseController with a
per-write deadline (writeTimeout, 30s): a stalled flush now returns
os.ErrDeadlineExceeded, the error plain http.Flusher.Flush swallows, and
the relay abandons the stream so the deferred cancel + src.Close release
the follow. SetWriteDeadline and rc.Flush are best-effort: a writer
without deadline support (httptest recorder; some HTTP/2 origins) ignores
the deadline and behaves exactly as before, so the guard degrades
gracefully.
This closes the leak the per-principal stream cap only bounded the blast
radius of. Verified by a deterministic test with a deadline-aware
ResponseWriter whose flush blocks until the deadline; the test times out
(fails closed) if the guard is removed.
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.
Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.
This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
The three felis-api http.Servers (internal, external, https) were built with
only Addr and Handler, leaving ReadHeaderTimeout, IdleTimeout, and ReadTimeout
at zero. A zero ReadHeaderTimeout is a Slowloris hole — a client trickling
header bytes pins a connection indefinitely — and a zero IdleTimeout lets
kept-alive connections accumulate (gosec G112).
Route all three listeners through a newAPIServer factory that sets a 10s
ReadHeaderTimeout and a 120s IdleTimeout. WriteTimeout and ReadTimeout are
left unset on purpose: the external and https faces stream Server-Sent Events
(console / build logs) for the lifetime of a client attachment, and a
WriteTimeout would sever a healthy long-lived stream. Slowloris is closed by
ReadHeaderTimeout, which bounds only the header phase.
withRequestID honored any inbound X-Request-Id verbatim, and that value is
echoed on the response, embedded in the error envelope, and persisted into
audit_logs.request_id. An unvalidated caller-supplied id is therefore an
audit-integrity vector: an arbitrarily long value bloats the audit row, and a
stray control byte (CR/LF) could smuggle a forged entry into a log sink.
Accept an inbound id only when it is well-formed — non-empty, at most 64
bytes, and restricted to a log-safe charset ([A-Za-z0-9._-]) — otherwise mint
a fresh server id. A rejected request loses its inbound trace link, which is
strictly better than storing attacker-controlled text in the audit trail.
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.
Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.
The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
The passkey login/assertion HTTP handler stays deferred after its design
checkpoint; capture the reasoning in the handler header so the decision is
durable in the repo rather than only in task notes.
- RP boundary (resolved): felis-api is the app-login relying party (panel.*);
the WebAuthn security gate lives at the Cloudflare Access edge. Spec §14 ties
WebAuthn/posture to admin.* (Access) while panel.* is plain app login, so
there is neither a spec-required assertion handler nor a backend step-up
consumer for one.
- Identifier (blocking): a from-zero login needs a unique, human-typable handle
to resolve an account, but users.email is nullable and non-unique and a
player's username is their Minecraft uuid. Username-first assertion has
nothing to key on; re-link stays the returning-player door.
Discoverable (usernameless) credentials are the future enabler; the adapter
crypto is already verified so that slice inherits correct crypto.
Build the assertion (login) half of the WebAuthn ceremony crypto in the
internal/passkey adapter, Oracle-verified against a virtual authenticator.
- BeginLogin/FinishLogin over go-webauthn BeginLogin/ValidateLogin,
username-first (allowCredentials scoped to the known user's bound
passkeys). Discoverable/usernameless login stays out of scope: the
enrolled credentials are non-resident and the challenge store is
user-keyed (migration 0007), so it would need a future migration.
- WebAuthnCredentials() now populates the stored COSE public key and
signature counter (assertion validation needs both to verify the
signature and detect clones); enrollment ignores them, so the change
is backward-compatible and the enrollment tests guard it.
- VerifiedAssertion seam output: which credential signed plus the raw
signature counter. Clone/regression policy is deliberately NOT here —
the counter is a ceremony fact and the future handler, which holds the
previously stored counter, decides reject/warn.
Scope: crypto adapter only. The login HTTP handlers, session minting,
and the panel.* relying-party boundary/tier decision remain a deferred
slice (no unauthenticated login route is added). BeginLogin/FinishLogin
live on the concrete adapter, not the api.PasskeyVerifier interface,
which grows only when a handler consumes them.
Tests (virtualwebauthn): a real enrollment chained into a real assertion
exercises the COSE public-key decode path and surfaces the advanced
signature counter, plus origin-mismatch and unbound-credential rejection.
The admin API persists the auto-update maintenance window as lowercase
JSON {"start","end"} (platform_settings key "update_window"), but
updates.Window had no json tags, so it marshaled/unmarshaled with
capitalized keys. The natural decode the update runner will use --
json.Unmarshal(stored, &updates.Window{}) -- would therefore miss every
key and silently yield the zero Window. That fails closed (a zero window
Contains nothing, so notify-only, never a rogue apply), so it is safe but
a latent silent-zero trap for the not-yet-built runner.
Add json:"start"/json:"end" to updates.Window so the obvious decode is
correct by construction; value time.Time treats a stored null as a no-op,
so a cleared/never-set window still decodes to the zero Window. Nothing
in the package serialized Window before, so this changes no existing
behavior.
Guarded by a cross-package contract test in internal/api that marshals
the real api.updateWindow DTO and unmarshals it into updates.Window --
asserting the interval survives (Contains(mid) is true) and that an empty
window decodes to the fail-closed zero Window -- so the two shapes cannot
drift apart silently.
Two admin-tier routes read and set a single platform-wide maintenance
window for the auto-update subsystem (decision core internal/updates):
GET /api/v1/updates/window
PUT /api/v1/updates/window
The window is stored as JSON {"start","end"} (RFC3339, or null when
unset) under the platform_settings key "update_window", reusing the
existing GetSetting/SetSetting KV seam -- no new Repo method, no
migration. Pointer times keep "unset" (null) distinct from a real
instant on both decode and encode; a never-set and an explicitly
cleared window both read back as {null,null}.
Validation mirrors the core's fail-closed Window: a window is either
fully set (both ends, end strictly after start) or fully cleared (both
null). A half-set, inverted, or empty-interval body is 400 and is never
persisted. Reads treat only a missing key as unset (ErrNotFound -> 200
nulls); any other store error 500s rather than fail open.
This is API + PERSISTENCE ONLY. Nothing consumes the stored window yet
-- the runner, the ReleaseSource/Notifier/Applier executors, and the
scheduler CronJob remain INTEGRATION-ONLY. Setting a window changes no
behavior until those land; it is the durable input they will read.
Nothing here force-updates ("不要强制自动更新").
Adds POST /api/v1/auth/bind, the one public pre-account entrypoint of the
player console (console.<root_domain>). An account-less player redeems the
one-time Bind Code minted in the in-game Login Lobby; in a single step the
platform creates a role=user player, links it to the verified in-game UUID,
and mints a host-only felis_session. Login is thus not forced at the edge
while operations stay app-authenticated.
The operator console (op.console.<root_domain>) is unaffected and stays
behind Zero Trust: a code whose UUID resolves to a staff (role=admin)
account is refused with 403 (ErrPlayerBindForbidden) without consuming the
code, so the public door provably never yields an admin principal — the
session it mints carries ViaAdminAccess=false and is host-only to console,
never sent to op.console.
Repo layer: new RedeemPlayerBindCode on the Repo interface, implemented on
PGRepo (single tx: resolve code, create-or-fetch the player, consume) and
the test fake. The returning-player branch is idempotent and is a deliberate
standing "log in via the game" door, not just first-time onboarding.
Honest labeling:
- ORACLE-VERIFIED (Go): account/session logic — role=user, refuse-staff,
idempotent create-or-fetch, single-use code, and the op.console redline
(player session rejected on admin routes). Covered by handlers_onboard_test
and the OpenAPI parity gate.
- INTEGRATION-dependent: the endpoint's security rests on the Bind Code having
been minted against an online-mode-Yggdrasil-authenticated UUID, a
precondition that lives in velocity/Java and is not verifiable from this
repo (CODE-ONLY). The Go layer proves the logic, not that identity guarantee.
- No app-level attempt cap: rate-limiting is deferred to the edge as for the
public /auth/login; the ~1e12 keyspace, single use and short TTL make a
blind app-level cap non-critical.
Introduce internal/updates: a pure, I/O-free engine that decides what
should happen to each tracked platform component (Felis control-plane,
k3s, cloudflared, Velocity) given its current version, the latest
discovered upstream, its policy, and the current time.
Updates are never force-applied. A component is Pinned (Minecraft, left
alone), Notify (a SysAdmin is told and applies out of band), or Scheduled
(Felis may apply, but only inside a maintenance window the SysAdmin set).
The load-bearing invariants are unit-tested: a pinned component never
changes, a downgrade is never proposed, a prerelease is never
auto-applied, and an apply happens only inside the window.
Version parsing tolerates the real feeds (leading v, k3s +k3s1 build
suffix, calendar versions, prerelease tails) and orders by SemVer
precedence. ReleaseSource/Notifier/Applier are declared as integration
seams and exercised via fakes; this package ships no network, SMTP, or
kubectl, and deliberately has no blind k3s-upgrade applier.
After the Cloudflare tunnel connector is installed and the origin has rolled
out, applyCloudflareEdge now fences the panel NodePort so the origin is
reachable only over loopback -- the hop the host-side connector uses -- and
never from a public interface. This closes the Access-bypass hole where a direct
https://<node-ip>:<nodeport>/ with the right Host header reached the origin
behind Cloudflare Access.
The fence is an nftables table hooked at prerouting priority -300 (raw), before
kube-proxy's NodePort DNAT (dstnat, -100), so it catches the packet on its
original destination port; a filter/INPUT rule would miss the DNAT'd, then
FORWARDed NodePort packet. Loopback is accepted first, so the connector origin
hop is untouched; the inet family fences a public IPv6 NodePort too.
It is gated on the connector actually serving (verifyConnectorServing polls
`cloudflared tunnel info`): fencing a dead tunnel would sever the only web path
to a still-up origin. If serving cannot be confirmed the port is left open (its
pre-tunnel state) and the failure is surfaced loudly. unfenceOriginNodePort is
the on-host break-glass reversal. The nft/cloudflared calls are INTEGRATION-ONLY;
the ruleset shape and the conn-count gate are pure and unit-tested.
KNOWN-LIMITATION: targets nftables; firewalld-native coordination is not yet
handled (a firewalld reload can flush the standalone table).
Setup previously called runner.StartConnector (`cloudflared service install`)
while the TUI applyCloudflareEdge separately installs cloudflared-felis.service
for the same tunnel from the same config -- two managed services serving one
tunnel from one connector config.
Drop StartConnector from cfsetup: running a connector is a host-specific side
effect (systemd/launchd/Windows service) that belongs with the caller, not in
this host- and domain-agnostic package whose documented side effects are tunnel
creation, DNS routing, and the Access app/policy calls. installCloudflaredService
in the host layer stays the single connector installer, so the routed-but-dead
1033 is still closed; RouteDNS --overwrite-dns still closes the stale-DNS 1033.
Setup created the tunnel, routed DNS, and wrote config.yml, but nothing
installed or started a connector for it. A one-click run therefore left the
tunnel routed-but-dead: every web hostname returned Cloudflare error 1033
(tunnel has no connector) even though the config was correct on disk.
Add a StartConnector step to the Runner seam, invoked right after the config
is written (and gated on ConfigPath, so a caller wanting only the Access
config is not forced to install a service). The ExecRunner implementation
runs `cloudflared --config <path> service install`, which installs and starts
a managed system service (systemd/launchd/Windows), and is idempotent on an
already-installed service. The orchestration — connector started, and only
after its config exists — is unit-tested against the fake Runner; the actual
service install is INTEGRATION-ONLY.
Together with the RouteDNS --overwrite-dns fix, this closes both distinct
paths to a 1033 half-state from a fresh setup: a stale DNS binding and a
missing connector.
RouteDNS ran `cloudflared tunnel route dns` without --overwrite-dns and
swallowed the resulting "record already exists" error as success. When a
hostname already had a CNAME from an earlier tunnel that was deleted and
recreated, the record stayed bound to the dead tunnel: the setup reported
the hostname "routed" while it kept returning Cloudflare error 1033 (the
tunnel it pointed at has no connector).
Pass --overwrite-dns so the record is repointed at the tunnel just created,
making the route idempotent and correct on every re-run, and drop the
now-unnecessary "already exists" swallow. INTEGRATION-ONLY (ExecRunner
shells out to the real cloudflared binary).
Construct the go-webauthn verifier at the composition root and attach
it to the API when auth.panel_hostname is configured (RP id = panel
hostname, origin = https://<panel hostname>, display name Felis). When
the hostname is unset or the verifier fails to build it stays nil and
the passkey ceremony routes report 503, matching the existing
nil-when-unconfigured subsystem pattern. An admin passkey, if ever
added, is a separate relying party on the admin host and is
intentionally not wired here.
Wrap github.com/go-webauthn/webauthn behind the api.PasskeyVerifier
seam so the api package stays free of go-webauthn types. The adapter
covers the credential-creation ceremony only (BeginRegistration /
CreateCredential); the login/assertion path is a deferred slice.
Ceremony state crosses the seam as opaque marshaled SessionData, the
attestation as an io.Reader, and the verified result as a plain
VerifiedCredential. SessionData carries no expiry so the challenge
row's TTL stays the single liveness authority. New rejects an empty
RP id or origin list so a misconfigured deployment fails at
construction rather than minting unverifiable challenges.
Tests drive a real relying party against a virtual authenticator
(descope/virtualwebauthn): a full creation round-trip plus adversarial
guards proving origin-mismatch and user-mismatch are rejected and
already-bound credentials are excluded.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14), web app face.
An already-authenticated principal binds a passkey to their own account and
manages the credentials they have bound; email-OTP stays the fallback factor.
- four account routes: POST register/begin mints a credential-creation
challenge, POST register/finish verifies the attestation against the
server-stashed SessionData and binds the credential, GET/DELETE credentials
list and unbind the caller's OWN passkeys. App-tier, principal-scoped (the
body never names a user).
- PasskeyVerifier seam keeps go-webauthn out of this package: ceremony state
crosses as opaque bytes, attestation as an io.Reader, result as a plain
VerifiedCredential. A nil verifier makes begin/finish report 503 so the
authenticated boundary is exercised before the real verifier is wired in.
- the view never leaks the public key; credential_id collisions map to 409.
- OpenAPI: the four paths plus the PasskeyCredential schema, keeping the
served-routes parity gate green.
Scope: ENROLLMENT only. The passkey login/assertion path (proving a passkey
from an unauthenticated state) is deferred; every ceremony here rides on a
known principal.
Tests: handler + challenge state machine against a fake repo and a fake
verifier (no real attestation crypto, no SQL). The decisive assertion is the
session-data round-trip -- the finish body carries no challenge, so the only
path for the stashed blob into FinishRegistration is store-stash then consume,
proving the challenge is server-held and never client-echoed. Also covers
supersede-on-begin, single-use, expiry, 503-unavailable, 409-already-bound,
owner-scoped list/delete, and external-only face separation.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14). Adds the data
layer an already-authenticated principal needs to bind and manage passkeys:
- migration 0007: webauthn_credentials (one bound passkey per row, public
attestation material only) and webauthn_challenges (server-stashed ceremony
state between begin and finish, single-use via consumed_at). Both rows are
bound to a known user_id; there is no usernameless login lookup, since the
assertion/login path is a deferred slice.
- PasskeyCredential type and five Repo methods (create/consume challenge,
create/list/delete credential) with the PG semantics the handlers rely on:
supersede-prior-live on begin, expiry-before-consume single-use on finish,
credential_id UNIQUE -> ErrConflict, owner-scoped delete -> ErrNotFound.
- ErrPasskeyChallengeInvalid sentinel for a missing/expired/consumed ceremony.
cooldownLimiter began as the wake-only throttle; the OTP-start hardening
reused it via the atomic reserve/release. Its type comment still called it
a per-server wake limiter and justified the per-replica behaviour as
"acceptable because the operator reconcile is idempotent" -- true for wake,
false for OTP, whose every admitted send is a non-idempotent email.
Rewrite the comment to describe the shared per-key limiter and record the
honest KNOWN-LIMITATION: the atomic reserve/release closes the
intra-replica concurrent burst, but the in-memory map throttles per
replica, so cross-replica bounding still needs a shared store. No
behaviour change.
The email-OTP resend cooldown checked the window with a peek (allowed)
and only recorded it after delivery. For OTP that throttle is the sole
defense and each admitted send is a real, non-idempotent email, so a
burst of truly concurrent starts all passed the peek before any recorded
and every one mailed: N concurrent starts bombed a mailbox with N codes.
Add an atomic reserve/release pair to cooldownLimiter: reserve checks and
records the window in one critical section under the mutex, so a
concurrent burst yields exactly one winner; release rolls a reservation
back only if it is still the current one, so a slow failing caller never
clobbers a newer holder. handleEmailOTPStart now reserves both the
principal and the recipient key up front and defers a rollback that frees
both windows on any mint, create, or delivery error — preserving the old
"a failed send does not consume the cooldown" property, now race-free.
The wake path keeps allowed→record: its real gate is the running cap and
its side effect (SetDesiredState) is idempotent, so the peek gap is
harmless there.
Tests: a frozen-clock gate-mailer fires 8 concurrent starts for one
victim from one principal and asserts exactly one mail and one 202; a
flaky-mailer test proves a failed delivery releases the window so an
immediate retry in the same instant is admitted.
A Running, ready server always reported 0/0 players: markRunningReady
never wrote Status.Players, and markStopped only cleared it. The panel
therefore showed an empty tally for live servers.
Extend the readiness probe to also sample the player count. Prober.Probe
now returns a PlayerCount{Online, Max}: RconProber still gates readiness
on Dial+auth, then runs a best-effort `list` and parses the vanilla
reply ("There are N of a max of M players online"). A failed or
unparseable tally is swallowed (0/0) so it never blocks readiness. The
reconciler threads the count into markRunningReady, which writes
Status.Players; markStopped still resets it to zero.
handleEmailOTPStart minted and mailed a code on every call, so an
authenticated caller could drive unbounded mail to any address they
typed — an email-bomb primitive against arbitrary mailboxes.
Add a separate otpLimiter (its own sync.Once and map, distinct from the
wake limiter) and throttle each send on two keys before anything is
minted: the caller (user:<id>) and the recipient (email:<lower>). A
refused send mints no code and mails nothing; both cooldowns are
recorded only after delivery succeeds, mirroring the wake path so a
failed mint or delivery never consumes the throttle. The two-key design
stops both one account fanning out across addresses and many accounts
converging on one mailbox.
A wake refused by the §9.1 running-server cap returns 503, but the
per-server cooldown was recorded before the cap check ran. A player
held because the cluster was momentarily full would then also have to
wait out the wake cooldown once a slot freed, even though their refused
wake never actually flipped desiredState.
Split cooldownLimiter.allow into allowed (peek, no record) and record
(commit). Both wake paths now consult allowed for the 429, then call
record only after SetDesiredState succeeds — so neither a 503
at_capacity nor a SetDesiredState error consumes the cooldown. The
split is safe against the running cap, which counts CRD truth via
ListServers and is independent of the limiter.
`docker build` failed twice over because .dockerignore excluded two trees the
image actually needs. The panel stage's `COPY panel/ ./` hit `"/panel": not
found`, and even past that the Go build would fail: the root felis package
//go:embeds deploy/bootstrap.sh and deploy/crd/*.yaml, which `COPY . .` dropped
along with the excluded deploy/.
The stale header comment claimed only internal/store/migrations was embedded,
which is what licensed the over-broad exclusions. Rewrite it to name all three
embedded trees (migrations, deploy assets, panel static) and warn against
re-adding panel/, deploy/, or internal/ without re-checking the go:embed list.
Tighten node_modules -> **/node_modules so a working-tree build no longer drags
panel/node_modules over the Linux modules npm ci installs in the panel stage.
When a staff account already exists, the break-glass console now opens on a
thin top-level menu (menuModel) where account operations are peers rather than
tails of one wizard: provision/reset the Owner, or add an Operator. A fresh
machine with no Owner skips the menu and goes straight to Owner bootstrap, since
minting an Operator first would create a staff account the login gate rejects.
The Operator path reuses ownerModel via a bgOperation discriminator. It is
insert-only (performAddOperator -> InsertOperator), wraps a duplicate username as
api.ErrConflict and routes back to the provision form for a retry rather than
tearing down, and deliberately never flips the global local_auth toggle the way
the Owner thread does. The post-exit summary and audit trail distinguish the two
outcomes (isOperator); only the Owner provision claims local-password login was
enabled.
Tests cover the operator-model defaults, path selection (insert vs upsert and
the local-auth gate), conflict-retry versus generic teardown, isOperator
propagation, and the root menu routing for both fresh and admin-present
machines.
Add docs/troubleshooting.md covering the common failure modes the spec
implies, grounded in the actual control-plane code paths:
- Stuck Starting (PodNotReady / RconSecretUnavailable / RconNotReachable)
and the deliberate absence of a Starting->Failed timeout.
- Failed reachable only via InvalidSpec on a malformed spec.storage.size,
plus the stale status.endpoint=direct caveat after a failure.
- Routing via status.endpoint direct/fallback and the empty fallbackServer
pitfall; wake 403/429/503 gate order.
- online-mode coupling and Velocity's offline-mode routing refusal.
- Cloudflare Access 401/403, nil-Keyfunc fail-closed, audience checks,
the absence of an issuer check, and local-session gating.
- Internal service-token (FELIS_SERVICE_TOKEN) rejection path.
- link/claim error codes, Kaniko build denials (SA-by-absence RBAC,
default-deny egress, internal-registry push gate), and the registry
DNS contract.
- Reaper backup-before-delete invariant and false-delete vectors.
- Unimplemented idle auto-stop, permanently-zero players.online, the
inert CRD fields, and the always-survives world PVC behaviour.
Each item is labelled with its evidence grade (GO-TESTED / CODE-ONLY /
INTEGRATION-ONLY / INERT) so operators know what is verified versus
asserted.