Commit Graph
100 Commits
Author SHA1 Message Date
flyemoji 694e3cb800 feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.

The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.

Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.

RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.

Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.

That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.

Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.

Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.

Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
2026-07-21 00:12:37 +09:00
flyemoji b3989fa4af fix(mail): prove SMTP deliverability before saving, and stop losing the relay
A live install passed the SMTP setup screen and then failed every one-time
code with a bare `internal error`. Four separate defects had to line up for
that, and each is fixed here.

The relay was configured with `from = noreply@<domain-A>` on an account
authenticated as `<user>@<domain-B>`. Providers that validate sender identity
— Fastmail among them — answer MAIL FROM with an unconditional 250 and only
refuse at end-of-DATA. Ping stopped at NOOP, so it never saw the refusal: the
wizard reported success, wrote the config, rolled felis-api, and every OTP
afterwards died at w.Close().

Ping now runs the same transaction a real code takes — connect, (STARTTLS,)
AUTH, MAIL FROM, RCPT TO, DATA — delivering one self-test message to the From
address, and SendOTP and Ping share deliver() so the check cannot drift from
the thing it checks. The self-test recipient cannot cause a false negative:
an authenticated submission relay accepts RCPT for any destination by
definition, while the sender identity it does validate is exactly what we
want tested. The setup screen now says a message will be sent, names the
address it went to, and warns that From must be an address the account is
allowed to send as.

A relay refusal also answered 500 `internal`, which reads as a broken panel
and sends the operator hunting through handler code instead of their [smtp]
block. It is now 502 `mail_undeliverable`, mapped inside deliverOTP so all
four doors that mail a code (onboarding, email login, op-login, migrate
step-up) answer alike. The relay's own text stays out of the response — it
can name the SMTP account, and these routes are reachable by any signed-in
player — and goes to the log instead.

writeError logged nothing when it collapsed an unmapped error to 500, so an
operator holding an `internal error` had nothing to grep for and diagnosis
degraded into guessing against a live install. It now logs the method, path,
wrapped chain and the same request_id the caller is shown.

Finally, write_felis_toml regenerated the config wholesale and never emitted
[smtp], so re-running the installer — the documented way to update felis-api —
silently erased a working relay and reverted OTP delivery to the no-Mailer
path, logging codes instead of sending them. It now carries the block forward,
cached on first read because the host toml is clobbered before the pod toml is
written. Same defect family as the root_domain loss fixed in ecbeb20: a
generated file holding a hand-set value with no carry-forward.

Tests cover the case a MAIL FROM probe cannot see: a fake relay that answers
250 to MAIL FROM and 550 at end-of-DATA must fail both Ping and SendOTP, and
the 502 must carry a distinct machine code without leaking the relay's text.
2026-07-21 00:11:40 +09:00
flyemoji 32be3e17b7 docs(readme): give an install command that works against a private repo
The documented one-liner fetches bootstrap.sh from raw.githubusercontent.com
unauthenticated, which 404s for as long as this repository stays private -- so
the single command the README exists to provide did not work for anyone.

The authenticated form goes through the contents API with the raw media type,
matching what github_api already does, and hands the token to curl over stdin
via --config rather than -H. argv is world-readable through /proc, and a token
on the command line would leak to any local user during the install; bootstrap
avoids that in its own fetches for the same reason and the README should not
teach the opposite.

sudo -E, because the installer needs that same token to resolve and download the
release. Without it sudo drops the variable and the run fails later, at the
release lookup, for a reason the operator has no way to connect to this command.

The public one-liner stays first: it is what this becomes once the repository is
public, and the note is scoped to the current state.

Also records that re-running the installer is how felis-api moves to a newer
release, that it now keeps the installed root domain, and that it does not keep
the channel.
2026-07-20 21:20:09 +09:00
flyemoji ecbeb20761 fix(bootstrap): reuse the installed root domain instead of re-deriving it
detect_node_ip recomputed FELIS_ROOT_DOMAIN from scratch on every run and fell
back to <node-ip>.nip.io. Nothing read the domain back out of the felis.toml an
earlier run wrote, so it survived only as long as the operator kept passing the
same environment.

That made re-running the installer destructive on any install with a real
domain, and re-running it is not optional: it is the only way to move felis-api
to a newer release, which is what `felis update` points operators at. A bare
re-run rewrote root_domain, panel_hostname and admin_hostname to nip.io names
while ensure_panel_tls_cert returned early on the certificate it had already
written, leaving the console serving a cert for hostnames it no longer answered
to -- with no re-domain flow to recover through.

Precedence is now explicit FELIS_ROOT_DOMAIN, then the domain the last run
persisted, then the nip.io default. First installs are unaffected. Deliberate
re-domains still work, because there is no other route to one, but they now warn
that the write-once certificate is not reissued and that the proxy and login
config carry the old name too.

Secrets were never exposed to this: load_or_make_secrets has always sourced
secrets.env before generating anything. The domain was the one piece of install
identity with no read-back.

The channel is deliberately left alone. FELIS_VERSION_BOOTSTRAP is not persisted
either, but defaulting a re-run to the release channel installs a working build
rather than breaking one, so cmd/felis/update.go states that instead. Its
warning about the domain went with the bug and would now be false.

Verified against the shipped function text: the ladder holds for a fresh host,
a re-run with and without the variable set, a re-domain, and a felis.toml whose
root_domain is missing or empty. Reverting the one line reproduces the nip.io
overwrite.
2026-07-20 21:20:08 +09:00
flyemoji f7815629bf fix(update): say that setup cannot move felis-api to a newer release
`felis update` offers `sudo felis setup` for every planner-backed target and
closed with a trailer calling setup idempotent. That is true for velocity --
install_velocity re-resolves the newest build of the pinned minor on each run --
and misleading for felis-api, which both --panel and --plugins resolve to.

setup hands bootstrap the binary it is itself running and takes the
bootstrap_from_tui arm, which skips the release lookup. The run rebuilds the
image and rolls the deployment off that SAME binary: it reports success and
leaves the version exactly where it was. An operator following this guidance to
apply a felis-api update would watch it appear to work and then see the same
version reported again.

Only the bootstrap installer moves felis-api, and naming it is where this gets
dangerous, so the warning ships with it. The installer is not an updater. Every
run re-derives FELIS_ROOT_DOMAIN through detect_node_ip and defaults it to
<node-ip>.nip.io; nothing reads the domain back out of the felis.toml an earlier
run wrote. A bare re-run -- which is exactly what README documents, with no
environment at all -- rewrites root-domain, panel-hostname and admin-hostname to
nip.io names, while ensure_panel_tls_cert returns early on the certificate it
already wrote and keeps serving the old hostnames. The console then fails to
match its own certificate, and there is no re-domain flow to recover with.
Persisted secrets are not at risk: load_or_make_secrets sources secrets.env
before it generates anything.

felis update stays report-only, so no command changed; only the claim about what
the offered one accomplishes, and the conditions under which the alternative is
safe to run.

Tested four ways, because the scoping and the warning are both the point:
--panel carries the caveat AND names FELIS_ROOT_DOMAIN, --velocity keeps the
ordinary trailer without either, and --mc, which offers no command at all, gets
neither.
2026-07-20 21:20:04 +09:00
flyemoji 8b25114fb4 fix(update): say that setup cannot move felis-api to a newer release
`felis update` offers `sudo felis setup` for every planner-backed target and
closed with a trailer calling setup idempotent. That is true for velocity --
install_velocity re-resolves the newest build of the pinned minor on each run --
and misleading for felis-api, which both --panel and --plugins resolve to.

setup hands bootstrap the binary it is itself running and takes the
bootstrap_from_tui arm, which skips the release lookup. The run rebuilds the
image and rolls the deployment off that SAME binary: it reports success and
leaves the version exactly where it was. An operator following this guidance to
apply a felis-api update would watch it appear to work and then see the same
version reported again.

The report now says so, scoped to runs that actually offered a felis-api
target, and points at the bootstrap installer -- the path that resolves and
downloads a release. felis update stays report-only, so no command changed;
only the claim about what the offered one accomplishes.

Tested three ways, because the scoping is the whole point: --panel carries the
caveat, --velocity keeps the ordinary trailer without it, and --mc, which
offers no command at all, gets neither.
2026-07-20 19:53:55 +09:00
flyemoji 8675cda001 fix(setup): refuse --dev rather than silently installing the release channel
`felis setup --dev` promised "install the dev channel (main HEAD)" and
installed release: the flag only exported FELIS_CHANNEL, a variable nothing in
the tree reads. deploy/bootstrap.sh reads FELIS_VERSION_BOOTSTRAP.

Renaming the variable would have been a worse bug than the dead one, because it
would look wired. setup runs bootstrap with FELIS_BOOTSTRAP_FROM_TUI=1, and on
that arm every reader of FELIS_VERSION_BOOTSTRAP is unreachable: the channel
case and its validation live in resolve_install_ref, which the TUI path skips
outright, and use_release_binary is only consulted by the elif that
`if bootstrap_from_tui` already short-circuited. setup re-images the host from
the felis binary it is itself running; there is no channel to pick.

So the flag refuses, exits 2 and names FELIS_VERSION_BOOTSTRAP=dev on the
installer, which is the mechanism that does work. Refusing beats defaulting:
the operator asked for dev, and release is the one answer they did not want.
The refusal precedes the root check, or an unprivileged operator gets told
about sudo instead of about the channel.

channelName had no other caller and goes with it. Nothing else referenced
--dev -- no doc, no script, no test -- so this removes a promise the tree only
ever made to itself.
2026-07-20 19:53:32 +09:00
flyemoji e5ea51c0db fix(bootstrap): wrap the downloaded binary in the same base CI ships
build_image_from_binary built on distroless/base-debian12 while the repo
Dockerfile's final stage uses distroless/static-debian12, so the image an
install runs did not match the image CI publishes.

Every binary that can reach HOST_BIN traces back to the Dockerfile's
CGO_ENABLED=0 build -- the downloaded CI asset, the binary the TUI is already
running, and the one build_image_from_source docker-cp's out of the image it
just built. None link glibc, so base-debian12 bought nothing and only widened
the runtime surface.

This mattered little while build_image_from_binary was the rare fallback.
659c8e5 made the release channel download a binary and wrap it here, which
makes this the image most installs actually run.
2026-07-20 19:53:04 +09:00
flyemoji 646d514a65 test(updates): pin the ordering of a dev build's own version stamp
659c8e5 chose "+g<sha>" over "-g<sha>" for the dev channel's stamp on the
grounds that "+" is SemVer build metadata, ignored for precedence, so a build
some commits past v1.2.3 still reads as v1.2.3 rather than as something older.
That claim is load-bearing and was asserted only in prose: if metadata ever
counted for ordering, every dev install would report an upgrade available onto
a release it already contains.

Parse does handle it -- build metadata is stripped before the prerelease tail,
so the "-earlyAccess+g<sha>" order a real build stamps splits correctly too --
but nothing tested it. The existing coverage compares "+k3s1" against "+k3s2",
which is metadata on BOTH sides; the case that matters here is metadata on one
side only, against the bare tag.

Four pairs in the compare table, plus "v0.0.0+g<sha>" in the parse table for a
build pinned to a ref with no tag behind it. Verified by mutation: disabling the
"+" split in Parse fails these cases specifically.

They sit at the end of the table rather than beside the other build-metadata
pairs. The entries are wide and carry no trailing comment, and inserting them
mid-table splits the contiguous comment block, which makes gofmt rewrite the
alignment of seven untouched lines.
2026-07-20 19:05:19 +09:00
flyemoji d9ef5e4ffd chore(deps): tidy the module files
Plain `go mod tidy` output, no hand edits, so the files match what the tool
produces from the current import graph:

  - github.com/minio/minio-go/v7 becomes direct. internal/submit/s3store.go has
    imported it since 598f3d3, which added the S3 upload backend without tidying.
  - golang.org/x/crypto becomes indirect. Nothing imports it directly any more —
    the passwordless migration removed the last caller.
  - github.com/clipperhouse/stringish drops out of the graph entirely.
  - go.sum loses 26 stale lines.

No version changes and no behaviour change; vet and the full test suite pass
unchanged.
2026-07-20 18:53:25 +09:00
flyemoji c4300cb005 feat(updater): authenticate GitHub polling and track the real release repo
felis-api's coord was the placeholder "felis/felis", which resolves against
nothing on real GitHub. It is now MliroLirrorsIngenuity/Felis — the same slug
deploy/bootstrap.sh clones from — so update reporting for the control plane
itself is live rather than parked.

That repo is private today, so the github source gained an optional token, read
from FELIS_GITHUB_TOKEN: the variable bootstrap already needs, so an operator
sets one value once. It comes from the environment and is never compiled in. A
constant would be committed to the very repository it protects, ship inside
every felis binary where strings(1) recovers it, reach every node the image is
imported onto, and need a rebuild and a redeploy to rotate.

Empty stays the correct posture for the other tracked components — k3s and
cloudflared are public — and an empty token sends no Authorization header at
all rather than an empty one.

GitHub answers 404, not 401 or 403, for a private repo the caller cannot see,
so "no token" and "no stable release published yet" arrive as the same status.
On an unauthenticated 404 the error now names both causes and the variable that
fixes the actionable one. With a token already set that hint would be wrong, so
it is suppressed.

Tests pin both halves: the Bearer header is sent only when the token is set,
and the diagnostic names the variable only when it is not.

doc.go's CAVEATS bullet still described the coord as a placeholder and the
component as "dark at runtime". Both were true only until this change; it now
records the real condition, which is that the component resolves like the others
but needs a credential while the repo is private.
2026-07-20 18:53:12 +09:00
flyemoji 659c8e5e9f feat(bootstrap): install the published release build instead of compiling on the host
deploy/bootstrap.sh now resolves the newest published GitHub release, downloads
the binary CI built for that tag, and builds a thin image around it. Compiling
on the target host becomes the fallback and the opt-in, not the default.

The panel is not a separate artifact. The Dockerfile copies panel/dist into
internal/panel/static before the go build, so the control plane — panel and
backend — ships as ONE file. The release channel therefore downloads exactly
one asset, felis-linux-<arch>, and needs no registry, no Go toolchain and no
checkout on the host.

The Minecraft game stack (limbo, lobby, the Velocity plugin) is still always
built locally. game_stack_source now keys on HAVE_PREBUILT_BINARY — the same
flag build_image uses — so on any prebuilt path it unpacks the tar embedded in
that binary instead of trusting a checkout an earlier install left behind.
Trusting the checkout would build the plugin from an old commit against a
freshly downloaded control plane: a silent version skew across the plugin/API
boundary.

Channels:

    (default)                      newest published release, downloaded
    FELIS_VERSION_BOOTSTRAP=dev    clone main and compile
    FELIS_REF=<ref>                pins the tree, forces the source path

The download is best-effort. A tag whose assets are not uploaded yet, an
architecture with no published asset, or an asset that fails validation each
warn and fall back to compiling THE SAME TAG from source — never a different
commit.

The ref is resolved right after install_base, the first point curl exists and
well before docker and k3s, so a missing FELIS_GITHUB_TOKEN or an unpublished
release costs the operator seconds instead of a k3s install they then have to
unwind. It is skipped on exactly the paths that never consume the result: the
TUI, which rebuilds the binary it is already running, and FELIS_SKIP_FETCH,
which builds whatever is staged. Resolving anyway would set FELIS_VERSION to
the newest tag and stamp a staged tree as that release.

The asset is staged next to HOST_BIN rather than in TMPDIR. Validation EXECUTES
it, and /tmp is noexec on CIS-hardened images, where the exec dies 126, the
check reads it as a bad asset, and every such host silently falls back to the
full on-host compile this path exists to avoid. It also keeps a private-repo
artifact out of a world-readable 1777 directory.

git_auth, which supplies the token to git for a private-repo clone, passes an EMPTY
credential.helper before the inline one. credential.helper is multi-valued: a bare
`-c credential.helper=...` APPENDS to whatever the host has configured rather than
replacing it, and an empty value is git's documented list reset. Without it, on a host
with a persistent helper (Git for Windows ships `manager` at SYSTEM scope) two things
go wrong. Git runs `credential approve` automatically after a successful clone and
feeds every helper in the list, so a `store` helper writes the PAT to
~/.git-credentials in cleartext — the token outlives the install, in a file bootstrap
never created and never cleans up. And because the inline helper is LAST, a
pre-existing helper answers `fill` first, so a stale cached credential can win and the
clone authenticates as the wrong account — surfacing as exactly the 404-on-private-repo
the surrounding code works hard to explain. Reproduced both against a real clone, and
confirmed the reset closes both.

internal/panel parses the new stamp. The dev channel now emits "<tag>+g<sha>", which
matched neither describeSuffix ("-N-g<sha>") nor releaseTag, so a dev build fell through
to the default case and the version badge rendered the entire stamp as the release with
no commit. A devSuffix case handles it; the git-describe case stays for hand-rolled
`-ldflags "-X main.version=$(git describe)"` builds. Table test covers both forms plus
the release, dirty and unstamped cases.

CRD application no longer branches on the install path: it is always
`felis bootstrap-assets crd`. That output is byte-identical to deploy/crd/ —
bootstrap_asset.go embeds that very file — and needs no checkout, so one source
replaces a branch whose two arms had to be kept in agreement by hand.

Dockerfile gains a FELIS_VERSION build arg wired into -X main.version, declared
after `go mod download` so a version bump does not invalidate that layer. Both
build stages are pinned to $BUILDPLATFORM so a multi-platform buildx run never
emulates them: the panel's output is architecture-independent and the Go stage
cross-compiles via TARGETARCH. The final stage stays on the target platform and
is COPY-only, which BuildKit performs without QEMU.

.github/workflows/release.yml publishes on a vX.Y.Z tag: vet, tests, then one
buildx run producing both architectures through the repo Dockerfile. Not a bare
`go build` — internal/panel/static holds a tracked placeholder index.html so the
//go:embed compiles without node, which means a direct build succeeds and
quietly ships a release whose panel is that placeholder.

The stamp is asserted end to end, because it fails silently: an unstamped binary
reports "dev", which the updater refuses to compare, disabling update reporting
for every install built from that release. The arm64 artifact is checked by ELF
machine type rather than by running it — runners have binfmt registered, so
executing an amd64 binary misnamed arm64 would succeed.

Prerelease tags are flagged explicitly. The trigger glob is v*, gh does not read
semver out of a tag name, and an RC published as a full release becomes
/releases/latest — the single endpoint the default channel installs from and
`felis update` polls.

No SHA256SUMS. A checksum fetched over the same TLS session, with the same
credential, from the same host as the binary adds no trust root; signing is the
real answer and is a separate decision.

Not verified: the download -> validate -> image -> k3s path has never run on a
host against a real published release, because no tag exists yet. The shell
logic around it is verified out of tree; the network and exec behaviour is not.
2026-07-20 18:53:01 +09:00
flyemoji d146f1ccd7 fix(bootstrap): keep the install alive on a host with only felis-api
restart_existing_control_plane ended in an and-list per deployment:

    [ "$had_api" = "1" ] && kube ... rollout restart deployment/felis-api
    [ "$had_operator" = "1" ] && kube ... rollout restart deployment/felis-operator

As the LAST command of a function, an and-list whose test is false returns 1,
and that becomes the function's exit status. The call site is bare, so under
`set -Eeuo pipefail` the installer dies there — after the bundle has been
applied and before the rollout wait, leaving a half-finished upgrade and no
message naming the cause.

It fires on any host carrying one control-plane deployment but not the other:
felis-api present without felis-operator restarts the api, then exits 1 on the
second test. Both present, or neither, happened to work, which is why it
survived.

Rewritten as explicit `if` statements, which return 0 when the test is false.
Verified out of tree against all four had_api/had_operator combinations.
2026-07-20 18:30:19 +09:00
flyemoji fe4c92c1c5 feat(files): add the server file editor
Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.

felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.

What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.

Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.

The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.

Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:

  * A write accepts arbitrary bytes at any path, so an owner can place a
    loadable plugin jar. This is deliberate — it is what a hosting panel's
    file manager does, scoped to a server the caller already owns and
    already drives through /command — but it is the one owner-tier route
    that lands executable code in a backend pod, since images are
    admin-only and modpack submissions need an admin verdict.
  * config/paper-global.yml is refused on read. felis-lobby's entrypoint
    writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
    identical on every backend, so reading it from a server you own would
    hand you the handshake key for everyone else's. It is the only path in
    the mount that is not the caller's own data, and therefore the only
    denial. The comparison is on the cleaned path, or ./config/... would
    walk straight through it.

Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.

The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
2026-07-20 14:33:29 +09:00
flyemoji 05cb8f6320 feat(cli): report component updates and make the router a data table
Add `felis update`, which reports which platform components have newer
versions available, and route `felis version`, which shipped implemented but
unreachable.

That bug is why the subcommand router is now a map rather than a switch.
cmdVersion existed with nothing dispatching to it and no usage line, so
`felis version` fell through to "unknown command" and no test noticed — a
switch offers no way to enumerate what it routes, so the usage text and the
router could not be compared. As data, they can: a test now walks the
Commands: block and the table in both directions, failing an entry added to
one without the other. bootstrap-assets stays deliberately undocumented and
is listed as such, which makes its absence a decision rather than an
oversight.

The host gatherer answers the two seams NewSysGatherer leaves nil, for the
one caller that can satisfy them without a cluster client. felis-api is
answered from the running binary's own build stamp rather than the
Deployment's image tag: deploy/bootstrap.sh builds the image from the same
checkout it installs /usr/local/bin/felis from and stamps both with one git
describe, so it is the same artifact, and it is the identity `felis version`
reports. Reading the Deployment answers a slightly different question — what
is rolled out — and stays the right seam for the in-cluster path.

Velocity is read from the jar's own META-INF/MANIFEST.MF
Implementation-Version, which is what the proxy reports about itself at
runtime, because bootstrap installs the jar under a fixed name with no
version in it. The filename extractor remains only as a fallback for a
hand-placed velocity-3.5.1.jar.

An unstamped local build reports "dev" and is refused with an actionable
message rather than being treated as 0.0.0, which would make every release
upstream look like an upgrade. The panel and the plugin jars have no version
of their own on purpose: they are embedded in or built alongside the felis
binary, so the felis version is theirs.
2026-07-20 14:32:58 +09:00
flyemoji f36d5b87f6 feat(images): mark platform-curated images and seed the lobby
The create-server form has no way to tell a user which of the whitelisted
images is a sensible starting point. Add 'recommended' as a third
image_whitelist.source alongside 'built' and 'external', and seed it with
the one image that has earned it.

The marker is presentation only. ImageAdmitted still turns solely on
enabled, so a recommended row is admitted by exactly the rule that governs
every other row and carries no extra privilege; a test pins both halves,
because the failure modes are silent and opposite — make admission
source-aware and the curated images vanish from the form, or let curation
bypass the disable switch and an admin who pulled an image finds it still
creatable.

Only one image is seeded, and the restraint is the point. Velocity runs
proxy-wide modern forwarding, so a backend that cannot verify the signed
handshake rejects every login the proxy sends it. The operator injects
FELIS_FORWARDING_SECRET into every backend but cannot make an image consume
it. An arbitrary public Minecraft image therefore passes admission, builds,
schedules, reports Ready — and then refuses every join, with nothing in the
server's status explaining why. Exactly two images read that variable,
deploy/limbo and deploy/lobby; limbo is the login gate and is nonsense as a
base for a user's server, which leaves lobby. The list grows when Felis
ships another forwarding-aware image, not before.

AdmitBuiltImage now preserves a 'recommended' source through its ON CONFLICT
path. Rebuilding a curated tag is the expected way to patch it, and that
rebuild arrives through this exact path, so a blind SET source = 'built'
would demote the curation on the first rebuild with nothing in the request
saying so. AddExternalImage deliberately does not preserve it: an admin
POSTing the ref is an explicit, named re-admission, and the 201 body reports
the Image it constructed without re-reading the row, so a sticky source
there would report a value the database does not hold.

The migration is idempotent via ON CONFLICT DO NOTHING, so an admin who
disabled or re-pointed the row does not have that decision undone on the
next apply.
2026-07-20 14:31:14 +09:00
flyemoji d26acc20ae feat(api): let in-game staff manage any server without claiming it
The web face has always granted staff the run of the fleet (isOwnerOrAdmin
passes an admin for stop/command/console/access on any node), but the
internal face explicitly had "no admin tier": a linked administrator in game
could only wake servers they owned or that autostartPolicy permitted. The
only way to manage another player's (or an unclaimed) server from inside the
game was to claim it — seizing ownership and burning the admin's own quota.

Give authorizeWakeByUUID the admin tier on the same trust anchor the op-login
approve already uses: verified online-mode UUID -> account link -> stored
role. A linked staff member now wakes ANY node under any policy (so
`/felis go` works fleet-wide without claiming); the owner bypass and the
policy gates are unchanged, and an unlinked UUID still fails safe.

Centralize the staff-role rule while at it: staffRole(role) in auth.go
(admin, plus owner as its superset) now backs Principal.IsAdmin, the session
ViaAdminAccess grading, the op-login approve gate and the new wake tier.
That also fixes a real hole in the approve gate, which required role=admin
exactly: an Owner manually promoted to role='owner' per migration 0011's
upgrade note would have been refused by their own in-game approval door.

The lobby menu still renders "Claim & Start" on ownerless tiles — claiming
becomes optional for staff rather than the only entry — so the velocity
plugin needs no change.
2026-07-20 11:10:23 +09:00
flyemoji f0b79e9edd feat(mail): deliver email one-time codes over SMTP and add the setup email screen
Felis never actually sent mail: OTP codes for onboarding, email login and
op-login were only written to the felis-api log behind a "demo has no SMTP"
limitation, and the Settings/SMTP flow those comments promised was never
built. Combined with the bootstrap Owner's address being recorded unverified
(87279a1), op-login start always took the anti-enumeration neutral branch and
minted a fake request_id, so the in-game approve inevitably answered "No
pending operator sign-in with that code".

Give the codes a real delivery path, configured in felis.toml rather than a
web settings page so config keeps a single source of truth:

- config: new [smtp] table (host, port defaulting to 587, from, username,
  password_ref). Validation requires a plausible from address and a sane
  port; the password itself never enters the config file.
- internal/mail (new): stdlib net/smtp mailer implementing the api.OTPMailer
  seam. Port 465 dials implicit TLS, other ports upgrade via STARTTLS when
  advertised; AUTH only when a username is configured (PlainAuth itself
  refuses plaintext, so the password cannot leak to a TLS-less relay).
  Ping() proves reachability and credentials without sending mail. The
  message shape (CRLF, Q-encoded bilingual subject) is pinned by test.
- platform: felis-smtp Secret constants and an optional FELIS_SMTP_PASSWORD
  env var on the felis-api Deployment, mirroring felis-uploads-s3.
- cmd/felis api: construct the real mailer when [smtp] is configured; keep
  the log fallback otherwise and say so at startup. Warn when a username is
  set but the credentials env is empty.
- setup TUI: "e" on the summary/status screen opens the email form (host,
  port, from, optional auth). Apply order: Ping preflight, [smtp] into both
  host and pod config files, felis-smtp Secret piped to kubectl via stdin,
  config Secret, felis-api rollout. A failed preflight leaves the install
  untouched. SMTP is deliberately not a wizard rail step: first-run stays
  mail-less by design, and the passkey minted at onboarding is the pre-SMTP
  owner credential.

Also make PGRepo.UserByEmail match case-insensitively (lower(email) =
lower($1)), honoring the interface contract and the users_verified_email_
unique partial index; the fake repo already matched with EqualFold.

Existing installs need the felis-api Deployment manifest re-applied (e.g. a
bootstrap re-run) before the new env var exists; a rollout restart alone
cannot add it.
2026-07-20 10:24:52 +09:00
flyemoji 7860152f57 feat(auth)!: go fully passwordless and fix cross-check review findings
Remove password authentication everywhere; the only session doors are
passkey (WebAuthn), email OTP, in-game bind codes, QR scan-login, and
op-login vouching. Remediates the 33-finding cross-check review across
backend, CLI, panel, plugins, and docs.

Backend/CLI:
- Drop password routes and fields from account/user/onboard/auth
  handlers; align tests (new account subtests, naming reserves
  "console", op-login/onboard/qr-login test updates).
- Add migrations 0016_op_login.sql and 0017_drop_password.sql.
- Thread panel/admin hostnames from hostcfg through api.go,
  setup_panel.go, tui_root.go and tui_preflight.go instead of
  hardcoding; bootstrap.sh writes panel-hostname/admin-hostname
  into felis.toml.
- Reword breakglass and TUI copy for passwordless flows.

Panel:
- Delete the ChangePassword page and all password UI; align
  login/auth/api/types with the passwordless contract; add the
  migration and op-login approval flows.
- i18n: convert ImageBuildPage durations/status badges and
  ServerLuckPerms strings to translation keys; drop 72 orphan keys
  per locale; unify the title as "Felis - Console".

Plugins (all six rebuilt):
- Velocity waiting router returns 503 at_capacity during wake;
  MOTD/control-channel copy and config comments.
- Paper zh menu title; Limbo bind-code TTL 600s with panel_url
  preference; unified /link lines in fabric/forge/neoforge; shared
  link-client javadoc contract fixes.

Docs: openapi.yaml, sequence-diagrams.md, deploy/limbo/README.md and
plugins/README.md aligned with the implementation.

BREAKING CHANGE: migration 0017 irreversibly drops
users.password_hash and users.must_change_password; password login
cannot be restored after migrating.
2026-07-20 04:47:32 +09:00
flyemoji c96b36a41f fix(bootstrap): retry transient Fill failures when resolving build jars
A single HTTP 502 from fill.papermc.io aborted the entire bootstrap. Build
resolution used a one-shot curl, so one gateway blip from an upstream that
flaps was indistinguishable from a permanent failure, and the run died before
Docker, k3s or any game server was provisioned.

Pass --retry 5 --retry-delay 2 to the build-resolution fetches. 502/503/504
are already in curl's built-in transient set, so the tool had solved this; the
flags were simply never passed. papermc_latest_jar is shared by the Paper and
the Velocity resolve, so hardening it once covers both callers. The
LOOHP/Limbo CI metadata fetch feeds the same step and gets the same treatment.

Deliberately no --retry-connrefused. It only adds ECONNREFUSED to a set that
already covers this incident, and it needs curl 7.52.0 while the yum (el7)
path the script supports ships 7.29.0, where an unrecognised long option is a
parse error rather than a warning:

  # centos:7, curl 7.29.0
  $ curl -fsSL --retry 5 --retry-delay 2 --retry-connrefused https://example.com
  curl: option --retry-connrefused: is unknown

Under set -Eeuo pipefail that exits 2 and trips the || die, so both hardened
fetches would hard-fail on a host where they used to work, each naming a cause
that is not the real one. A comment above papermc_latest_jar records this so
the flag does not come back.

The failure message was actively misleading. "no Paper build for Minecraft
26.2 (the login gate speaks only that protocol)" reads as "that Minecraft
version is unsupported", sending the reader after a version-pinning problem
that does not exist: Paper 26.2 build 60 resolved fine minutes later. Say what
is actually known instead, that the build likely exists and Fill is flapping.

Verified against a local always-502 server: curl now issues 6 requests
(1 initial + 5 retries) over 10.1s before giving up, where it previously
issued 1 and died. Verified on centos:7 that this flag set is accepted, and
against the live Fill v3 API that the resolve still returns a jar URL.

Known and deliberately unchanged: no fetch sets --max-time, so an upstream
that accepts a connection and never answers still blocks forever. --retry does
not cover that, as it fires only once a request completes with a failure. The
remaining single-shot downloads (cloudflared, the Docker GPG key and repo
list, k3s, the Temurin JRE, the Velocity jar, the Go toolchain) keep their
existing no-retry shape rather than widen this diff on a script that is about
to provision a live host.
2026-07-19 19:33:48 +09:00
flyemoji 87279a1366 fix(setup): record email unverified so onboarding works without SMTP
At bootstrap there is no SMTP, so the old /setup flow was unreachable: it
requested an emailed OTP that could never arrive. Setup now records the
Owner's email address unverified (no OTP round-trip) and requires a passkey,
deferring SMTP configuration to a later Settings page. Setup completes on
email-recorded + passkey-enrolled, and the lockdown lifts on the passkey, not
on email_verified: a passkey is the Owner's only pre-SMTP login credential
(email-OTP login refuses admin accounts).

The record-email endpoint (POST /account/email) now clears email_verified in
the same write. Only VerifyEmailOTP, which proves control of the address, may
set that flag; recording a fresh unproven address must never leave a stale
email_verified=true asserting a proof the user never gave. The change strictly
tightens the invariant, so no existing reader breaks.

Remove the dead ErrEmailTaken path and its documented 409: no migration puts a
unique index on users.email and the codebase does not enforce email
uniqueness, so the unique-violation branch was unreachable and the 409 an
impossible response.

The /setup route (Setup.tsx, setEmail helper, setup i18n copy) is rewritten to
match: record-email, mandatory passkey, no skip-for-now. The SMTP settings
page and post-setup configure-SMTP nudge are deferred.
2026-07-17 01:41:07 +09:00
flyemoji 93190e7a5b fix(cfsetup): regenerate missing tunnel credentials on re-bootstrap
cloudflared writes the tunnel credentials JSON only at `tunnel create`. An
idempotent re-run against a tunnel that already exists — or a reset +
re-bootstrap where the old box's ~/.cloudflared was wiped but the
Cloudflare-side tunnel survived — finds no local credentials file, and the
connector crash-loops with "Tunnel credentials file doesn't exist". A tunnel
that never comes up leaves op.console unreachable, so the one-time setup link
minted just before it ages out (30-min TTL) unredeemed.

CreateTunnel now resolves the tunnel id on both paths (fresh create and
already-exists) and routes through ensureCredentials, which re-fetches the
token with `cloudflared tunnel token --cred-file` (authenticating via cert.pem,
preserving the same id / DNS / Access) when the file is absent. The secret is
written to the file, not stdout, and the file is chmod 0600 so it is not left
world-readable next to cert.pem.

The self-heal is unconditional on re-bootstrap: Setup gates on Pre.check()
(cert.pem present) before CreateTunnel, so the token re-fetch always has its
cert.pem authority.
2026-07-16 23:23:42 +09:00
flyemoji 5fbb1db4ec feat(setup): add the op.console owner onboarding wizard
The /setup route redeems the one-time token from `felis setup`, then walks
the Owner through email-OTP verification and passkey enrollment before
handing off to the console. It sits outside RequireAuth — the visitor
arrives without a session and the redeem is what mints one — and is
reload-safe: a spent token resumes from the surviving session via
/auth/setup/status.

Adds the Setup page and its /setup route, the setup API client methods
(redeem/status), and the en-US/zh-CN onboarding strings.
2026-07-16 18:03:58 +09:00
flyemoji 60732a6283 feat(operator): gate op.console to staff and land owner setup there
The operator console (op.console.<root>) requires internal permission
verification on top of Zero-Trust: a passkey is not access. requireExternal
now refuses any non-admin principal arriving on the admin host, before any
handler, so op.console is staff-only at the door rather than per-route —
including on the passwordless demo face where Cloudflare Access is not in
front. The gate is inert on the player console (console.<root>).

Owner first-run setup is staff onboarding, so `felis setup` mints the
one-time setup URL on op.console.<root>/setup (was console.<root>). The
passkey verifier lists both console and op.console in RPOrigins so the
one-time binding asserts on either face under the shared console.<root>
RP-ID.

Session admin-access now includes role=owner, not only admin: the owner is
a superset of admin, so excluding it left IsOwner() unreachable through a
passwordless session. No path assigns role=owner yet — this is forward
consistency.

The bootstrap summary now names console.<root> the player panel and
op.console.<root> the operator console where the Owner runs setup, fixing
text that told operators not to run setup there.

Tests: op.console door gate (non-admin refused, player console unaffected,
admin passes) and owner session admin-access; the setup-bind default-host
test follows the move to op.console.
2026-07-16 18:03:30 +09:00
flyemoji 26b62e91cd feat(bootstrap): federate LittleSkin by default and let Velocity own its runtime dir
Point Velocity's authlib (mojang.sessionserver) at the felis-api hasJoined multiplexer so
a full install federates Mojang plus the configured [[auth_source]] set out of the box,
not just the standalone `felis nano`, and ship a default LittleSkin auth_source in the
generated felis.toml (delete the block for a Mojang-only server). Velocity now owns
velocity.toml and its working tree — it migrates the config version and extracts
localizations on start — while the jar and forwarding secret stay root-owned read-only and
ReadWritePaths widens to VELOCITY_DIR; the felis-api-internal ClusterIP lookup is factored
into a felis_internal_ip helper shared by the link config and the sessionserver override.
Re-include plugins/velocity in the docker context because the felis binary now embeds all
four plugin trees for `felis bootstrap-assets game-stack`.
2026-07-16 13:27:04 +09:00
flyemoji 9ea35304e3 fix(setup): source the console host from the panel hostname, not op.console
The owner setup URL and the limbo login link were built from the admin host
(op.console.<root>, with an op.console.localhost fallback) and a hardcoded
console.<root>, so an operator who set a custom panel_hostname got an unreachable setup
link and a wrong login target. Thread the resolved panel host (defaultPanelHostname)
through performSetupMCBind, the MC-bind TUI, and the login system-server env
(new FELIS_PANEL_HOSTNAME); the limbo plugin prefers it and keeps console.<root> only as
the fallback for an older operator whose env predates it. This also matters for security:
the only wired WebAuthn verifier is scoped to the panel host, so passkey enrollment must
land on the panel face, never op.console.

While here, the limbo login handler checks link status before minting a bind code: an
already-linked player is sent straight to the lobby instead of being shown a useless code.
2026-07-16 13:27:02 +09:00
flyemoji b5cd4501e5 fix(passkey): serve flat WebAuthn options to register and username login
go-webauthn marshals CredentialCreation/CredentialAssertion as {"publicKey": {...}},
but the panel's register (Account.tsx) and username-first login (Login.tsx) read the
options flat (options.challenge, options.user.id), so base64urlToBytes(undefined) threw
"Cannot read properties of undefined (reading 'replace')" and neither ceremony could
start. Strip the envelope in the register-begin and username-login-begin handlers via a
small unwrapPublicKey helper; discoverable login keeps the envelope because it reads
options.publicKey.* plus a top-level options.login_id. The begin tests now feed a wrapped
body and assert the handlers return it flat, so they genuinely exercise the unwrap.
2026-07-16 13:26:56 +09:00
flyemoji be0c4c41f4 feat(bootstrap): install authenticated game stack 2026-07-14 02:58:21 +09:00
flyemoji dab8fc214b feat(setup): bind owner through login gate 2026-07-14 02:57:15 +09:00
flyemoji 5dc8eb92a8 feat(operator): secure system server workloads 2026-07-14 02:56:19 +09:00
flyemoji 688c86da18 feat(proxy): enforce login-first routing 2026-07-14 02:55:11 +09:00
flyemoji 16b7ad2756 feat(runtime): add authenticated system backends 2026-07-14 02:54:13 +09:00
flyemoji fd062882ed feat(nano): give a Mojang player's name back to them, by prefixing the squatter
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.

The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.

A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.

The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.

prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.

Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.

Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:

  premium FLYEMOJ1     -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
  LittleSkin FLYEMOJ1  -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
  both online at once, zero "already connected" rejections
  LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3

The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.

Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.

BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
2026-07-13 13:00:54 +09:00
flyemoji b323975ddb fix(nano): -Dmojang.sessionserver takes the full hasJoined URL, not the base
d417efc got this backwards, in both the code comment and the installer summary.
It claimed authlib appends /session/minecraft/hasJoined itself, so the property
should be given the base URL only. Velocity does not work that way, and a real
login says so: pointed at http://127.0.0.1:8081, a Mojang login arrives at nano
as

    GET /?username=FLYEMOJ1&serverId=-23ae0b50...

with no path at all. Velocity appends the query string to the property verbatim
and issues the request itself; authlib is not in the loop. nano has no route on
/, so it answers 404 and Velocity kicks the player with authservers_down.
Velocity's own default for the property is the full URL,
https://sessionserver.mojang.com/session/minecraft/hasJoined, which is the same
thing said another way.

With the full endpoint URL the same account logs straight in, so both the
nano.go header and summary_nano now print

    -Dmojang.sessionserver=http://127.0.0.1:8081/session/minecraft/hasJoined

and note that the flag belongs between `java` and `-jar`.

Verified against Velocity 3.5.1 + Paper 26.2 on the deploy host: a Mojang login
reaches the backend with its real Mojang UUID unchanged, and a LittleSkin login
under the same username reaches it as UUIDv3(felisAuthNS, "littleskin:"+id) --
two different players on the backend, which is the point.
2026-07-13 11:05:45 +09:00
flyemoji d417efc8bf fix(nano): make the installer work on EL10 and stop serving hasJoined to the world
Deploying `felis nano` to a real Rocky Linux 10 host surfaced four defects that
no local check could see. Fixed together because they all sit on the same path
from `curl|bash` to a running felis-nano.service.

* docker killed the nano install on EL10. `acquire_nano_binary` pulled in
  docker purely to build the binary; on Rocky 10.2 the docker-ce el10 rpms
  install but dockerd refuses to start, so the install died at
  `systemctl enable --now docker`. nano needs one static binary, not an image,
  so the docker dependency is gone: fetch_source -> install_go_toolchain
  (pinned FELIS_GO_VERSION, default 1.26.4, amd64/arm64) -> build_nano_binary.

* the built binary could not be exec'd by systemd (203/EXEC). The Go linker
  renames its output out of $TMPDIR, and a same-filesystem rename carries the
  source SELinux label, so `go build -o /usr/local/bin/felis` produced a binary
  labelled user_tmp_t rather than bin_t. root is unconfined and could run it by
  hand, which is what made this look fine, but the DynamicUser service could
  not. build_nano_binary now stages the output and installs it as a fresh file
  so the policy type transition labels it bin_t, with restorecon as a belt.

* re-running the installer did not converge. `systemctl enable --now` is a
  no-op on an already-active unit, so a rebuilt binary was installed while the
  old process kept running. Now enable + restart.

* the Velocity wiring comment in cmd/felis/nano.go was wrong. authlib appends
  /session/minecraft/hasJoined itself, so -Dmojang.sessionserver takes the base
  URL only, as the installer has always printed.

Also bind to loopback by default. hasJoined is unauthenticated by protocol --
authlib speaks the vanilla sessionserver dialect and sends no token -- so a
public bind is an open auth relay: anyone can point their own proxy at it and
spend this host's egress IP on Mojang until Mojang rate-limits it and the
operator's own players stop getting in. It is not an identity bypass (a caller
still needs a serverId hash bound to their own server key, which the upstream
Yggdrasil validates), but it is someone else's traffic on your address.
FELIS_NANO_LISTEN and the -listen flag now default to 127.0.0.1:8081, which a
same-host Velocity reaches unchanged; serving an off-host proxy is an explicit
opt-in. configure_nano_firewall no longer opens a port for a loopback bind, and
summary_nano prints the real bind address plus the relay warning.

Verified on the target host: installs with no docker present, service active,
binary labelled bin_t, `ss` shows LISTEN 127.0.0.1:8081, an external request is
unreachable, and an in-host request returns 204 with the login logged.

BREAKING CHANGE: felis nano defaults to 127.0.0.1:8081 instead of 0.0.0.0:8081.
A Velocity proxy on another machine must now set FELIS_NANO_LISTEN (or -listen)
to a reachable address, and should allow that port only from the proxy's IP.
2026-07-13 10:30:18 +09:00
flyemoji 87aa9ea625 feat(deploy): bootstrap can install Felis-nano only, chosen at the start prompt
bootstrap.sh now asks up front whether to install the full Felis control
plane or only Felis-nano, and grows a parallel install path for the
nano-only case.

- prompt_install_mode() runs right after OS detection and reads /dev/tty
  (so it works under `curl ... | sudo bash`) offering [1] Felis / [2]
  Felis-nano, default full. FELIS_INSTALL_MODE=full|nano skips the prompt
  for non-interactive runs; no tty falls back to full.
- main_nano() installs only what nano needs: the felis binary (reusing
  the embedded-binary / docker-build acquisition), a template felis.toml
  carrying a commented [[auth_source]] example (Mojang-only until edited),
  a felis-nano.service unit running `felis nano -config ... -listen ...`
  under DynamicUser hardening, and a firewalld port-open for the listen
  port. None of the k3s / Postgres / migrate / bundle steps run.
- write_nano_config is idempotent (leaves any existing config untouched)
  and its template is valid as-is. summary_nano prints the hasJoined
  endpoint and the Velocity -Dmojang.sessionserver flag, offering the
  127.0.0.1 form when the proxy is on the same host.

Verified on WSL: `bash -n` clean; the emitted template loads via
config.LoadNano and the exact systemd ExecStart command serves 204 on a
miss ("Mojang + 0 third-party source(s)"); a duplicate-tag config still
exits non-zero citing "unique". Not exercised: a full main_nano run,
systemd activation of the unit, and shellcheck (unavailable in this env).
2026-07-13 02:39:50 +09:00
flyemoji d177428fb0 feat(felis): add felis nano — Yggdrasil hasJoined multiplexer without a control plane
`felis nano` serves the vanilla sessionserver protocol
(GET /session/minecraft/hasJoined) as a federating multiplexer over
Mojang plus any number of third-party Yggdrasil roots, with no k3s,
Postgres, or panel — a MultiLogin-style auth front-end delivered as a
subcommand of the single felis binary rather than a separate build.

- config.LoadNano reads only [[auth_source]] blocks; it skips the
  database.url / root_domain / archive requirements the full server
  needs. Zero sources is valid (Mojang-only).
- Mojang is prepended in code (Identity:true), never from config, so it
  is always the sole identity root. Third-party profiles are rewritten
  to canonical = UUIDv3(felisAuthNS, tag+":"+nativeID).
- validateAuthSources rejects unknown keys, duplicate tags, and
  scheme-less URLs — a malformed nano config fails loud at load.
- Reuses api.HasJoinedHandler with a stub Repo (no blacklist backend);
  a rejected login is a 204, matching the vanilla sessionserver.
- nano.go binds the -listen flag and ignores [server] listen in config.

Verified on WSL (go1.26.4): go build/vet/test ./... green; a runtime
smoke against the template config returns 204 on a miss and logs
"Mojang + 0 third-party source(s)"; a duplicate-tag config exits
non-zero citing "unique".
2026-07-13 02:39:41 +09:00
flyemoji ecea20ee7c feat(nano): configure hasJoined auth sources via [[auth_source]], Mojang-anchored
Step 2 of Felis-nano: a [[auth_source]] array-of-tables (tag + full hasJoined
url, config order = priority) supplies the multiplexer's third-party Yggdrasil
roots; cmd/felis prepends Mojang as the sole code-owned identity anchor and
wires them into API.AuthSources. With no sources configured the endpoint stays
inert (204s), unchanged from step 1.

The config deliberately has no identity/trusted field: Mojang is the only source
whose self-asserted UUIDs are trusted verbatim, so no misconfiguration can
reopen the impersonation hole the per-source UUID rewrite closes. An identity=
key is an unknown key and Load rejects it. Validate adds two fail-fast guards:
unique tags (namespace collision) and a scheme-qualified url (else the source is
silently dead, never validating any login).
2026-07-12 02:27:02 +09:00
flyemoji ed8fa0cdbf docs(changes): record the Felis-nano hasJoined resolver and reconcile the audit ledger row 2026-07-12 02:10:49 +09:00
flyemoji ff550c41ef feat(nano): federating hasJoined multiplexer with per-source UUID namespacing 2026-07-12 02:09:47 +09:00
flyemoji 667c6d33d2 docs(changes): record the adversarial input-validation audit (sink-first negative-path) 2026-07-08 12:38:40 +09:00
flyemoji b7b4a3be45 docs(changes): record the round-2 backup/restore mutation audit and index the owner-gate test
Round-2 mutation audit of the backup/restore data-safety surface (7 fail-open gates pinned, 1 coverage gap found and closed by the owner-gate test in 85b8a92). Reverts the Pending section and indexes c67a4d3 + 85b8a92 into the committed ledger.
2026-07-08 05:59:13 +09:00
flyemoji 85b8a92a0e test(api): pin restore's owner gate against a superseded former owner
handleRestoreBackup's owner-or-admin gate was not pinned by any test: the former-owner gate backstopped every non-owner case the suite exercised, so a broken owner gate would not redden. Add the mirror of the former-owner test — a released former owner (still the backup's former_owner, no longer the current owner) must get 403 — the sole subtest that fails when the owner gate is disabled. Found by the round-2 backup/restore mutation audit; production code unchanged.
2026-07-08 05:58:00 +09:00
flyemoji c67a4d3f35 docs(changes): close §B4 with the S3 archive backend deferred by design
The operational break-glass ops (Owner provision, OP create, halt, Sync backup) are built and oracle-verified. The fourth §B4 line item — the tarS3 archive backend — is recorded as a deliberate deferral, not a silent gap: config.Load fail-closes store=tarS3 (frozen by TestLoadRejectsUnimplementedArchiveStore), tarLocal is the tested baseline every backup/restore path uses today, and offsite/cross-cluster DR is opt-in future work. Supersedes the "S3 remains open" note in the phase-2b peer doc.
2026-07-08 05:56:59 +09:00
flyemoji 729bd7ba0f docs(changes): index the mutation audit and two lagging ledger rows 2026-07-07 23:45:50 +09:00
flyemoji 4626ab580e docs(changes): mutation-audit the ledger's "unit-tested" safety claims
Break each load-bearing safety gate the change ledger names as "unit-tested" and
confirm the specific test turns red — passing proves GREEN, not that the test would
catch a regression. All 18 fail-open crown-jewel gates across the subsystems (auto-update
pin/no-downgrade/no-prerelease/window, passkey clone-refuse, modpack CAS + Trivy scan-gate,
cfsetup fail-closed + NodePort fence conn-count, SSE cap, OTP atomic reserve, idle stop,
startup/readiness timeout, /readyz deps, naming reservation, service-token login-pod-only)
are mutation-proven to pin behaviour. Every documented "unit-tested" claim is reconciled
one-for-one; the fence conn-count gate, previously mis-classified as integration-only, is
corrected and verified. No code changed — read-and-verify only.
2026-07-07 23:45:12 +09:00
flyemoji 5a7cd5abe0 docs(changes): fold 346ec68 cloudflare-edge walkthrough into its detail doc
A completeness re-check of the ledger backfill (in-scope pre-fad48ff
feat/fix/refactor commits vs SHAs actually cited in detail docs, excluding
the auto-generated ledger table) surfaced one backend functional commit with
no detail-doc home: 346ec68 refactor(deploy) — an 868-insertion rework of the
cloudflare-edge TUI walkthrough plus a tested cfsetup integration-runner path.
The "improved cloudflare walkthrough" subject undersold a behaviour change, so
it is folded into the existing cloudflare-tunnel-access-edge detail doc and its
INDEX row rather than left orphaned. Backend detail-doc coverage of the
pre-ledger history is now complete (0 backend orphans).
2026-07-07 21:05:24 +09:00
flyemoji 096d59716e docs(changes): backfill detail docs for pre-ledger functional commits
Retroactively author 15 grouped detail docs covering the backend
functional (feat/fix) commits made before the change ledger was
established (fad48ff), closing the ledger's detail-doc axis for the
pre-convention history. Each doc groups a feature's constituent commits,
lists their SHAs with subjects, and carries a backfill note stating it
was reconstructed from git history on 2026-07-07 and not independently
re-verified (current tree green at 9911b8c).

Add a Detail docs section to INDEX.md linking every detail doc (the 6
existing + 15 backfill) to the commit(s) it covers, so a doc is findable
from the index without a column on the auto-generated ledger table. Catch
the table up with the missing 9911b8c row.

Scope: backend (Go/Java/K8s) only, per the ledger's stated convention
that frontend/panel commits are the collaborator's UI work; non-functional
commits (docs/style/chore/refactor) keep their table row without a
dedicated detail doc.
2026-07-07 20:55:51 +09:00
flyemoji 9911b8cd41 docs(changes): record the break-glass backup console peer (§B4 Sync phase 2b) 2026-07-07 19:34:00 +09:00
flyemoji fc748d3462 feat(breakglass): add "back up a world now" console peer (§B4 Sync)
Adds a break-glass console operation that snapshots a stopped world by
calling the felis-api internal face while the API is alive, rather than
rendering the backup Job locally: the Job needs felis-api deployment
coordinates the console does not hold.

The peer resolves the felis-api-internal ClusterIP Service + service
token from the control namespace, POSTs the internal backup endpoint
with the operator os_user for audit attribution, and maps 409/503/404
to friendly outcome cards. Core decision logic lives in backupnow.go
(unit-tested against a fake client + httptest); tui_backupnow.go is the
untested bubbletea glue mirroring tui_halt.go.
2026-07-07 19:32:13 +09:00
flyemoji 4918d93412 docs(changes): record the felis-api internal-face ClusterIP Service fix
Detail doc + ledger row for 2ba9948: the separate felis-api-internal ClusterIP
Service that gives the login pod (and the break-glass console) a routable 8081.
2026-07-07 11:24:38 +09:00
flyemoji 2ba994889e fix(platform): front the felis-api internal face on its own ClusterIP Service
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.

Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).

Manifest-level fix; the live packet path is pending real-cluster verification.
2026-07-07 11:24:09 +09:00
flyemoji 73195ca48f docs(changes): record the internal-face backup endpoint (§B4 Sync phase 2a)
Detail doc + ledger row for f2fc57c: the internal (service-token) twin of the
on-demand world backup endpoint that the break-glass console peer will call.
2026-07-07 10:48:51 +09:00
flyemoji f2fc57cad9 feat(api): add internal-face break-glass world backup endpoint (§B4 Sync)
Add POST /api/v1/internal/servers/{name}/backup so the on-node break-glass
console can snapshot a stopped world while felis-api is alive. It goes through
the API (not direct-to-CRD like halt) because rendering the backup Job needs
deployment coordinates (FELIS_IMAGE, FELIS_BACKUP_PVC) only felis-api holds.

Service-token auth (no Principal); the middleware IS the authorization, since
the operator already has root on the node. Refactor the RWO stopped-gate,
optional-Backuper 503, async hand-off and audit+202 into a shared enqueueBackup
tail so the external (owner/admin) and internal (break-glass) faces cannot
diverge on the security-critical stopped-gate. The internal audit is attributed
to break-glass/internal so a console-initiated backup is distinguishable from an
owner self-service one.
2026-07-07 10:48:01 +09:00
flyemoji c3b4d7b957 docs(changes): backfill 7a7c0d5 into the change ledger
Move the on-demand world backup entry from Pending into the committed
ledger and record its SHA in the detail doc.
2026-07-07 10:05:58 +09:00
flyemoji 7a7c0d53ab feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
2026-07-07 10:04:30 +09:00
flyemoji fad48ff21d docs(changes): establish the change ledger for functional changes
Add docs/changes/ — a durable, in-repo map of every functional change and
the commit that records it, independent of git log. INDEX.md carries the
convention (each functional change gets a dated detail doc plus a ledger
row) and the full oldest-first ledger, regenerable losslessly from git.
Seed detail docs for the two changes just landed: the break-glass halt op
(c2ee21a) and the /felis migrate command (c1aa38b).
2026-07-06 23:52:12 +09:00
flyemoji c1aa38bac1 feat(velocity): add /felis migrate to open an account migration (§B3 inherit)
Add the in-game /felis migrate command that a player runs to open an
account migration, the entry point for handing their owned servers to
another account (§B3 inherit, scenario A). The command posts the player's
Mojang-verified UUID to the existing handleMigrateStart backend, which puts
the account into migrate mode; the player then finishes on the web console
(prove identity, name the receiving account, redeem a one-time code).

Mirrors the existing /felis claim path: requires a real player past login
limbo, acts on the caller's account rather than the current server, expects
201 Created affirming started=true (a 201 without it is a contract breach,
not a refusal), and maps the handler refusals (404 not_linked, 409
account_retired) to player-facing guidance. On success it points the player
at https://console.<root_domain>, derived from config, never hardcoded.
Compile-verified against velocity-api:3.3.0-SNAPSHOT via the podman gradle
toolchain. Closes the code-only gap named in handlers_account_migrate.go.
2026-07-06 23:50:57 +09:00
flyemoji c2ee21ae05 feat(breakglass): add halt-a-server op to the recovery console (§B4)
Add a root-gated "Halt a running server" operation to the break-glass
console. The operator picks a server from the live fleet and the console
flips that MinecraftServer CRD's spec.desiredState to Stopped via a
spec-only merge patch, disjoint from the operator's status writes, so it
cannot race or clobber reconciliation. It is the panel-independent
emergency stop for when the box still has root plus a kubeconfig.

System servers (login/lobby) are allowed but flagged: a system tag in the
picker and an explicit WARNING in the post-exit summary, since halting
login takes the shared auth front door down with no fallback. Audit is
best-effort so a halt still works with the audit sink down. Already-stopped
is a distinct no-op. The core (halt.go) is unit-tested against a real
controller-runtime fake client that applies the patch.
2026-07-06 23:50:09 +09:00
flyemoji fdb6efbd88 feat(account): migrate a live account's owned servers to a new account (§B3 inherit)
Old account runs /felis migrate in-game to open a migration, proves control via a
fresh web step-up (passkey forced when enrolled, else email-OTP), names the target
and mints a one-time code. The target redeems it while authenticated AS that target:
in one transaction the source's owned servers re-point to the target and the source
is retired (sessions revoked, disabled, soft-deleted), which also spends the code so
it cannot be replayed. Only server ownership moves; the mc_uuid link and web
credentials stay with the source, so migrate is not a credential-theft primitive.

- 0015 migration: account_migrations state machine (initiated -> confirmed ->
  code_issued -> redeemed), one live migration per source
- Repo/PGRepo: Start/ForSource/Confirm/IssueCode/Redeem
- 8 routes (1 internal /felis side, 7 web) with openapi parity
- passkey step-up runs the same clone-signal (sign-count) check as the login door
- code bound to the named target at issue and at redeem

Quota is grandfathered at redeem: no per-target quota re-check when servers move.
2026-07-05 20:50:48 +09:00
flyemoji 9e1df12975 feat(passkey): advance sign_count, reject clone-warned assertions
Both login doors (username-first and discoverable) now run a shared applyAssertionCounter after a verified assertion. A signature-counter regression — go-webauthn's CloneWarning, the possible-cloned-authenticator signal — is refused fail-closed with the same opaque passkey_login_invalid envelope any other finish failure returns (no clone oracle to a prober) and audited distinctly as auth.passkey_clone_rejected under the resolved account. A clean assertion advances the stored sign_count to the asserted value and stamps last_used_at, before any session is minted.

Counter-less/synced authenticators report 0 and never warn, so they pass through and simply re-stamp 0; the check gates only counter-keeping hardware authenticators, where a rollback is the meaningful signal. Email-OTP and username-first passkey remain fallbacks, so a rejected clone is never bricked.

Adds Repo.AdvanceCredentialSignCount (pgrepo UPDATE by credential_id) and surfaces CloneWarning from the internal/passkey adapter's FinishLogin/FinishDiscoverableLogin. Proven by real-crypto adapter tests (a counter regression still verifies but flags CloneWarning), handler tests (advance-and-stamp on success, fail-closed on clone), and a symmetric test on each door so both call sites of the shared helper are covered.
2026-07-05 16:02:30 +09:00
flyemoji 0dbd557a7a fix(store): renumber discoverable-login migration 0013 -> 0014
A resource-cache migration (0013_resource_cache.sql) was merged onto main concurrently and also claimed version 0013. LoadMigrations rejects any duplicate migration version, so the app would refuse to boot with both files present.

Renumber the discoverable-login migration to 0014. The two migrations touch disjoint objects (0013 ALTERs servers to add cached_* columns; this one CREATEs webauthn_discoverable_challenges), so their relative order does not matter, and the table name is unchanged -- no Go reference moves.

Renaming a just-published migration is safe here because neither version has been applied to a persistent database yet: there is no schema_migrations row for version 13 to reconcile. This is a pre-application renumber, not a history rewrite of an already-applied migration.
2026-07-05 04:31:45 +09:00
flyemoji 154002edf3 docs(auth): cite MultiLogin reference for UUID-keyed reclaim split
Anchor the username-collision reclaim's UUID-keyed, proxy-detected design to
the multi-Yggdrasil reference: CaaMoe/MultiLogin v6 binds identity as
serviceId+online-UUID via "identity cards" that decouple the in-game name from
online identity — keyed by UUID, never by name. Note that §B3's Mojang-priority
reclaim goes beyond the common "protect the first-bound name" behavior by
evicting a squatter once the genuine Mojang owner appears and stashing the
squatter's data for the code-only inherit path.
2026-07-05 04:05:56 +09:00
flyemoji ec468baef9 feat(auth): add discoverable (usernameless) passkey login
A from-zero login door: the browser calls navigator.credentials.get() with an
empty allowCredentials, the authenticator returns an assertion carrying the
resident credential's userHandle, and the server resolves the account from that
handle alone — nothing is typed or client-named.

Routes (both Public):
  POST /api/v1/auth/passkey/login/discoverable/begin
  POST /api/v1/auth/passkey/login/discoverable/finish

Begin stashes the ceremony SessionData server-side keyed by an opaque login_id
under a global cap; finish consumes it single-use, hands the
authenticator-revealed userHandle to a UserByID resolver, and mints a session
only for the account the assertion actually verified to. Every finish rejection
— no live challenge, expired, bad assertion, unresolvable handle — collapses to
one passkey_login_invalid envelope, so finish is never an existence/state
oracle. SignCount is surfaced but not yet consumed, exactly as the
username-first door, so the from-zero path offers no clone-detection bypass.

The discoverable VERIFY path is Oracle-verified end to end against a virtual
authenticator (internal/passkey): it resolves the account from the signed
userHandle, fails closed when the handle names no account, and rejects an
assertion signed by a credential not bound to the resolved user — the
impersonation guard unique to usernameless login. Enrollment now requests a
resident key (authenticatorSelection.residentKey=preferred), the only
server-side half a unit test can pin.

Whether an authenticator actually stores a resident key is a device property no
test can reach, so this door is INERT for a credential until its owner enrolls a
NEW passkey against these options; "preferred" (not "required") preserves the
no-lockout fallback to username-first + email-OTP.
2026-07-05 04:05:56 +09:00
flyemoji 7db57b9fff feat(updater): add VersionGatherer extraction core and CLI gather seam
Give the Runner a way to read each component's CURRENT version so it can be
compared against the release sources already wired. Three pure extractors turn
raw system text into an updates.Version, each fail-closed:

  - versionFromCLI      — a `<tool> --version` banner   (k3s, cloudflared)
  - versionFromImageRef — a container image tag         (felis-api)
  - versionFromJarName  — a proxy jar filename          (velocity)

sysGatherer routes each Topology component to the right extractor over an
injected seam; every path is exercised with a fake runner, mirroring how the
release sources are proven against httptest.

The load-bearing case is k3s: its Git tag "v1.36.2+k3s1" parses stable, but a
registry cannot store '+', so the same build ships as image tag "v1.36.2-k3s1",
which parses as a prerelease unless repaired. versionFromImageRef normalizes
"-k3sN"/"-rke2rN" back to "+", so an image read and a CLI read agree instead of
the image masquerading as a prerelease and being barred from comparison.

Honest runtime state after this slice — a green suite is not "the updater runs
against real infra": only the CLI seam (execRunner) is wired, so of the four
tracked components just cloudflared is live end to end (gatherable AND
Scheduled/appliable). k3s is CLI-gatherable but Notify-only. felis-api and
velocity are NOT yet runtime-gatherable: their producing seams — a k8s read of
the control-plane Deployment image, and an off-cluster jar inspection — are left
nil, so both surface an explicit "gather seam not wired" error rather than a
wrong version. felis-api self-update is therefore not functional yet.

Remaining integration (tracked in doc.go): the two producing seams, the concrete
Notifier (SMTP + in-game), the Applier (image bump, cloudflared swap), the
`felis update` CLI + CronJob entry point, and the runtime append of the Pinned
Minecraft fleet.
2026-07-05 04:05:56 +09:00
flyemoji 7d27640c07 feat(updater): add GitHub Releases source and route felis-api/k3s/cloudflared
Give RoutingSource its second upstream so every non-pinned component now
resolves a real latest-stable: Velocity via PaperMC (already wired), and
felis-api, k3s and cloudflared via the GitHub REST API.

github.go queries /repos/{repo}/releases/latest (one request, rate-limit
friendly) and fails closed: a transport error, a non-200 status (404 = no
stable release), an undecodable body, a draft/prerelease flag, or an
unparseable / prerelease-parsing tag all return an error, never a zero
version. It sends the User-Agent GitHub requires (a UA-less request is
403'd) and tolerates the two live tag styles -- cloudflared's CalVer
"2026.6.1" and k3s's v-prefixed, build-tagged "v1.36.2+k3s1" -- while
String() keeps the raw tag for the report.

source.go routes sourceGitHub to it and drops the errGitHubNotWired stub;
velocity still routes to PaperMC.

Tests: github_test.go covers both tag styles, the User-Agent gate, and
fail-closed on 404 / prerelease-flag / unparseable tag, with fixtures
captured from api.github.com on 2026-07-05. runner_test.go now drives
PaperMC and GitHub through dual httptest servers end to end with no source
degrading to an error.

doc.go re-tiers the verification boundary: both release sources are now
built and live-grounded; the VersionGatherer's version-extraction core is
the next verifiable slice (logic over an exec seam, not pure I/O); the
genuine I/O remainder is the Notifier, Applier and felis update CLI/CronJob.
felis-api's coord is still a placeholder slug, so that component is dark at
runtime until a real repository is configured.
2026-07-05 01:03:32 +09:00
flyemoji 9896fe16c3 docs(updater): correct PaperMC UA/fixture overclaims, re-tier the boundary
An out-of-band curl of the live Fill v3 endpoint contradicted two claims the
previous commit shipped and surfaced a mis-tiering:

- User-Agent is NOT enforced: fill.papermc.io/v3/projects/velocity returned
  HTTP 200 to a bare curl UA. The comments claimed a generic UA "is refused"
  and the API "REQUIRES" a contact UA. Reword to what is true — PaperMC's usage
  policy asks for a descriptive UA and may block generic ones, but sending it is
  etiquette/defensive here, not a gate Felis depends on.
- The test fixture's shape was invented, not captured: the real "versions"
  object groups the entire 3.x line under a single key "3.0.0", not the
  per-minor keys the fixture used. Replace it with the real body (keys and
  version strings as returned). The key-agnostic parser already produced the
  right answer, and an independent max-stable check confirms 3.4.0.
- Re-tier doc.go: the GitHub Releases source is verifiable-here (the same
  httptest-testable shape as PaperMC), not integration remainder. It is why
  3 of 4 components report "latest unknown" today and is the next verifiable
  slice — the release-source work is only ~half done until it exists.

No production logic changed. WSL oracle: build + vet clean, internal/updater
10/10, full tree go test RC=0 (19 ok, 0 fail).
2026-07-05 00:28:36 +09:00
flyemoji 96b3cc901c feat(updater): wire updates.Run to a caller with PaperMC v3 release discovery
internal/updates is a pure, fakes-tested decision core with no production caller,
so nothing could produce its "版本号状态" report. Add internal/updater as that caller:

- topology: the fixed platform components and their user-set policies (felis-api
  and cloudflared Scheduled+manageable; k3s Notify, high-blast-radius single node;
  velocity Notify, off-cluster and unmanageable). Minecraft is pinned by ABSENCE,
  never force-tracked here, appended from the live fleet at runtime.
- PaperMC Fill v3 release source: the v2 API (api.papermc.io) was retired
  2026-07-01 and returns HTTP 410, so this targets fill.papermc.io/v3, sends the
  required non-generic User-Agent, and returns the newest STABLE version, filtering
  the -SNAPSHOT/rc prereleases the plan would otherwise suppress. Its test fixture
  is captured from the live v3 response shape (2026-07-04).
- RoutingSource: the single ReleaseSource updates.Run requires, dispatching
  velocity to PaperMC and returning errGitHubNotWired for the GitHub-backed
  components so they degrade to "latest unknown" honestly, never a fabricated one.
- Runner: gather current versions (seam) -> assemble Components -> updates.Run ->
  Report; report-only when notifier and applier are nil.

Verification boundary: the parse/plan/compose logic is unit-tested (httptest +
fakes, fixture grounded in the live v3 shape). Live network/TLS/User-Agent
enforcement, the GitHub Releases source, the concrete version gatherer, the
notifier and applier, and the felis update CLI/CronJob remain integration work,
enumerated in doc.go.
2026-07-04 23:41:46 +09:00
flyemoji 0a2accd245 chore: stop tracking Autohand-generated AGENTS.md
Added by the tooling in commit 0c1cc59, not authored guidance. Untrack and gitignore it: the file claims precedence over CLAUDE.md and tells agents to run go fmt, which rewrites the CRLF working tree. The file stays on disk (git rm --cached) so local tooling keeps it, but it is no longer tracked or committed.
2026-07-04 23:11:25 +09:00
flyemoji c20b12c655 refactor(api): drop dead password-era ResetMailer, reconcile passkey-unbind docs
The passwordless migration left ResetMailer (SendPasswordReset) and its API field with zero callers and no wiring; the web console authenticates via email-OTP and passkey only. Remove both, plus the now-orphaned context import that the interface was the last user of in handlers_users.go.

Reconcile the DeleteAllPasskeyCredentialsForUser docs in repo.go and pgrepo.go: they claimed there was no production caller, but 2f22027 wired the owner-tier DELETE /users/{id}/passkeys. Both now note that a complete authenticator remediation pairs the unbind with a session revoke (unbinding alone leaves the live hijacked session; revoking alone leaves a re-enrollable credential), and the OpenAPI operation carries the same guidance in a new description. Reword the stale local-password test-fake header, since the passwordless fakes carry no must_change_password field.

No behavior change. gofmt, build, and the full test tree are green; OpenAPI parity and passkey-unbind tests pass; a grep confirms ResetMailer/SendPasswordReset are gone from the Go tree.
2026-07-04 21:47:13 +09:00
flyemoji 4f59d5128a feat(auth): add owner-tier passkey-unbind remediation endpoint
Add DELETE /api/v1/users/{id}/passkeys (owner-only) to unbind every passkey a
target account holds — the authenticator remediation that stops a passkey planted
or retained via a transiently-hijacked session from surviving as a standing login
foothold. It wires the previously-uncalled DeleteAllPasskeyCredentialsForUser and
is deliberately not a lockout: the account re-enters via the email-OTP door
(players) or op-login's in-game approval (staff), then re-enrolls. Documented in
the OpenAPI, so the served/documented parity gate covers it.

Remove RevokeUserSessionsExcept: a change-password-era orphan with no callers
since the passwordless migration. Its keep-one ("log out my other devices")
semantics is inherently self-service, and no such slice is on the roadmap; the
admin remediation path already uses RevokeAllUserSessions.
2026-07-04 21:47:12 +09:00
flyemoji 3b43f05a83 refactor(api): drop dead login concurrency limiter and reconcile passwordless comments
The passwordless migration (b330d77) removed the password-login route, leaving
concurrencyLimiter — its bcrypt concurrency cap — with no caller, and scattered
stale "local-password" / "change-password" references through the surviving auth
code's comments.

- Remove the dead concurrencyLimiter (type + newConcurrencyLimiter + acquire):
  no caller, no struct field, no test. Reword the one streamLimiter doc that
  contrasted against it.
- Realign comments in repo.go, pgrepo.go, session.go, util.go to the passwordless
  reality: staff lookups feed email-OTP / passkey / setup redeem, not a password
  compare; RevokeUserSessionsExcept and DeleteAllPasskeyCredentialsForUser are
  retained (uncalled) for the P5 account-remediation path (#78); "local sessions"
  no longer implies a password.

Comments and dead code only; no behavior change. Full WSL test tree green.
2026-07-04 21:47:12 +09:00
flyemoji 0c1cc598c1 feat(auth): migrate console login to passwordless
Replace console password auth with a passwordless surface — the pre-session
login doors plus an identifier-first discovery endpoint — and remove the
password paths.

- Login doors (Public, pre-session): email-OTP, passkey assertion, op.console
  login with in-game approval, and setup-token redeem.
- /api/v1/auth/options: identifier-first discovery reporting which console
  methods an email can use. The single sanctioned existence oracle; methods
  are computed with no role branch, so staff and player accounts in the same
  credential state return byte-identical bodies (staffness invisible by
  construction).
- Remove password auth: drop StaffUser.PasswordHash and the /auth/login,
  /auth/change-password and /users/{id}/reset-password endpoints (and test).
- Data layer: UserByEmail, verified-email uniqueness, setup-token store
  (migration 0012).
- Reconcile docs/openapi.yaml with the served surface; the method/path/face/
  tier parity gate (TestOpenAPIMatchesServedRoutes) passes.
- felis TUI: in-game MC bind, owner/break-glass OP provisioning, version.
- Velocity /felis command suite.

Consolidates the accumulated backend migration work; the frontend (panel/)
is left untouched. Full Go tree green on WSL (go build ./... && go test ./...).
2026-07-04 21:47:12 +09:00
flyemoji b84debf872 feat(deploy): one-shot demo bring-up wrapper
demo-up.sh collapses bootstrap -> build+import the limbo/lobby images -> wire [velocity] login_image/lobby_image into felis.host.toml -> felis setup into a single command, ending in the interactive Owner-creation TUI (the only step it cannot automate). Prefers prebuilt tars under deploy/images, else builds on the host, auto-resolving the LOOHP/Limbo CI jar and the latest stable Paper jar (all overridable by env); SKIP_BOOTSTRAP/SKIP_SETUP toggles for reruns.
2026-07-03 01:26:15 +09:00
flyemoji d9e866fcdb fix(deploy): make the lobby image actually build
The lobby image had never been built and two defects blocked it: the felis image .dockerignore excluded plugins/* and only re-included limbo/shared, so the lobby Dockerfile's COPY plugins/paper landed empty; and the plugin stage used eclipse-temurin:21-jdk, which ships no gradle (and the tree vendors no wrapper), failing with 'gradle: not found'. Re-include plugins/paper and build the paper plugin on gradle:8.14-jdk21, matching the limbo image. Verified: both images build and boot (limbo /healthz 200 on 25565; lobby reaches 'Done' with felis-paper enabled).
2026-07-03 01:26:15 +09:00
flyemoji c7315e44b6 feat(deploy): login-limbo and lobby images with game-port pinning
deploy/limbo assembles LOOHP/Limbo from its loose CI artifacts plus the felis-limbo plugin (and the shared link core), with an entrypoint that pins server-port to the operator's GamePort (25565) on every start. deploy/lobby carries the Paper + felis-paper hub image. .dockerignore re-includes plugins/limbo and plugins/shared so the plugin image build sees them.
2026-07-02 19:38:38 +09:00
flyemoji 241fe21f8a feat(limbo): felis-limbo in-game login flow over the shared account-link client
The login limbo now performs the onboarding inside Limbo: on join it checks the collision blacklist, mints a bind code, opens a book linking the player to console.<root_domain> (guiding them to the system browser), polls link-status, and transfers to the lobby via BungeeCord Connect — fail-closed on blacklist, mint/transport error, or window elapse. FelisApiClient gains linkStatus/isBlacklisted on the existing internal transport.
2026-07-02 19:38:38 +09:00
flyemoji a63f49dcb3 feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey
WebAuthn is unusable inside the WeChat/QQ in-app WebViews, so a document navigation carrying those UAs is served a bilingual 'open in your system browser' interstitial instead of the passkey-centric SPA. API/config/health/asset requests pass through, and an ack cookie (ua_ack) lets a determined user or false-positive continue. Backend-only; the SPA is untouched.
2026-07-02 19:38:38 +09:00
flyemoji f554d525d4 feat(cli): provision login/lobby system servers with login env and token replica
setup builds the always-on, reaper-exempt login/lobby MinecraftServers (create-if-absent), bakes the login limbo's non-secret config (internal API URL, root domain, lobby name) into spec.env, and replicates the felis-service-token Secret from the control namespace into the minecraft namespace so the operator's namespace-local secretKeyRef on the login pod resolves.
2026-07-02 19:38:37 +09:00
flyemoji 3fdb3d032e feat(platform): internal API base-URL helper and single-sourced token secret
InternalAPIBaseURL builds the felis-api internal-face DNS from SAAPI and the internal port for cross-namespace callers (the login limbo). The service-token Secret name/key now reference the shared naming constants so the Deployment wiring and the operator's login-pod injection cannot drift.
2026-07-02 19:38:37 +09:00
flyemoji dc23cb54d2 feat(operator): system-server pod readiness probe and login service-token env
buildStatefulSet gates readiness on an HTTP probe when HealthHTTPPort is set (exposing it as a named container port). buildEnv injects FELIS_SERVICE_TOKEN into the login server only — keyed off the reserved name so it can never leak into a user pod — sourced from a Secret via secretKeyRef, never inlined into the CRD.
2026-07-02 19:38:37 +09:00
flyemoji 159107b4e3 feat(api): HTTP readiness knob on MinecraftServer and login-gate fallback default
StartupSpec.HealthHTTPPort/Path switch pod readiness from plain-TCP to an HTTP GET for RCON-less loaders (LOOHP/Limbo) that report 'started' only after the first tick. User servers now default FallbackServer to the login gate, never the lobby, so a stopped/starting backend keeps authentication in front of a fresh connection.
2026-07-02 19:38:37 +09:00
flyemoji 9ef817f2b3 feat(naming): system-server names, validation, and service-token identifiers
SystemLoginServer/SystemLobbyServer plus ValidateSystemServerName (format rule without the reservation check) let the platform provision the reserved login/lobby names users can never claim. ServiceTokenSecretName/Key are the one source of truth for the internal-API credential Secret, shared by the platform renderer and the operator's login-pod injection.
2026-07-02 19:38:37 +09:00
flyemoji 9bed51b67f feat(config): add [velocity] login_image/lobby_image for system servers
setup provisions the always-on login/lobby system services only when these image refs are set; empty means skip-and-say-so (the same fail-loud stance manifests takes), since no official LOOHP/Limbo image exists and a deployment must build its own.
2026-07-02 19:38:37 +09:00
flyemoji 54bc6ef211 fix(api): clear bound passkeys on password change to close a takeover foothold
handleChangePassword revoked other sessions but never cleared webauthn_credentials, and enrollment needs no step-up. A passkey planted through a transiently-hijacked session needs no password, so it survived the reset + session-revoke as a standing login foothold. Add DeleteAllPasskeyCredentialsForUser and call it in the change-password remediation so every passkey is unbound alongside the session revoke. Removing zero rows is a successful no-op. Email-OTP remains the fallback factor, so this never locks anyone out; the user re-enrolls a passkey afterward if they want one.
2026-07-02 06:55:21 +09:00
flyemoji 7278cd7c6a feat(passkey): require and record user verification at enrollment
Enrollment set no AuthenticatorSelection, so user verification defaulted to preferred (not enforced), and the UV/backup flags the ceremony reported were discarded. Set UserVerification=required so a bound passkey always proves possession AND user (a UV-incapable device falls back to email-OTP), and capture user_verified/backup_eligible/backup_state through VerifiedCredential -> PasskeyCredential -> webauthn_credentials (migration 0009) so a future login path can enforce UV per credential. Adds a negative test proving a presence-only authenticator is rejected, and asserts the roundtrip records UV=true.
2026-07-02 06:55:21 +09:00
flyemoji 20e31fb08f fix(store): cascade-delete passkeys and challenges on user removal
webauthn_credentials.user_id and webauthn_challenges.user_id referenced users(id) with the default ON DELETE NO ACTION, so a future user-delete would either fail or leave orphaned auth material. Recreate both FKs ON DELETE CASCADE: a bound passkey and a pending challenge are ephemeral and must not outlive the account. Scoped to the passkey tables only, not blanket, so retention-bearing child data (world_backups) is not swept away with an account.
2026-07-02 06:55:21 +09:00
flyemoji 99532759b2 fix(api): bound webauthn_challenges growth by superseding all prior rows
The supersede DELETE in CreatePasskeyChallenge filtered consumed_at IS NULL, so it only reaped the prior LIVE challenge; the row that each finish stamps consumed_at on was left behind. A begin->finish loop therefore accumulated one dead row per cycle, unbounded. Drop the consumed_at clause so a fresh begin reaps ALL prior rows for (user, purpose), bounding the table at one row per (user, purpose) with zero net growth per cycle. Deleting an already-consumed row is safe: it has been redeemed and nothing reads it. The fake mirrors the widened supersede.
2026-07-02 06:55:21 +09:00
flyemoji cdbb5abc35 fix(api): record credential id in passkey-register audit event
handlePasskeyRegisterFinish logged an empty target for account.passkey.registered, while the delete half logs the credential id. An operator auditing the log could see that a passkey was bound but not which one. Pass cred.ID as the audit target so bind and unbind are symmetric, and tighten the enrollment test to assert both halves name the credential id.
2026-07-02 06:55:21 +09:00
flyemoji 6368ab1914 fix(api): coalesce MyServers owned flag so ownerless rows do not 500
The MyServers query lists both a user's own servers and unclaimed (owner_id IS NULL) servers, but computed owned as s.owner_id = $1. For an ownerless row that comparison is SQL NULL, which fails to scan into the Go bool and 500s the whole listing. Wrap it in COALESCE(..., false) so an ownerless row reports owned=false while still surfacing as claimable.
2026-07-02 06:55:17 +09:00
flyemoji 8f41a003b0 fix(api): clear the SSE write deadline on return so it can't leak to a reused connection
The per-write deadline that severs a stalled SSE reader was never cleared on
return. Server.WriteTimeout is deliberately unset -- a WriteTimeout would sever
a healthy long-lived stream -- and with it unset net/http never resets the
connection write deadline between keep-alive requests. So the deadline the last
writeChunk left set leaks onto the next request that reuses the pooled
connection and fails its first write for no reason. Clear it to the zero value
on return via a deferred rc.SetWriteDeadline; best-effort, a no-op on writers
without deadline support.

Also record honestly at the header flush that the connect-time stall stays
bounded only by the per-principal stream cap, not severed by this guard -- only
the mid-stream stall is closed. Adds a test pinning the clear (fails closed:
neutering the deferred clear leaves a +writeTimeout deadline set on return).
2026-07-01 23:06:26 +09:00
flyemoji 2c56d17989 docs(api): record the quota-claim TOCTOU as a KNOWN-LIMITATION (audit #4)
QuotaAvailable and ClaimServer run as two separate statements, so the
count read is not serialized against a concurrent claim's UPDATE: two
claims by one user for two different ownerless servers can both pass the
gate and both succeed, leaving the user one server over quota. It is low
severity — quota over-provisioning under a deliberate burst, not an
authorization, ownership, or isolation break, since each server is still
claimed atomically via UPDATE ... WHERE owner_id IS NULL.

Closing it requires Postgres transaction semantics (advisory-xact-lock on
the user, or SERIALIZABLE with retry) folding the gate into a single repo
method — verifiable only against a real Postgres, not the hermetic
fakeRepo suite. Documented at QuotaAvailable with back-references from the
two claim gates (handleClaim and the internal UUID claim) rather than
patched blind.
2026-07-01 22:37:36 +09:00
flyemoji d6e3189629 fix(api): bound SSE relay writes with a deadline to sever stalled readers
relayLogStream copied a pod-log follow to the client with a plain
flusher.Flush per event. On a client that stays connected but stops
reading (its TCP receive window shut), net/http buffers the small
"data:" line and only touches the socket at Flush, which then blocks
forever inside the write. The select's <-ctx.Done() branch is never
reached, because r.Context() cancels on an actual disconnect, not on a
stall, so the relay goroutine and its upstream apiserver follow leak for
the life of the process.

Route every event's write+flush through http.ResponseController with a
per-write deadline (writeTimeout, 30s): a stalled flush now returns
os.ErrDeadlineExceeded, the error plain http.Flusher.Flush swallows, and
the relay abandons the stream so the deferred cancel + src.Close release
the follow. SetWriteDeadline and rc.Flush are best-effort: a writer
without deadline support (httptest recorder; some HTTP/2 origins) ignores
the deadline and behaves exactly as before, so the guard degrades
gracefully.

This closes the leak the per-principal stream cap only bounded the blast
radius of. Verified by a deterministic test with a deadline-aware
ResponseWriter whose flush blocks until the deadline; the test times out
(fails closed) if the guard is removed.
2026-07-01 22:30:40 +09:00
flyemoji 3c1d64749f fix(api): cap concurrent SSE streams per principal
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.

Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.

This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
2026-07-01 22:04:19 +09:00
flyemoji c6c0772a7a fix(api): set read/idle timeouts on the felis-api listeners
The three felis-api http.Servers (internal, external, https) were built with
only Addr and Handler, leaving ReadHeaderTimeout, IdleTimeout, and ReadTimeout
at zero. A zero ReadHeaderTimeout is a Slowloris hole — a client trickling
header bytes pins a connection indefinitely — and a zero IdleTimeout lets
kept-alive connections accumulate (gosec G112).

Route all three listeners through a newAPIServer factory that sets a 10s
ReadHeaderTimeout and a 120s IdleTimeout. WriteTimeout and ReadTimeout are
left unset on purpose: the external and https faces stream Server-Sent Events
(console / build logs) for the lifetime of a client attachment, and a
WriteTimeout would sever a healthy long-lived stream. Slowloris is closed by
ReadHeaderTimeout, which bounds only the header phase.
2026-07-01 21:12:03 +09:00
flyemoji 164ac447ef fix(api): validate inbound X-Request-Id before echo and audit persist
withRequestID honored any inbound X-Request-Id verbatim, and that value is
echoed on the response, embedded in the error envelope, and persisted into
audit_logs.request_id. An unvalidated caller-supplied id is therefore an
audit-integrity vector: an arbitrarily long value bloats the audit row, and a
stray control byte (CR/LF) could smuggle a forged entry into a log sink.

Accept an inbound id only when it is well-formed — non-empty, at most 64
bytes, and restricted to a log-safe charset ([A-Za-z0-9._-]) — otherwise mint
a fresh server id. A rejected request loses its inbound trace link, which is
strictly better than storing attacker-controlled text in the audit trail.
2026-07-01 21:07:08 +09:00
flyemoji 7a51c1d9c3 fix(api): bound concurrent login bcrypt to shed CPU-pin floods
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.

Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.

The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
2026-07-01 21:01:28 +09:00
flyemoji f34711c174 docs(api): record passkey login-handler deferral rationale
The passkey login/assertion HTTP handler stays deferred after its design
checkpoint; capture the reasoning in the handler header so the decision is
durable in the repo rather than only in task notes.

- RP boundary (resolved): felis-api is the app-login relying party (panel.*);
  the WebAuthn security gate lives at the Cloudflare Access edge. Spec §14 ties
  WebAuthn/posture to admin.* (Access) while panel.* is plain app login, so
  there is neither a spec-required assertion handler nor a backend step-up
  consumer for one.
- Identifier (blocking): a from-zero login needs a unique, human-typable handle
  to resolve an account, but users.email is nullable and non-unique and a
  player's username is their Minecraft uuid. Username-first assertion has
  nothing to key on; re-link stays the returning-player door.

Discoverable (usernameless) credentials are the future enabler; the adapter
crypto is already verified so that slice inherits correct crypto.
2026-07-01 19:15:44 +09:00
flyemoji e035142abc feat(passkey): add WebAuthn login/assertion crypto adapter
Build the assertion (login) half of the WebAuthn ceremony crypto in the
internal/passkey adapter, Oracle-verified against a virtual authenticator.

- BeginLogin/FinishLogin over go-webauthn BeginLogin/ValidateLogin,
  username-first (allowCredentials scoped to the known user's bound
  passkeys). Discoverable/usernameless login stays out of scope: the
  enrolled credentials are non-resident and the challenge store is
  user-keyed (migration 0007), so it would need a future migration.
- WebAuthnCredentials() now populates the stored COSE public key and
  signature counter (assertion validation needs both to verify the
  signature and detect clones); enrollment ignores them, so the change
  is backward-compatible and the enrollment tests guard it.
- VerifiedAssertion seam output: which credential signed plus the raw
  signature counter. Clone/regression policy is deliberately NOT here —
  the counter is a ceremony fact and the future handler, which holds the
  previously stored counter, decides reject/warn.

Scope: crypto adapter only. The login HTTP handlers, session minting,
and the panel.* relying-party boundary/tier decision remain a deferred
slice (no unauthenticated login route is added). BeginLogin/FinishLogin
live on the concrete adapter, not the api.PasskeyVerifier interface,
which grows only when a handler consumes them.

Tests (virtualwebauthn): a real enrollment chained into a real assertion
exercises the COSE public-key decode path and surfaces the advanced
signature counter, plus origin-mismatch and unbound-credential rejection.
2026-07-01 18:45:26 +09:00
flyemoji 7464fa700b fix(updates): tag Window JSON so the persisted maintenance window round-trips
The admin API persists the auto-update maintenance window as lowercase
JSON {"start","end"} (platform_settings key "update_window"), but
updates.Window had no json tags, so it marshaled/unmarshaled with
capitalized keys. The natural decode the update runner will use --
json.Unmarshal(stored, &updates.Window{}) -- would therefore miss every
key and silently yield the zero Window. That fails closed (a zero window
Contains nothing, so notify-only, never a rogue apply), so it is safe but
a latent silent-zero trap for the not-yet-built runner.

Add json:"start"/json:"end" to updates.Window so the obvious decode is
correct by construction; value time.Time treats a stored null as a no-op,
so a cleared/never-set window still decodes to the zero Window. Nothing
in the package serialized Window before, so this changes no existing
behavior.

Guarded by a cross-package contract test in internal/api that marshals
the real api.updateWindow DTO and unmarshals it into updates.Window --
asserting the interval survives (Contains(mid) is true) and that an empty
window decodes to the fail-closed zero Window -- so the two shapes cannot
drift apart silently.
2026-07-01 18:26:42 +09:00