Commit Graph
159 Commits
Author SHA1 Message Date
Lemon-miaow 2a55a0d265 test(pgint): verify the business stores against a real Postgres
The hermetic suites encode the store contracts against fakes; PGRepo drifted
behind them three times (attempt accounting, a missing JOIN, a missing FOR
UPDATE) while every unit test stayed green. This harness replays the real
embedded migrations onto a throwaway database — its name must contain "pgint"
or the harness refuses to run — and exercises the SQL directly: sessions, the
onboarding email-OTP lifecycle (supersede/expiry/lockout), the pre-session
login consume, the op-login state machine, link and bind-code redemption,
submissions, and builds with the image admission round trip.

Run it after touching SQL under internal/api/pgrepo.go, internal/submit, or
internal/build; CONTRIBUTING.md carries the one-liner.
2026-09-23 03:18:56 +08:00
Lemon-miaow 02fd2de502 fix(build): three drill-driven fixes so the lane actually completes on a starter node
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:

- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
  Pod never scheduled on the platform's own starter node: FailedScheduling /
  Insufficient memory, Pending forever. Requests are now a small floor
  (250m / 512Mi, never above a configured cap) while the limits stay the
  safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
  the source owner; a 65532-owned context (the distroless felis image uid)
  fails that under the pod's dropped capabilities ('copying dockerfile:
  chown /kaniko/Dockerfile: operation not permitted'). The fetch container
  now extracts as root — the uid Kaniko already runs as — so the copy
  succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
  step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
  renders --db-repository, and docs/troubleshooting.md §8e now carries the
  verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
  internal registry; --insecure already covers its plain HTTP).

Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
2026-09-22 23:01:25 +08:00
Lemon-miaow f79e5ebb5e feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:

- submit: derived context refs become the internal-face URL
  /api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
  gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
  route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
  image's new fetch-context entrypoint) that streams the blob with the
  namespace-local service-token Secret — never mounted into Kaniko — and
  extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
  reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
  build namespace gets the token Secret through the existing replica mechanism
  (bootstrap.sh + felis setup); the build egress lock opens exactly the control
  namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
  escapes/symlinks/non-gzip).

Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
2026-09-22 22:45:09 +08:00
Lemon-miaow 0c8e29b05a fix(platform): give every control-plane Deployment real probes (#8 follow-up)
The api, operator and registry Deployments shipped with no liveness/readiness
probes at all: a wedged process stayed 'Running' forever, and the operator had
no health listener to probe in the first place. Kaniko build evidence on a
fresh install showed the only cluster-wide red after a disk-pressure pass was
Deployment status that never reflected health.

- felis-api: readiness /readyz (DB + K8s API round-trip) and liveness /healthz
  on the internal face (:8081), the only listener that serves both endpoints;
  liveness deliberately avoids /readyz so a DB blip cannot restart the api.
- felis-operator: new --health-probe-bind-address (:8081) with controller-
  runtime's /healthz + /readyz (registered ping checks; an unregistered handler
  map would 404), plus the matching container port and probes.
- registry: /v2/ probes on the pinned port, so a broken storage backend stops
  reading as 'Running'.

Tests pin paths, ports, and that each probe targets a declared container port.
2026-09-22 22:22:37 +08:00
Lemon-miaow 87a9f4eb25 feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
2026-09-22 22:03:48 +08:00
Lemon-miaow 0a2d654e68 fix(platform): control plane runs system-cluster-critical, so eviction refuses it (#8)
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.

Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
2026-09-22 21:55:54 +08:00
Lemon-miaow fe310743a2 fix(platform): give the control plane a PriorityClass eviction shield (#8)
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
2026-09-22 21:08:56 +08:00
Lemon-miaow fd33fd05e1 fix(install): backups exist on a default install; retention resolves real world dirs (#6)
Three faces of one gap, all on the supported install path:

- Backup/restore answered 503 out of the box: nothing ever rendered the
  archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
  PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
  'felis manifests' names it by default (--backup-pvc= is the explicit
  no-store shape); bootstrap passes it through so the generated felis.toml
  [archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
  now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
  env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
  nothing deleted).
- Even when enabled it could not find a world on a stock install:
  resolveWorldDir now also resolves the exact local-path directory
  <pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
  glob, so a stale deleted PV's bytes can't be archived in place of the
  current world). Reaper Role gains persistentvolumeclaims:get (weaker
  than the delete it already held).

README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
2026-09-22 20:48:07 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow a2df2f242b fix(operator): re-probe RCON every 2s while Starting
The readiness gate is status-driven; at a 5s re-probe cadence the observed
'container Ready but API still 409 not_running' window was 6~10s. Halving the
cadence halves the worst case; probes still only run while unreachable.
2026-09-22 20:31:53 +08:00
Lemon-miaow 2a8f897e61 fix(api): a session-store outage answers 503, not 401
Resolving a session cookie failed identically whether the credential was
missing or Postgres was unreachable: local_auth_enabled read errors fell into
the fail-closed 'disabled' branch and SessionUser errors into 'invalid
session', both surfacing as 401 'authentication required' — a lie that reads
as 'log in again' during an outage. Split the enabled-read into
(enabled, error), tag non-ErrNotFound store failures with errAuthBackend, and
map that to a new 503 auth_unavailable in requireExternal. Fail-closed is
unchanged: missing setting / bad value / missing session stay 401.
2026-09-22 20:30:43 +08:00
Lemon-miaow c839454a1f fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
another namespace either (live drill: 'error looking up service account
minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
RoleBinding subject to the Minecraft namespace alongside the CronJob.
2026-09-22 20:18:36 +08:00
Lemon-miaow e4f2cff532 fix(manifests): render the retention reaper CronJob into the Minecraft namespace
A Pod can only mount PVCs from its own namespace; the CronJob referenced the
minecraft-namespace backup PVC while being rendered under ControlNamespace, so
it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim
felis-backups not found'. The reaper Role/RoleBinding were already
minecraft-scoped (the objects it touches live there), so the CronJob was the
odd one out. The minecraft felis-config replica (felis setup, backup Job fix)
supplies its config mount.
2026-09-22 20:14:19 +08:00
Lemon-miaow 0414913bc7 fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
6-way concurrent redeem of one code 500'd on users_username_key (each request
generated a fresh user id but the same uuid-derived username), plus the rarer
two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
locks its code row and handles the conflict.

- SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
  as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
- INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
  cross-code racers converge on the winner's row (role checked, staff still
  refused) instead of a unique-violation 500.
- account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.

Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
db clean; zero unmapped errors.
2026-09-22 20:08:38 +08:00
Lemon-miaow dcc3b7403e fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake)
The interface doc promised 'each joined to its staff username', the fake and
the pending handler both project username and created_at, but the PG query
selected neither — live internal /op-login/pending returned username:"" and
created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
tests can't see PG-only regressions.
2026-09-22 20:04:32 +08:00
Lemon-miaow 52549f7b3a fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
The PG implementation was a single UPDATE ... WHERE code_hash that returned
ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
the pre-session email-login door (and the op-login finish / migration confirm
doors) fell through to writeError's unmapped-error 500, and attempts were never
charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
intended contract; only the PG side had drifted.

Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
UPDATE the newest live row, expiry + attempt cap before the hash compare,
mismatch charges one attempt and returns ErrOTPInvalid without consuming,
match consumes and commits. Verified live on the VM: 5 wrong guesses return
400 and stop at attempts=5 (correct code then also refused, unconsumed);
fresh code redeems; replay returns 400.
2026-09-22 19:51:04 +08:00
Lemon-miaow 9309ff5a7f chore: apply the missed S1016 conversions in handlers_users
The gofmt/staticcheck commit staged handlers_user.go (singular) for the
formatting fix but missed this sibling for its two struct-literal-to-
conversion cleanups.
2026-09-22 17:57:03 +08:00
Lemon-miaow e690b058db fix(restore): wait for the tracking finalizer before recreating
Live verification of the previous commit showed the immediate retry STILL
stranded: deleting a finished Job leaves it terminating (job-tracking
finalizer), so the re-Create collided with the dying object and was
mapped to ErrAlreadyExists a second time. Poll until the name actually
frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck
finalizer ever outlives the budget. Fake-client tests pin both the
replace-finished and coalesce-in-flight branches.
2026-09-22 17:54:01 +08:00
Lemon-miaow 90ccbfede4 fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
2026-09-22 17:50:04 +08:00
Lemon-miaow a05edc934c chore: gofmt the tree, clear staticcheck, add a CI gofmt gate
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
2026-09-22 17:49:54 +08:00
flyemoji 99c31c1d4e fix(config): refuse plaintext auth-source urls to public hosts
An auth_source url could be http:// to any host. Anyone on the path
to a public root, or anyone who can spoof its DNS name, can then
answer hasJoined with a 200 and log in as any player of that source,
including a third-party account linked to staff. The player's IP also
travels in cleartext. Mojang logins are unaffected, since that source
is built in over https.

Config load now refuses http:// unless the host is localhost or a
loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12,
192.168/16, fc00::/7), so a root on the same host or the LAN still
works without TLS. The decision is made on the literal host because
nothing is resolved at load time, so a LAN root named by hostname
needs its IP address or https. The error says what to change.

The new test covers public names and addresses, link-local, 0.0.0.0
and the first address past 172.16/12 (all refused over http, all
accepted over https), and the loopback and private forms that stay
allowed. It fails on the old check.
2026-09-22 14:24:41 +09:00
flyemoji e9f74f3f0f fix(nano): stop trusting an expired free name while mojang is failing
When the premium-name lookup failed, isPremiumName fell back to any
cached answer, however old. An expired "free" is exactly the answer
that may have stopped being true: someone can buy the name after it
was last seen free. For as long as api.mojang.com kept failing (429,
5xx, a timeout), a third-party player holding that name kept it on
every reconnect, and the Velocity registry, keyed on the name, turned
its new owner away as already connected. A hostile source could drive
the host into Mojang's rate limit on purpose to hold names that way.

A failed lookup now always counts as taken, so the player is renamed
with the source's prefix. An expired "taken" already gave that answer,
so only the stale "free" case changes. The cost is cosmetic: during an
outage an ordinary third-party player may get a prefix they do not
need, and their data follows the UUID, not the name.

A new test gives the cache a free entry past its TTL and has Mojang
answer 429. It fails on the old fallback. The two comments that
described the fallback now describe the fail-closed rule.
2026-09-22 14:17:53 +09:00
flyemoji c2a5645c55 fix: keep internal section numbers out of runtime messages
Four messages that reach an operator or an API client cited sections
of a specification nobody outside the project can read:

- the unimplemented archive store error from config load
- the running-server cap refusal, from both the user wake and the
  internal wake
- the missing memory ceiling guard, in the API and in felis apply

The references are gone and the wording is otherwise unchanged. Each
message still says what went wrong and, where there is one, what to
do about it. The test for the archive store message checks for the
tarLocal remediation, which is still there.
2026-09-22 13:44:38 +09:00
flyemoji 1ebd73a309 docs(nano): describe the hasjoined path as it works
The comments around hasJoined still described an authlib client that
is not in the path. Velocity reads -Dmojang.sessionserver and sends
the request itself, and it turns a 204 into its online-mode-only kick,
not authlib's "failed to verify username". The route comment in api.go
also offered "a thin login hook" as an alternative that does not
exist.

Other comments had drifted from the code:

- The [[auth_source]] doc said an empty list ships the multiplexer
  off. Mojang is always prepended, so an empty list means Mojang is
  the only source.
- The premium-name cache said Mojang does not recycle names. A name
  frees up when its owner renames away. The day-long "taken" TTL still
  holds, because a stale "taken" costs a third-party player only a
  prefix.
- The cache bound claimed entries come only from players who
  authenticated somewhere. Any third-party source that validates a
  login adds one, so a hostile source can force the map to clear. That
  costs repeat lookups, or a fail-closed prefix while Mojang is
  unreachable, never an identity.

The rewrite rationale now states what it costs a backend operator. A
chat-session key that a third-party source signed over its native UUID
cannot verify against the canonical UUID, so chat from those players
can only be accepted unsigned.

In the tests, comments that repeated their subtest names are gone.
2026-09-22 13:42:03 +09:00
flyemoji e0ad78af98 test(config): make the identity-key test fail when the key is accepted
TestLoadRejectsAuthSourceIdentityKey is the guard against a config line
identity = true making a third-party source's UUIDs trusted as-is. Its
fixture had no prefix, so Load failed on the prefix rule and the test
passed on that error. With the unknown-key check in decodeConfig
disabled, the test still passed.

The fixture now carries a valid prefix, the error must mention unknown
keys and identity, and LoadNano is checked alongside Load. With the
unknown-key check disabled, both loaders now fail the test; the old
version of the test passes against the same change.
2026-09-22 13:36:29 +09:00
flyemoji 30b4e1dfb2 test(nano): cover the bar-list error, bad identity id and ip relay
Three paths in handleHasJoined had no test that fails when they break:

- A bar-list lookup error answers 500. Logging it and carrying on would
  admit a reclaimed squatter during a database outage.
- An identity (Mojang) id that does not parse answers 204. Ignoring the
  parse error would emit the nil UUID for every such login, so they all
  share one player's data.
- The ip parameter is relayed to each source. Dropping it turns off the
  sources' check that the session is used from the player's own address.

One subtest each. Mutants that ignore the bar-list error, ignore the id
parse error, or stop appending ip each fail their subtest.
2026-09-22 13:35:50 +09:00
flyemoji 942e9a5ff8 test(nano): cover the premium-name cache rules
isPremiumName decides on every third-party login whether the player
keeps their name, and none of its rules had a test that fails when the
rule breaks: treating a 429 or 5xx from api.mojang.com as "free",
swapping the free and taken TTLs, flipping the freshness comparison,
answering "free" from an expired taken entry during an outage, or
dropping the clear-at-4096 bound. Each of those leaves a squatter
holding a name its owner has bought, or grows the cache without limit,
with CI green.

TestPremiumNameCache drives isPremiumName against a stub that answers
with a fixed status and counts lookups, and seeds cache entries at chosen
ages. Five mutants of handlers_hasjoined.go, one per rule above, each
fail at least one subtest. It does not test an expired "free" entry
during an outage; what that case should return is still open.
2026-09-22 13:34:43 +09:00
flyemoji 3f7274d29f test(nano): pin the auth namespace and one rewritten uuid as literals
The rewrite test computed its expected UUID from felisAuthNS itself, so
a change to the namespace seed moved both sides together and still
passed. Such a change gives every third-party player a new UUID on next
login, orphaning their playerdata and account links and letting any
squatter barred by the old UUID back in.

The test now also compares felisAuthNS and the rewrite of
littleskin:<Notch's id> against fixed strings, 07228eae-77f6-500e-
9dc0-436afbc87c27 and b63bcc1c611432eeb7b3af3a15012e48. Both were
computed independently with Python's uuid5/uuid3, not read back from
the code. Prefixing the seed with https:// fails the test.
2026-09-22 13:32:45 +09:00
flyemoji a7fe525bfc test(api): keep the package's tests off the live mojang profile api
mojangProfileAPI defaults to https://api.mojang.com, and only the tests
that call stubMojangNames or setProfileAPI swap it out. A new test that
reaches a third-party login without doing so would query the real
service: its result then depends on network access and on whether
someone owns the name that day, and the shared premium cache can carry
that answer into later tests.

A TestMain now points the lookup at an address nothing listens on
before any test runs, so a forgotten stub always takes the same
fail-closed path. Tests that stub it restore this address, not the live
one, when they finish.
2026-09-22 13:31:53 +09:00
flyemoji 8e9c8ca4e6 fix(nano): say that [server] listen is ignored instead of defaulting it
LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset,
and a test pinned that value, but felis nano never reads it: it binds
the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An
operator moving nano off loopback by writing [server] listen in its
config got connection refused from the proxy and no hint that the key
did nothing.

LoadNano no longer sets the default, and nano prints a line naming the
ignored value and the address it actually binds whenever the key is
set. It is a warning rather than a load error so a full felis.toml
copied onto a nano host keeps starting. The assertion that pinned the
unused default is removed along with it.

The new test runs cmdNano against a config that sets [server] listen
and one that does not, with an unbindable -listen so it returns after
loading. The first must warn and the second must not; with the old
default restored, the second prints a warning about 0.0.0.0:8080.
2026-09-22 13:30:23 +09:00
flyemoji 1905cac950 fix(config): refuse auth-source tags padded with whitespace
A third-party player's UUID is hashed from the source tag byte for byte,
so the tag is a permanent namespace: change it and every player of that
source comes back as someone new, with their playerdata, permissions,
account links and reclaim bans left behind. Nothing said so, and a tag
with a stray leading or trailing space, which nobody can see in the
file, loaded as a brand new namespace.

Such a tag is now rejected at load, and the AuthSourceConfig doc states
that the tag is permanent, case included. The charset stays otherwise
open: tightening it would force existing installs to rename, which is
the very thing that rekeys their players.

The new test loads a tag with a trailing space, a leading space and a
trailing tab through LoadNano; all three loaded before this change.
2026-09-22 13:24:23 +09:00
flyemoji 72a2750461 fix(config): refuse mojang as an auth-source tag
Mojang is prepended in code as the first, identity source, and the
config templates say not to list it. Nothing enforced that. A listed
tag = "mojang" loaded, and nano's startup list printed it as if Mojang
had been pointed at that url, while the real Mojang was still asked
first. The listed entry was a separate third-party source: asked again
on every login that got past Mojang, adding up to five seconds when its
url was Mojang's own and it answered 204 each time.

Any case of "mojang" is now rejected at load with a message saying
Mojang is built in and must not be listed. The duplicate-tag check could
not catch this because the built-in source never passes through it.

The new test loads "mojang" and "Mojang" through LoadNano; both loaded
before this change.
2026-09-22 13:23:37 +09:00
flyemoji 2c74080b78 fix(config): reject auth-source urls the resolver cannot query
The url check only looked for an http:// or https:// prefix. Several
shapes passed it and then left the source dead at login time: no host
("https://"), a bad port, surrounding whitespace (sent as %20 and
answered 404), and any query or fragment. The resolver appends
"?username=…&serverId=…" to the url as a string, so an existing query
swallows those parameters and a fragment hides them from the request
entirely. Each loaded green, and every login from that source failed.

The url is now parsed and must be http or https with a host, no query,
no fragment and no surrounding whitespace. Load and LoadNano share the
check. The shipped LittleSkin default and plain http:// endpoints, such
as a same-host root on loopback, still load.

The new test feeds each rejected shape to LoadNano. Against the previous
prefix check, six of the seven load; only ftp:// was refused.
2026-09-22 13:23:02 +09:00
flyemoji 3338d6f0fe fix(nano): drop oversized hasJoined parameters before asking sources
username, serverId and ip were forwarded to every configured source at
whatever length the caller sent, up to the megabyte net/http allows in a
request line. Velocity never sends more than a 16-character name, a
41-character signed SHA-1 serverId and a textual IP address, so only a
direct caller reaches those sizes, and each such request cost one
oversized upstream call per source.

Any of the three over 64 bytes is now answered 204 before a source is
asked, the same as a missing username or serverId. 64 bytes still
leaves room for a 16-character name in multi-byte UTF-8.

The subtest behind this points a source that validates anything at the
handler and sends missing and oversized fields, expecting 204 and zero
upstream requests, then a well-formed login that gets 200. It replaces
the old missing-username case, whose only source was unreachable, so
the test passed even with the guard removed. Dropping the length check
now fails it on the long username; dropping the whole guard fails it on
the first missing field.
2026-09-22 13:21:43 +09:00
flyemoji a4779186a4 fix(nano): refuse hasJoined requests that declare a body
A GET to hasJoined with a Content-Length and no body held its
connection indefinitely. The handler returned, but net/http tries to
drain an unread body before it writes the answer, and nothing bounds
that wait: ReadHeaderTimeout ends with the headers. One such request
per socket pins a goroutine and a descriptor on nano or on felis-api's
internal face.

Velocity never sends a body, so any request that declares one, including
a chunked one, now gets a 400 with Connection: close, which skips the
drain and releases the connection once the answer is written.

The new subtest writes that request over a raw socket and waits three
seconds for an answer. Before the change it times out with no response
at all; now it reads a 400 marked close.
2026-09-22 13:19:55 +09:00
flyemoji ff81295aa9 fix(nano): report failing sources instead of treating them as a no
A source that timed out, answered 5xx or 429, redirected, or sent a 200
without a usable profile was skipped exactly like one that answered 204.
With nobody else validating, the login got a 204 and Velocity told the
player their account is offline-mode. Nothing was logged, so a dead or
mistyped source URL, or an http:// root that now redirects to https since
redirects stopped being followed, failed every one of its players with
no trace.

Each such failure now logs the source tag and the cause; for a 3xx it
names the Location to configure instead. When no source validates and at
least one failed, the answer is 503, which Velocity reports as the auth
servers being down and logs with the status. A source answering 204 is
still a plain no, and a validating source still wins regardless of
failures before it.

The new subtest puts a 503 source, a redirecting source and an
unreachable one each behind a Mojang that answers 204, and expects 503.
Against the previous handler every case returns 204.
2026-09-22 13:18:33 +09:00
flyemoji 1dd62a9bdc fix(nano): cap upstream response headers at 16 KiB
The hasJoined and name-lookup clients limited the body to 64 KiB but
left headers at the transport default of 1 MiB. A configured root could
answer with a megabyte of headers and stall the body, holding a few MiB
of heap per in-flight login for the full five seconds; enough parallel
logins take down the host, and every source's logins with it.

Both clients now share a transport with MaxResponseHeaderBytes set to
16 KiB. Real roots come nowhere near it: Mojang's sessionserver sends
338 bytes of headers, LittleSkin 752, api.mojang.com 327. A source over
the cap fails the request and the resolver moves on to the next one.

The new subtest puts a source with 64 KiB of headers and a valid profile
ahead of an honest one and expects the honest player. Without the cap
the padded source wins.
2026-09-22 13:02:48 +09:00
flyemoji 2180e77cf5 chore: drop tool-name markers from source comments
Seventeen comments opened with a tag naming the tool that wrote them.
The tag goes and each comment keeps its reasoning, now starting as a
plain sentence. None of the reasoning changes.

The AGENTS.md note in .gitignore drops the story of how the file got
into the tree and keeps the one fact a reader needs: its advice to run
go fmt is destructive on this CRLF working tree.

Comments only; no code, build or test changes.
2026-09-22 12:57:19 +09:00
flyemoji 4e98ae6e56 fix(config): reject auth-source tags that contain a colon
A third-party player's canonical UUID is UUIDv3 over tag+":"+nativeID,
and the native id is whatever the source answers. Tags were only
checked for being non-empty and unique, so both "guild" and "guild:eu"
could be configured. The "guild" root could then answer hasJoined with
id "eu:X" and receive exactly the UUID of "guild:eu"'s player X, along
with their playerdata, permissions and account links. Real native ids
are 32 hex digits, so only the shorter tag's source can do this, and
only when the operator has configured such a pair; when they have, it
is a full impersonation.

Reject a ':' in a tag at load. With colon-free tags the join is
unambiguous: two different (tag, id) pairs can no longer produce the
same input, since equal inputs force equal tags and duplicate tags are
already refused. The tag is deliberately not narrowed any further.
It is a permanent UUID namespace, and forcing an operator to rename a
tag that has no colon would move every one of its players to a new
UUID. The hash input and the native id are left exactly as they were,
so no existing player's UUID changes.

Load and LoadNano share validateAuthSources; the new test runs both
against the guild / guild:eu pair and fails on the previous config.go.
2026-09-22 12:53:35 +09:00
flyemoji 1976fca809 fix(nano): always relay properties as an array
sessionProfile tagged properties with omitempty, so an upstream answer
of "properties": [] (or null, or no key at all) reached Velocity with
no properties key. A Yggdrasil root may legitimately answer that way for
a player without a skin. Velocity 3.5.1's GameProfile deserializer
passes the missing key on as null and ImmutableList.copyOf throws, so
that player hangs at login with nothing logged, even though the same
answer sent straight to Velocity is accepted. Mojang always sends
textures, which is why the premium path and the hardware runs never hit
it.

Drop omitempty and replace a nil slice with an empty one before the
response is written. Removing omitempty alone is not enough: a nil
slice marshals as null, which Velocity rejects the same way.

The new subtest feeds the relay [], null and a missing key and expects
"properties":[] every time. The previous handler fails all three.
2026-09-22 12:52:04 +09:00
flyemoji 28d3638952 fix(nano): stop following redirects from upstream Yggdrasil roots
authHTTPClient kept net/http's default redirect policy, so a configured
third-party root that answered hasJoined with a 3xx made this host fetch
whatever URL it named, up to ten hops. That is a blind SSRF into
anything the host can reach, and it includes the multiplexer's own
listener: a root that redirects back to /session/minecraft/hasJoined
re-enters the handler, which queries Mojang and every source again and
gets redirected again, until the outer 5s client timeout fires. With a
50ms Mojang stub, one login produced 86 nested handler calls and 86
Mojang requests from this host's egress IP. The loopback default does
not help, because the redirect target is resolved from this host.

Return the 3xx as the response instead. resolveHasJoined already skips
any non-200 answer and closes its body, so a redirecting source is now
treated like one that is down, and the next source gets its turn. The
same probe now makes one handler call and one Mojang request.
Neither Mojang's nor LittleSkin's hasJoined redirects.

The new subtest puts a redirecting root ahead of an honest one and
checks that the redirect target is never contacted and the honest
source's player is returned. The pre-fix handler fails it.
2026-09-22 12:50:56 +09:00
flyemoji 8fe255e38f fix(api): relay Mojang logins when no auth source is configured
felis api wired the hasJoined multiplexer only when felis.toml had at
least one [[auth_source]]. With none, the source list stayed nil and
every hasJoined answer was a 204. That was harmless while nothing
pointed at the route, but the installer now starts Velocity with
-Dmojang.sessionserver aimed at felis-api unconditionally, and the
generated felis.toml tells the operator to delete the LittleSkin block
for a Mojang-only server. Doing exactly that turned every login away,
premium accounts included, and felis-api logged nothing about it.

Always build the list through authSourcesFromConfig, which prepends
Mojang in code, so an empty config is a Mojang-only relay. felis nano
already behaves this way with the same file.

The new test pins authSourcesFromConfig itself: Mojang first, the only
Identity source, and still present when nothing is configured. Marking
a configured source Identity makes it fail. The call site in cmdAPI is
now a single unconditional assignment and has no unit test of its own.
2026-09-22 12:46:00 +09:00
flyemoji 800a9042a1 test: use placeholder domains in setup and system-server tests
Three tests carried the maintainer's production root domain, a personal
mailbox and the public IP of a live demo host as fixture values. None of
them needs the value to be real: the re-domain test only needs two
different roots, and the setup flow only needs a well-formed address.

Swap them for the placeholders the rest of the suite already uses
(mc.example.net, [email protected]), and move the "before" root in the
re-domain test to 203.0.113.10.nip.io. That address is from the RFC 5737
documentation range, so the stale install the test models still has an
IP-derived hostname, which is the case the refresh exists for.
2026-09-22 12:43:54 +09:00
flyemoji afdbfac7a8 docs: index the deferred integration seams and correct two stale markers
INTEGRATION-ONLY and KNOWN-LIMITATION are grep-able, but the grep answers the
wrong question. Thirty-four Go sites share the two markers and they carry four
different meanings: "declared, nothing implements it" reads exactly like
"implemented, only its I/O is unreachable from here", and neither reads
differently from a limitation that was accepted on purpose and is not coming
back. docs/deferred-seams.md sorts them, following the bucketed shape
internal/updater/doc.go already uses for its own package rather than starting a
second convention.

Sorting them turned up two markers that had outlived the condition they describe.

config.go called the modpack upload transport a deferred integration after both
backends had shipped -- LocalContextStore and S3ContextStore, selected in
cmd/felis by the shape of user_uploads_context, with the uploads PVC mounted and
the felis-uploads-s3 Secret rendered. What is still deferred is the far end:
Kaniko reading that context from inside the build Pod.

updater/doc.go listed the `felis update` CLI and the off-cluster Velocity jar read
under REMAINING INTEGRATION. Both exist -- cmd/felis/update.go, and
gatherer_host.go, which answers Velocity from the installed jar's manifest and
felis-api from the running binary's build stamp. The two nil seams that bullet
also names are real, but they belong to the in-cluster gatherer only, so the
bullet now says which caller has what and which is still empty.

The index also records the collision that makes a naive grep misleading:
docs/troubleshooting.md uses [INTEGRATION-ONLY] for something else, defined in its
own opening at :19 -- the symptom is produced by the kubelet, kaniko or a live
handshake, so it cannot be reproduced from the repository. Those twelve marks say
where a failure comes from, not that something is unbuilt, and are excluded.

Both code changes are comments. Every file:line the index cites was checked
against the line it points at.
2026-07-28 17:27:01 +09:00
flyemoji 23792d6251 fix(crd): remove spec.storage.retainOnDelete rather than leave it inert
The field validated, shipped in the CRD, and reached no controller. The world
PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate,
StatefulSet deletion does not cascade to template PVCs, and no finalizer exists
anywhere in the operator. So setting it true described what already happened,
and setting it false did nothing at all. False is the worse half: it reads as a
request to delete a world, and was silently ignored.

This departs from spec v4.1 §5, which asks for
"删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half
was ever built. Restoring that line means adding a finalizer whose other listed
duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs,
so the only work it would newly do is delete worlds, on a path that does not
pass the reaper's verified-backup check. The reaper is the one thing in the
system allowed to destroy a world and it earns that by proving a backup first.
A second door without that check is not an improvement.

The spec is a frozen versioned document, so it is left alone and the departure
is recorded in troubleshooting.md §13, beside the behaviour it explains. §12
loses its inert row and its opening sentence, which existed to introduce this
one field: every field in that table is now read by a controller.

Deployed installs need nothing. A CR still carrying retainOnDelete keeps
working, because a v1 CRD prunes unknown keys on the next write and the
behaviour the field claimed to control was never conditional.

go build, go vet and go test ./... pass on Linux with zero failures; the CRD
still parses and storage keeps size and storageClassName.
2026-07-28 17:27:01 +09:00
flyemoji 5d4f3063a9 feat(operator): make any Paper image joinable behind the forwarding proxy
Velocity modern forwarding is proxy-WIDE. A backend that cannot verify the signed
handshake does not degrade -- it rejects every login the proxy forwards. Until now the
only backends that could verify it were the two images Felis builds itself
(deploy/limbo, deploy/lobby), which read FELIS_FORWARDING_SECRET in their own
entrypoints. An arbitrary Paper image a user brings does not, so it passed admission,
started, reported Ready, and was UNJOINABLE. The platform's answer was to recommend the
lobby image as a base for a user's own world (0018_recommended_images.sql), which was
never a good base -- it carries the /menu plugin whose job is to TRANSFER a joining
player away, the exact opposite of a server you mean to stay on.

The fix configures forwarding from OUTSIDE the image instead of requiring it inside.
The operator now injects a root `felis init-forwarding` initContainer into every user
server; it writes the proxies.velocity block into config/paper-global.yml and forces
online-mode=false in server.properties on the /data PVC before the main container
starts. The image needs no forwarding logic of its own, so the joinable set stops being
"images that self-configure forwarding" and becomes every Paper-family image the
platform runs.

buildStatefulSet gates the injection on the ABSENCE of the system-role label: the
Felis-built system servers already consume the secret in their entrypoints and the
login gate is a limbo, not Paper. It is also gated on a non-empty felis image name --
the operator Deployment passes its own image as FELIS_IMAGE, and an operator without it
skips the injection rather than failing, because a cluster whose proxy is not in modern
mode has nothing to configure.

The init runs as root deliberately. The world volume's ownership comes from the storage
provisioner and the main container runs as whatever UID its image declares, so root is
the only UID that can reliably write these files; it then chmods them 0666/0777 so that
non-root main container can rewrite them on boot. The privilege is bounded -- the init
exits before the server container starts and the server container keeps its own UID.
The alternative, an fsGroup on the pod, is noted in the code as the upgrade path if the
init ever stops running as root.

The writer merges rather than overwrites, both because Paper expands paper-global.yml to
its full default tree on first boot and because the panel file editor may edit either
file between boots. It sets proxies.velocity.* and the single online-mode key and leaves
every other setting alone. It is a no-op on an empty secret, for the same reason the env
var is optional: a proxy that is not in modern mode provisions no Secret, and wedging
every server's init on a missing optional value would be worse than the status quo.

felis-paper (deploy/paper) is the platform's plain-Paper expression of that base and
0019 seeds it recommended: same PAPER_JAR_URL the lobby build already resolves, no /menu
plugin, no forwarding gate, and a correctly-escaped RCON channel so the console, the
online-player list and permission commands work out of the box. 0018's row is left in
place -- an admin who kept it can keep it; this only adds the better default beside it.

Three fixes ride along, each of which the 1.8 path hit in practice.

bootstrap pins ViaVersion's serverside-blockconnections off. ConnectionData.init() only
builds its block-connection provider when Via's lowest supported protocol is below 1.13;
under modern forwarding the Velocity injector reports 393, so init() returns early,
blockConnectionProvider stays null, and the first 1.12.2->1.13 chunk rewrite dereferences
it -- a 1.8 client takes an NPE on the first chunk it is sent and never finishes joining.
Every call site is behind isServersideBlockConnections(), so switching it off skips all
of them, at a cosmetic pre-1.13 cost: fences and glass panes stop drawing connected.
ViaVersion ships the option ON, so a fresh install shipped that NPE. Seeding a file with
this one key suffices -- Config#loadConfig parses the bundled default as the base map and
merges the on-disk file over it, so every other option stays current across version
bumps. The absence of "Loading block connection mappings" in the log is NOT evidence this
worked: init() gates on the protocol version too, and that half fails on its own, so the
line is missing either way. The config value is the only evidence, which is what the test
asserts.

The Velocity unit gains -Dfelis.legacy-forwarding.servers=legacy18. A protocol-47 backend
sits behind ViaVersion, which strips modern forwarding's login-plugin-message when it
down-translates the proxy->backend pipeline to 47 -- the packet is registered from 1.13
and has nowhere to go. Only the handshake address field survives Via, so the Felis fork
forwards the named servers BungeeCord-style while every other backend keeps modern+secret
untouched. v1 hardcodes the one legacy backend; rendering the list from the MinecraftServer
CRs is the upgrade path.

deploy/lobby's set_prop escapes the value before substituting it. The RCON password is
operator-provisioned arbitrary bytes, and a '|', '\' or '&' in one corrupts a bare
`sed s|...|...|` and silently kills the key -- taking the console, the online-player list
and permission commands with it. deploy/paper was written with the escaping, so the lobby
gets the same rather than leaving the sibling caller broken.

Verified: the full Go suite passes on Windows and on Fedora 44 (go1.26.4), where
TestWriteForwardingFileModes actually runs its POSIX mode assertions instead of skipping.
The new tests cover the initContainer's image, root UID, world mount and secret env; the
merge preserving unrelated config trees; the properties upsert including the commented-key
case; and the bootstrap script both writing the Via key and still calling the function
that writes it.

Not verified: the initContainer has never run in a real cluster, and the felis-paper
image is code-only here as the other game-stack images are -- no Go CI builds them.

The ViaVersion pin is the one piece with live evidence, and that evidence is what it was
written from. Before it, a client was cut within a second of "logged in with entity id"
on legacy18 while the proxy logged the NPE above -- REMAP OF LEVEL_CHUNK chained into
Protocol1_8To1_9's MAP_BULK_CHUNK. It was applied by hand to the running proxy on
2026-07-24 at 14:47 and only then written back into bootstrap. At 14:48:14 the same
player joined real Paper 1.8.8 through the fork, issued commands, approved an op-login
from in-game at 14:50:39, and held the connection until 15:30:09 -- 42 minutes.

Neither session says which client version it was. The proxy never logged a protocol
number. It bounds above at 1.16.4, from the viabackwards "(1.17->1.16.4) ... for 1.16
players and below" warning that fired for that player on the lobby leg, and no lower --
Via floors every handshake to the proxy's 393, so anything from 47 up is admissible.
Reading Protocol1_8To1_9 in the stack as a client-version tell is backwards: that chain
runs on the BACKEND leg, up-translating the 47 server's chunks to the floor. What the
NPE proves is that the pin was load-bearing, not who was holding the mouse.

That is one hand-run session on one host, and it is not a cell. The 393->47 leg has one
now, in Felis-Legacy -- FL-009 puts a genuine protocol-47 client on a stock Paper 1.8.8
behind this proxy and flips this same option: on it, cut 0.2s after JoinGame with the
fault above; off, holds. No automated test in THIS repository exercises the leg.
2026-07-28 09:13:58 +09:00
flyemoji 37ab87f07c docs(mail): say plainly what a green SMTP self-test does not prove
The Ping comment claimed the self-test could not produce a false negative and
left the impression it therefore proved deliverability. It does not, and the
distinction is the whole trap: a relay that gates sender identity at
end-of-DATA gates it on the way OUT. Fastmail answers 250 for
[email protected] addressed to the account's own mailbox and 551 5.7.1 "Not
authorised to send from this header address" for that same From addressed to
anyone else — and only the account's exact authorized identity passes the
second one, another local-part on the same domain is refused too.

So a green Ping means connect, TLS, AUTH and message shape are good, and
nothing more; the operator still has to have authorized From as a sending
identity with their provider, and the first real OTP is what proves they did.
Addressing the self-test somewhere external would not fix it either — the only
mailbox an operator can reliably check is usually inside the same account — so
the honest move is to scope the claim rather than buy false confidence with a
bigger probe.
2026-07-22 14:40:58 +09:00
flyemoji a8c4077202 fix(api): log the panic value and stack behind the opaque 500
withRecover turned a panicking handler into a 500 envelope and threw the panic
away. The client is meant to get an opaque "internal error" — that part is
right — but nothing was written server-side, so a recovered panic was an
untraceable 500: an operator holding "internal error" has no message, no
stack and no request to grep for, and diagnosis degrades into guessing against
a live install. That is what it cost during the email-OTP report.

The panic value, a stack and the method+path are now logged first, keyed by
the same request_id writeError already stamps on unmapped errors, so the
client envelope and the server log can be joined. The test pins all three
markers plus the unchanged 500/"panic" response, because a silent recover
looks exactly like a working one from the outside.
2026-07-22 14:40:42 +09:00
flyemoji c4c964578e fix(setup): stop the forced-onboarding gate trapping players who have no email
setup_required is what the SPA polls to decide whether the onboarding wall is
still owed, and it disagreed with the middleware that actually enforces the
wall. requireOnboarded lifts on a verified email OR an enrolled passkey;
setup_required answered `u.Email == "" || !hasPasskey`. A console-tier player
joins through the bind-code door with no email at all — by design, there is no
SMTP at that point — so the email term never clears and the SPA keeps them on
the setup screen forever, even after they enroll the passkey that already
unlocked the API for them.

The predicate now lives in one place (setupRequired) and both endpoints call
it, so the next edit to the unlock condition cannot drift them apart again.
Keying it on EmailVerified rather than email presence is the deliberate part:
presence is exactly the term that trapped the no-email player, and it was also
wrong on its own terms — an unverified address is not an authentication
factor, so it was never what the lockdown could safely lift on.

Also lands the regression test for the mechanism behind the live claim-403
report: /me/servers answers 200 for a bind-onboarded player (which is why the
dashboard renders the 认领 button at all) while claim, wake and status all
answer 403 with code "setup_required" — i.e. the refusal comes from
requireOnboarded before the handler, not from isOwnerOrAdmin inside it, which
would have said "forbidden". Enrolling a passkey and changing nothing else
lifts all three, which isolates the gate as the sole cause. The backend authz
is correct; the button that leads a locked-down player into a 403 is the
frontend's to hide.
2026-07-22 14:40:28 +09:00
flyemoji 694e3cb800 feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.

The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.

Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.

RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.

Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.

That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.

Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.

Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.

Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
2026-07-21 00:12:37 +09:00