Commit Graph
9 Commits
Author SHA1 Message Date
Lemon-miaow c57daaf861 feat(cli): add felis converge for fields a newer desired spec never delivered (tracker #1)
Provisioning is create-if-absent, so a field the desired spec gained after an
install (spec.rcon, spec.startup.healthHTTPPort, a derived env key) never
reaches the existing login/lobby CR while every re-run of setup reports
success — the reported 'configuration updates never reach an installed
deployment' symptom. converge is the explicit pass: it fills exactly the
zero-valued whitelist fields and the derived env (including a missing key,
which refreshDerivedEnv deliberately never adds), and never overwrites a
non-zero value. The timing stays with the operator because enabling RCON or
the HTTP readiness gate on a pre-listener image would wedge that server in
Starting until it was marked Failed.

Tests: fills predated fields while operator edits survive / non-zero values
left alone / absent + foreign + unset-image guards. usage table updated so the
router-parity test passes; troubleshooting gains §12b.
2026-09-24 10:35:14 +08:00
Lemon-miaow 328e570309 fix(setup,install): refresh the workload namespace's felis-config mirror (#51) 2026-09-23 20:07:20 +08:00
Lemon-miaow 8e7c7bbf24 fix(reaper): deliver pre-reap warnings for real — and never fake a delivery
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:

- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
  `warning suppressed — no warner wired` and does NOT stamp; a delivery error
  logs and retries on the next daily run (bounded by the warning window). The
  stamps are no longer burned by notices nobody received.

- A real channel: mail.SendNotice (the second and last message shape the mail
  package sends) plus a mailWarner that resolves the owner's VERIFIED email
  and mails the notice through the configured [smtp] relay. `felis reaper`
  wires it when [smtp] is set (same password_ref convention as felis-api) and
  prints exactly what happens when it is not.

Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.

Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
2026-09-23 03:47:19 +08:00
Lemon-miaow f79e5ebb5e feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:

- submit: derived context refs become the internal-face URL
  /api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
  gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
  route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
  image's new fetch-context entrypoint) that streams the blob with the
  namespace-local service-token Secret — never mounted into Kaniko — and
  extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
  reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
  build namespace gets the token Secret through the existing replica mechanism
  (bootstrap.sh + felis setup); the build egress lock opens exactly the control
  namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
  escapes/symlinks/non-gzip).

Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
2026-09-22 22:45:09 +08:00
flyemoji a0064adfdd fix(setup): stop the login gate sending players to the old console after a re-domain
Changing root_domain updated felis.toml and the panel, but the login gate kept
pointing players at the hostname it was created with. FELIS_ROOT_DOMAIN and
FELIS_PANEL_HOSTNAME are baked into the login MinecraftServer at provisioning time,
FelisLimboPlugin reads them to build the link an unauthenticated player is told to
open, and ensureSystemServers is create-if-absent — so nothing in the install ever
rewrote them. On the demo host the CR still carried
console.159.223.32.51.nip.io hours after the domain had moved to
mc.flyemoji.network: every joining player was handed a link that bypasses the
tunnel, hits the node directly and trips a certificate warning, on the one screen
someone with no account is guaranteed to see.

setup now converges these values on an existing system service instead of skipping
it. Create-if-absent stays the rule for everything else, and the comment on it is
still true — an operator's edits to a system service must survive a re-run. These
three names are the exception because they are not the operator's to own: they are a
copy of config that is wrong the moment config changes, and there is no other writer
who could notice.

The convergence is deliberately narrow. Only a name already present with a different
value is rewritten, so env the operator added by hand is untouched and the rest of
the spec — image, memory, storage — is not read at all. A derived name that is
absent from the live object is left absent rather than added back: a deliberate
removal and drift look identical from here, and re-adding it would mean fighting the
operator on every run. The outcome string reports the refresh so a setup run does not
silently rewrite the front door.

This closes one surface of a re-domain, not the whole of it. The write-once panel
certificate at deploy/bootstrap.sh keeps its old SANs, and so do the Velocity config
and the forwarding material; the warning in bootstrap.sh that says so is still
accurate. What changes is that the surface players actually walk through now catches
up when setup is re-run.

Checks: a login gate built with the old domain converges onto the new one and says
so; an env var the operator added and a hand-raised javaMemory both survive that
same run; and an install whose config already matches reports no refresh, so a
routine setup does not read like a re-domain. The middle one is the one worth having
— converging config must not turn into a licence to clobber the edits
create-if-absent exists to protect.
2026-07-21 00:47:04 +09:00
flyemoji 694e3cb800 feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.

The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.

Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.

RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.

Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.

That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.

Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.

Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.

Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
2026-07-21 00:12:37 +09:00
flyemoji 9ea35304e3 fix(setup): source the console host from the panel hostname, not op.console
The owner setup URL and the limbo login link were built from the admin host
(op.console.<root>, with an op.console.localhost fallback) and a hardcoded
console.<root>, so an operator who set a custom panel_hostname got an unreachable setup
link and a wrong login target. Thread the resolved panel host (defaultPanelHostname)
through performSetupMCBind, the MC-bind TUI, and the login system-server env
(new FELIS_PANEL_HOSTNAME); the limbo plugin prefers it and keeps console.<root> only as
the fallback for an older operator whose env predates it. This also matters for security:
the only wired WebAuthn verifier is scoped to the panel host, so passkey enrollment must
land on the panel face, never op.console.

While here, the limbo login handler checks link status before minting a bind code: an
already-linked player is sent straight to the lobby instead of being shown a useless code.
2026-07-16 13:27:02 +09:00
flyemoji dab8fc214b feat(setup): bind owner through login gate 2026-07-14 02:57:15 +09:00
flyemoji f554d525d4 feat(cli): provision login/lobby system servers with login env and token replica
setup builds the always-on, reaper-exempt login/lobby MinecraftServers (create-if-absent), bakes the login limbo's non-secret config (internal API URL, root domain, lobby name) into spec.env, and replicates the felis-service-token Secret from the control namespace into the minecraft namespace so the operator's namespace-local secretKeyRef on the login pod resolves.
2026-07-02 19:38:37 +09:00