Commit Graph
7 Commits
Author SHA1 Message Date
Lemon-miaow f650bf892a fix(operator): idle auto-stop couldn't write — patch the spec, and grant the patch
Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:

- The stop used a whole-object Update while the same reconcile loop writes
  status; that risks clobbering a concurrent status write. Switch to the
  reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
  markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
  closed with 403 (visible in the operator log as 'cannot update resource
  "minecraftservers"'). Grant patch and pin it in the RBAC scope test.

With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
2026-09-23 04:08:58 +08:00
Lemon-miaow fd33fd05e1 fix(install): backups exist on a default install; retention resolves real world dirs (#6)
Three faces of one gap, all on the supported install path:

- Backup/restore answered 503 out of the box: nothing ever rendered the
  archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
  PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
  'felis manifests' names it by default (--backup-pvc= is the explicit
  no-store shape); bootstrap passes it through so the generated felis.toml
  [archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
  now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
  env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
  nothing deleted).
- Even when enabled it could not find a world on a stock install:
  resolveWorldDir now also resolves the exact local-path directory
  <pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
  glob, so a stale deleted PV's bytes can't be archived in place of the
  current world). Reaper Role gains persistentvolumeclaims:get (weaker
  than the delete it already held).

README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
2026-09-22 20:48:07 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow c839454a1f fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
another namespace either (live drill: 'error looking up service account
minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
RoleBinding subject to the Minecraft namespace alongside the CronJob.
2026-09-22 20:18:36 +08:00
Lemon-miaow 90ccbfede4 fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
2026-09-22 17:50:04 +08:00
flyemoji 694e3cb800 feat(rcon): provision per-server RCON so the console, player list and permissions work
A server created through the panel never had RCON. CreateServer built a
MinecraftServerSpec without a Rcon block at all, so the field took its zero value
and every downstream consumer read Enabled=false. Nothing failed loudly: the
operator skips the probe when RCON is off and marks the server Ready on pod
readiness alone, so the panel showed "运行中" for a server the control plane could
not talk to. Everything that rides the write channel (spec §8 写=RCON) was dead —
the online-player list returned nothing because Status.Players is only ever
sampled by the probe, and console writes answered 503 ErrConsoleUnavailable
because internal/api/console.go refuses when Enabled is false.

The whole RCON machinery already existed — builders gate the service port,
container port, preStop save-and-stop hook and the RCON_* env on Spec.Rcon,
the reconciler probes and reports, console.go dials, the NetworkPolicy opens
25575 to {api, operator}. The only thing missing was that nobody ever turned it
on or created a password. This wires the three layers that were absent.

Provisioning lives in the operator, not in felis-api. felis-api holds secrets:get
and not create, and giving it create solely to mint a password it immediately
stops caring about (console.go re-reads the Secret at command time) would widen
the API's powers for nothing. The operator already reads every Secret in the
namespace, so adding create there grants no read it did not have. It also makes
provisioning declarative: a Secret deleted by hand comes back on the next pass, a
controller reference garbage-collects it with the server so no delete path has to
remember it, and a server that predates RCON only needs spec.rcon filled in for
the password to appear. The name comes from naming.RconSecretName so felis-api,
`felis setup` and the operator cannot drift apart on it.

RCON is enabled per system service rather than by default, because enabling it on
a backend that serves no RCON listener is destructive rather than merely useless:
the operator gates readiness on the probe, so such a server never leaves Starting
and is eventually marked Failed. The login limbo is exactly that backend
(LOOHP/Limbo has no RCON) and it is the front door, so it stays off; the lobby
runs Paper and is administered through the panel like any other server, so it is
on.

Paper only reads RCON settings from server.properties, so the operator's injected
RCON_PASSWORD did nothing on its own — felis-lobby's entrypoint now writes the
three keys on every boot. Rewriting them each time makes the copy in the world
volume derived state rather than the source of truth, so an owner who edits them
through the panel's file editor cannot lock the control plane out of their own
server. Without a password it sets enable-rcon=false and warns rather than
refusing to start: unlike the forwarding secret, a missing RCON password degrades
the server rather than making it unsafe.

That password landing in server.properties is a §286 exposure (RCON 密码绝不下发
前端), since server.properties is readable through the file editor. It is redacted
on read rather than the file being denied outright the way config/paper-global.yml
is: the forwarding secret is cluster-wide material that merely happens to sit in
the volume, whereas server.properties is the single most-edited config an owner
has, and hiding one line should not cost them MOTD, difficulty and view-distance.
The write path is deliberately left alone — the boot-time rewrite restores the
real value, which is what makes redacting rather than denying safe here.

Also guards idle auto-stop on Rcon.Enabled. Status.Players is only meaningful
when the probe ran; with RCON off it keeps its zero value, which that branch would
have read as "empty" and used to stop a server full of people. AutoStopEnabled is
not currently settable through any path, so this is a latent footgun rather than a
live bug, but it is one line and the alternative is discovering it in production.

Checks: the operator provisions a missing Secret with a 32-hex-char password and a
controller reference, and does not rotate an existing one; idle auto-stop stays
inert without RCON; the editor redacts rcon.password from the world root's
server.properties while leaving the rest of the file (and a plugin's own nested
copy) intact; login has RCON off and lobby has it on with the shared secret name;
CreateServer sets the block. That last one departs from K8sCluster being
integration-tested against a live cluster: this defect was a struct literal
missing a field, it shipped, and a fake client is enough to pin a struct literal.

Existing servers are NOT migrated by this change — CreateServer only covers new
ones and ensureSystemServers is create-if-absent, so a `felis setup` re-run will
not touch an existing lobby. A deployed install additionally needs the
felis-lobby image rebuilt and re-imported for the entrypoint change, and its pods
recreated, before the RCON keys reach server.properties.
2026-07-21 00:12:37 +09:00
flyemoji 47fcd90f75 feat(platform): add node orchestration and the felis entrypoint
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them.
2026-06-26 23:32:38 +09:00