Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:
- The stop used a whole-object Update while the same reconcile loop writes
status; that risks clobbering a concurrent status write. Switch to the
reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
closed with 403 (visible in the operator log as 'cannot update resource
"minecraftservers"'). Grant patch and pin it in the RBAC scope test.
With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
Three faces of one gap, all on the supported install path:
- Backup/restore answered 503 out of the box: nothing ever rendered the
archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
'felis manifests' names it by default (--backup-pvc= is the explicit
no-store shape); bootstrap passes it through so the generated felis.toml
[archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
nothing deleted).
- Even when enabled it could not find a world on a stock install:
resolveWorldDir now also resolves the exact local-path directory
<pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
glob, so a stale deleted PV's bytes can't be archived in place of the
current world). Reaper Role gains persistentvolumeclaims:get (weaker
than the delete it already held).
README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.
The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them.