Use shared action buttons and default to the login space. Allow staff to save startup-only experience settings while running and request a durable restart through the existing operator flow.
Grant the API PVC list permission needed to detect retained worlds before server creation, and show internal failures with a request ID. Cover permission, maintenance, restart and UI behavior with regression checks.
Configure connection and storage before creating or resuming the one-time Owner login. Remove Minecraft prerequisites from setup and preserve established login credentials.
Let staff preview and confirm roles from configured authentication sources using the existing account-link storage and game UUID mapping. Retain in-game code proof for players, add client-version and lobby guidance, and support NodePort passkey origins.
Two changes to the registry Deployment, both prerequisite to GC-durable images:
- Dedicated resource template: the control plane's 256Mi memory limit was a
live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.
- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
reach the Service VIP (live stack: "Empty reply"), so the node-side pull
path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
the plain-HTTP registry off every other interface.
Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.
Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.
Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.
Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:
- The stop used a whole-object Update while the same reconcile loop writes
status; that risks clobbering a concurrent status write. Switch to the
reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
closed with 403 (visible in the operator log as 'cannot update resource
"minecraftservers"'). Grant patch and pin it in the RBAC scope test.
With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:
- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
`warning suppressed — no warner wired` and does NOT stamp; a delivery error
logs and retries on the next daily run (bounded by the warning window). The
stamps are no longer burned by notices nobody received.
- A real channel: mail.SendNotice (the second and last message shape the mail
package sends) plus a mailWarner that resolves the owner's VERIFIED email
and mails the notice through the configured [smtp] relay. `felis reaper`
wires it when [smtp] is set (same password_ref convention as felis-api) and
prints exactly what happens when it is not.
Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.
Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:
- submit: derived context refs become the internal-face URL
/api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
image's new fetch-context entrypoint) that streams the blob with the
namespace-local service-token Secret — never mounted into Kaniko — and
extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
build namespace gets the token Secret through the existing replica mechanism
(bootstrap.sh + felis setup); the build egress lock opens exactly the control
namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
escapes/symlinks/non-gzip).
Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
The api, operator and registry Deployments shipped with no liveness/readiness
probes at all: a wedged process stayed 'Running' forever, and the operator had
no health listener to probe in the first place. Kaniko build evidence on a
fresh install showed the only cluster-wide red after a disk-pressure pass was
Deployment status that never reflected health.
- felis-api: readiness /readyz (DB + K8s API round-trip) and liveness /healthz
on the internal face (:8081), the only listener that serves both endpoints;
liveness deliberately avoids /readyz so a DB blip cannot restart the api.
- felis-operator: new --health-probe-bind-address (:8081) with controller-
runtime's /healthz + /readyz (registered ping checks; an unregistered handler
map would 404), plus the matching container port and probes.
- registry: /v2/ probes on the pinned port, so a broken storage backend stops
reading as 'Running'.
Tests pin paths, ports, and that each probe targets a declared container port.
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.
Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
Three faces of one gap, all on the supported install path:
- Backup/restore answered 503 out of the box: nothing ever rendered the
archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
'felis manifests' names it by default (--backup-pvc= is the explicit
no-store shape); bootstrap passes it through so the generated felis.toml
[archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
nothing deleted).
- Even when enabled it could not find a world on a stock install:
resolveWorldDir now also resolves the exact local-path directory
<pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
glob, so a stale deleted PV's bytes can't be archived in place of the
current world). Reaper Role gains persistentvolumeclaims:get (weaker
than the delete it already held).
README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
another namespace either (live drill: 'error looking up service account
minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
RoleBinding subject to the Minecraft namespace alongside the CronJob.
A Pod can only mount PVCs from its own namespace; the CronJob referenced the
minecraft-namespace backup PVC while being rendered under ControlNamespace, so
it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim
felis-backups not found'. The reaper Role/RoleBinding were already
minecraft-scoped (the objects it touches live there), so the CronJob was the
odd one out. The minecraft felis-config replica (felis setup, backup Job fix)
supplies its config mount.
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.
The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.