- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
build_mem_limit overrides; empty keeps the compiled-in defaults. An
air-gapped or mirrored install has no route to gcr.io/aquasec (the
build egress policy allows only DNS + registry + package mirrors), so
builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
impossible (PVCs cannot cross namespaces) and that the s3 lane also
lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
13b rewritten (verified eviction refusal, 5m pressure-transition,
image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
volume (worlds + config + plugins + cache) and a restore rolls all of
it back — OpenAPI/README wording updated to match (same-tag images are
still watched for regressions by the openapi parity gate).
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.
Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.
Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
The field validated, shipped in the CRD, and reached no controller. The world
PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate,
StatefulSet deletion does not cascade to template PVCs, and no finalizer exists
anywhere in the operator. So setting it true described what already happened,
and setting it false did nothing at all. False is the worse half: it reads as a
request to delete a world, and was silently ignored.
This departs from spec v4.1 §5, which asks for
"删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half
was ever built. Restoring that line means adding a finalizer whose other listed
duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs,
so the only work it would newly do is delete worlds, on a path that does not
pass the reaper's verified-backup check. The reaper is the one thing in the
system allowed to destroy a world and it earns that by proving a backup first.
A second door without that check is not an improvement.
The spec is a frozen versioned document, so it is left alone and the departure
is recorded in troubleshooting.md §13, beside the behaviour it explains. §12
loses its inert row and its opening sentence, which existed to introduce this
one field: every field in that table is now read by a controller.
Deployed installs need nothing. A CR still carrying retainOnDelete keeps
working, because a v1 CRD prunes unknown keys on the next write and the
behaviour the field claimed to control was never conditional.
go build, go vet and go test ./... pass on Linux with zero failures; the CRD
still parses and storage keeps size and storageClassName.
Sections 11 and 12 told the operator that idle auto-stop, both startup
budgets, and the player tally are read by nobody. All four are read, and
§1 repeated the same claim in its strongest form: "the operator has no
start timeout ... loops forever".
spec.idle.autoStopEnabled reconciler.go:175
spec.idle.emptySecondsBeforeStop reconciler.go:175
spec.startup.timeoutSeconds reconciler.go:479, called at :126
spec.startup.readinessTimeoutSeconds reconciler.go:490, called at :157
status.players.online reconciler.go:413 (markRunningReady)
The repository already contained the disproof:
TestReconcileRunning_StartupTimeoutConvertsToFailed and
TestReconcileRunning_ReadinessTimeoutConvertsToFailed both assert the
escalation §1 said does not exist. The test §1 cited,
TestReconcileRunning_RconProbeFailureStaysStarting, only asserts that a
single failed probe does not flap the phase; that was read as "forever".
§11 was the costly one, because it misdiagnosed a configuration problem
as a missing feature. Idle auto-stop and the player tally both hang off
spec.rcon.enabled -- the tally is a by-product of the RCON readiness
probe (prober.go:63 runs `list`), and reconciler.go:175 carries the RCON
condition explicitly so a never-sampled zero cannot stop a server full of
people. Following the old text, an operator whose RCON was never enabled
would conclude the feature was unwritten and stop. The section now opens
with the jsonpath that reads spec.rcon.enabled.
Also separated the prober's fixed 5s dial timeout (prober.go:45) from
spec.startup.readinessTimeoutSeconds, which §1c conflated: the former
bounds one probe, the latter is a deadline for the whole start measured
from status.startRequestedAt.
spec.storage.retainOnDelete is the one field still genuinely inert, so
the [INERT] legend and §12 stay -- §12 now records the condition each
read field depends on instead of claiming none of them are read.
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.
Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).
Manifest-level fix; the live packet path is pending real-cluster verification.
Add docs/troubleshooting.md covering the common failure modes the spec
implies, grounded in the actual control-plane code paths:
- Stuck Starting (PodNotReady / RconSecretUnavailable / RconNotReachable)
and the deliberate absence of a Starting->Failed timeout.
- Failed reachable only via InvalidSpec on a malformed spec.storage.size,
plus the stale status.endpoint=direct caveat after a failure.
- Routing via status.endpoint direct/fallback and the empty fallbackServer
pitfall; wake 403/429/503 gate order.
- online-mode coupling and Velocity's offline-mode routing refusal.
- Cloudflare Access 401/403, nil-Keyfunc fail-closed, audience checks,
the absence of an issuer check, and local-session gating.
- Internal service-token (FELIS_SERVICE_TOKEN) rejection path.
- link/claim error codes, Kaniko build denials (SA-by-absence RBAC,
default-deny egress, internal-registry push gate), and the registry
DNS contract.
- Reaper backup-before-delete invariant and false-delete vectors.
- Unimplemented idle auto-stop, permanently-zero players.online, the
inert CRD fields, and the always-survives world PVC behaviour.
Each item is labelled with its evidence grade (GO-TESTED / CODE-ONLY /
INTEGRATION-ONLY / INERT) so operators know what is verified versus
asserted.