# Felis Troubleshooting Checklist This file is the spec §28 #23 deliverable: a故障排查清单 (troubleshooting checklist) for the failure modes the platform actually produces. Every symptom below is traced to a concrete signal — a `status.conditions` reason, an HTTP error code, a log string, or a manifest name — so an operator can map what they see to the code path that emitted it. ## How to read this document Each entry is **symptom → likely cause → where to look → fix**. Signals are graded for how far the in-repo Go test suite proves the behaviour: - **[GO-TESTED]** — a hermetic `*_test.go` exercises this exact path; the string/code is asserted in CI. - **[CODE-ONLY]** — the code path and string exist and are real, but no unit test drives them (notably the Velocity Java plugin, which is not compiled or tested in this repo). - **[INTEGRATION-ONLY]** — the symptom is produced by the kubelet, kaniko, containerd, Postgres, or the network, not by Felis Go code; you will see it in `kubectl describe` / pod logs, never in `MinecraftServer.status`. - **[INERT]** — the configuration field exists in the CRD but no controller reads it. Tuning it does nothing. §12 lists the one field this still applies to, alongside the fields that *are* read and the condition each depends on. The operator never invents the parent domain; routing identity is `spec.subdomain` under the deployment zone. Examples below use `` / `registry..svc:` placeholders rather than any concrete host. --- ## 1. Server is stuck in `Starting` and never becomes `Running` `MinecraftServer.status.phase` stays `Starting`. A start that never succeeds is requeued every 5s until one of the two startup budgets expires, then escalated to `Failed` — `StartupTimeout` if the pod never passed TCP readiness, `ReadinessTimeout` if the RCON probe never succeeded. Both default to **300s** when `spec.startup.timeoutSeconds` / `spec.startup.readinessTimeoutSeconds` are unset or `0`, and both are measured from `status.startRequestedAt`. [GO-TESTED: `TestReconcileRunning_StartupTimeoutConvertsToFailed`, `TestReconcileRunning_ReadinessTimeoutConvertsToFailed`.] So `Starting` seen *once* is normal and `TestReconcileRunning_RconProbeFailureStaysStarting` asserts exactly that — a single failed probe must not flap the phase. `Starting` seen for longer than the budget means the reconcile loop is not running at all; check the operator's own logs before tuning anything. Either way the underlying cause is diagnosed from pod state, not from `MinecraftServer.status`. First, read the condition reason: ``` kubectl get minecraftserver -o jsonpath='{.status.conditions}' ``` `markStarting` writes the same reason to both `Ready=False` and `RconReached=False`. The reason is exactly one of: | `status.conditions[].reason` | Meaning | Requeue | |---|---|---| | `PodNotReady` | Pod not TCP-ready yet (`status.readyReplicas < 1`) | 5s | | `RconSecretUnavailable` | RCON secret missing or malformed | 10s | | `RconNotReachable` | RCON dial/auth failed | 5s | [GO-TESTED for the reason set.] ### 1a. `PodNotReady` — pod never goes ready The operator cannot tell *why* the pod is not ready; **a broken image, an unbound PVC, and a backend that simply has not finished booting all surface as the identical `PodNotReady` signal.** [GO-TESTED that the reason is emitted; [INTEGRATION-ONLY] for the underlying pod cause.] You must drop to the pod: ``` kubectl get pod -l app.kubernetes.io/name= kubectl describe pod # look at Events + container State ``` - **`ImagePullBackOff` / `ErrImagePull`** → `spec.image` is wrong, the tag does not exist, or the registry is unreachable. `spec.image` is copied verbatim into the container with **zero validation** by the operator. Fix the image reference, or see §7 (registry reachability) and §6 (build push target). - **PVC `Pending`** → `kubectl get pvc -l app.kubernetes.io/name=`. A nonexistent `spec.storage.storageClassName`, or a request larger than any class can satisfy, leaves the PVC unbound. The operator does **not** error on this (only a malformed *quantity* errors — §2); it waits in `Starting` indefinitely. Fix the StorageClass name or capacity. [INTEGRATION-ONLY.] - **Container crash-looping before the readiness port opens** → check container logs; this is a backend/entrypoint problem, not a Felis problem. ### 1b. `RconSecretUnavailable` — RCON secret missing or malformed The probe needs the RCON password from `spec.rcon.secretRef`. The message is the verbatim error: [CODE-ONLY for these branches] - `rcon.secretRef.name and .key are required when rcon is enabled` — you enabled `spec.rcon.enabled` but left `secretRef.name` or `secretRef.key` empty. - `secret "" has no key ""` — the Secret exists but lacks the named key. Fix: create the Secret with the referenced key, or correct `secretRef`. Verify: ``` kubectl get secret -o jsonpath='{.data.}' | base64 -d | wc -c ``` ### 1c. `RconNotReachable` — RCON dial or auth failed The probe is a TCP connect **plus** RCON auth handshake, then immediate close — **no command is ever run; a successful auth IS the entire readiness gate.** The message is the verbatim dial error: - `rcon: authentication failed` → the password in the Secret does not match the backend's `rcon.password`. Reconcile the two. [GO-TESTED that this maps to `RconNotReachable`.] - `connection refused` / `i/o timeout` → the backend has not opened the RCON port yet, RCON is disabled in `server.properties`, or `spec.rcon.port` (default 25575) is wrong. [INTEGRATION-ONLY for the live handshake.] The per-probe timeout is a fixed 5s in code (`prober.go:45`, shortened further if the reconcile context has a nearer deadline). It is **not** derived from `spec.startup.readinessTimeoutSeconds`, which is the deadline for the whole start, not for one probe — see §12. --- ## 2. Server is in `Failed` There is exactly **one** path to `PhaseFailed`: `markFailed(server, "InvalidSpec", err)`, reached only when the StatefulSet cannot be built. In practice this means **a malformed `spec.storage.size`** (an unparseable resource quantity), surfaced as: ``` invalid storage size "": ``` `status.conditions` will show `Ready=False` and `Provisioned=False`, both with reason `InvalidSpec`. Fix the quantity (e.g. `10Gi`, not `10 GB`) and the server leaves `Failed` on the next reconcile. [GO path is real; the specific branch is [CODE-ONLY] — the reconciler test fixture uses a valid size.] Two caveats when a server has been in `Failed`: - A pod-not-ready or unreachable-RCON server is **never** `Failed`; it is `Starting` (§1). If you see `Failed`, it is a spec problem, not a runtime one. - `markFailed` does **not** reset `status.endpoint`. A server that fails after having been `Running` keeps a stale `endpoint.mode=direct`. The proxy should treat any non-`Running` phase as "do not route direct" rather than trusting a lingering `direct` endpoint (§4). --- ## 3. Players can't join / get sent to the wrong place Routing is driven by `status.endpoint`: - `markRunningReady` is the **only** writer of `endpoint.mode=direct` (`address=`), and only while `phase=Running` and RCON-ready. - `markStarting`, `markStopping`, `markStopped` all write `endpoint.mode=fallback`, `address=spec.fallbackServer`. So the proxy should route `direct` **only** when `phase=Running`; otherwise it gets a `fallback` endpoint. [GO-TESTED for the direct/fallback toggle via `markRunningReady`/`markStopped`.] Two pitfalls: 1. **`spec.fallbackServer` is empty** → the fallback endpoint `address` is `""`, so during `Starting`/`Stopping`/`Stopped` the proxy has no lobby to park the player in. Set `spec.fallbackServer` to a registered Velocity server name. 2. **Stale `direct` after `Failed`** → see §2; the proxy must not honour a `direct` endpoint unless `phase=Running`. ### 3a. Wake-on-join is refused, slow, or rate-limited When a player joins a stopped server, the proxy parks them in the fallback and calls the internal wake API. The authorization gate order is **autostartPolicy → per-server cooldown → global running cap**. Map the API result: | HTTP | Code | Cause | Fix | |---|---|---|---| | `403` | `forbidden` | `autostartPolicy=allowlist` and UUID not allowlisted, or `ownerOnly` and caller is not owner | Add the UUID / claim the server / set `autostartPolicy=public` | | `429` | (cooldown) | Wake retried within the 30s per-server `WakeCooldown` | Wait out the cooldown | | `503` | `at_capacity` | Global `MaxRunningServers` cap reached | Stop another server or raise the cap | [GO-TESTED: `handlers_internal_wake_test.go`, cooldown, running-cap shape.] The operator's RCON probe — **not** the wake call — is the authoritative readiness gate; the proxy polls `GET /api/v1/internal/servers/{name}/status` every ~2s and teleports when `ready=true`. The Velocity-side consumption of these codes (`403` → "You're not allowed to start «server»"; `429` → re-queue; other → "Couldn't start … Try again shortly.") lives in the Java plugin and is **[CODE-ONLY]** — the codes it reacts to are produced by the Go-tested `authorizeWakeByUUID` / cooldown limiter, so grade the two halves separately. --- ## 4. Routing is disabled even though servers are up (online-mode coupling) Wake/claim/allowlist semantics trust **Mojang-verified online-mode UUIDs**. If the proxy runs `online-mode=false`, those identities are spoofable, so the Velocity plugin **refuses to activate routing**: ``` Felis routing DISABLED: the proxy is in offline mode (online-mode=false). Domain autostart and the allowlist trust Mojang-verified UUIDs; refusing to route on spoofable identities. /link remains available. Set online-mode=true to enable routing. ``` [CODE-ONLY — `FelisVelocityPlugin.onProxyInitialize`.] `/link` still works (account binding does not depend on routing), but no domain autostart happens. The summary line prints `routing: disabled (offline mode)` or `(no root-domain set)`. Fix: set `online-mode=true` on the proxy, or configure the root domain if the log says `no root-domain set`. On the server side, `spec.autostartPolicy` and `spec.onlineMode` are documented as only meaningful when the proxy enforces `online-mode=true`. Setting them does not by itself make an offline proxy safe — the proxy guard is the enforcement point. --- ## 5. Web panel returns 401 / 403 (Zero-Trust / Cloudflare Access) The external face accepts either a Cloudflare Access JWT (`Cf-Access-Jwt-Assertion` header) **or** a local session cookie. The error envelope is always `{"error":{"code","message","request_id"}}`. [GO-TESTED.] - **`401 unauthorized`** — not authenticated: no/invalid Access JWT and no valid session. [GO-TESTED.] - **`403 forbidden`** — authenticated but not permitted (e.g. a non-admin principal hitting an admin route; `IsAdmin()` requires `role=admin` **and** arrival via the admin Access audience/host). [GO-TESTED.] ### 5a. Every external request 401s on a fresh deploy The Access verifier is wired **fail-closed**: `Keyfunc` (the JWKS key function) is `nil` until deployment wiring supplies it. With a nil Keyfunc, **every** JWT verification fails, and startup logs: ``` felis api: external face fails closed (Access JWKS key function not configured) ``` [INTEGRATION-ONLY — the live JWKS path is a deployment point.] This is intended: the panel rejects all callers until JWKS is configured. Fix by wiring the Access JWKS key function for `cfg.Auth.AccessJWTAud`. ### 5b. Token rejected with audience error ``` token audience does not include "" ``` The JWT's `aud` claim does not contain the configured `cfg.Auth.AccessJWTAud` (or the admin audience for admin routes). [GO-TESTED.] Confirm the Access application audience matches `cfg.Auth.AccessJWTAud`. **Trust-model note for operators:** verification is **expiration-required + audience + signing-key (JWKS)**. There is **no `iss` (issuer) check** anywhere in the verifier. Trust rests entirely on the audience claim plus the JWKS signing key. When documenting or auditing the trust boundary, do not assume issuer is validated — it is not. ### 5c. Local-password login fails or is silently rejected Local sessions use the `felis_session` cookie (HttpOnly, Secure, SameSite=Lax, 12h TTL, host-only). They are gated by the `local_auth_enabled` row in `platform_settings`, read live per request and **fail-closed** (missing or unparseable → treated as disabled). Symptoms: - Cookie present but login rejected with `local auth disabled` → the `local_auth_enabled` setting is false/absent. A present cookie under disabled local-auth is **rejected outright**, not fallen through to the JWT path. - `invalid session: …` → bad/forged session hash. Fix: set `local_auth_enabled=true` in `platform_settings` if local password auth is intended. [GO-TESTED for the session/QR-login logic.] --- ## 6. Internal API rejects Velocity / proxy callers (service-token) The internal face (`--internal-addr :8081`, routes under `/api/v1/internal/...`) is **never** Zero-Trust; it authenticates a single service token via `Authorization: Bearer `, compared in constant time. In-cluster it is reached through the ClusterIP Service `felis-api-internal` (port 8081), which is separate from the external NodePort `felis-api` (443) precisely so the no-Zero-Trust face is never exposed on a node. On the control-plane node the break-glass console reaches it by resolving that Service's ClusterIP and dialing `:8081`. - **Internal calls fail to *connect* (not 401)** → the `felis-api-internal` Service is missing or its selector no longer matches the api pods. `kubectl -n felis get svc felis-api-internal` must show a ClusterIP with 8081; a bare `felis-api` name serves only 443 and every internal call would hang/refuse. - **All internal calls 401** → the token is unset or wrong. The API reads env `FELIS_SERVICE_TOKEN`. If unset, startup logs: ``` felis api: warning: FELIS_SERVICE_TOKEN unset — internal face will reject all callers ``` and wires an empty token, which rejects **everyone** (no bypass). [GO-TESTED for the constant-time compare / empty-token rejection.] In-cluster, the token's source of truth is the Secret `felis-service-token` (key `token`), injected as `FELIS_SERVICE_TOKEN` on the API Deployment. Fix: ``` kubectl get secret felis-service-token -o jsonpath='{.data.token}' | base64 -d ``` Ensure the proxy is configured with the identical value. --- ## 7. Account-link and claim API errors Codes from `handlers_account.go` / `handlers_internal.go`. [GO-TESTED.] | HTTP | Code | When | |---|---|---| | `400` | `bad_request` | Missing `mc_uuid`, empty code, or invalid `auth_source` (must be `mojang`/`thirdparty`) | | `400` | `invalid_code` | Link code unknown or expired (10-min TTL, 8-symbol code) | | `409` | `already_linked` | That MC UUID is already linked to **another** user | | `412` | `not_linked` | Claim/owner op by a caller with no verified account link | | `403` | `quota_exceeded` | Claim would exceed the user's server quota | | `409` | `already_claimed` | Atomic `UPDATE … WHERE owner_id IS NULL` affected 0 rows | | `404` | `not_found` | Unknown server/resource | Flow reminder: the **code is minted in-game** on the internal face (`POST /api/v1/internal/account/link/code`, proves the UUID) and **verified on the web** external face (`POST /api/v1/account/link/verify`, proves the user). A `409 already_linked` rolls the transaction back and **preserves** the code so a different user can still use it. The claim's race-safety is the single conditional `UPDATE` under READ COMMITTED — `1` row → `200`, `0` rows → `409`. The real SQL execution is [INTEGRATION-ONLY] (no sqlmock/dockertest in repo); the handler logic is [GO-TESTED] via an in-memory fake repo. --- ## 8. Image build fails (Kaniko, spec §14/§16) ### 8a. Build push rejected at submission with `400` Pre-build validation rejects any push target that is not the internal registry: ``` must target the internal registry ".svc:>", not "" ``` [GO-TESTED via the `validate` gate.] Fix the image reference to push to `cfg.Registry.URL` (the internal registry — §9). ### 8b. Build Job's ServiceAccount can do nothing (RBAC "denial" by design) The build/restore Job SAs (`felis-build`, `felis-restore`, namespace `felis-build`) have **no Role and no RoleBinding anywhere** — isolation is the *absence* of permissions (spec §16/§21). If you see the build SA denied a namespaced API operation, **that is correct, not a misconfiguration.** [GO-TESTED that the rendered manifests give these SAs no Role.] Do not "fix" it by granting the build SA permissions. ### 8c. Build hangs then fails fetching base image / packages The build namespace runs a default-deny egress NetworkPolicy (`felis-build-egress`); the **only** allowed egress is the package-mirror CIDRs from `--package-cidr`, which **defaults to none** (fail-closed, no internet). [GO-TESTED for the netpol shape.] A kaniko run that hangs pulling a base image or OS package from a non-allowlisted host is the egress policy doing its job — the hang/failure text comes from kaniko/containerd and is [INTEGRATION-ONLY]. Fix: add the mirror CIDR via `--package-cidr`, or pre-stage the base image in the internal registry. ### 8d. Build reaches `Failed` phase `reconcileBuilds` polls the Job; a Job reaching `Failed` is surfaced via `writeBuildError` (JobPhase→Failed). [GO-TESTED for the mapping.] The underlying cause — a kaniko build error or the **Trivy CRITICAL-CVE gate** failing the build before push (spec §16) — is in the Job's pod logs and is [INTEGRATION-ONLY]. Inspect: ``` kubectl logs -n felis-build job/ ``` ### 8e. Build Pods never start: executor images and air-gapped installs The build Job runs Kaniko and Trivy from external registries by default (`gcr.io/kaniko-project/executor:latest`, `aquasec/trivy:latest`). On a box whose build namespace cannot reach those registries (the egress policy allows only DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the build stays `building` until its deadline. Point the overrides at images **in the internal registry** — the one pull source that survives an image GC (a bare node-containerd import does not: kubelet's image GC collects unused images under disk pressure, and an air-gapped box then has nothing to restore them from) — in `felis.toml`: ```toml [registry] url = "registry.felis.svc:5000" build_namespace = "felis-build" kaniko_image = "registry.felis.svc:5000/mirror/kaniko-executor:v1.24.0" trivy_image = "registry.felis.svc:5000/mirror/trivy:0.74.0" trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2" trivy_java_db_repository = "registry.felis.svc:5000/mirror/trivy-java-db:1" build_cpu_limit = "2" build_mem_limit = "4Gi" ``` Mirror the executor images into the registry once. On the node itself, push through the loopback hostPort the registry Deployment binds (docker treats `127.0.0.1` as insecure by default; the installer leaves the daemon stopped, so `sudo systemctl start docker` first): ```sh docker pull gcr.io/kaniko-project/executor:v1.24.0 # any versions you pin docker pull aquasec/trivy:0.74.0 docker pull mirror.gcr.io/aquasec/trivy-java-db:1 docker tag gcr.io/kaniko-project/executor:v1.24.0 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0 docker tag aquasec/trivy:0.74.0 127.0.0.1:5000/mirror/trivy:0.74.0 docker tag mirror.gcr.io/aquasec/trivy-java-db:1 127.0.0.1:5000/mirror/trivy-java-db:1 docker push 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0 docker push 127.0.0.1:5000/mirror/trivy:0.74.0 docker push 127.0.0.1:5000/mirror/trivy-java-db:1 ``` From another machine, port-forward the registry instead (`kubectl -n felis port-forward svc/registry 5000:5000`) and push to `localhost:5000/...` — the registry keys a repository by the path after the host, so pushes through either door land in the same place the build Pods will pull from. Put them in **both** `/etc/felis/felis.host.toml` (host-side CLI) and `/etc/felis/felis.pod.toml` (the file rendered into the API's `felis-config` Secret — the two differ only in the database URL; the setup screens re-render the Secret from the pod file, so edits made only through `kubectl` on the live Secret are lost at the next reconfigure). A Deployment restart alone is NOT enough — the API Pod mounts the Secret, never the host file. Re-render the Secret from the pod file, then roll `felis-api`: ```sh kubectl -n felis create secret generic felis-config \ --from-file=felis.toml=/etc/felis/felis.pod.toml --dry-run=client -o yaml | kubectl apply -f - kubectl -n felis rollout restart deployment/felis-api ``` Unset fields keep the defaults. `trivy_db_repository` is not optional on an egress-locked box. Trivy fetches its vulnerability DB from `mirror.gcr.io`/`ghcr.io` unless told otherwise, and the build egress policy denies those hosts — so the scan step fails closed (`failed to download vulnerability DB`) and NO build ever completes, even though Kaniko pushed the image. Mirror the DB into the internal registry once: ``` # On the node (docker treats 127.0.0.1 as insecure by default), or through the # port-forward above: # docker pull mirror.gcr.io/aquasec/trivy-db:2 # docker tag mirror.gcr.io/aquasec/trivy-db:2 127.0.0.1:5000/mirror/trivy-db:2 # docker push 127.0.0.1:5000/mirror/trivy-db:2 ``` The Job's Trivy container already runs with `--insecure`, so the internal registry's plain HTTP works for the DB pull exactly as it does for the scanned image. Re-mirror the tag periodically (Trivy refreshes the DB several times a day upstream; a stale mirror only means stale CVE data, never a failed gate). `trivy_java_db_repository` is the same story one step lazier: Trivy downloads the Java DB on demand the first time it scans an image containing Java artifacts — every real modpack — and that download fails closed too. Mirror `mirror.gcr.io/aquasec/trivy-java-db:1` alongside the vulnerability DB (commands above); the Java DB refreshes far less often than the vulnerability DB, so a one-off mirror is usually fine. --- ## 9. Registry push/pull failures (spec §15) The in-cluster registry is Deployment/Service/PVC named `registry` in the control namespace (or `--registry-namespace`): - **Push/pull target (the exact string to match):** `registry..svc:` (port from `--registry-port`, default `5000`; e.g. `--felis-image registry.felis.svc:5000/felis:v1`). A *wrong* push URL is caught at build time by the validate gate (§8a, `400`). An *unreachable* registry at runtime (wrong DNS/port, PVC unbound, missing default StorageClass) surfaces as kaniko push or kubelet pull errors — [INTEGRATION-ONLY], **not** a Felis-emitted string. - **Storage:** PVC is RWO, `10Gi`, mounted at `/var/lib/registry`, **no `storageClassName`** → binds the cluster default class. If the cluster has no default StorageClass the PVC stays `Pending` and the registry never starts. - **Node-side pulls:** containerd cannot dial the Service VIP (the live stack answered "Empty reply"), so the registry Deployment binds a loopback hostPort (`127.0.0.1:`) and the installer writes a `/etc/rancher/k3s/registries.yaml` mirror relaying `registry..svc:` onto it. That pair is what lets kubelet re-pull a garbage-collected image; both halves must survive together (remove either and every pull after an image GC fails). - **Memory:** the registry's limit is 2Gi, deliberately larger than the other control-plane pods' 256Mi — a live 475MB-layer push OOM-killed the 256Mi template mid-upload (audit #46). Very large layers need headroom here, not more CPU. - **Selector quirk worth knowing:** the registry Service selector is only `name + component=registry` — it deliberately lacks the `part-of=felis-control-plane` label, so the registry is *invisible* to the RCON-peer NetworkPolicy selector. This is intended isolation, not a bug; do not "fix" it by adding the label. Registry manifest rendering is [GO-TESTED]; actual serving is [CODE-ONLY/INTEGRATION-ONLY]. --- ## 10. World reaper: false-deletes and skipped backups (spec §18) The reaper is a **run-once daily CronJob batch**, not an operator controller. It reaps a world only when `now - last_active_at > 15d` (`inactive_15d`); the 15-day deadline is **hard-fixed in code** (only `warn_before` / `retention` / `max_local_bytes` are configurable from `felis.toml [archive]`). ### What a "backup" contains A backup tars the server's ENTIRE data volume — the same volume the server mounts at `/data`: world folders, `server.properties`, plugins/mods, configs, jars, libraries, logs and cache, not just the `world/` directory. A restore replaces the volume's contents with the archive (files added since the backup are pruned), so a restore also rolls config/plugin changes back. Sizes are dominated by libraries/cache on stock Paper servers (~170MB for a fresh instance before any world growth) — do not size the archive PVC as if only world data were stored. ### The backup-before-delete invariant The reap sequence (all [GO-TESTED] hermetically) preserves the world unless a **confirmed, DB-recorded backup exists**: 1. `ensureCapacity` (only if `max_local_bytes > 0`) → store full ⇒ world **preserved** (not deleted). 2. `Archiver.Archive` fails ⇒ world **preserved**, PVC untouched. 3. `InsertBackup` (DB) fails ⇒ the orphan archive is deleted, PVC **untouched**. 4. **Only then** `DeletePVC` → `ReleaseWorld` → `Stop` (cosmetic) → audit → `felis_reaper_worlds_deleted_total++`. So a missing backup never results in a deleted world. [GO-TESTED: `TestReapArchiveFailurePreservesWorld`, `TestReapInsertBackupFailurePreservesWorld`, `TestReapIdleWorldFullSequence`.] ### Exemptions (world never reaped) - `spec.reaperExempt=true` → skipped entirely (system servers). [GO-TESTED `TestReapExemptServerNeverTouched`.] - CRD missing → logs `reaper: CRD missing, skipping`, skipped. - Idle `≤ 15d` → not yet eligible. ### Pre-reap warnings (the `warn_before` offsets) An OWNED server inside a warning window gets an email notice (`3d`/`1d` before the deadline, `warn_before` from `[archive]`) to the owner's **verified** email — the same `[smtp]` relay felis-api uses. The `warned_3d_at` / `warned_1d_at` stamps record a **delivered** notice: - No `[smtp]` configured (or owner has no verified address): the run logs `reaper: warning suppressed — no warner wired` / a delivery error and does NOT stamp. Nothing is falsely recorded as sent, and the day SMTP is configured the pending warning can still go out. - Delivery failure (relay down): logged and retried on the next daily run — bounded by the warning window, since the reap removes the candidate anyway. - `warned=` in the run output counts DELIVERED notices, not attempts. The reaper runs in the minecraft namespace and reads the **mirrors** of `felis-smtp` and `felis-config` there (a `secretKeyRef` is namespace-local). The installed `felis setup`'s "configure email" screen refreshes both mirrors when it applies, so configuring SMTP after install is enough; a manual edit of the control-namespace Secret alone is not. [GO-TESTED: the delivered/retried/ suppressed matrix in `internal/reaper`; live-drilled end to end against a local SMTP sink.] ### Genuine false-delete risk vectors - **Stale `last_active_at`.** The keep-alive is `RecordJoin`, called from the internal `join-event` handler. **If join events are not delivered to the API, the activity clock never resets** and an actively-played world becomes reap-eligible after 15 days. Verify join events are flowing (§3a) — this is the most important reaper check. [INTEGRATION-ONLY for the live Postgres write.] - **Unowned servers are still reaped.** A server with `owner_id=""` gets **no pre-deletion warning** (`maybeWarn` skips unowned), but is still reaped at 15d. [GO-TESTED `TestReapUnownedServerStillReaped`.] Claim or exempt servers you want to keep. - `DeletePVC` is idempotent (missing PVC is not an error), so a re-run will not fail on already-reaped worlds; and `Stop` failure is only logged, so a reaped world's `MinecraftServer` may not be flipped to `Stopped`. Only `TarLocal` (tar+gzip) archiving is implemented; VolumeSnapshot/Longhorn backends return `not implemented in this build`. The live PVC delete / Postgres store paths are [INTEGRATION-ONLY]. ### Where worlds are read from (hostPath resolution) The CronJob mounts `--worlds-host-path` read-only at `/worlds`; the resolver runs `cmd/felis/reaper.resolveWorldDir`: it looks for `/`, then for the stock local-path directory `/__` derived from the live PVC's `spec.volumeName` (never a glob — a leftover directory of a deleted PV must not stand in for the world the PVC currently binds). Pointing the flag at k3s's storage root (`/var/lib/rancher/k3s/storage`) is therefore the supported way to enable retention on a stock install. Two deployment facts the resolver cannot fix: - **Permissions.** The reaper Pod runs as **root** and carries `DAC_OVERRIDE`: worlds are written by the game image's own UID (root for every Paper image we ship), and Paper saves `level.dat` mode-0600, so any fixed non-root identity (the previous uid-1000 convention, and the ACL setup that went with it) could neither walk the tree nor read the files — every archive failed `open …/level.dat: permission denied` and the same defect failed on-demand backups/restores. Root is the same identity the game container itself runs as (see the operator's forwarding-init note); `DAC_OVERRIDE` extends the archive to game images with a different UID. If a world is still **preserved** while a reap was expected, it is now a different cause: check the run's ERROR logs for the resolver's `lstat` messages before suspecting permissions. - **Node placement.** Multi-node clusters: the world's directory exists only on the node holding its volume, and the CronJob sets no `nodeSelector`, so add one (single-node starters are pinned implicitly). --- ## 11. Idle auto-stop never fires; player count always shows 0 Both are implemented, and both hang off the same switch: **`spec.rcon.enabled`**. Check it first. ```sh kubectl get minecraftserver -o jsonpath='{.spec.rcon.enabled}' ``` The player tally is a by-product of the RCON readiness probe — `prober.go:63` runs `list` on the same connection that just authenticated, and `parseListReply` extracts the tally from `There are (\d+) of a max of (\d+) players online`. With RCON disabled the probe never runs, `players` keeps its zero value, and `markRunningReady` (`reconciler.go:413`) writes that zero into `status.players.online`. So a permanent 0 means "never sampled", not "nobody online". Idle auto-stop (`reconciler.go:175`) reads that same tally, which is why it carries the RCON condition explicitly: ```go if server.Spec.Rcon.Enabled && server.Spec.Idle.AutoStopEnabled && server.Spec.Idle.EmptySecondsBeforeStop > 0 { ``` The comment above it says why: with RCON off the zero tally "would read as 'empty' and use to stop a server full of people". So the guard is deliberate — enabling `spec.idle.*` without RCON is a no-op by design, not a missing feature. Both fields set and still nothing happens? Then the probe is failing rather than disabled: the server would be stuck in `Starting` with `RconNotReachable` (`reconciler.go:156`), which is §1's symptom, not this one. This path used to fail even with everything configured correctly, through three stacked defects proven and fixed on a live cluster (auditfix21/22): the `emptySince` stamp was pruned by a missing CRD status field, a quiescent empty server produced no watch events to re-check the timer, and the Role lacked the `minecraftservers:patch` grant the stop write needs. If auto-stop ever looks dead again, check these three in order (each is now pinned by a test): ```sh # ① The stamp must persist — should print a timestamp, not an empty string, # a few seconds after a server goes Ready with zero players. kubectl get minecraftserver -o jsonpath='{.status.emptySince}' # ② The operator must be able to write spec.desiredState (403 in the operator # log = missing patch grant on Role felis-operator). kubectl auth can-i patch minecraftservers -n --as=system:serviceaccount::felis-operator # ③ A wake-up must be scheduled: while empty, expect whatever you set # as emptySecondsBeforeStop to elapse and the box to flip to Stopped without # any external action. ``` While players are online the operator re-probes on a 30s cadence so it notices the moment the last one leaves; while empty it schedules a wake-up exactly at the deadline. Quiet operator logs on an idle server are normal — the action is the scheduled wake-up, not a stream of reconciles. Note the reaper's `last_active_at` (§10) is a *different* subsystem (Postgres business layer, bumped by join events) — it keeps worlds alive against the reaper, but it does **not** auto-stop empty running servers. --- ## 12. A configuration field seems to be ignored Every field below is read by a controller. What varies is the condition that decides whether setting it does anything. | Field | What you might expect | Reality | |---|---|---| | `spec.startup.timeoutSeconds` | Start budget before `Failed` | Read by `startupTimedOut` (`reconciler.go:479`), called at `:126`. `0` or unset falls back to **300s**, then `markFailed("StartupTimeout")` | | `spec.startup.readinessTimeoutSeconds` | First-probe budget | Read by `readinessTimedOut` (`reconciler.go:490`), called at `:157`. `0` or unset falls back to **300s**, then `markFailed("ReadinessTimeout")`. Not to be confused with the prober's own 5s dial timeout (`prober.go:45`) | | `spec.idle.autoStopEnabled` | Auto-stop empty servers | Read at `reconciler.go:175` — but gated on `spec.rcon.enabled`, since the player tally comes from the RCON probe (§11) | | `spec.idle.emptySecondsBeforeStop` | Empty grace period | Same branch. Must be `> 0`; the guard treats `0` as "off", not "stop immediately" | Both startup budgets are measured from the same `status.startRequestedAt`, so `readinessTimeoutSeconds` is not a budget *after* pod readiness — it is a deadline for the whole start, applied on the RCON-probe branch. --- ## 13. World PVC survives after I deleted the MinecraftServer This is expected. The world PVC is a StatefulSet `VolumeClaimTemplate`. There is **no `persistentVolumeClaimRetentionPolicy` and no finalizer** anywhere in the operator. Deleting the `MinecraftServer` garbage-collects the StatefulSet, but StatefulSet deletion does **not** cascade to its template PVCs, and nothing else cleans them up. So the world PVC **always survives** server deletion. The **only** code that deletes a world PVC is the reaper, and only after a verified backup (§10). To reclaim a world PVC manually: ``` kubectl get pvc -l app.kubernetes.io/name= kubectl delete pvc # irreversible — the world is gone ``` `spec.storage.retainOnDelete` sat in the CRD and reached no controller. Spec v4.1 §5 asks for it — 「删除:finalizer 清 Service/STS/ConfigMap,PVC 按 `retainOnDelete`」 — and neither half was ever built: there is no finalizer, and nothing read the field. It was removed rather than implemented, which is a deliberate departure from that line, recorded here because the spec is a frozen document and still says otherwise. The reasoning is that implementing it buys a second path that deletes a world — one that skips the reaper's verified-backup check — in order to restore a finalizer whose other listed duties (Service, StatefulSet, ConfigMap) ownerReference GC already performs. A CR still carrying the field keeps working: the API server prunes the unknown key on its next write, and nothing above changes, because retention was never conditional in the first place. --- ## 13b. Node runs out of disk: what survives, and how to recover A full disk is the most destructive failure this stack sees: kubelet evicts game pods (the control plane is protected below), and its image GC then collects images nothing is running. The images have a pull source now — the in-cluster registry — so they come back without an operator re-import; freeing space is what completes the recovery. **Eviction.** Every control-plane pod (api, operator, reaper, registry) runs under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's node-pressure eviction refuses to touch those pods — the log shows *"Eviction manager: cannot evict a critical pod"* for each of them — while game-server pods at the default priority 0 are evicted first. A drill that filled the disk to 1.7G free saw exactly this: login/lobby evicted, the whole control plane still Running (before the fix the same drill evicted the api, operator and registry too, and the image-GC stage below followed). User-defined PriorityClasses cannot substitute: the API caps them at 1e9, below kubelet's critical threshold. The built-in class allows preemption (its policy is fixed), so a control-plane pod that cannot fit may preempt a game pod — deliberate: the management plane must be placeable. **The pressure condition clears slowly.** After you free space, the node can stay `DiskPressure:True` for up to ~5 minutes (`--eviction-pressure-transition-period` defaults to 5m, to stop the condition flapping); pods that need scheduling wait for it. This is the bulk of the "recovery takes minutes" observation, not a stuck node. **The images may be gone — they come back on their own.** If pods were evicted, the kubelet can garbage-collect their images (unused > 2 minutes under imagefs pressure). Every image this platform runs is ALSO hosted in the in-cluster registry: the installer builds each one as `registry..svc:5000/felis/…` and mirrors it there, and the node's containerd is configured (a registries.yaml mirror onto the registry's loopback hostPort) to relay those refs back through it. So a GC'd image is re-pulled on the next attempt with no operator action — delete the stuck pod to force an immediate retry (or wait out the backoff), and the workload converges. If a pull does NOT come back: 1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are `k3s ctr images ls -q` and the world/backup PVCs under `/var/lib/rancher/k3s/storage`). 2. Check the registry: `kubectl -n felis get pods -l app.kubernetes.io/component=registry` and, on the node, `curl -s http://127.0.0.1:5000/v2/` (expect `{}`). 3. Check the mirror file: `/etc/rancher/k3s/registries.yaml` must map `registry.felis.svc:5000` to `http://127.0.0.1:5000`. Missing or changed: re-run the installer (it rewrites the file and restarts k3s only when the content changed). 4. Re-mirror a tag the registry does not have (hand-built images were never pushed): `sudo systemctl start docker` (the installer leaves the daemon stopped), then `docker tag 127.0.0.1:5000/: && docker push 127.0.0.1:5000/:`. For an image that is in neither place, the old fallback still stands: re-run the installer (it rebuilds/re-imports from the local Docker store AND mirrors into the registry), or for a single image `docker save felis: | k3s ctr images import -`. The Docker store remains a deliberate second copy on the node; treat it as the recovery path, not as free space. --- ## 14. Metrics for diagnosis (spec §23) All four mandated metrics have real producers; scrape them when triaging: - `felis_servers_total` — managed server count. - `felis_start_duration_seconds` — histogram, observed once per start when readiness is first reached (`ReadySignalAt − StartRequestedAt`). A start that never completes (§1) contributes **nothing** here — absence of observations is itself the signal that starts are hanging. - `felis_image_build_failures_total` — increments on build Job failure (§8d). - `felis_reaper_worlds_deleted_total` — increments only after a world PVC is actually deleted post-backup (§10); a spike here means worlds crossed the 15d idle line — cross-check that join events are flowing (§10 risk vectors). ### Scraping The series come from two processes: - `felis-operator` pod `:8080/metrics` — `felis_servers_total`, `felis_start_duration_seconds` (no Service; scrape pod-scoped, e.g. a PodMonitor targeting port `metrics`). - `felis-api` internal face `:8081/metrics` (Service `felis-api-internal`) — `felis_image_build_failures_total`. Unauthenticated like the probes; ClusterIP-only, and the external face never serves it. - `felis_reaper_worlds_deleted_total` is produced inside the one-shot reaper CronJob, which exits long before any scrape interval — without a pushgateway it has no scrape path. Read the reaper Pod log or the `world_backups` table for deletions instead. ### Alert rules `deploy/alerts/` ships ready-made rules: build failures, slow starts, node disk/memory thresholds, and the kubelet `DiskPressure` condition. - Plain Prometheus: add `felis-alerts.yaml` to `rule_files`. Check and unit-test it standalone with `promtool check rules felis-alerts.yaml` and `promtool test rules felis-alerts_test.yml` (the tests pin exactly when each alert fires). - kube-prometheus-stack / prometheus-operator: `kubectl apply -f felis-prometheusrule.yaml` (adjust its `release:` label to your stack's ruleSelector). --- ## 15. Control-plane upgrades, and rolling back a bad one There is no in-place updater: an upgrade is re-running the installer (`curl -fsSL | sudo bash`), which rebuilds/re-imports the image and re-applies the bundle. (`sudo felis setup` is not this path; on a completed install it only opens the config console.) The channel is not persisted across the re-run, so pass `FELIS_VERSION_BOOTSTRAP=dev` on a host that tracks main. Two properties of the control plane matter when you do: - Both Deployments use strategy **Recreate** (single replica, no leader election: two overlapping instances would fight over the same cluster). An upgrade takes the panel/API down for the rollout window — seconds normally, longer if the new image still has to be imported. - If the new pod cannot start (bad tag, missing image), the installer's rollout wait fails after 180s and prints `kubectl describe` diagnostics: you see `ErrImagePull`/`ImagePullBackOff` there instead of a silent hang. Roll back with: ``` kubectl -n felis rollout undo deploy/felis-api kubectl -n felis rollout status deploy/felis-api ``` (the same for `felis-operator` and `registry`). `rollout undo` returns to the previous ReplicaSet, whose image is normally still on the node; if the image GC collected it, the registry re-serves it automatically (§13b) for every tag the installer built — only hand-built tags need a manual re-mirror. ## Quick reference: symptom → section | Symptom | Section | |---|---| | Stuck `Starting`, never `Running` | §1 | | `Starting` with `PodNotReady` (image? PVC? boot?) | §1a | | RCON secret/auth/port errors | §1b, §1c | | Phase `Failed` | §2 | | Players land in lobby / wrong place | §3, §4 | | Wake refused / rate-limited (403/429/503) | §3a | | Routing disabled, offline-mode | §4 | | Panel 401/403; fails-closed; audience error | §5 | | Local password login rejected | §5c | | Internal callers 401 (service token) | §6 | | Link/claim 400/409/412/403/404 | §7 | | Build push 400 / SA denied / egress hang / Failed / executor ImagePullBackOff | §8, §8e | | Registry push/pull unreachable | §9 | | World deleted unexpectedly / backup skipped | §10 | | Idle auto-stop not firing; player count 0 | §11 | | A config field seems ignored | §12 | | PVC left behind after delete | §13 | | Node out of disk; pods evicted / ImagePullBackOff | §13b | | Which metric to scrape | §14 | | Upgrade / roll back a bad control-plane image | §15 |