# Felis Troubleshooting Checklist This file is the spec §28 #23 deliverable: a故障排查清单 (troubleshooting checklist) for the failure modes the platform actually produces. Every symptom below is traced to a concrete signal — a `status.conditions` reason, an HTTP error code, a log string, or a manifest name — so an operator can map what they see to the code path that emitted it. ## How to read this document Each entry is **symptom → likely cause → where to look → fix**. The `kubectl` commands run as root on the node (`sudo -i`, or `sudo k3s kubectl …`): the admin kubeconfig `/etc/rancher/k3s/k3s.yaml` is readable by root only (§13c). Signals are graded for how far the in-repo Go test suite proves the behaviour: - **[GO-TESTED]** — a hermetic `*_test.go` exercises this exact path; the string/code is asserted in CI. - **[PG-TESTED]** — a `-tags pgint` test in `internal/pgint` drives it against a real Postgres with the shipped migrations (CONTRIBUTING.md has the command). - **[CODE-ONLY]** — the code path and string exist and are real, but no unit test drives them (notably the Velocity Java plugin, which is not compiled or tested in this repo). - **[INTEGRATION-ONLY]** — the symptom is produced by the kubelet, kaniko, containerd, Postgres, or the network, not by Felis Go code; you will see it in `kubectl describe` / pod logs, never in `MinecraftServer.status`. - **[INERT]** — the configuration field exists in the CRD but no controller reads it. Tuning it does nothing. **No CRD field carries this status today**; the last one, `spec.storage.retainOnDelete`, was removed rather than implemented (§13 records why). A field whose change seems ignored is almost always a condition instead — §12 lists the fields that *are* read and the condition each depends on. The operator never invents the parent domain; routing identity is `spec.subdomain` under the deployment zone. Examples below use `` / `registry..svc:` placeholders rather than any concrete host. --- ## 0. First look: `felis status`, `felis doctor`, `felis support-bundle` Three read-only commands, all run as root on the node, give the state of the whole host before any of the sections below. They change nothing and mail nothing, so they are safe to run at any time. **`sudo felis status`** prints the platform at a glance: the release, the node and its kubelet, each control-plane Deployment (ready replicas, image, pod restarts), the game proxy (the `felis-velocity` unit and whether the game port accepts connections), every server with its desired state, phase, players and newest world backup, the newest database bundle and the off-site copy, the watched disks and memory, and the alerts the watchdog has open. A part that is down reads as such and the rest still prints: with k3s stopped the cluster line says `unreachable (...)` and the servers `unknown while the cluster is unreachable`; with PostgreSQL down the backup column reads `?` and a line under the table says why. [GO-TESTED: `TestStatusReport`, `TestStatusWithPartsDown`, `TestStatusClusterEdges`, `TestStatusBackups`, `TestStatusWatchdog`] **`sudo felis doctor`** runs every check `felis watchdog` runs, with the settings `felis-watchdog.service` gives it, plus what only the host shows: a Felis unit that failed or a long-running one (`k3s`, `felis-velocity`) that stopped, a timer that no longer fires, alerts that reach no one (no `[smtp]` relay, no owner with a verified address), an unusable heartbeat URL. It prints one line per area, `✓` fine, `!` warnings, `✗` something critical, `-` not checked and why (the off-site copy on an install without one), each finding with where to look next: ```text felis doctor on felis-1 at 2026-09-27 12:00 UTC checks run as /etc/systemd/system/felis-watchdog.service runs them (config /etc/felis/felis.host.toml) ✓ configuration ✓ Kubernetes cluster ✗ PostgreSQL critical postgres: PostgreSQL is unreachable: sign-in, the panel and server management fail → k3s kubectl -n felis get pods -l app.kubernetes.io/component=postgres; k3s kubectl -n felis logs deploy/felis-postgres --tail=100 (...) - off-site copy: not checked, not configured ... 1 problem(s): 1 critical, 0 warning(s) ``` It exits 1 when it found anything and 0 otherwise, so a script can run it. It never mails, pings the heartbeat or touches the watchdog's state: an alert it shows is mailed by the watchdog's own next run, on the watchdog's delays (§14). A missing or unreadable `felis-watchdog.service` is itself a critical finding, and the checks then run with the watchdog's defaults. [GO-TESTED: `TestDoctorReportsByArea`, `TestDoctorMailsPingsAndSavesNothing`, `TestDoctorWithoutUnitOrConfig`, `TestUnitFindings`, `TestPrintDoctorReport`] **`sudo felis support-bundle`** collects what someone helping needs into one file, `/var/lib/felis/support/felis-support--