chore: work the tracker items that need no cluster #19

Merged
FLYEMOJ1 merged 10 commits from chore/issue-sweep into main 2026-07-29 01:52:15 +09:00
Showing only changes of commit 3af5cc360c - Show all commits
+60 -29
View File
@@ -20,8 +20,8 @@ graded for how far the in-repo Go test suite proves the behaviour:
containerd, Postgres, or the network, not by Felis Go code; you will see it in
`kubectl describe` / pod logs, never in `MinecraftServer.status`.
- **[INERT]** — the configuration field exists in the CRD but no controller
reads it. Tuning it does nothing. These are the highest-value traps and are
collected in §10.
reads it. Tuning it does nothing. §12 lists the one field this still applies
to, alongside the fields that *are* read and the condition each depends on.
The operator never invents the parent domain; routing identity is
`spec.subdomain` under the deployment zone. Examples below use
@@ -32,14 +32,21 @@ concrete host.
## 1. Server is stuck in `Starting` and never becomes `Running`
`MinecraftServer.status.phase` stays `Starting`. The single most important fact:
**the operator has no start timeout.** A start that never succeeds loops in
`Starting`, requeued every 5s, **forever** — it is never auto-escalated to
`Failed`. [GO-TESTED: `TestReconcileRunning_RconProbeFailureStaysStarting`
asserts the phase stays `Starting`.] Do **not** reach for
`spec.startup.timeoutSeconds` / `spec.startup.readinessTimeoutSeconds` — both are
[INERT] (§10). A hung start must be diagnosed from pod state, not from
`MinecraftServer.status`.
`MinecraftServer.status.phase` stays `Starting`. A start that never succeeds is
requeued every 5s until one of the two startup budgets expires, then escalated
to `Failed` — `StartupTimeout` if the pod never passed TCP readiness,
`ReadinessTimeout` if the RCON probe never succeeded. Both default to **300s**
when `spec.startup.timeoutSeconds` / `spec.startup.readinessTimeoutSeconds` are
unset or `0`, and both are measured from `status.startRequestedAt`.
[GO-TESTED: `TestReconcileRunning_StartupTimeoutConvertsToFailed`,
`TestReconcileRunning_ReadinessTimeoutConvertsToFailed`.]
So `Starting` seen *once* is normal and
`TestReconcileRunning_RconProbeFailureStaysStarting` asserts exactly that — a
single failed probe must not flap the phase. `Starting` seen for longer than the
budget means the reconcile loop is not running at all; check the operator's own
logs before tuning anything. Either way the underlying cause is diagnosed from
pod state, not from `MinecraftServer.status`.
First, read the condition reason:
@@ -111,8 +118,10 @@ message is the verbatim dial error:
port yet, RCON is disabled in `server.properties`, or `spec.rcon.port`
(default 25575) is wrong. [INTEGRATION-ONLY for the live handshake.]
The per-probe timeout is a fixed 5s in code — it is **not** derived from
`spec.startup.readinessTimeoutSeconds` ([INERT]).
The per-probe timeout is a fixed 5s in code (`prober.go:45`, shortened further if
the reconcile context has a nearer deadline). It is **not** derived from
`spec.startup.readinessTimeoutSeconds`, which is the deadline for the whole
start, not for one probe — see §12.
---
@@ -463,14 +472,35 @@ store paths are [INTEGRATION-ONLY].
## 11. Idle auto-stop never fires; player count always shows 0
**Idle auto-stop is entirely unimplemented in the operator.** `spec.idle.*`
([INERT], §12) is read by no controller, and no idle controller is registered.
An empty `Running` server stays `Running`.
Both are implemented, and both hang off the same switch: **`spec.rcon.enabled`**.
Check it first.
Relatedly, `status.players.online` is **permanently 0**: the only writer of
`status.players` zeroes it on stop, and the RCON prober only dials + closes — it
never runs `list`. Any panel reading `status.players.online` will always show
empty. Do not build alerting on it, and do not expect "empty server" automation.
```sh
kubectl get minecraftserver <name> -o jsonpath='{.spec.rcon.enabled}'
```
The player tally is a by-product of the RCON readiness probe — `prober.go:63`
runs `list` on the same connection that just authenticated, and `parseListReply`
extracts the tally from `There are (\d+) of a max of (\d+) players online`. With
RCON disabled the probe never runs, `players` keeps its zero value, and
`markRunningReady` (`reconciler.go:413`) writes that zero into
`status.players.online`. So a permanent 0 means "never sampled", not "nobody
online".
Idle auto-stop (`reconciler.go:175`) reads that same tally, which is why it
carries the RCON condition explicitly:
```go
if server.Spec.Rcon.Enabled && server.Spec.Idle.AutoStopEnabled && server.Spec.Idle.EmptySecondsBeforeStop > 0 {
```
The comment above it says why: with RCON off the zero tally "would read as
'empty' and use to stop a server full of people". So the guard is deliberate —
enabling `spec.idle.*` without RCON is a no-op by design, not a missing feature.
Both fields set and still nothing happens? Then the probe is failing rather than
disabled: the server would be stuck in `Starting` with `RconNotReachable`
(`reconciler.go:156`), which is §1's symptom, not this one.
Note the reaper's `last_active_at` (§10) is a *different* subsystem (Postgres
business layer, bumped by join events) — it keeps worlds alive against the
@@ -478,22 +508,23 @@ reaper, but it does **not** auto-stop empty running servers.
---
## 12. Inert configuration fields (highest-value traps)
## 12. A configuration field seems to be ignored
These CRD fields exist and validate, but **no controller reads them.** Setting
them has **no effect**. Verified by grep: each appears only in the type
definition and its deepcopy, never in a controller.
One CRD field exists and validates but is read by no controller; the rest of
this table records fields that *are* read, together with the condition that
decides whether setting them does anything.
| Field | What you might expect | Reality |
|---|---|---|
| `spec.startup.timeoutSeconds` | Start budget before `Failed` | **[INERT]** — no Starting→Failed timeout exists; a hung start loops forever (§1) |
| `spec.startup.readinessTimeoutSeconds` | First-probe budget | **[INERT]** — probe timeout is a fixed 5s in code |
| `spec.idle.autoStopEnabled` | Auto-stop empty servers | **[INERT]** — idle auto-stop unimplemented (§11) |
| `spec.idle.emptySecondsBeforeStop` | Empty grace period | **[INERT]** |
| `spec.startup.timeoutSeconds` | Start budget before `Failed` | Read by `startupTimedOut` (`reconciler.go:479`), called at `:126`. `0` or unset falls back to **300s**, then `markFailed("StartupTimeout")` |
| `spec.startup.readinessTimeoutSeconds` | First-probe budget | Read by `readinessTimedOut` (`reconciler.go:490`), called at `:157`. `0` or unset falls back to **300s**, then `markFailed("ReadinessTimeout")`. Not to be confused with the prober's own 5s dial timeout (`prober.go:45`) |
| `spec.idle.autoStopEnabled` | Auto-stop empty servers | Read at `reconciler.go:175` — but gated on `spec.rcon.enabled`, since the player tally comes from the RCON probe (§11) |
| `spec.idle.emptySecondsBeforeStop` | Empty grace period | Same branch. Must be `> 0`; the guard treats `0` as "off", not "stop immediately" |
| `spec.storage.retainOnDelete` | Keep/drop PVC on delete | **[INERT]** — world PVCs **always** survive server deletion; only the reaper ever deletes a world PVC (§13) |
If a runbook tells someone to "tune `startup.timeoutSeconds`" for a hung start,
it is wrong — use `kubectl describe pod` (§1a) instead.
Both startup budgets are measured from the same `status.startRequestedAt`, so
`readinessTimeoutSeconds` is not a budget *after* pod readiness — it is a
deadline for the whole start, applied on the RCON-probe branch.
---