docs(troubleshooting): correct four fields the runbook documents as inert
Sections 11 and 12 told the operator that idle auto-stop, both startup budgets, and the player tally are read by nobody. All four are read, and §1 repeated the same claim in its strongest form: "the operator has no start timeout ... loops forever". spec.idle.autoStopEnabled reconciler.go:175 spec.idle.emptySecondsBeforeStop reconciler.go:175 spec.startup.timeoutSeconds reconciler.go:479, called at :126 spec.startup.readinessTimeoutSeconds reconciler.go:490, called at :157 status.players.online reconciler.go:413 (markRunningReady) The repository already contained the disproof: TestReconcileRunning_StartupTimeoutConvertsToFailed and TestReconcileRunning_ReadinessTimeoutConvertsToFailed both assert the escalation §1 said does not exist. The test §1 cited, TestReconcileRunning_RconProbeFailureStaysStarting, only asserts that a single failed probe does not flap the phase; that was read as "forever". §11 was the costly one, because it misdiagnosed a configuration problem as a missing feature. Idle auto-stop and the player tally both hang off spec.rcon.enabled -- the tally is a by-product of the RCON readiness probe (prober.go:63 runs `list`), and reconciler.go:175 carries the RCON condition explicitly so a never-sampled zero cannot stop a server full of people. Following the old text, an operator whose RCON was never enabled would conclude the feature was unwritten and stop. The section now opens with the jsonpath that reads spec.rcon.enabled. Also separated the prober's fixed 5s dial timeout (prober.go:45) from spec.startup.readinessTimeoutSeconds, which §1c conflated: the former bounds one probe, the latter is a deadline for the whole start measured from status.startRequestedAt. spec.storage.retainOnDelete is the one field still genuinely inert, so the [INERT] legend and §12 stay -- §12 now records the condition each read field depends on instead of claiming none of them are read.
This commit is contained in:
1 file changed
+60
-29
+60
-29
@@ -20,8 +20,8 @@ graded for how far the in-repo Go test suite proves the behaviour:
|
||||
containerd, Postgres, or the network, not by Felis Go code; you will see it in
|
||||
`kubectl describe` / pod logs, never in `MinecraftServer.status`.
|
||||
- **[INERT]** — the configuration field exists in the CRD but no controller
|
||||
reads it. Tuning it does nothing. These are the highest-value traps and are
|
||||
collected in §10.
|
||||
reads it. Tuning it does nothing. §12 lists the one field this still applies
|
||||
to, alongside the fields that *are* read and the condition each depends on.
|
||||
|
||||
The operator never invents the parent domain; routing identity is
|
||||
`spec.subdomain` under the deployment zone. Examples below use
|
||||
@@ -32,14 +32,21 @@ concrete host.
|
||||
|
||||
## 1. Server is stuck in `Starting` and never becomes `Running`
|
||||
|
||||
`MinecraftServer.status.phase` stays `Starting`. The single most important fact:
|
||||
**the operator has no start timeout.** A start that never succeeds loops in
|
||||
`Starting`, requeued every 5s, **forever** — it is never auto-escalated to
|
||||
`Failed`. [GO-TESTED: `TestReconcileRunning_RconProbeFailureStaysStarting`
|
||||
asserts the phase stays `Starting`.] Do **not** reach for
|
||||
`spec.startup.timeoutSeconds` / `spec.startup.readinessTimeoutSeconds` — both are
|
||||
[INERT] (§10). A hung start must be diagnosed from pod state, not from
|
||||
`MinecraftServer.status`.
|
||||
`MinecraftServer.status.phase` stays `Starting`. A start that never succeeds is
|
||||
requeued every 5s until one of the two startup budgets expires, then escalated
|
||||
to `Failed` — `StartupTimeout` if the pod never passed TCP readiness,
|
||||
`ReadinessTimeout` if the RCON probe never succeeded. Both default to **300s**
|
||||
when `spec.startup.timeoutSeconds` / `spec.startup.readinessTimeoutSeconds` are
|
||||
unset or `0`, and both are measured from `status.startRequestedAt`.
|
||||
[GO-TESTED: `TestReconcileRunning_StartupTimeoutConvertsToFailed`,
|
||||
`TestReconcileRunning_ReadinessTimeoutConvertsToFailed`.]
|
||||
|
||||
So `Starting` seen *once* is normal and
|
||||
`TestReconcileRunning_RconProbeFailureStaysStarting` asserts exactly that — a
|
||||
single failed probe must not flap the phase. `Starting` seen for longer than the
|
||||
budget means the reconcile loop is not running at all; check the operator's own
|
||||
logs before tuning anything. Either way the underlying cause is diagnosed from
|
||||
pod state, not from `MinecraftServer.status`.
|
||||
|
||||
First, read the condition reason:
|
||||
|
||||
@@ -111,8 +118,10 @@ message is the verbatim dial error:
|
||||
port yet, RCON is disabled in `server.properties`, or `spec.rcon.port`
|
||||
(default 25575) is wrong. [INTEGRATION-ONLY for the live handshake.]
|
||||
|
||||
The per-probe timeout is a fixed 5s in code — it is **not** derived from
|
||||
`spec.startup.readinessTimeoutSeconds` ([INERT]).
|
||||
The per-probe timeout is a fixed 5s in code (`prober.go:45`, shortened further if
|
||||
the reconcile context has a nearer deadline). It is **not** derived from
|
||||
`spec.startup.readinessTimeoutSeconds`, which is the deadline for the whole
|
||||
start, not for one probe — see §12.
|
||||
|
||||
---
|
||||
|
||||
@@ -463,14 +472,35 @@ store paths are [INTEGRATION-ONLY].
|
||||
|
||||
## 11. Idle auto-stop never fires; player count always shows 0
|
||||
|
||||
**Idle auto-stop is entirely unimplemented in the operator.** `spec.idle.*`
|
||||
([INERT], §12) is read by no controller, and no idle controller is registered.
|
||||
An empty `Running` server stays `Running`.
|
||||
Both are implemented, and both hang off the same switch: **`spec.rcon.enabled`**.
|
||||
Check it first.
|
||||
|
||||
Relatedly, `status.players.online` is **permanently 0**: the only writer of
|
||||
`status.players` zeroes it on stop, and the RCON prober only dials + closes — it
|
||||
never runs `list`. Any panel reading `status.players.online` will always show
|
||||
empty. Do not build alerting on it, and do not expect "empty server" automation.
|
||||
```sh
|
||||
kubectl get minecraftserver <name> -o jsonpath='{.spec.rcon.enabled}'
|
||||
```
|
||||
|
||||
The player tally is a by-product of the RCON readiness probe — `prober.go:63`
|
||||
runs `list` on the same connection that just authenticated, and `parseListReply`
|
||||
extracts the tally from `There are (\d+) of a max of (\d+) players online`. With
|
||||
RCON disabled the probe never runs, `players` keeps its zero value, and
|
||||
`markRunningReady` (`reconciler.go:413`) writes that zero into
|
||||
`status.players.online`. So a permanent 0 means "never sampled", not "nobody
|
||||
online".
|
||||
|
||||
Idle auto-stop (`reconciler.go:175`) reads that same tally, which is why it
|
||||
carries the RCON condition explicitly:
|
||||
|
||||
```go
|
||||
if server.Spec.Rcon.Enabled && server.Spec.Idle.AutoStopEnabled && server.Spec.Idle.EmptySecondsBeforeStop > 0 {
|
||||
```
|
||||
|
||||
The comment above it says why: with RCON off the zero tally "would read as
|
||||
'empty' and use to stop a server full of people". So the guard is deliberate —
|
||||
enabling `spec.idle.*` without RCON is a no-op by design, not a missing feature.
|
||||
|
||||
Both fields set and still nothing happens? Then the probe is failing rather than
|
||||
disabled: the server would be stuck in `Starting` with `RconNotReachable`
|
||||
(`reconciler.go:156`), which is §1's symptom, not this one.
|
||||
|
||||
Note the reaper's `last_active_at` (§10) is a *different* subsystem (Postgres
|
||||
business layer, bumped by join events) — it keeps worlds alive against the
|
||||
@@ -478,22 +508,23 @@ reaper, but it does **not** auto-stop empty running servers.
|
||||
|
||||
---
|
||||
|
||||
## 12. Inert configuration fields (highest-value traps)
|
||||
## 12. A configuration field seems to be ignored
|
||||
|
||||
These CRD fields exist and validate, but **no controller reads them.** Setting
|
||||
them has **no effect**. Verified by grep: each appears only in the type
|
||||
definition and its deepcopy, never in a controller.
|
||||
One CRD field exists and validates but is read by no controller; the rest of
|
||||
this table records fields that *are* read, together with the condition that
|
||||
decides whether setting them does anything.
|
||||
|
||||
| Field | What you might expect | Reality |
|
||||
|---|---|---|
|
||||
| `spec.startup.timeoutSeconds` | Start budget before `Failed` | **[INERT]** — no Starting→Failed timeout exists; a hung start loops forever (§1) |
|
||||
| `spec.startup.readinessTimeoutSeconds` | First-probe budget | **[INERT]** — probe timeout is a fixed 5s in code |
|
||||
| `spec.idle.autoStopEnabled` | Auto-stop empty servers | **[INERT]** — idle auto-stop unimplemented (§11) |
|
||||
| `spec.idle.emptySecondsBeforeStop` | Empty grace period | **[INERT]** |
|
||||
| `spec.startup.timeoutSeconds` | Start budget before `Failed` | Read by `startupTimedOut` (`reconciler.go:479`), called at `:126`. `0` or unset falls back to **300s**, then `markFailed("StartupTimeout")` |
|
||||
| `spec.startup.readinessTimeoutSeconds` | First-probe budget | Read by `readinessTimedOut` (`reconciler.go:490`), called at `:157`. `0` or unset falls back to **300s**, then `markFailed("ReadinessTimeout")`. Not to be confused with the prober's own 5s dial timeout (`prober.go:45`) |
|
||||
| `spec.idle.autoStopEnabled` | Auto-stop empty servers | Read at `reconciler.go:175` — but gated on `spec.rcon.enabled`, since the player tally comes from the RCON probe (§11) |
|
||||
| `spec.idle.emptySecondsBeforeStop` | Empty grace period | Same branch. Must be `> 0`; the guard treats `0` as "off", not "stop immediately" |
|
||||
| `spec.storage.retainOnDelete` | Keep/drop PVC on delete | **[INERT]** — world PVCs **always** survive server deletion; only the reaper ever deletes a world PVC (§13) |
|
||||
|
||||
If a runbook tells someone to "tune `startup.timeoutSeconds`" for a hung start,
|
||||
it is wrong — use `kubectl describe pod` (§1a) instead.
|
||||
Both startup budgets are measured from the same `status.startRequestedAt`, so
|
||||
`readinessTimeoutSeconds` is not a budget *after* pod readiness — it is a
|
||||
deadline for the whole start, applied on the RCON-probe branch.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in new issue
Block a user