Commit Graph
16 Commits
Author SHA1 Message Date
Lemon-miaow 72553cb414 docs(troubleshooting): the installer leaves docker stopped — start it before manual pushes
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
2026-09-23 19:35:13 +08:00
Lemon-miaow fa0e8d7d97 docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
2026-09-23 19:03:11 +08:00
Lemon-miaow 43df08b52a feat(alerts): ship Felis alert rules with promtool unit tests; document scraping & rules (troubleshooting §14) 2026-09-23 16:17:51 +08:00
Lemon-miaow ae6e9256c6 docs(troubleshooting): 8e — apply build-image overrides through the config Secret (restart alone does not) 2026-09-23 15:56:44 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00
Lemon-miaow 089d4f3a80 docs(audit): idle auto-stop postmortem — ledger #25 and self-checks in §11
Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
2026-09-23 04:14:11 +08:00
Lemon-miaow 8e7c7bbf24 fix(reaper): deliver pre-reap warnings for real — and never fake a delivery
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:

- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
  `warning suppressed — no warner wired` and does NOT stamp; a delivery error
  logs and retries on the next daily run (bounded by the warning window). The
  stamps are no longer burned by notices nobody received.

- A real channel: mail.SendNotice (the second and last message shape the mail
  package sends) plus a mailWarner that resolves the owner's VERIFIED email
  and mails the notice through the configured [smtp] relay. `felis reaper`
  wires it when [smtp] is set (same password_ref convention as felis-api) and
  prints exactly what happens when it is not.

Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.

Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
2026-09-23 03:47:19 +08:00
Lemon-miaow 02fd2de502 fix(build): three drill-driven fixes so the lane actually completes on a starter node
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:

- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
  Pod never scheduled on the platform's own starter node: FailedScheduling /
  Insufficient memory, Pending forever. Requests are now a small floor
  (250m / 512Mi, never above a configured cap) while the limits stay the
  safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
  the source owner; a 65532-owned context (the distroless felis image uid)
  fails that under the pod's dropped capabilities ('copying dockerfile:
  chown /kaniko/Dockerfile: operation not permitted'). The fetch container
  now extracts as root — the uid Kaniko already runs as — so the copy
  succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
  step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
  renders --db-repository, and docs/troubleshooting.md §8e now carries the
  verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
  internal registry; --insecure already covers its plain HTTP).

Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
2026-09-22 23:01:25 +08:00
Lemon-miaow 87a9f4eb25 feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
2026-09-22 22:03:48 +08:00
Lemon-miaow 0a2d654e68 fix(platform): control plane runs system-cluster-critical, so eviction refuses it (#8)
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.

Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
2026-09-22 21:55:54 +08:00
Lemon-miaow fe310743a2 fix(platform): give the control plane a PriorityClass eviction shield (#8)
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
2026-09-22 21:08:56 +08:00
Lemon-miaow 2b87a5a13b fix(install): grant the reaper traverse on the worlds root (#6)
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.

Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
2026-09-22 21:03:17 +08:00
flyemoji 23792d6251 fix(crd): remove spec.storage.retainOnDelete rather than leave it inert
The field validated, shipped in the CRD, and reached no controller. The world
PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate,
StatefulSet deletion does not cascade to template PVCs, and no finalizer exists
anywhere in the operator. So setting it true described what already happened,
and setting it false did nothing at all. False is the worse half: it reads as a
request to delete a world, and was silently ignored.

This departs from spec v4.1 §5, which asks for
"删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half
was ever built. Restoring that line means adding a finalizer whose other listed
duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs,
so the only work it would newly do is delete worlds, on a path that does not
pass the reaper's verified-backup check. The reaper is the one thing in the
system allowed to destroy a world and it earns that by proving a backup first.
A second door without that check is not an improvement.

The spec is a frozen versioned document, so it is left alone and the departure
is recorded in troubleshooting.md §13, beside the behaviour it explains. §12
loses its inert row and its opening sentence, which existed to introduce this
one field: every field in that table is now read by a controller.

Deployed installs need nothing. A CR still carrying retainOnDelete keeps
working, because a v1 CRD prunes unknown keys on the next write and the
behaviour the field claimed to control was never conditional.

go build, go vet and go test ./... pass on Linux with zero failures; the CRD
still parses and storage keeps size and storageClassName.
2026-07-28 17:27:01 +09:00
flyemoji 3af5cc360c docs(troubleshooting): correct four fields the runbook documents as inert
Sections 11 and 12 told the operator that idle auto-stop, both startup
budgets, and the player tally are read by nobody. All four are read, and
§1 repeated the same claim in its strongest form: "the operator has no
start timeout ... loops forever".

  spec.idle.autoStopEnabled       reconciler.go:175
  spec.idle.emptySecondsBeforeStop  reconciler.go:175
  spec.startup.timeoutSeconds     reconciler.go:479, called at :126
  spec.startup.readinessTimeoutSeconds  reconciler.go:490, called at :157
  status.players.online           reconciler.go:413 (markRunningReady)

The repository already contained the disproof:
TestReconcileRunning_StartupTimeoutConvertsToFailed and
TestReconcileRunning_ReadinessTimeoutConvertsToFailed both assert the
escalation §1 said does not exist. The test §1 cited,
TestReconcileRunning_RconProbeFailureStaysStarting, only asserts that a
single failed probe does not flap the phase; that was read as "forever".

§11 was the costly one, because it misdiagnosed a configuration problem
as a missing feature. Idle auto-stop and the player tally both hang off
spec.rcon.enabled -- the tally is a by-product of the RCON readiness
probe (prober.go:63 runs `list`), and reconciler.go:175 carries the RCON
condition explicitly so a never-sampled zero cannot stop a server full of
people. Following the old text, an operator whose RCON was never enabled
would conclude the feature was unwritten and stop. The section now opens
with the jsonpath that reads spec.rcon.enabled.

Also separated the prober's fixed 5s dial timeout (prober.go:45) from
spec.startup.readinessTimeoutSeconds, which §1c conflated: the former
bounds one probe, the latter is a deadline for the whole start measured
from status.startRequestedAt.

spec.storage.retainOnDelete is the one field still genuinely inert, so
the [INERT] legend and §12 stay -- §12 now records the condition each
read field depends on instead of claiming none of them are read.
2026-07-28 13:49:07 +09:00
flyemoji 2ba994889e fix(platform): front the felis-api internal face on its own ClusterIP Service
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.

Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).

Manifest-level fix; the live packet path is pending real-cluster verification.
2026-07-07 11:24:09 +09:00
flyemoji ac02c69612 docs(troubleshooting): add operator failure-mode checklist
Add docs/troubleshooting.md covering the common failure modes the spec
implies, grounded in the actual control-plane code paths:

- Stuck Starting (PodNotReady / RconSecretUnavailable / RconNotReachable)
  and the deliberate absence of a Starting->Failed timeout.
- Failed reachable only via InvalidSpec on a malformed spec.storage.size,
  plus the stale status.endpoint=direct caveat after a failure.
- Routing via status.endpoint direct/fallback and the empty fallbackServer
  pitfall; wake 403/429/503 gate order.
- online-mode coupling and Velocity's offline-mode routing refusal.
- Cloudflare Access 401/403, nil-Keyfunc fail-closed, audience checks,
  the absence of an issuer check, and local-session gating.
- Internal service-token (FELIS_SERVICE_TOKEN) rejection path.
- link/claim error codes, Kaniko build denials (SA-by-absence RBAC,
  default-deny egress, internal-registry push gate), and the registry
  DNS contract.
- Reaper backup-before-delete invariant and false-delete vectors.
- Unimplemented idle auto-stop, permanently-zero players.online, the
  inert CRD fields, and the always-survives world PVC behaviour.

Each item is labelled with its evidence grade (GO-TESTED / CODE-ONLY /
INTEGRATION-ONLY / INERT) so operators know what is verified versus
asserted.
2026-06-30 13:38:37 +09:00