Commit Graph
612 Commits
Author SHA1 Message Date
Lemon-miaow 4d4cdd6ea7 fix(api,panel): refuse file operations on a server without a world volume (#58)
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
2026-09-23 22:17:08 +08:00
Lemon-miaow 70c988e702 fix(panel): render the console command's reply (#57)
sendCommand returns the RCON reply and the component's own contract says it
"displays its plain-text reply", but the reply was discarded and the pod log
does not echo command output, so a sent command produced no visible result at
all (verified live with 'list'). The last command echo + reply now render
above the prompt, terminal-style.
2026-09-23 22:08:19 +08:00
Lemon-miaow 0790f8dfd3 fix(panel): surface the RCON reply for player-access mutations (#56)
The whitelist/ban/kick mutations reported a canned success message and threw
away the server's reply, so a refused command still read as done: live, the
vanilla server answers "That player does not exist" for a name it has never
seen (any player who has not joined yet), while the panel said the player had
been whitelisted/banned. api.ts documents these replies as "surfaced verbatim
as confirmation"; now they are. The localized string stays as the fallback
for a silent server.
2026-09-23 22:08:19 +08:00
Lemon-miaow 9c3a1c5be1 docs(audit): thirty-fifth batch — NetworkPolicy live enforcement matrix green (whitelist closed loop), velocity refresh loop verified 2026-09-23 21:43:23 +08:00
Lemon-miaow cbb0f11288 docs(audit): thirty-fourth batch — felis nano sweep (red 13 -> green 60/60), #55 bad-source ladder stall 2026-09-23 21:30:53 +08:00
Lemon-miaow 9dad61f508 fix(nano): screen unusable profiles per source instead of stopping the ladder (#55)
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.

Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.

Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
2026-09-23 21:21:21 +08:00
Lemon-miaow 971ae01caf docs(audit): thirty-third batch — /updates window API+UI sweep green, #53 repo move, #54 update guidance red/green
Records: (1) the /updates maintenance-window API sweep (unset null, 400s/415 for the invalid set, write->read-back->survives API pod restart, 401/403 auth) and the CDP panel sweep (status transitions, validation copy, clear, audit x3, zero console errors); (2) #53 red/green with the live fix53 binary ('MliroLirrorsIngenuity/Felis' -> 'FelisMC/Felis' in the 404 line) plus the token'd check (old path 301 / new path 200) and the fact the new home has no stable release yet; (3) #54 red/green with the live fix54 binary (installer one-liner + single trailer replace 'run: sudo felis setup'), the setup-vs-installer evidence, and the doc/test synchronization. Reachability table gains #53 (2) and #54 (2, docs); stats 54 total; repro-entry notes updated to the new remote, host binary v0.0.0+fix54 and the batch's artifacts.
2026-09-23 21:04:02 +08:00
Lemon-miaow 397a400d57 fix(update): point the apply guidance at the installer, not felis setup (#54)
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.

Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
2026-09-23 21:01:32 +08:00
Lemon-miaow 2c6739ad76 fix(updater,install,docs): follow the move to FelisMC/Felis (#53)
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.

Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
2026-09-23 20:58:34 +08:00
Lemon-miaow cec9a98305 docs(audit): thirty-second batch — S3 storage wizard sweep, #52 mirror lag red/green, cleanup 2026-09-23 20:39:09 +08:00
Lemon-miaow de7fb2c936 fix(setup): converge the workload felis-config mirror on every apply path (#52)
felis setup's in-TUI applies (storage / connection / edge) refreshed only the
control-namespace felis-config Secret; the workload-namespace mirror kept the
render from the previous run's startup pass until the next setup or installer
run. Found live: after 's -> Local' the minecraft copy still carried
user_uploads_context = s3://felis-wizard-uploads while the control copy and
both tomls were local. The 'configure email' path already overwrote both
mirrors, so storage/connection were the odd ones out.

Move the mirror refresh into applyFelisConfigSecret — the single choke point
every apply path calls — best-effort with a warning, since a control-plane
default install may not have the workload namespace at all. The smtp helper
drops its now-duplicate felis-config block.
2026-09-23 20:31:48 +08:00
Lemon-miaow 01988305a8 docs(audit): thirty-first batch — first-install walkthrough, #51 replica refresh red/green 2026-09-23 20:15:46 +08:00
Lemon-miaow 328e570309 fix(setup,install): refresh the workload namespace's felis-config mirror (#51) 2026-09-23 20:07:20 +08:00
Lemon-miaow cf5a790ea8 docs(audit): thirtieth batch — #50 smtp carry hoard, healed by two live re-runs 2026-09-23 19:57:00 +08:00
Lemon-miaow 4d3c85fd06 fix(bootstrap): [smtp] carry stops hoarding the auth_source comment block (#50) 2026-09-23 19:45:26 +08:00
Lemon-miaow 28fe7c43a9 docs(audit): twenty-ninth batch — image durability live drills; #46–#49
- registry hosting + loopback pull path landed (a9b275a/13d64e0/fa0e8d7) and
  drilled live: three installer re-runs, then GC simulations on the control
  plane (rolled felis-api pulled back in 25ms) and a game image (lobby-0,
  182MB in 10ms).
- #46 registry OOM (475MB-layer push killed the 256Mi template; dmesg evidence)
  fixed and re-verified: oom-kill count unchanged across a full rebuild+push.
- #47 AppleDouble ._*.sql embedding broke felis migrate on a Mac-staged tree;
  .dockerignore fix probed live with a planted junk file.
- #48 per-image docker start/stop tripped systemd start-limit-hit mid-batch;
  one wrap per batch, re-run mirrors all four.
- #49 installer re-runs silently reverted operator [registry]/[archive] config;
  carry-forward landed + live-verified into host toml, pod toml and the Secret,
  and the carried pins drove a successful POST /images/build.
- reachability table extended to #49 (① 20 | ② 19 | ③ 3+ | ④ 4 | 决策 3).
2026-09-23 19:35:13 +08:00
Lemon-miaow 72553cb414 docs(troubleshooting): the installer leaves docker stopped — start it before manual pushes
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
2026-09-23 19:35:13 +08:00
Lemon-miaow 20a95da487 docs(update): the plugins note is a rebuild + registry re-mirror now, not a node re-import
The felis-paper/felis-limbo jars are baked into the lobby/limbo images; with
the images hosted in the in-cluster registry, the extra step is pushing the
rebuilt image there (which is also what survives an image GC), not a bare
containerd import. The installer re-run does both.
2026-09-23 19:33:04 +08:00
Lemon-miaow c7e585e21d fix(bootstrap): mirror the image batch under ONE docker start/stop
Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.

push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
2026-09-23 19:23:44 +08:00
Lemon-miaow 5fa8b7412e fix(build): keep macOS ._*/.DS_Store junk out of the image
A Mac-staged tree (BSD tar materializes extended attributes as ._<name>
sidecars) went through the docker build and one landed in
internal/store/migrations/ — //go:embed-ed into the binary, where every
`felis migrate` then died with 'migration "._0004..." has a non-numeric
version'. Observed live wiring up the auditfix42 image: the installer's own
run_migrations failed on it. Exclude the sidecars and .DS_Store from the
build context; deploy/*.yaml and plugins/ have the same exposure.
2026-09-23 19:19:04 +08:00
Lemon-miaow b8e554dac7 fix(bootstrap): carry the operator's [archive] keys across re-runs too
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.

persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
2026-09-23 19:09:57 +08:00
Lemon-miaow 765a8923a4 fix(bootstrap): re-runs keep the operator's [registry] overrides
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.

Found while landing the registry-hosting work, which depends on those same
keys surviving.

- persisted_registry_block carries the operator-owned [registry] keys and the
  [registry.s3] subtable forward, same first-readable-file rule as
  persisted_smtp_block; url/build_namespace stay installer-owned (they must
  match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
  as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
  idempotence (a second re-run writes a byte-identical file).
2026-09-23 19:05:54 +08:00
Lemon-miaow fa0e8d7d97 docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
2026-09-23 19:03:11 +08:00
Lemon-miaow 13d64e0000 feat(bootstrap): host every built image in the internal registry — GC-durable pulls
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:

- Every image the installer builds is now a registry ref
  (registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
  containerd under that exact name (first boot needs no registry round-trip)
  and mirrored into the registry after deploy_bundle (push_image_to_registry:
  push endpoint 127.0.0.1:5000, and only the path after the host matters to the
  registry — a push there lands where kubelet's mirrored pull looks). A ref
  outside the registry is warned about, not silently unmirrored.

- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
  registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
  registry Deployment binds (node containerd cannot dial the Service VIP — live
  drill: "Empty reply"). k3s regenerates containerd config only at agent start,
  so a CONTENT change restarts k3s and an identical file (every re-run)
  restarts nothing.

- import_registry_image caches registry:2 into containerd so the registry
  Deployment can start on a box that cannot reach Docker Hub.

- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
  'felis-paper:demo') at the registry refs — a user server created from those
  rows must not strand when GC collects the bare tag. Only recommended rows
  still holding the old seed are touched; enabled is preserved; a pre-existing
  target row wins over a duplicate.

bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
2026-09-23 19:02:58 +08:00
Lemon-miaow a9b275abbb fix(platform): registry OOM (audit #46) + loopback hostPort — the node-side pull path
Two changes to the registry Deployment, both prerequisite to GC-durable images:

- Dedicated resource template: the control plane's 256Mi memory limit was a
  live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
  mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
  same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.

- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
  reach the Service VIP (live stack: "Empty reply"), so the node-side pull
  path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
  http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
  the plain-HTTP registry off every other interface.

Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
2026-09-23 18:55:23 +08:00
Lemon-miaow 92c06ac8dc docs(audit): twenty-eighth batch — #45 blind-review fix live-verified (context download, byte-exact); reachability 1-45 2026-09-23 17:01:19 +08:00
Lemon-miaow 168a37542b feat(api,panel): reviewer context download for submissions; dockerfile field documented as audit-only (audit #45) 2026-09-23 16:41:55 +08:00
Lemon-miaow edd9d63f5e docs(audit): twenty-seventh batch — alert module live drill (real build failure -> pending -> firing), cleanup, #44 reachability 2026-09-23 16:35:23 +08:00
Lemon-miaow 43df08b52a feat(alerts): ship Felis alert rules with promtool unit tests; document scraping & rules (troubleshooting §14) 2026-09-23 16:17:51 +08:00
Lemon-miaow 94f71eea19 feat(api): serve felis_* metrics on the internal face (build-failure counter's only scrape path) 2026-09-23 16:17:50 +08:00
Lemon-miaow 17ede3c6aa docs(audit): twenty-sixth batch ledger — setup wizard re-run screens, #44, build-pin drift incident & re-verify 2026-09-23 16:04:20 +08:00
Lemon-miaow ae6e9256c6 docs(troubleshooting): 8e — apply build-image overrides through the config Secret (restart alone does not) 2026-09-23 15:56:44 +08:00
Lemon-miaow abb5910d2f fix(cli): setup re-run keeps its already-set-up framing after connect/storage reconfigure 2026-09-23 15:56:44 +08:00
Lemon-miaow 5450ec786f docs(audit): reachability grading for findings #1-#43 (who actually hits each one) 2026-09-23 15:45:19 +08:00
Lemon-miaow 24a6ab3d1e docs(audit): twenty-fifth batch ledger — breakGlass console screens & backup/restore gates (#39–#43, live-verified) 2026-09-23 07:50:27 +08:00
Lemon-miaow ac3a557566 fix(cli): Sync picker hides system servers; keep the two 409 refusals apart
Two defects from the live Sync drill:

- The picker listed the system servers (login/lobby), which the backup API can
  never accept (reserved names, no servers row): the pick died in name
  validation with a raw "server name is reserved" error. backupPickable now
  filters them out; the halt picker keeps them on purpose (break-glass retains
  full power over system servers).
- backupErrorFromResponse mapped every 409 to the stopped gate, so the new
  world-volume refusal would have displayed the wrong reason. The 409 arm now
  keys on the body's error code; a code-less body still reads as the stopped
  gate.

Live (auditfix38): the picker shows only user servers; a world-less pick shows
the API's own "no world volume yet — start it once" text; the not_stopped text
is unchanged.
2026-09-23 07:49:07 +08:00
Lemon-miaow 508a1c02da fix(api): refuse backup/restore before a missing world volume
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
2026-09-23 07:49:00 +08:00
Lemon-miaow 55d515d41f fix(provisioning): keep the Owner seat single; clash on the operator name stays retryable
Two defects live-drilled in the break-glass staff provisioning:

- An Owner reset that typed any username other than the occupied seat took
  UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
  existing seat — possibly the compromised account the reset was meant to
  replace — live; every owner row is undeletable through the panel, so the tier
  could never converge back to one. provisionOwner now refuses with
  ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
  form); bootstrap still mints, and the seat's own username still resets in
  place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
  console keys its rename prompt off api.ErrConflict — the "choose another
  name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
  the contract; PGRepo had drifted). Map the unique violation to ErrConflict
  and pin it in pgint.

Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
2026-09-23 07:48:54 +08:00
Lemon-miaow f6dbfd3625 docs(audit): twenty-fourth batch ledger — live S3 upload-channel drill (0 defects, reverted clean) 2026-09-23 07:05:42 +08:00
Lemon-miaow f378953982 docs(audit): twenty-third batch ledger — reaper node pin (#38 + multi-node gap) 2026-09-23 07:00:33 +08:00
Lemon-miaow daf760220b fix(cli): pin the reaper to its storage node; drop the stale uid-1000 note
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
2026-09-23 06:58:29 +08:00
Lemon-miaow a31eca65c3 docs(audit): twentieth–twenty-second batch ledger — build outcome visibility, files page, fleet system services 2026-09-23 06:53:40 +08:00
Lemon-miaow 2f90851c03 fix(panel): mark platform system services read-only in the fleet table
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
2026-09-23 06:51:45 +08:00
Lemon-miaow 0a36b3fda9 feat(panel): add the server files page for the world-volume repair lever
The backend could list/read/write a stopped server's world volume since the
file-editor slice, but the panel had no entry, so the one repair path for a
server that will not boot (a wrong line in server.properties) was API-only.
New /servers/:name/files page: breadcrumb browser, editor dialog with the
base64 []byte codec, binary files open read-only, the stopped gate is owned
up front (with a stop action) instead of letting every call 409, and a
doorway card on the console. i18n files namespace + wire-shape tests.
2026-09-23 06:42:45 +08:00
Lemon-miaow 72c4aa3895 fix(submissions): surface each linked build's outcome to the submitter
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
2026-09-23 06:31:13 +08:00
Lemon-miaow 4933c075b0 docs(audit): nineteenth-batch ledger — passkey unbind panel entry 2026-09-23 06:25:01 +08:00
Lemon-miaow 11ac4f50e6 feat(panel): expose owner passkey unbind in the user danger zone
DELETE /users/{id}/passkeys shipped as the owner-tier remediation for a
lost or compromised authenticator, but nothing in the panel reached it.
Add the danger-zone action with a confirm dialog; the account keeps its
other doors (email OTP, in-game op-login re-enrollment), so this severs
a credential without locking anyone out. Wire-shape test pins the call.
2026-09-23 06:24:45 +08:00
Lemon-miaow 35d93d7612 docs(audit): eighteenth-batch ledger — #35 world-executor identity defect and the backups-page completion 2026-09-23 06:20:44 +08:00
Lemon-miaow 97a64c8a33 feat(panel): add back up now and recent operations to the backups page
The backups page could list and restore archives but not create one,
and nothing surfaced backup/restore Job outcomes — a failed 202 was
visible only through kubectl. Add a Back up now action (enabled only on
a stopped server, the backend's own gate; a raced 409 is surfaced in
its words) and a Recent operations card fed by GET /servers/{name}/jobs
that shows running/succeeded/failed with the Job's failure message,
re-reads on an interval while a Job is running, and persists across
reloads. Wire-shape tests pin both endpoints.
2026-09-23 06:20:16 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00