Commit Graph
186 Commits
Author SHA1 Message Date
Lemon-miaow 4d4cdd6ea7 fix(api,panel): refuse file operations on a server without a world volume (#58)
Live: the files page against a server whose world claim does not exist (never
started, or reaped) created a Job whose Pod stayed Pending on
FailedScheduling (persistentvolumeclaim not found) until the executor's 90s
wait expired — a 90s spinner answered by a misleading 504 files_timeout, for
a request that is knowably impossible. Backup and restore have refused this
shape with 409 no_world_volume since the #42 round; the file routes now run
the same gate before any Job is created, and the panel maps the code to a
localized message (it previously fell back to the English server text).
2026-09-23 22:17:08 +08:00
Lemon-miaow 9dad61f508 fix(nano): screen unusable profiles per source instead of stopping the ladder (#55)
A configured Yggdrasil root answering 200 with a name outside the Minecraft charset (or an identity UUID that does not parse) was rejected one layer up in the handler: a silent 204 with no log line, and because the rejection returned instead of continuing, every source behind the broken one was unreachable for that login. The resolver already treats the same class (200 without a usable profile, non-200, unreachable) as skip + log + failed; the name/UUID screens lived above it and silently stopped the ladder instead.

Live on the audit box, a single sloppy root produced 204s with no trace anywhere, and [bad root, valid root] answered 204 where the valid root would have admitted the login; nothing else in the nano matrix (60 checks across input validation, canonical rewrite, premium rename, failure modes, failover, log discipline, properties relay) was red.

Screen both shapes inside resolveHasJoined, before a 200 can win: identity ids must parse, third-party names must match the charset. A bad answer is logged ('unusable profile name' / 'unparseable profile id'), skipped, and counted as failed — 503 when nothing else validates, and later sources get their turn. The handler's guards stay as the last line before anything leaves (comments updated).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix55): the five bad-name cases and the two failover cases all pass; matrix rerun 60/60.
2026-09-23 21:21:21 +08:00
Lemon-miaow 2c6739ad76 fix(updater,install,docs): follow the move to FelisMC/Felis (#53)
The repository moved to FelisMC/Felis, but the felis-api release coordinate, the installer's default FELIS_REPO_URL, the PaperMC user-agent strings and both READMEs still named MliroLirrorsIngenuity/Felis. Live on the audit box, 'felis update' reported 'github: MliroLirrorsIngenuity/Felis releases/latest returned HTTP 404 -- ...', pointing operators at a coordinate that no longer exists; the old path keeps answering today only because GitHub still 301s the transfer (verified with a read token against api.github.com: old path 301, new path 200), and if that redirect is ever retired every install and every update check breaks with it.

Replace the coordinate in the six tracked files: the updater topology and both test fixtures, the bootstrap default URL and user-agent strings, and README.md/README_EN.md. Green live: the same command now reports 'github: FelisMC/Felis releases/latest returned HTTP 404 -- ...' (still 404 because the new home has published no stable release yet -- a release-process fact, not a code bug).

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean.
2026-09-23 20:58:34 +08:00
Lemon-miaow 13d64e0000 feat(bootstrap): host every built image in the internal registry — GC-durable pulls
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:

- Every image the installer builds is now a registry ref
  (registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
  containerd under that exact name (first boot needs no registry round-trip)
  and mirrored into the registry after deploy_bundle (push_image_to_registry:
  push endpoint 127.0.0.1:5000, and only the path after the host matters to the
  registry — a push there lands where kubelet's mirrored pull looks). A ref
  outside the registry is warned about, not silently unmirrored.

- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
  registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
  registry Deployment binds (node containerd cannot dial the Service VIP — live
  drill: "Empty reply"). k3s regenerates containerd config only at agent start,
  so a CONTENT change restarts k3s and an identical file (every re-run)
  restarts nothing.

- import_registry_image caches registry:2 into containerd so the registry
  Deployment can start on a box that cannot reach Docker Hub.

- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
  'felis-paper:demo') at the registry refs — a user server created from those
  rows must not strand when GC collects the bare tag. Only recommended rows
  still holding the old seed are touched; enabled is preserved; a pre-existing
  target row wins over a duplicate.

bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
2026-09-23 19:02:58 +08:00
Lemon-miaow a9b275abbb fix(platform): registry OOM (audit #46) + loopback hostPort — the node-side pull path
Two changes to the registry Deployment, both prerequisite to GC-durable images:

- Dedicated resource template: the control plane's 256Mi memory limit was a
  live-bite bug (#46) — pushing a 475MB layer OOM-killed the registry
  mid-upload (dmesg oom-kill, oom_score_adj 989) and the push failed; the
  same push completes in 2s with 2Gi. Registry limits are now 1 CPU / 2Gi.

- The container port carries hostPort 127.0.0.1:5000. Node containerd cannot
  reach the Service VIP (live stack: "Empty reply"), so the node-side pull
  path is a registries.yaml mirror rewriting registry.<ns>.svc:5000 onto
  http://127.0.0.1:5000, which lands on this hostPort. Loopback-only keeps
  the plain-HTTP registry off every other interface.

Tests pin both: exactly one port with hostIP 127.0.0.1, and a memory limit
>= 2Gi (exceeding the control-plane template) with the #46 evidence cited.
2026-09-23 18:55:23 +08:00
Lemon-miaow 168a37542b feat(api,panel): reviewer context download for submissions; dockerfile field documented as audit-only (audit #45) 2026-09-23 16:41:55 +08:00
Lemon-miaow 94f71eea19 feat(api): serve felis_* metrics on the internal face (build-failure counter's only scrape path) 2026-09-23 16:17:50 +08:00
Lemon-miaow 508a1c02da fix(api): refuse backup/restore before a missing world volume
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
2026-09-23 07:49:00 +08:00
Lemon-miaow 55d515d41f fix(provisioning): keep the Owner seat single; clash on the operator name stays retryable
Two defects live-drilled in the break-glass staff provisioning:

- An Owner reset that typed any username other than the occupied seat took
  UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
  existing seat — possibly the compromised account the reset was meant to
  replace — live; every owner row is undeletable through the panel, so the tier
  could never converge back to one. provisionOwner now refuses with
  ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
  form); bootstrap still mints, and the seat's own username still resets in
  place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
  console keys its rename prompt off api.ErrConflict — the "choose another
  name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
  the contract; PGRepo had drifted). Map the unique violation to ErrConflict
  and pin it in pgint.

Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
2026-09-23 07:48:54 +08:00
Lemon-miaow daf760220b fix(cli): pin the reaper to its storage node; drop the stale uid-1000 note
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
2026-09-23 06:58:29 +08:00
Lemon-miaow 2f90851c03 fix(panel): mark platform system services read-only in the fleet table
login/lobby carry reserved names, so every per-server route rejects them —
yet the cockpit offered claim/stop/wake and a console link on their rows,
each answering 400 bad_name. The fleet view now marks them (system:true,
shared naming.IsSystemServer) and the panel renders a plain label instead
of dead actions.
2026-09-23 06:51:45 +08:00
Lemon-miaow 72c4aa3895 fix(submissions): surface each linked build's outcome to the submitter
/me/submissions (and the admin queue) now attach build_status/build_error by
a read-only Builder.Get — until now a failed build was visible only on the
admin-tier /images/build routes, so the person who submitted the modpack
never learned the build died. A missing build row renders as "no outcome";
any other lookup failure surfaces instead of being swallowed. The panel's
My Submissions page renders the outcome in the expanded row, localised.
2026-09-23 06:31:13 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00
Lemon-miaow bb9798e32c fix(api): in-game identity resolution and link takeover ignore dead accounts
UserByMCUUID now resolves only live accounts: claim, menu, wake
authorization, op-login vouch and the QR link-status poll treat a
disabled or soft-deleted link holder exactly like an unlinked UUID
instead of a retired identity. VerifyLinkCode lets a soft-deleted
link be taken over by a fresh in-game code (the deleted account is
gone, e.g. a migrated source), while a disabled holder still 409s so
the lockout is not bypassable; failed attempts still do not consume
the code. Fake repo and pgint coverage pin both branches.
2026-09-23 05:40:29 +08:00
Lemon-miaow 58535890c4 fix(api): dead accounts cannot log in, hold sessions, or keep identity assets 2026-09-23 05:28:32 +08:00
Lemon-miaow 1918da29be fix(api): the quota/link admin sub-resources require a live user (404, not FK 500) 2026-09-23 04:56:06 +08:00
Lemon-miaow 36b954d347 docs(backupjob): the backup Job name is not deterministic anymore
The comment described a deterministic-name collision that the unique random
suffix made near-impossible; align it with Backuper.Backup and jobspec's
contract (ErrAlreadyExists survives only as the defensive no-op).
2026-09-23 04:41:01 +08:00
Lemon-miaow 30857df5b2 fix(cfsetup): upsert the Access policy — never swallow already-exists over a broader rule set
CreateAccessPolicy treated a Cloudflare "policy_already_exists" as idempotent
success and kept whatever policy was there. On a re-run with a changed
identity — or against a hand-made broader policy — op.console would stay
guarded by something weaker than the fail-closed body this package builds and
guards, while Setup reported success. The fail-closed validation only ever ran
on the policy we built, never on the one that stayed live.

Now it upserts by name: lookup, PUT the guarded body over the existing policy,
POST only when absent (a racing POST re-looks up and PUTs). apiPost/apiPut
share one apiWrite; three httptest cases pin update-over-existing, create-when-
absent, and the race fallback.
2026-09-23 04:30:53 +08:00
Lemon-miaow 82b5a606f7 fix(operator): end the start attempt on success — stale anchor caused false StartupTimeout
Found live while validating the RCON-secret heal: a server that had already
recovered to Ready was marked Failed(StartupTimeout) minutes later, the moment
an unrelated pod rollout briefly dropped readyReplicas. The anchor
(status.startRequestedAt) was never cleared on success, so its 300s budget
kept ticking under a healthy server and any later blip spent it.

markRunningReady now clears the anchor: every start-or-recovery attempt gets
its own budget. It also flips ConditionProvisioned back to True — markFailed
sets it False and nothing ever reset it, leaving a permanent failure flag on
recovered servers that every conditions consumer would read.

Unit tests pin both: anchor cleared on Ready, Provisioned recovers from
Failed to Running.
2026-09-23 04:23:59 +08:00
Lemon-miaow 56f3abdb36 fix(operator): heal a deleted RCON Secret instead of locking the server out
Deleting the per-server RCON Secret used to leave a running pod authenticating
with the lost password while the operator re-minted a fresh one and probed
with it: the RCON gate failed forever (live: 96s+ of RconNotReachable, headed
for ReadinessTimeout) and nothing re-triggered a pod restart — the server only
came back when the pod was deleted by hand.

Two changes pair up:
- Owns(&corev1.Secret{}) so the deletion is noticed at all (a quiet Running
  server emits no other events; the Secret is controller-owned, so the watch
  maps it back to the CR).
- The pod template now carries a fingerprint of the current password
  (RconSecretAnnotation). Re-creation changes the fingerprint, the
  StatefulSet rolls, and the new pod picks the new password up; while the
  Secret is untouched the value is stable so no spurious rolls.

Unit tests pin stability across reconciles and the change-on-recreation roll.
2026-09-23 04:20:49 +08:00
Lemon-miaow f650bf892a fix(operator): idle auto-stop couldn't write — patch the spec, and grant the patch
Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:

- The stop used a whole-object Update while the same reconcile loop writes
  status; that risks clobbering a concurrent status write. Switch to the
  reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
  markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
  closed with 403 (visible in the operator log as 'cannot update resource
  "minecraftservers"'). Grant patch and pin it in the RBAC scope test.

With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
2026-09-23 04:08:58 +08:00
Lemon-miaow 1c89a5eeeb fix(operator): wake up for idle auto-stop — the timer had no driver
EmptySince was stamped and then never revisited: player joins/leaves do not
touch the CRD, RCON is only probed inside Reconcile, and a steady Running
server produces no watch events (its status update goes out unchanged and is a
no-op). Live, an empty server with a 30s grace sat Running for minutes with
zero reconciles in the log — the auto-stop existed only on paper.

reconcileRunning now returns a RequeueAfter for idle-enabled servers: exactly
at the deadline while empty, or a 30s probe cadence while occupied so the
moment the last player leaves is noticed. New unit tests pin all three:
deadline requeue, occupied cadence, and no requeue when disabled.
2026-09-23 04:05:31 +08:00
Lemon-miaow 8e7c7bbf24 fix(reaper): deliver pre-reap warnings for real — and never fake a delivery
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:

- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
  `warning suppressed — no warner wired` and does NOT stamp; a delivery error
  logs and retries on the next daily run (bounded by the warning window). The
  stamps are no longer burned by notices nobody received.

- A real channel: mail.SendNotice (the second and last message shape the mail
  package sends) plus a mailWarner that resolves the owner's VERIFIED email
  and mails the notice through the configured [smtp] relay. `felis reaper`
  wires it when [smtp] is set (same password_ref convention as felis-api) and
  prints exactly what happens when it is not.

Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.

Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
2026-09-23 03:47:19 +08:00
Lemon-miaow bb68fefe04 fix(quota): make the claim gate atomic, and stop zeroing the storage cache
Two defects in the §9.3 quota path, both invisible to the hermetic suite:

- Audit #4's TOCTOU was real and documented: QuotaCheck and ClaimServer were
  separate statements, so two concurrent claims by one user for two different
  ownerless servers both read count < max_servers and both won. The gate now
  lives inside ClaimServer, in the SAME transaction as the ownership write,
  under pg_advisory_xact_lock(hashtext(user_id)) — the aggregate read, the
  four-dimension re-check (shared with QuotaCheck via one helper so the two
  cannot drift), and the UPDATE are one serialized decision. The loser gets
  ErrQuotaExceeded, which both claim handlers map to the same 403 the
  sequential path gives; the server row is additionally taken FOR UPDATE so
  same-server races still resolve to exactly one winner.

- The server PATCH path called UpdateServerResources(..., 0) for storage even
  though a resources patch cannot change storage. The cached columns are the
  ONLY input to the quota aggregate, so every resource patch silently dropped
  that server's storage contribution from its owner's cap. The handler now
  reads the current spec and passes storage through.

Red-then-green: the new pgint test drives two real concurrent claims against
max_servers=1 (before: both win; now: exactly one win + one gated 403, and the
DB shows one owned row); the hermetic suite pins the 403 mapping and the
storage-preserving cache write.
2026-09-23 03:37:59 +08:00
Lemon-miaow e0d23780d8 fix(auth): make the owner role real — provisioning, staff doors, panel guards
Found live while verifying the admin email-edit fix: the Owner account could
not load /api/v1/users at all. Root cause: migration 0011 adds the 'owner'
role and gates every user-administration route on it, but NOTHING ever wrote
it. break-glass (UpsertOwner), the setup MC-bind (CompleteOwnerSetup), and the
re-provision path all forced 'admin', so in a fresh install the entire
owner tier — list/create/edit/disable/delete users, quotas, sessions — was
unreachable. The role was a dead letter in the other direction too: staff
predicates that predate the role did not know it.

- UpsertOwner and CompleteOwnerSetup now write role='owner'; the username-
  conflict arm re-asserts it, which is also the documented pre-0011 promotion
  path ("re-provision via break-glass"). InsertOperator stays plain 'admin'.
- Staff doors learn the role: op-login start/finish admit the Owner; the
  player email door refuses it like any staff account; the in-game approver
  check already used staffRole.
- Reclaim protection: IsProtectedAdminLink (and the break-glass bootstrap
  switch AdminExists) count admin OR owner — the Owner must never be displaced
  by a Mojang-priority reclaim.
- Panel guards make migration 0011's claim true now that owner rows exist: an
  owner can never be demoted, deleted, or disabled through the API (only the
  local break-glass console resets the identity); username/email edits still
  work.

Tests: pgint pins both provisioning paths, the protected-link predicate and
the reset/promote semantics; hermetic suites cover the owner-admitting staff
door, the owner-refusing player door, the three panel guards, and break-glass
attribution.
2026-09-23 03:30:29 +08:00
Lemon-miaow d1ec40f738 fix(users): an admin email edit must clear the stale verification
UpdateUser wrote a new address but kept email_verified, so patching a verified
account asserted a proof of an address nobody had proven — and the
pre-session login mails and resolves on exactly that flag, so a typo'd edit
could hand the account's sign-in codes to the wrong mailbox.

Changing the address now clears the flag in the same write; a no-op patch that
passes the same value keeps it. The fake mirrors the semantics, and the pgint
suite pins both halves (same value keeps proof, new value drops it).
2026-09-23 03:19:55 +08:00
Lemon-miaow b6ef27cd2d fix(auth): enforce the verified-email uniqueness that email login assumes
The design has claimed since migration 0010 that at most one account can hold
a PROVEN email address, with ErrEmailTaken as the 409 a second verifier sees.
Neither half ever shipped: no migration created users_verified_email_unique,
and VerifyEmailOTP had no guard at all — the sentinel was defined but never
returned, so two accounts could both verify one address. The damage is not
cosmetic: the pre-session login resolves accounts BY verified email, so the
duplicate decided which identity a mailed sign-in code belonged to.

- Migration 0020 creates the partial unique index (lower(email) WHERE
  email_verified) the comments have been citing — the database-level backstop.
- VerifyEmailOTP now refuses the take-over with ErrEmailTaken BEFORE consuming
  the code (the address, not the code, is the problem), charges no attempt,
  and maps a lost cross-user race (unique violation) to the same answer.
- The verify handler answers 409 email_taken instead of a generic 500.

Covered by the pgint suite (sequential double-verify refused with the code
still live, a direct duplicate write still loses to the index, the refused
account can still prove its own address) and a hermetic 409 case.
2026-09-23 03:19:35 +08:00
Lemon-miaow 2a55a0d265 test(pgint): verify the business stores against a real Postgres
The hermetic suites encode the store contracts against fakes; PGRepo drifted
behind them three times (attempt accounting, a missing JOIN, a missing FOR
UPDATE) while every unit test stayed green. This harness replays the real
embedded migrations onto a throwaway database — its name must contain "pgint"
or the harness refuses to run — and exercises the SQL directly: sessions, the
onboarding email-OTP lifecycle (supersede/expiry/lockout), the pre-session
login consume, the op-login state machine, link and bind-code redemption,
submissions, and builds with the image admission round trip.

Run it after touching SQL under internal/api/pgrepo.go, internal/submit, or
internal/build; CONTRIBUTING.md carries the one-liner.
2026-09-23 03:18:56 +08:00
Lemon-miaow 02fd2de502 fix(build): three drill-driven fixes so the lane actually completes on a starter node
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:

- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
  Pod never scheduled on the platform's own starter node: FailedScheduling /
  Insufficient memory, Pending forever. Requests are now a small floor
  (250m / 512Mi, never above a configured cap) while the limits stay the
  safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
  the source owner; a 65532-owned context (the distroless felis image uid)
  fails that under the pod's dropped capabilities ('copying dockerfile:
  chown /kaniko/Dockerfile: operation not permitted'). The fetch container
  now extracts as root — the uid Kaniko already runs as — so the copy
  succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
  step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
  renders --db-repository, and docs/troubleshooting.md §8e now carries the
  verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
  internal registry; --insecure already covers its plain HTTP).

Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
2026-09-22 23:01:25 +08:00
Lemon-miaow f79e5ebb5e feat(build): make the user-modpack build lane read its context (closes the last functional gap)
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:

- submit: derived context refs become the internal-face URL
  /api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
  gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
  route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
  image's new fetch-context entrypoint) that streams the blob with the
  namespace-local service-token Secret — never mounted into Kaniko — and
  extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
  reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
  build namespace gets the token Secret through the existing replica mechanism
  (bootstrap.sh + felis setup); the build egress lock opens exactly the control
  namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
  escapes/symlinks/non-gzip).

Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
2026-09-22 22:45:09 +08:00
Lemon-miaow 0c8e29b05a fix(platform): give every control-plane Deployment real probes (#8 follow-up)
The api, operator and registry Deployments shipped with no liveness/readiness
probes at all: a wedged process stayed 'Running' forever, and the operator had
no health listener to probe in the first place. Kaniko build evidence on a
fresh install showed the only cluster-wide red after a disk-pressure pass was
Deployment status that never reflected health.

- felis-api: readiness /readyz (DB + K8s API round-trip) and liveness /healthz
  on the internal face (:8081), the only listener that serves both endpoints;
  liveness deliberately avoids /readyz so a DB blip cannot restart the api.
- felis-operator: new --health-probe-bind-address (:8081) with controller-
  runtime's /healthz + /readyz (registered ping checks; an unregistered handler
  map would 404), plus the matching container port and probes.
- registry: /v2/ probes on the pinned port, so a broken storage backend stops
  reading as 'Running'.

Tests pin paths, ports, and that each probe targets a declared container port.
2026-09-22 22:22:37 +08:00
Lemon-miaow 87a9f4eb25 feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
2026-09-22 22:03:48 +08:00
Lemon-miaow 0a2d654e68 fix(platform): control plane runs system-cluster-critical, so eviction refuses it (#8)
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.

Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
2026-09-22 21:55:54 +08:00
Lemon-miaow fe310743a2 fix(platform): give the control plane a PriorityClass eviction shield (#8)
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
2026-09-22 21:08:56 +08:00
Lemon-miaow fd33fd05e1 fix(install): backups exist on a default install; retention resolves real world dirs (#6)
Three faces of one gap, all on the supported install path:

- Backup/restore answered 503 out of the box: nothing ever rendered the
  archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
  PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
  'felis manifests' names it by default (--backup-pvc= is the explicit
  no-store shape); bootstrap passes it through so the generated felis.toml
  [archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
  now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
  env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
  nothing deleted).
- Even when enabled it could not find a world on a stock install:
  resolveWorldDir now also resolves the exact local-path directory
  <pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
  glob, so a stale deleted PV's bytes can't be archived in place of the
  current world). Reaper Role gains persistentvolumeclaims:get (weaker
  than the delete it already held).

README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
2026-09-22 20:48:07 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow a2df2f242b fix(operator): re-probe RCON every 2s while Starting
The readiness gate is status-driven; at a 5s re-probe cadence the observed
'container Ready but API still 409 not_running' window was 6~10s. Halving the
cadence halves the worst case; probes still only run while unreachable.
2026-09-22 20:31:53 +08:00
Lemon-miaow 2a8f897e61 fix(api): a session-store outage answers 503, not 401
Resolving a session cookie failed identically whether the credential was
missing or Postgres was unreachable: local_auth_enabled read errors fell into
the fail-closed 'disabled' branch and SessionUser errors into 'invalid
session', both surfacing as 401 'authentication required' — a lie that reads
as 'log in again' during an outage. Split the enabled-read into
(enabled, error), tag non-ErrNotFound store failures with errAuthBackend, and
map that to a new 503 auth_unavailable in requireExternal. Fail-closed is
unchanged: missing setting / bad value / missing session stay 401.
2026-09-22 20:30:43 +08:00
Lemon-miaow c839454a1f fix(manifests): reaper ServiceAccount lives in (and binds from) the Minecraft namespace
Follow-up to the CronJob placement fix: a Pod cannot USE a ServiceAccount from
another namespace either (live drill: 'error looking up service account
minecraft/felis-reaper: serviceaccount not found'). Move the SA and its
RoleBinding subject to the Minecraft namespace alongside the CronJob.
2026-09-22 20:18:36 +08:00
Lemon-miaow e4f2cff532 fix(manifests): render the retention reaper CronJob into the Minecraft namespace
A Pod can only mount PVCs from its own namespace; the CronJob referenced the
minecraft-namespace backup PVC while being rendered under ControlNamespace, so
it could never schedule — live drill: FailedScheduling 'persistentvolumeclaim
felis-backups not found'. The reaper Role/RoleBinding were already
minecraft-scoped (the objects it touches live there), so the CronJob was the
odd one out. The minecraft felis-config replica (felis setup, backup Job fix)
supplies its config mount.
2026-09-22 20:14:19 +08:00
Lemon-miaow 0414913bc7 fix(api): serialise RedeemPlayerBindCode — concurrent redeem 500s become clean 400s/idempotent converges
6-way concurrent redeem of one code 500'd on users_username_key (each request
generated a fresh user id but the same uuid-derived username), plus the rarer
two-codes-one-uuid race. Same drift family as VerifyLinkCode, which already
locks its code row and handles the conflict.

- SELECT ... FOR UPDATE the code row: same-code racers serialise; losers exit
  as ErrLinkCodeInvalid (400 invalid_code), no user row is attempted.
- INSERT users ... ON CONFLICT (username) DO NOTHING + re-read by username:
  cross-code racers converge on the winner's row (role checked, staff still
  refused) instead of a unique-violation 500.
- account_links ON CONFLICT (mc_uuid) DO NOTHING for the same race.

Verified live: same-code x6 = 1x200 + 5x400; two-codes x2 = 2x200 same user;
db clean; zero unmapped errors.
2026-09-22 20:08:38 +08:00
Lemon-miaow dcc3b7403e fix(api): fill ListPendingOpLogins username/created_at (PG lagged the interface+fake)
The interface doc promised 'each joined to its staff username', the fake and
the pending handler both project username and created_at, but the PG query
selected neither — live internal /op-login/pending returned username:"" and
created_at:0001-01-01. Same drift class as ConsumeLoginEmailOTP: fake-based
tests can't see PG-only regressions.
2026-09-22 20:04:32 +08:00
Lemon-miaow 52549f7b3a fix(api): ConsumeLoginEmailOTP honesty — wrong/expired/consumed codes are ErrOTPInvalid 400, not a 500
The PG implementation was a single UPDATE ... WHERE code_hash that returned
ErrNotFound on zero rows: every wrong, expired, replayed or superseded code on
the pre-session email-login door (and the op-login finish / migration confirm
doors) fell through to writeError's unmapped-error 500, and attempts were never
charged so otpMaxAttempts/ErrOTPLocked could not trigger. The fake repo and the
Repo interface ("SAME code lifecycle as VerifyEmailOTP") already documented the
intended contract; only the PG side had drifted.

Mirror VerifyEmailOTP's transaction without its users write: SELECT ... FOR
UPDATE the newest live row, expiry + attempt cap before the hash compare,
mismatch charges one attempt and returns ErrOTPInvalid without consuming,
match consumes and commits. Verified live on the VM: 5 wrong guesses return
400 and stop at attempts=5 (correct code then also refused, unconsumed);
fresh code redeems; replay returns 400.
2026-09-22 19:51:04 +08:00
Lemon-miaow 9309ff5a7f chore: apply the missed S1016 conversions in handlers_users
The gofmt/staticcheck commit staged handlers_user.go (singular) for the
formatting fix but missed this sibling for its two struct-literal-to-
conversion cleanups.
2026-09-22 17:57:03 +08:00
Lemon-miaow e690b058db fix(restore): wait for the tracking finalizer before recreating
Live verification of the previous commit showed the immediate retry STILL
stranded: deleting a finished Job leaves it terminating (job-tracking
finalizer), so the re-Create collided with the dying object and was
mapped to ErrAlreadyExists a second time. Poll until the name actually
frees (bounded, ~10s) and surface a 'retry shortly' error if a stuck
finalizer ever outlives the budget. Fake-client tests pin both the
replace-finished and coalesce-in-flight branches.
2026-09-22 17:54:01 +08:00
Lemon-miaow 90ccbfede4 fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
2026-09-22 17:50:04 +08:00
Lemon-miaow a05edc934c chore: gofmt the tree, clear staticcheck, add a CI gofmt gate
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
2026-09-22 17:49:54 +08:00
flyemoji 99c31c1d4e fix(config): refuse plaintext auth-source urls to public hosts
An auth_source url could be http:// to any host. Anyone on the path
to a public root, or anyone who can spoof its DNS name, can then
answer hasJoined with a 200 and log in as any player of that source,
including a third-party account linked to staff. The player's IP also
travels in cleartext. Mojang logins are unaffected, since that source
is built in over https.

Config load now refuses http:// unless the host is localhost or a
loopback or private IP address (127.0.0.0/8, ::1, 10/8, 172.16/12,
192.168/16, fc00::/7), so a root on the same host or the LAN still
works without TLS. The decision is made on the literal host because
nothing is resolved at load time, so a LAN root named by hostname
needs its IP address or https. The error says what to change.

The new test covers public names and addresses, link-local, 0.0.0.0
and the first address past 172.16/12 (all refused over http, all
accepted over https), and the loopback and private forms that stay
allowed. It fails on the old check.
2026-09-22 14:24:41 +09:00
flyemoji e9f74f3f0f fix(nano): stop trusting an expired free name while mojang is failing
When the premium-name lookup failed, isPremiumName fell back to any
cached answer, however old. An expired "free" is exactly the answer
that may have stopped being true: someone can buy the name after it
was last seen free. For as long as api.mojang.com kept failing (429,
5xx, a timeout), a third-party player holding that name kept it on
every reconnect, and the Velocity registry, keyed on the name, turned
its new owner away as already connected. A hostile source could drive
the host into Mojang's rate limit on purpose to hold names that way.

A failed lookup now always counts as taken, so the player is renamed
with the source's prefix. An expired "taken" already gave that answer,
so only the stale "free" case changes. The cost is cosmetic: during an
outage an ordinary third-party player may get a prefix they do not
need, and their data follows the UUID, not the name.

A new test gives the cache a free entry past its TTL and has Mojang
answer 429. It fails on the old fallback. The two comments that
described the fallback now describe the fail-closed rule.
2026-09-22 14:17:53 +09:00
flyemoji c2a5645c55 fix: keep internal section numbers out of runtime messages
Four messages that reach an operator or an API client cited sections
of a specification nobody outside the project can read:

- the unimplemented archive store error from config load
- the running-server cap refusal, from both the user wake and the
  internal wake
- the missing memory ceiling guard, in the API and in felis apply

The references are gone and the wording is otherwise unchanged. Each
message still says what went wrong and, where there is one, what to
do about it. The test for the archive store message checks for the
tarLocal remediation, which is still there.
2026-09-22 13:44:38 +09:00