Provisioning is create-if-absent, so a field the desired spec gained after an
install (spec.rcon, spec.startup.healthHTTPPort, a derived env key) never
reaches the existing login/lobby CR while every re-run of setup reports
success — the reported 'configuration updates never reach an installed
deployment' symptom. converge is the explicit pass: it fills exactly the
zero-valued whitelist fields and the derived env (including a missing key,
which refreshDerivedEnv deliberately never adds), and never overwrites a
non-zero value. The timing stays with the operator because enabling RCON or
the HTTP readiness gate on a pre-listener image would wedge that server in
Starting until it was marked Failed.
Tests: fills predated fields while operator edits survive / non-zero values
left alone / absent + foreign + unset-image guards. usage table updated so the
router-parity test passes; troubleshooting gains §12b.
A logged-in user could file submissions without bound and stream a 1 GiB
context per submission. The only limits were the single-blob size cap and the
5 GiB uploads PVC (platform/workloads.go); nothing counted a user's rows or
bytes, so one account could fill the volume and every other user's upload
would start failing.
- Create: per-user pending_review cap (default 5) — the review queue cannot
be parked full of one account's rows. Check-then-insert, documented soft.
- UploadContext: per-user stored-context budget (default 2 GiB) charged
against the blob store's REAL sizes (new Blobs.Size on local/S3 stores), so
the sum cannot drift from the volume; the write is capped at the remaining
budget, so the excess is refused before it is persisted, and a re-upload is
charged only for its new bytes.
- API: per-user create/upload throttles (30s/15s, cmd/felis-wired) on a
dedicated cooldown keyspace, reserve→release so a failed attempt never
burns the window and a burst collapses to one winner; ErrQuotaExceeded →
403 submission_quota_exceeded (distinct from the 400 an oversize blob
gets), 429 submission_cooldown for the throttles.
- Panel: zh/en copy for both codes; openapi documents 403/429 on the two
user routes; pgint covers the pending-queue count.
Unit tests: submit package (cap, budget boundary/exact-fit/replacement,
oversize-vs-quota split) and api handlers (quota 403 both paths, throttle
429 + recovery + failure-release). go vet/go test/gofmt clean; panel
vitest 118 + typecheck green.
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.
Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.
Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
felis setup's in-TUI applies (storage / connection / edge) refreshed only the
control-namespace felis-config Secret; the workload-namespace mirror kept the
render from the previous run's startup pass until the next setup or installer
run. Found live: after 's -> Local' the minecraft copy still carried
user_uploads_context = s3://felis-wizard-uploads while the control copy and
both tomls were local. The 'configure email' path already overwrote both
mirrors, so storage/connection were the odd ones out.
Move the mirror refresh into applyFelisConfigSecret — the single choke point
every apply path calls — best-effort with a warning, since a control-plane
default install may not have the workload namespace at all. The smtp helper
drops its now-duplicate felis-config block.
The felis-paper/felis-limbo jars are baked into the lobby/limbo images; with
the images hosted in the in-cluster registry, the extra step is pushing the
rebuilt image there (which is also what survives an image GC), not a bare
containerd import. The installer re-run does both.
Two defects from the live Sync drill:
- The picker listed the system servers (login/lobby), which the backup API can
never accept (reserved names, no servers row): the pick died in name
validation with a raw "server name is reserved" error. backupPickable now
filters them out; the halt picker keeps them on purpose (break-glass retains
full power over system servers).
- backupErrorFromResponse mapped every 409 to the stopped gate, so the new
world-volume refusal would have displayed the wrong reason. The 409 arm now
keys on the body's error code; a code-less body still reads as the stopped
gate.
Live (auditfix38): the picker shows only user servers; a world-less pick shows
the API's own "no world volume yet — start it once" text; the not_stopped text
is unchanged.
Two defects live-drilled in the break-glass staff provisioning:
- An Owner reset that typed any username other than the occupied seat took
UpsertOwner's insert arm and silently minted a SECOND owner row, leaving the
existing seat — possibly the compromised account the reset was meant to
replace — live; every owner row is undeletable through the panel, so the tier
could never converge back to one. provisionOwner now refuses with
ownerSeatTakenError naming the seat (recoverable: the TUI routes back to the
form); bootstrap still mints, and the seat's own username still resets in
place. PGRepo gains OwnerUsername for the guard.
- InsertOperator returned the raw driver error on a taken username while the
console keys its rename prompt off api.ErrConflict — the "choose another
name" leg died with SQLSTATE 23505 against real Postgres (the fake encoded
the contract; PGRepo had drifted). Map the unique violation to ErrConflict
and pin it in pgint.
Live (auditfix37): fresh username refused naming the seat; seat reset kept the
id/email with still exactly one owner; taken operator name returned to the form
with the retry note, and the retyped name succeeded (drill rows cleaned).
Two things in the same surface. --reaper-node is the supported multi-node
answer: the rendered CronJob's pod gets a kubernetes.io/hostname selector, so
it reads the hostPath on the node that actually holds the worlds instead of
possibly scheduling where it is empty (naming a node without
--worlds-host-path is fail-loud). And the render note still told operators to
grant uid-1000 traverse / setfacl after #35 moved every world executor to
root+DAC_OVERRIDE — it now states that fact instead of the obsolete ritual.
'smtpSecretManifest' hardcoded namespace=felis, so the 'configure email'
refresh of the minecraft-namespace copies failed before it began: kubectl
refuses a manifest whose namespace conflicts with -n (found live: 'the
namespace from the provided object "felis" does not match the namespace
"minecraft"'), and the felis-config mirror never ran at all because the
smtp apply returned early. A later SMTP change could therefore never reach
the reaper's pre-reap warnings — the exact failure the refresh was added to
close.
Render the Secret with the caller's namespace (felis for the control-plane
apply, the workload namespace for the mirror). Regression test pins both.
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:
- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
`warning suppressed — no warner wired` and does NOT stamp; a delivery error
logs and retries on the next daily run (bounded by the warning window). The
stamps are no longer burned by notices nobody received.
- A real channel: mail.SendNotice (the second and last message shape the mail
package sends) plus a mailWarner that resolves the owner's VERIFIED email
and mails the notice through the configured [smtp] relay. `felis reaper`
wires it when [smtp] is set (same password_ref convention as felis-api) and
prints exactly what happens when it is not.
Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.
Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
Found live while verifying the admin email-edit fix: the Owner account could
not load /api/v1/users at all. Root cause: migration 0011 adds the 'owner'
role and gates every user-administration route on it, but NOTHING ever wrote
it. break-glass (UpsertOwner), the setup MC-bind (CompleteOwnerSetup), and the
re-provision path all forced 'admin', so in a fresh install the entire
owner tier — list/create/edit/disable/delete users, quotas, sessions — was
unreachable. The role was a dead letter in the other direction too: staff
predicates that predate the role did not know it.
- UpsertOwner and CompleteOwnerSetup now write role='owner'; the username-
conflict arm re-asserts it, which is also the documented pre-0011 promotion
path ("re-provision via break-glass"). InsertOperator stays plain 'admin'.
- Staff doors learn the role: op-login start/finish admit the Owner; the
player email door refuses it like any staff account; the in-game approver
check already used staffRole.
- Reclaim protection: IsProtectedAdminLink (and the break-glass bootstrap
switch AdminExists) count admin OR owner — the Owner must never be displaced
by a Mojang-priority reclaim.
- Panel guards make migration 0011's claim true now that owner rows exist: an
owner can never be demoted, deleted, or disabled through the API (only the
local break-glass console resets the identity); username/email edits still
work.
Tests: pgint pins both provisioning paths, the protected-link predicate and
the reset/promote semantics; hermetic suites cover the owner-admitting staff
door, the owner-refusing player door, the three panel guards, and break-glass
attribution.
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:
- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
Pod never scheduled on the platform's own starter node: FailedScheduling /
Insufficient memory, Pending forever. Requests are now a small floor
(250m / 512Mi, never above a configured cap) while the limits stay the
safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
the source owner; a 65532-owned context (the distroless felis image uid)
fails that under the pod's dropped capabilities ('copying dockerfile:
chown /kaniko/Dockerfile: operation not permitted'). The fetch container
now extracts as root — the uid Kaniko already runs as — so the copy
succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
renders --db-repository, and docs/troubleshooting.md §8e now carries the
verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
internal registry; --insecure already covers its plain HTTP).
Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
A submitted modpack was durable but unreadable: the uploads PVC cannot cross
namespaces (felis-api mounts it; Kaniko runs in felis-build) and the s3 lane
handed the sandboxed build Pod no credentials, so NO user build could ever
consume its context. The transport is now the API itself:
- submit: derived context refs become the internal-face URL
/api/v1/internal/submissions/{id}/context (service-token gated), and Blobs
gains Open (local + s3) with an ErrBlobNotFound sentinel for the route's 404.
- api: serves that route on the internal face only (openapi.yaml updated; the
route-coverage test enforces it).
- build: an http(s) context renders a context-fetch initContainer (the felis
image's new fetch-context entrypoint) that streams the blob with the
namespace-local service-token Secret — never mounted into Kaniko — and
extracts it under a zip-slip guard into a size-limited emptyDir that Kaniko
reads read-only as --context=/context.
- platform/install: the api Deployment carries its own internal base URL; the
build namespace gets the token Secret through the existing replica mechanism
(bootstrap.sh + felis setup); the build egress lock opens exactly the control
namespace on the internal port.
- cmd/felis: fetch-context entrypoint (registered, documented, unit-tested for
escapes/symlinks/non-gzip).
Tests cover rendering, hardening, the s3/local Open paths, and the route's
404/503 mapping. Verified next on the real single-node cluster with Kaniko.
The api, operator and registry Deployments shipped with no liveness/readiness
probes at all: a wedged process stayed 'Running' forever, and the operator had
no health listener to probe in the first place. Kaniko build evidence on a
fresh install showed the only cluster-wide red after a disk-pressure pass was
Deployment status that never reflected health.
- felis-api: readiness /readyz (DB + K8s API round-trip) and liveness /healthz
on the internal face (:8081), the only listener that serves both endpoints;
liveness deliberately avoids /readyz so a DB blip cannot restart the api.
- felis-operator: new --health-probe-bind-address (:8081) with controller-
runtime's /healthz + /readyz (registered ping checks; an unregistered handler
map would 404), plus the matching container port and probes.
- registry: /v2/ probes on the pinned port, so a broken storage backend stops
reading as 'Running'.
Tests pin paths, ports, and that each probe targets a declared container port.
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
build_mem_limit overrides; empty keeps the compiled-in defaults. An
air-gapped or mirrored install has no route to gcr.io/aquasec (the
build egress policy allows only DNS + registry + package mirrors), so
builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
impossible (PVCs cannot cross namespaces) and that the s3 lane also
lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
13b rewritten (verified eviction refusal, 5m pressure-transition,
image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
volume (worlds + config + plugins + cache) and a restore rolls all of
it back — OpenAPI/README wording updated to match (same-tag images are
still watched for regressions by the openapi parity gate).
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.
Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
Three faces of one gap, all on the supported install path:
- Backup/restore answered 503 out of the box: nothing ever rendered the
archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
'felis manifests' names it by default (--backup-pvc= is the explicit
no-store shape); bootstrap passes it through so the generated felis.toml
[archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
nothing deleted).
- Even when enabled it could not find a world on a stock install:
resolveWorldDir now also resolves the exact local-path directory
<pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
glob, so a stale deleted PV's bytes can't be archived in place of the
current world). Reaper Role gains persistentvolumeclaims:get (weaker
than the delete it already held).
README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
The setup URL carries a 43-char token and overruns 80 columns; the TUI
renderer clipped it. Break it at the query '=' boundary (token on its own
line) with a shared wrapDisplayURL helper used by both the Owner wizard and
the mc-bind wizard; unit test pins the no-loss concatenation.
Without SetLogger, the first reconcile prints
'[controller-runtime] log.SetLogger(...) was never called; logs will not be
displayed' followed by a full stack trace (live-observed in felis-operator).
Route it through logr.FromSlogHandler(slog.Default()) so its messages are
ordinary stderr lines; go-logr/logr promoted to a direct dependency.
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.
The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
Nine files had drifted from gofmt and nothing checked; nine staticcheck
findings were live (three dead symbols, capitalization, a redundant
Sprintf, two literal-to-conversion sites, a nil test context). Fix all
of them and make CI fail on unformatted Go so this cannot re-drift.
Four messages that reach an operator or an API client cited sections
of a specification nobody outside the project can read:
- the unimplemented archive store error from config load
- the running-server cap refusal, from both the user wake and the
internal wake
- the missing memory ceiling guard, in the API and in felis apply
The references are gone and the wording is otherwise unchanged. Each
message still says what went wrong and, where there is one, what to
do about it. The test for the archive store message checks for the
tarLocal remediation, which is still there.
felis nano serves the same hasJoined handler as felis api, but behind
nanoStubRepo, which implements only the bar-list lookup and embeds a nil
Repo for everything else. Only the full-api path was tested, against a
complete fake store, so a second store call added to handleHasJoined
would pass CI and panic on every nano login. The loopback default of
-listen, the one thing keeping nano from being an open auth relay, was
not pinned either.
The default moves into a nanoDefaultListen constant, and two tests
cover the path. One serves a login through api.HasJoinedHandler with
nanoStubRepo and a fake identity source and expects the profile back.
The other requires the default to parse as a loopback IP. Taking the
bar-list method off the stub makes the first panic on the nil Repo;
defaulting to 0.0.0.0:8081 or :8081 fails the second.
LoadNano filled in [server] listen = "0.0.0.0:8080" when it was unset,
and a test pinned that value, but felis nano never reads it: it binds
the -listen flag, which the installer sets from FELIS_NANO_LISTEN. An
operator moving nano off loopback by writing [server] listen in its
config got connection refused from the proxy and no hint that the key
did nothing.
LoadNano no longer sets the default, and nano prints a line naming the
ignored value and the address it actually binds whenever the key is
set. It is a warning rather than a load error so a full felis.toml
copied onto a nano host keeps starting. The assertion that pinned the
unused default is removed along with it.
The new test runs cmdNano against a config that sets [server] listen
and one that does not, with an unbindable -listen so it returns after
loading. The first must warn and the second must not; with the old
default restored, the second prints a warning about 0.0.0.0:8080.
The installer and the config template tell the operator to run
systemctl restart felis-nano after editing the source list. nano had no
signal handling, so SIGTERM killed it mid-request: a login waiting on an
upstream had its connection reset, and Velocity disconnected that
player with "authentication servers are down". felis api already drains
on shutdown; nano did not.
nano now listens itself, serves until SIGINT or SIGTERM, then shuts the
server down gracefully with a 30-second limit. That outlasts the source
scan of any realistic list, at five seconds per source, and stays well
inside systemd's default 90-second stop timeout.
The new test holds a request inside the handler, cancels the serve
context, and checks that serveNano is still running 200 ms later, that
the held request then gets its answer, and that serveNano returns 0.
Replacing the graceful shutdown with Close fails it.
felis nano logged every request with the raw RequestURI and %s. That
text is the caller's: a right-to-left override reordered the line as
displayed, an invalid UTF-8 byte made journald store the entry as a
binary blob that journalctl -f shows as "[N blob data]", and a query
near net/http's one-megabyte limit became a one-megabyte log line.
The URI is now capped at 256 bytes, several times a real hasJoined
query, and printed with %q, so control, bidi and invalid bytes appear
escaped. The handler assembly moved into nanoHandler so the logged
handler can be tested on its own; cmdNano serves it unchanged.
The new test sends a query carrying U+202E, a 0x9b byte and 4 KiB of
padding, and expects a valid UTF-8 line with the override escaped and
no more than twice the cap. Restoring the old unquoted line fails it.
Seventeen comments opened with a tag naming the tool that wrote them.
The tag goes and each comment keeps its reasoning, now starting as a
plain sentence. None of the reasoning changes.
The AGENTS.md note in .gitignore drops the story of how the file got
into the tree and keeps the one fact a reader needs: its advice to run
go fmt is destructive on this CRLF working tree.
Comments only; no code, build or test changes.