cooldownLimiter began as the wake-only throttle; the OTP-start hardening
reused it via the atomic reserve/release. Its type comment still called it
a per-server wake limiter and justified the per-replica behaviour as
"acceptable because the operator reconcile is idempotent" -- true for wake,
false for OTP, whose every admitted send is a non-idempotent email.
Rewrite the comment to describe the shared per-key limiter and record the
honest KNOWN-LIMITATION: the atomic reserve/release closes the
intra-replica concurrent burst, but the in-memory map throttles per
replica, so cross-replica bounding still needs a shared store. No
behaviour change.
The email-OTP resend cooldown checked the window with a peek (allowed)
and only recorded it after delivery. For OTP that throttle is the sole
defense and each admitted send is a real, non-idempotent email, so a
burst of truly concurrent starts all passed the peek before any recorded
and every one mailed: N concurrent starts bombed a mailbox with N codes.
Add an atomic reserve/release pair to cooldownLimiter: reserve checks and
records the window in one critical section under the mutex, so a
concurrent burst yields exactly one winner; release rolls a reservation
back only if it is still the current one, so a slow failing caller never
clobbers a newer holder. handleEmailOTPStart now reserves both the
principal and the recipient key up front and defers a rollback that frees
both windows on any mint, create, or delivery error — preserving the old
"a failed send does not consume the cooldown" property, now race-free.
The wake path keeps allowed→record: its real gate is the running cap and
its side effect (SetDesiredState) is idempotent, so the peek gap is
harmless there.
Tests: a frozen-clock gate-mailer fires 8 concurrent starts for one
victim from one principal and asserts exactly one mail and one 202; a
flaky-mailer test proves a failed delivery releases the window so an
immediate retry in the same instant is admitted.
A Running, ready server always reported 0/0 players: markRunningReady
never wrote Status.Players, and markStopped only cleared it. The panel
therefore showed an empty tally for live servers.
Extend the readiness probe to also sample the player count. Prober.Probe
now returns a PlayerCount{Online, Max}: RconProber still gates readiness
on Dial+auth, then runs a best-effort `list` and parses the vanilla
reply ("There are N of a max of M players online"). A failed or
unparseable tally is swallowed (0/0) so it never blocks readiness. The
reconciler threads the count into markRunningReady, which writes
Status.Players; markStopped still resets it to zero.
handleEmailOTPStart minted and mailed a code on every call, so an
authenticated caller could drive unbounded mail to any address they
typed — an email-bomb primitive against arbitrary mailboxes.
Add a separate otpLimiter (its own sync.Once and map, distinct from the
wake limiter) and throttle each send on two keys before anything is
minted: the caller (user:<id>) and the recipient (email:<lower>). A
refused send mints no code and mails nothing; both cooldowns are
recorded only after delivery succeeds, mirroring the wake path so a
failed mint or delivery never consumes the throttle. The two-key design
stops both one account fanning out across addresses and many accounts
converging on one mailbox.
A wake refused by the §9.1 running-server cap returns 503, but the
per-server cooldown was recorded before the cap check ran. A player
held because the cluster was momentarily full would then also have to
wait out the wake cooldown once a slot freed, even though their refused
wake never actually flipped desiredState.
Split cooldownLimiter.allow into allowed (peek, no record) and record
(commit). Both wake paths now consult allowed for the 429, then call
record only after SetDesiredState succeeds — so neither a 503
at_capacity nor a SetDesiredState error consumes the cooldown. The
split is safe against the running cap, which counts CRD truth via
ListServers and is independent of the limiter.
`docker build` failed twice over because .dockerignore excluded two trees the
image actually needs. The panel stage's `COPY panel/ ./` hit `"/panel": not
found`, and even past that the Go build would fail: the root felis package
//go:embeds deploy/bootstrap.sh and deploy/crd/*.yaml, which `COPY . .` dropped
along with the excluded deploy/.
The stale header comment claimed only internal/store/migrations was embedded,
which is what licensed the over-broad exclusions. Rewrite it to name all three
embedded trees (migrations, deploy assets, panel static) and warn against
re-adding panel/, deploy/, or internal/ without re-checking the go:embed list.
Tighten node_modules -> **/node_modules so a working-tree build no longer drags
panel/node_modules over the Linux modules npm ci installs in the panel stage.
When a staff account already exists, the break-glass console now opens on a
thin top-level menu (menuModel) where account operations are peers rather than
tails of one wizard: provision/reset the Owner, or add an Operator. A fresh
machine with no Owner skips the menu and goes straight to Owner bootstrap, since
minting an Operator first would create a staff account the login gate rejects.
The Operator path reuses ownerModel via a bgOperation discriminator. It is
insert-only (performAddOperator -> InsertOperator), wraps a duplicate username as
api.ErrConflict and routes back to the provision form for a retry rather than
tearing down, and deliberately never flips the global local_auth toggle the way
the Owner thread does. The post-exit summary and audit trail distinguish the two
outcomes (isOperator); only the Owner provision claims local-password login was
enabled.
Tests cover the operator-model defaults, path selection (insert vs upsert and
the local-auth gate), conflict-retry versus generic teardown, isOperator
propagation, and the root menu routing for both fresh and admin-present
machines.
Add docs/troubleshooting.md covering the common failure modes the spec
implies, grounded in the actual control-plane code paths:
- Stuck Starting (PodNotReady / RconSecretUnavailable / RconNotReachable)
and the deliberate absence of a Starting->Failed timeout.
- Failed reachable only via InvalidSpec on a malformed spec.storage.size,
plus the stale status.endpoint=direct caveat after a failure.
- Routing via status.endpoint direct/fallback and the empty fallbackServer
pitfall; wake 403/429/503 gate order.
- online-mode coupling and Velocity's offline-mode routing refusal.
- Cloudflare Access 401/403, nil-Keyfunc fail-closed, audience checks,
the absence of an issuer check, and local-session gating.
- Internal service-token (FELIS_SERVICE_TOKEN) rejection path.
- link/claim error codes, Kaniko build denials (SA-by-absence RBAC,
default-deny egress, internal-registry push gate), and the registry
DNS contract.
- Reaper backup-before-delete invariant and false-delete vectors.
- Unimplemented idle auto-stop, permanently-zero players.online, the
inert CRD fields, and the always-survives world PVC behaviour.
Each item is labelled with its evidence grade (GO-TESTED / CODE-ONLY /
INTEGRATION-ONLY / INERT) so operators know what is verified versus
asserted.
Adversarial cross-check of the three §28 diagrams against the wake, claim
and link code paths surfaced two fidelity drifts:
- The wake/status/join-event lanes used abbreviated /internal/... paths;
the registered internal-face routes carry the /api/v1 prefix (api.go),
matching the convention the claim and link diagrams already use.
- The /link diagram showed the game posting {mc_uuid, auth_source}, but no
shipped in-game caller sends auth_source — LinkClient posts {mc_uuid}
and the API defaults auth_source to mojang server-side.
The claim diagram already matched the code (verbatim atomic UPDATE,
404/409/200 mapping) and is unchanged.
Wire the fourth mandated §23 metric to a real producer. The histogram
spans two reconcile passes, so anchor and observation must persist in
status:
- Add status.startRequestedAt, set once on the first Starting reconcile
of a start attempt and cleared on Stopped so the next start re-anchors.
- Observe felis_start_duration_seconds exactly when readiness is first
reached (ReadySignalAt - StartRequestedAt), guarded so a server that
reaches ready without a Starting pass records nothing.
- Mirror the field into the deepcopy and the structural CRD schema so the
apiserver does not prune it on patchStatus round-trips.
- Promote prometheus/client_golang and client_model to direct deps now
that the operator and its tests import them.
Tests drive a step clock through Starting -> Running asserting the exact
observed duration, and through Running -> Stopped asserting the metric is
observed once and the anchor clears.
A per-object reconcile cannot maintain felis_servers_total (spec §23): it
sees one server per call, so it could never Set a correct fleet-wide gauge
and inc/dec on transitions would drift on any missed event. Add a snapshot
producer instead.
metrics.SyncServerGauge Resets the GaugeVec then Sets one child per state,
so a state that drains to zero reports 0 rather than a stale last value.
operator.GaugeSyncer is a manager.Runnable that periodically Lists the
fleet and republishes from it, defaulting an unset desiredState to Stopped.
SyncOnce is exercised end-to-end against a fake client (List, default,
republish); the ticker loop in Start is the only untested I/O edge.
Wire the build subsystem to the felis_image_build_failures_total counter
(spec §23). It advances at the two terminal-failure producers: finishAt
(the Sync JobFailed/JobUnknown verdict — a kaniko failure or a CRITICAL
CVE from trivy's --exit-code 1) and Submit's job-creation bypass path,
which records its failure directly without going through finishAt.
Cancellations and successful builds are deliberately not counted.
A delta-asserting test exercises both Inc sites plus a successful-build
negative control that proves the StatusFailed guard discriminates rather
than firing on every terminal write, all over the existing in-memory
Store/Jobs fakes.
Introduce internal/metrics exposing the four metric families spec §23
mandates at minimum: felis_servers_total (gauge by desired state),
felis_start_duration_seconds (histogram with Minecraft cold-start
buckets), felis_image_build_failures_total and
felis_reaper_worlds_deleted_total (counters). Collectors are
package-level vars so any subsystem records without an import cycle;
Register wires them into a prometheus.Registerer and is idempotent.
Wire registration into the operator against controller-runtime's global
Registry, so /metrics on the manager's existing metrics endpoint carries
the felis_* families. Instrument the reaper to increment
felis_reaper_worlds_deleted_total in lockstep with Summary.WorldsReaped,
at the one point a world's PVC has actually been deleted.
Build the React consumer over the fail-closed view-mode logic so an admin
can view the app as each of the three homes (User/Admin/SysAdmin) and step
down to a plain User-Side home.
- ViewModeProvider holds the raw requested home (seeded from localStorage,
shape-checked only) and resolves it live against is_admin on every render,
so a demotion or transient /me failure collapses to the User home with no
flash, while an unentitled value is never stored or applied.
- RoleSwitcher renders only for admins (availableViewModes > 1); switching
re-gates the choice and navigates to the chosen home's root.
- AppShell drives its sidebar from sectionsForView(view, isAdmin), which only
ever narrows visibleSections — an admin viewing as a user sees a plain
user's sidebar and lands on the Dashboard at /.
- landingPathForView / viewModeLabelKey added to the logic layer (tested);
view_* and view_switch_label i18n keys added for en-US and zh-CN.
Placement note: the spec calls for a top-right avatar control, but the panel
has no desktop top bar, so the switcher lives in the sidebar foot beside the
user strip. Functionally complete; placement is not yet spec-parity.
Pure logic layer for the top-right avatar role-switcher: derive the home
a principal is in and may switch into, mirroring nav.ts/auth.ts so the
decision is unit-tested without a React renderer.
- ViewMode is derived from NavSection["id"], so the three switchable
homes (User/Admin/SysAdmin) can never drift from the nav sections.
- availableViewModes / effectiveViewMode resolve a requested view against
the live is_admin flag, failing closed: a non-admin or a demoted admin
always collapses to the User home.
- restoreViewMode re-gates a persisted (localStorage) choice on every
read, never trusting the stored value over the live flag, closing the
one escalation vector a client-side persona could open.
- sectionsForView composes the view ceiling on top of visibleSections, so
the switcher only ever narrows the sidebar, never widens access.
The avatar dropdown UI that consumes this lands as a separate increment.
Add the insert-only Operator-creation path to the break-glass console
(felis breakGlass). An Operator is an additional staff admin: role=admin
with must_change_password=true, identical in shape to the Owner, since
Felis has no separate operator DB role (migration 0003).
Unlike the Owner upsert, provisioning is insert-only -- a username already
taken returns ErrConflict (ON CONFLICT DO NOTHING + zero RowsAffected)
rather than silently resetting a live account, so adding an Operator can
never clobber the Owner's or another Operator's credential. A typed
password is used as-is; an empty one is replaced with a generated
one-time credential returned for display. Operator-add does not touch
local_auth_enabled -- that global gate belongs to the Owner thread alone.
Accountability is recorded best-effort under a break_glass.operator_create
audit action, written only after a successful provision.
The TUI menu router that reaches this path is deferred; this lands the
fully unit-testable logic layer (provisionOperator, performAddOperator,
auditAddOperator) with the PGRepo insert kept integration-only.
QR scan-to-login is a device-code grant where the QR encodes the existing
short-lived account-link code (spec §B3 player game-login). velocity mints a
code in-game, renders it as a QR, the player scans it on a phone already signed
in to the panel, and that web session's verify writes the durable account_links
row bound to that user. The only new verifiable surface that flow needs is the
completion poll velocity calls to learn the link landed and admit the player.
Add GET /api/v1/internal/account/link/status/{mc_uuid}: a read-only, internal
handleLinkStatus keyed by the verified mc_uuid velocity already holds. It reuses
the existing UserByMCUUID, so it adds no migration and no mutation to the
load-bearing VerifyLinkCode; ErrNotFound maps to {linked:false} (pending /
not-yet-scanned), a hit to {linked:true, user_id}. Keying on the public UUID and
not the scanned code means the read carries no guessing surface and needs no
attempt cap — the internal face already gates it to service callers, and the poll
consumes nothing so a velocity restart re-polls safely.
QR render, limbo collision routing, in-game admit, and the reclaim
inherit-disambiguation stay CODE-ONLY (Java/Velocity) and are labeled as such;
this endpoint reports link completion only.
Document the route in openapi.yaml (x-felis-face internal, x-felis-tier service)
so the parity gate holds, and cover it with a hermetic vertical that proves the
poll reflects the durable link only after the external verify and binds the
verifier's id, plus unknown-uuid, idempotency, and internal-only face separation.
- Convert CRLF to LF across Go, panel, and plugin files
- Add Cloudflare API token template URL to breakGlass TUI edge intro
- Verify API token in cfsetup before creating tunnel, DNS, or Access app
Add an optional edge-setup flow to the `felis breakGlass` sudo TUI,
reachable as an independent peer of Owner provisioning through a new
top-level menu (so reaching it never forces an Owner password reset).
The flow drives the operator's own Cloudflare consent (interactive
`cloudflared tunnel login`, suspending the alt-screen, plus an API
token) and then calls cfsetup to stand up a Tunnel routing the
configured admin and panel hosts and a fail-closed Access application.
It stays gated shut unless an admin hostname is configured and the
operator is logged in (edgeReady), and refuses empty or bare-domain
credentials before any side effect. On success the TUI surfaces the
issued Access aud and an explicit ACTION REQUIRED note; it never edits
felis.toml. The live cloudflared and Cloudflare API calls are
integration-only and exercised against a real account.
Add internal/cfsetup, the verifiable core of an optional one-click
Cloudflare Tunnel + Access provisioning flow for the SysAdmin edge
(spec §14). It is domain-agnostic (every FQDN is composed from the
configured root_domain) and IdP-agnostic (any valid Access JWT aud is
accepted, whichever IdP fronts it), so a SysAdmin who brings their own
domain or Zero-Trust scheme stays fully supported.
The load-bearing safety property is a fail-closed guard on the
recommended Access policy. validateFailClosed is an allowlist that
refuses any policy that could be public: a bypass/non-allow decision, an
empty include, an "everyone" include not narrowed by a constraining
require (include rules are OR, so "everyone" beside an identity is still
public), or any include rule it cannot positively recognize as a scoped
identity. Setup runs the guard before any side effect, so a public
policy aborts the run with nothing created.
The tunnel ingress routes only the web hostnames to the local panel
origin and terminates in the mandatory fail-shut 404 catch-all; the raw
game host is never proxied. Gating preconditions (cloudflared present,
tunnel login completed, API token) are hard checks with no side effects
on failure.
The actual cloudflared exec, DNS routing, and Access API calls live in
runner.go and are integration-only: they require the operator's own live
Cloudflare account and interactive browser consent, which cannot be
unit-tested. The policy guard, ingress generation, request bodies, and
gating are unit-tested.