handlePasskeyRegisterFinish logged an empty target for account.passkey.registered, while the delete half logs the credential id. An operator auditing the log could see that a passkey was bound but not which one. Pass cred.ID as the audit target so bind and unbind are symmetric, and tighten the enrollment test to assert both halves name the credential id.
The MyServers query lists both a user's own servers and unclaimed (owner_id IS NULL) servers, but computed owned as s.owner_id = $1. For an ownerless row that comparison is SQL NULL, which fails to scan into the Go bool and 500s the whole listing. Wrap it in COALESCE(..., false) so an ownerless row reports owned=false while still surfacing as claimable.
The per-write deadline that severs a stalled SSE reader was never cleared on
return. Server.WriteTimeout is deliberately unset -- a WriteTimeout would sever
a healthy long-lived stream -- and with it unset net/http never resets the
connection write deadline between keep-alive requests. So the deadline the last
writeChunk left set leaks onto the next request that reuses the pooled
connection and fails its first write for no reason. Clear it to the zero value
on return via a deferred rc.SetWriteDeadline; best-effort, a no-op on writers
without deadline support.
Also record honestly at the header flush that the connect-time stall stays
bounded only by the per-principal stream cap, not severed by this guard -- only
the mid-stream stall is closed. Adds a test pinning the clear (fails closed:
neutering the deferred clear leaves a +writeTimeout deadline set on return).
QuotaAvailable and ClaimServer run as two separate statements, so the
count read is not serialized against a concurrent claim's UPDATE: two
claims by one user for two different ownerless servers can both pass the
gate and both succeed, leaving the user one server over quota. It is low
severity — quota over-provisioning under a deliberate burst, not an
authorization, ownership, or isolation break, since each server is still
claimed atomically via UPDATE ... WHERE owner_id IS NULL.
Closing it requires Postgres transaction semantics (advisory-xact-lock on
the user, or SERIALIZABLE with retry) folding the gate into a single repo
method — verifiable only against a real Postgres, not the hermetic
fakeRepo suite. Documented at QuotaAvailable with back-references from the
two claim gates (handleClaim and the internal UUID claim) rather than
patched blind.
relayLogStream copied a pod-log follow to the client with a plain
flusher.Flush per event. On a client that stays connected but stops
reading (its TCP receive window shut), net/http buffers the small
"data:" line and only touches the socket at Flush, which then blocks
forever inside the write. The select's <-ctx.Done() branch is never
reached, because r.Context() cancels on an actual disconnect, not on a
stall, so the relay goroutine and its upstream apiserver follow leak for
the life of the process.
Route every event's write+flush through http.ResponseController with a
per-write deadline (writeTimeout, 30s): a stalled flush now returns
os.ErrDeadlineExceeded, the error plain http.Flusher.Flush swallows, and
the relay abandons the stream so the deferred cancel + src.Close release
the follow. SetWriteDeadline and rc.Flush are best-effort: a writer
without deadline support (httptest recorder; some HTTP/2 origins) ignores
the deadline and behaves exactly as before, so the guard degrades
gracefully.
This closes the leak the per-principal stream cap only bounded the blast
radius of. Verified by a deterministic test with a deadline-aware
ResponseWriter whose flush blocks until the deadline; the test times out
(fails closed) if the guard is removed.
Console and build-log relays hold a Server-Sent Event connection open for the
life of a client's attachment; a stalled reader pins the relay goroutine plus
its upstream kube-apiserver follow. Without a bound, one authenticated
principal could open these repeatedly and accumulate leaked control-plane
connections.
Add a per-principal stream cap (streamLimiter) enforced before either relay
opens its follow stream, returning 429 too_many_streams past the limit.
cmd/felis wires it to 16; zero disables it, matching the "zero disables"
idiom of the other levers.
This bounds the blast radius of the stalled-stream leak; it does not close the
leak itself -- the per-write deadline that severs a stalled stream is a
separate change.
The three felis-api http.Servers (internal, external, https) were built with
only Addr and Handler, leaving ReadHeaderTimeout, IdleTimeout, and ReadTimeout
at zero. A zero ReadHeaderTimeout is a Slowloris hole — a client trickling
header bytes pins a connection indefinitely — and a zero IdleTimeout lets
kept-alive connections accumulate (gosec G112).
Route all three listeners through a newAPIServer factory that sets a 10s
ReadHeaderTimeout and a 120s IdleTimeout. WriteTimeout and ReadTimeout are
left unset on purpose: the external and https faces stream Server-Sent Events
(console / build logs) for the lifetime of a client attachment, and a
WriteTimeout would sever a healthy long-lived stream. Slowloris is closed by
ReadHeaderTimeout, which bounds only the header phase.
withRequestID honored any inbound X-Request-Id verbatim, and that value is
echoed on the response, embedded in the error envelope, and persisted into
audit_logs.request_id. An unvalidated caller-supplied id is therefore an
audit-integrity vector: an arbitrarily long value bloats the audit row, and a
stray control byte (CR/LF) could smuggle a forged entry into a log sink.
Accept an inbound id only when it is well-formed — non-empty, at most 64
bytes, and restricted to a log-safe charset ([A-Za-z0-9._-]) — otherwise mint
a fresh server id. A rejected request loses its inbound trace link, which is
strictly better than storing attacker-controlled text in the audit trail.
The public /auth/login route runs a full-cost bcrypt compare on every
request — including the anti-enumeration dummy-hash compare for an unknown
user — with no bound on how many run at once. A flood of concurrent logins
therefore pins every core in bcrypt, starving the rest of the API.
Cap the simultaneous compares with a small non-blocking concurrency limiter
(a buffered-channel semaphore): a login that cannot take a slot is shed with
429 auth_busy before the compare, rather than piling more work onto the
scheduler. The slot guards only the hash and is released the instant the
compare returns. It is a concurrency cap, not a per-account lockout, so it
never fences out the one admin trying to break-glass in, and the 429 lands
before any credential distinction so it leaks nothing about the username.
The cap follows the existing "zero disables" lever idiom (WakeCooldown,
MaxRunningServers); cmd/felis wires it to the core count (floored at 4).
The passkey login/assertion HTTP handler stays deferred after its design
checkpoint; capture the reasoning in the handler header so the decision is
durable in the repo rather than only in task notes.
- RP boundary (resolved): felis-api is the app-login relying party (panel.*);
the WebAuthn security gate lives at the Cloudflare Access edge. Spec §14 ties
WebAuthn/posture to admin.* (Access) while panel.* is plain app login, so
there is neither a spec-required assertion handler nor a backend step-up
consumer for one.
- Identifier (blocking): a from-zero login needs a unique, human-typable handle
to resolve an account, but users.email is nullable and non-unique and a
player's username is their Minecraft uuid. Username-first assertion has
nothing to key on; re-link stays the returning-player door.
Discoverable (usernameless) credentials are the future enabler; the adapter
crypto is already verified so that slice inherits correct crypto.
Build the assertion (login) half of the WebAuthn ceremony crypto in the
internal/passkey adapter, Oracle-verified against a virtual authenticator.
- BeginLogin/FinishLogin over go-webauthn BeginLogin/ValidateLogin,
username-first (allowCredentials scoped to the known user's bound
passkeys). Discoverable/usernameless login stays out of scope: the
enrolled credentials are non-resident and the challenge store is
user-keyed (migration 0007), so it would need a future migration.
- WebAuthnCredentials() now populates the stored COSE public key and
signature counter (assertion validation needs both to verify the
signature and detect clones); enrollment ignores them, so the change
is backward-compatible and the enrollment tests guard it.
- VerifiedAssertion seam output: which credential signed plus the raw
signature counter. Clone/regression policy is deliberately NOT here —
the counter is a ceremony fact and the future handler, which holds the
previously stored counter, decides reject/warn.
Scope: crypto adapter only. The login HTTP handlers, session minting,
and the panel.* relying-party boundary/tier decision remain a deferred
slice (no unauthenticated login route is added). BeginLogin/FinishLogin
live on the concrete adapter, not the api.PasskeyVerifier interface,
which grows only when a handler consumes them.
Tests (virtualwebauthn): a real enrollment chained into a real assertion
exercises the COSE public-key decode path and surfaces the advanced
signature counter, plus origin-mismatch and unbound-credential rejection.
The admin API persists the auto-update maintenance window as lowercase
JSON {"start","end"} (platform_settings key "update_window"), but
updates.Window had no json tags, so it marshaled/unmarshaled with
capitalized keys. The natural decode the update runner will use --
json.Unmarshal(stored, &updates.Window{}) -- would therefore miss every
key and silently yield the zero Window. That fails closed (a zero window
Contains nothing, so notify-only, never a rogue apply), so it is safe but
a latent silent-zero trap for the not-yet-built runner.
Add json:"start"/json:"end" to updates.Window so the obvious decode is
correct by construction; value time.Time treats a stored null as a no-op,
so a cleared/never-set window still decodes to the zero Window. Nothing
in the package serialized Window before, so this changes no existing
behavior.
Guarded by a cross-package contract test in internal/api that marshals
the real api.updateWindow DTO and unmarshals it into updates.Window --
asserting the interval survives (Contains(mid) is true) and that an empty
window decodes to the fail-closed zero Window -- so the two shapes cannot
drift apart silently.
Two admin-tier routes read and set a single platform-wide maintenance
window for the auto-update subsystem (decision core internal/updates):
GET /api/v1/updates/window
PUT /api/v1/updates/window
The window is stored as JSON {"start","end"} (RFC3339, or null when
unset) under the platform_settings key "update_window", reusing the
existing GetSetting/SetSetting KV seam -- no new Repo method, no
migration. Pointer times keep "unset" (null) distinct from a real
instant on both decode and encode; a never-set and an explicitly
cleared window both read back as {null,null}.
Validation mirrors the core's fail-closed Window: a window is either
fully set (both ends, end strictly after start) or fully cleared (both
null). A half-set, inverted, or empty-interval body is 400 and is never
persisted. Reads treat only a missing key as unset (ErrNotFound -> 200
nulls); any other store error 500s rather than fail open.
This is API + PERSISTENCE ONLY. Nothing consumes the stored window yet
-- the runner, the ReleaseSource/Notifier/Applier executors, and the
scheduler CronJob remain INTEGRATION-ONLY. Setting a window changes no
behavior until those land; it is the durable input they will read.
Nothing here force-updates ("不要强制自动更新").
Adds POST /api/v1/auth/bind, the one public pre-account entrypoint of the
player console (console.<root_domain>). An account-less player redeems the
one-time Bind Code minted in the in-game Login Lobby; in a single step the
platform creates a role=user player, links it to the verified in-game UUID,
and mints a host-only felis_session. Login is thus not forced at the edge
while operations stay app-authenticated.
The operator console (op.console.<root_domain>) is unaffected and stays
behind Zero Trust: a code whose UUID resolves to a staff (role=admin)
account is refused with 403 (ErrPlayerBindForbidden) without consuming the
code, so the public door provably never yields an admin principal — the
session it mints carries ViaAdminAccess=false and is host-only to console,
never sent to op.console.
Repo layer: new RedeemPlayerBindCode on the Repo interface, implemented on
PGRepo (single tx: resolve code, create-or-fetch the player, consume) and
the test fake. The returning-player branch is idempotent and is a deliberate
standing "log in via the game" door, not just first-time onboarding.
Honest labeling:
- ORACLE-VERIFIED (Go): account/session logic — role=user, refuse-staff,
idempotent create-or-fetch, single-use code, and the op.console redline
(player session rejected on admin routes). Covered by handlers_onboard_test
and the OpenAPI parity gate.
- INTEGRATION-dependent: the endpoint's security rests on the Bind Code having
been minted against an online-mode-Yggdrasil-authenticated UUID, a
precondition that lives in velocity/Java and is not verifiable from this
repo (CODE-ONLY). The Go layer proves the logic, not that identity guarantee.
- No app-level attempt cap: rate-limiting is deferred to the edge as for the
public /auth/login; the ~1e12 keyspace, single use and short TTL make a
blind app-level cap non-critical.
Introduce internal/updates: a pure, I/O-free engine that decides what
should happen to each tracked platform component (Felis control-plane,
k3s, cloudflared, Velocity) given its current version, the latest
discovered upstream, its policy, and the current time.
Updates are never force-applied. A component is Pinned (Minecraft, left
alone), Notify (a SysAdmin is told and applies out of band), or Scheduled
(Felis may apply, but only inside a maintenance window the SysAdmin set).
The load-bearing invariants are unit-tested: a pinned component never
changes, a downgrade is never proposed, a prerelease is never
auto-applied, and an apply happens only inside the window.
Version parsing tolerates the real feeds (leading v, k3s +k3s1 build
suffix, calendar versions, prerelease tails) and orders by SemVer
precedence. ReleaseSource/Notifier/Applier are declared as integration
seams and exercised via fakes; this package ships no network, SMTP, or
kubectl, and deliberately has no blind k3s-upgrade applier.
After the Cloudflare tunnel connector is installed and the origin has rolled
out, applyCloudflareEdge now fences the panel NodePort so the origin is
reachable only over loopback -- the hop the host-side connector uses -- and
never from a public interface. This closes the Access-bypass hole where a direct
https://<node-ip>:<nodeport>/ with the right Host header reached the origin
behind Cloudflare Access.
The fence is an nftables table hooked at prerouting priority -300 (raw), before
kube-proxy's NodePort DNAT (dstnat, -100), so it catches the packet on its
original destination port; a filter/INPUT rule would miss the DNAT'd, then
FORWARDed NodePort packet. Loopback is accepted first, so the connector origin
hop is untouched; the inet family fences a public IPv6 NodePort too.
It is gated on the connector actually serving (verifyConnectorServing polls
`cloudflared tunnel info`): fencing a dead tunnel would sever the only web path
to a still-up origin. If serving cannot be confirmed the port is left open (its
pre-tunnel state) and the failure is surfaced loudly. unfenceOriginNodePort is
the on-host break-glass reversal. The nft/cloudflared calls are INTEGRATION-ONLY;
the ruleset shape and the conn-count gate are pure and unit-tested.
KNOWN-LIMITATION: targets nftables; firewalld-native coordination is not yet
handled (a firewalld reload can flush the standalone table).
Setup previously called runner.StartConnector (`cloudflared service install`)
while the TUI applyCloudflareEdge separately installs cloudflared-felis.service
for the same tunnel from the same config -- two managed services serving one
tunnel from one connector config.
Drop StartConnector from cfsetup: running a connector is a host-specific side
effect (systemd/launchd/Windows service) that belongs with the caller, not in
this host- and domain-agnostic package whose documented side effects are tunnel
creation, DNS routing, and the Access app/policy calls. installCloudflaredService
in the host layer stays the single connector installer, so the routed-but-dead
1033 is still closed; RouteDNS --overwrite-dns still closes the stale-DNS 1033.
Setup created the tunnel, routed DNS, and wrote config.yml, but nothing
installed or started a connector for it. A one-click run therefore left the
tunnel routed-but-dead: every web hostname returned Cloudflare error 1033
(tunnel has no connector) even though the config was correct on disk.
Add a StartConnector step to the Runner seam, invoked right after the config
is written (and gated on ConfigPath, so a caller wanting only the Access
config is not forced to install a service). The ExecRunner implementation
runs `cloudflared --config <path> service install`, which installs and starts
a managed system service (systemd/launchd/Windows), and is idempotent on an
already-installed service. The orchestration — connector started, and only
after its config exists — is unit-tested against the fake Runner; the actual
service install is INTEGRATION-ONLY.
Together with the RouteDNS --overwrite-dns fix, this closes both distinct
paths to a 1033 half-state from a fresh setup: a stale DNS binding and a
missing connector.
RouteDNS ran `cloudflared tunnel route dns` without --overwrite-dns and
swallowed the resulting "record already exists" error as success. When a
hostname already had a CNAME from an earlier tunnel that was deleted and
recreated, the record stayed bound to the dead tunnel: the setup reported
the hostname "routed" while it kept returning Cloudflare error 1033 (the
tunnel it pointed at has no connector).
Pass --overwrite-dns so the record is repointed at the tunnel just created,
making the route idempotent and correct on every re-run, and drop the
now-unnecessary "already exists" swallow. INTEGRATION-ONLY (ExecRunner
shells out to the real cloudflared binary).
Construct the go-webauthn verifier at the composition root and attach
it to the API when auth.panel_hostname is configured (RP id = panel
hostname, origin = https://<panel hostname>, display name Felis). When
the hostname is unset or the verifier fails to build it stays nil and
the passkey ceremony routes report 503, matching the existing
nil-when-unconfigured subsystem pattern. An admin passkey, if ever
added, is a separate relying party on the admin host and is
intentionally not wired here.
Wrap github.com/go-webauthn/webauthn behind the api.PasskeyVerifier
seam so the api package stays free of go-webauthn types. The adapter
covers the credential-creation ceremony only (BeginRegistration /
CreateCredential); the login/assertion path is a deferred slice.
Ceremony state crosses the seam as opaque marshaled SessionData, the
attestation as an io.Reader, and the verified result as a plain
VerifiedCredential. SessionData carries no expiry so the challenge
row's TTL stays the single liveness authority. New rejects an empty
RP id or origin list so a misconfigured deployment fails at
construction rather than minting unverifiable challenges.
Tests drive a real relying party against a virtual authenticator
(descope/virtualwebauthn): a full creation round-trip plus adversarial
guards proving origin-mismatch and user-mismatch are rejected and
already-bound credentials are excluded.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14), web app face.
An already-authenticated principal binds a passkey to their own account and
manages the credentials they have bound; email-OTP stays the fallback factor.
- four account routes: POST register/begin mints a credential-creation
challenge, POST register/finish verifies the attestation against the
server-stashed SessionData and binds the credential, GET/DELETE credentials
list and unbind the caller's OWN passkeys. App-tier, principal-scoped (the
body never names a user).
- PasskeyVerifier seam keeps go-webauthn out of this package: ceremony state
crosses as opaque bytes, attestation as an io.Reader, result as a plain
VerifiedCredential. A nil verifier makes begin/finish report 503 so the
authenticated boundary is exercised before the real verifier is wired in.
- the view never leaks the public key; credential_id collisions map to 409.
- OpenAPI: the four paths plus the PasskeyCredential schema, keeping the
served-routes parity gate green.
Scope: ENROLLMENT only. The passkey login/assertion path (proving a passkey
from an unauthenticated state) is deferred; every ceremony here rides on a
known principal.
Tests: handler + challenge state machine against a fake repo and a fake
verifier (no real attestation crypto, no SQL). The decisive assertion is the
session-data round-trip -- the finish body carries no challenge, so the only
path for the stashed blob into FinishRegistration is store-stash then consume,
proving the challenge is server-held and never client-echoed. Also covers
supersede-on-begin, single-use, expiry, 503-unavailable, 409-already-bound,
owner-scoped list/delete, and external-only face separation.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14). Adds the data
layer an already-authenticated principal needs to bind and manage passkeys:
- migration 0007: webauthn_credentials (one bound passkey per row, public
attestation material only) and webauthn_challenges (server-stashed ceremony
state between begin and finish, single-use via consumed_at). Both rows are
bound to a known user_id; there is no usernameless login lookup, since the
assertion/login path is a deferred slice.
- PasskeyCredential type and five Repo methods (create/consume challenge,
create/list/delete credential) with the PG semantics the handlers rely on:
supersede-prior-live on begin, expiry-before-consume single-use on finish,
credential_id UNIQUE -> ErrConflict, owner-scoped delete -> ErrNotFound.
- ErrPasskeyChallengeInvalid sentinel for a missing/expired/consumed ceremony.
cooldownLimiter began as the wake-only throttle; the OTP-start hardening
reused it via the atomic reserve/release. Its type comment still called it
a per-server wake limiter and justified the per-replica behaviour as
"acceptable because the operator reconcile is idempotent" -- true for wake,
false for OTP, whose every admitted send is a non-idempotent email.
Rewrite the comment to describe the shared per-key limiter and record the
honest KNOWN-LIMITATION: the atomic reserve/release closes the
intra-replica concurrent burst, but the in-memory map throttles per
replica, so cross-replica bounding still needs a shared store. No
behaviour change.
The email-OTP resend cooldown checked the window with a peek (allowed)
and only recorded it after delivery. For OTP that throttle is the sole
defense and each admitted send is a real, non-idempotent email, so a
burst of truly concurrent starts all passed the peek before any recorded
and every one mailed: N concurrent starts bombed a mailbox with N codes.
Add an atomic reserve/release pair to cooldownLimiter: reserve checks and
records the window in one critical section under the mutex, so a
concurrent burst yields exactly one winner; release rolls a reservation
back only if it is still the current one, so a slow failing caller never
clobbers a newer holder. handleEmailOTPStart now reserves both the
principal and the recipient key up front and defers a rollback that frees
both windows on any mint, create, or delivery error — preserving the old
"a failed send does not consume the cooldown" property, now race-free.
The wake path keeps allowed→record: its real gate is the running cap and
its side effect (SetDesiredState) is idempotent, so the peek gap is
harmless there.
Tests: a frozen-clock gate-mailer fires 8 concurrent starts for one
victim from one principal and asserts exactly one mail and one 202; a
flaky-mailer test proves a failed delivery releases the window so an
immediate retry in the same instant is admitted.
A Running, ready server always reported 0/0 players: markRunningReady
never wrote Status.Players, and markStopped only cleared it. The panel
therefore showed an empty tally for live servers.
Extend the readiness probe to also sample the player count. Prober.Probe
now returns a PlayerCount{Online, Max}: RconProber still gates readiness
on Dial+auth, then runs a best-effort `list` and parses the vanilla
reply ("There are N of a max of M players online"). A failed or
unparseable tally is swallowed (0/0) so it never blocks readiness. The
reconciler threads the count into markRunningReady, which writes
Status.Players; markStopped still resets it to zero.
handleEmailOTPStart minted and mailed a code on every call, so an
authenticated caller could drive unbounded mail to any address they
typed — an email-bomb primitive against arbitrary mailboxes.
Add a separate otpLimiter (its own sync.Once and map, distinct from the
wake limiter) and throttle each send on two keys before anything is
minted: the caller (user:<id>) and the recipient (email:<lower>). A
refused send mints no code and mails nothing; both cooldowns are
recorded only after delivery succeeds, mirroring the wake path so a
failed mint or delivery never consumes the throttle. The two-key design
stops both one account fanning out across addresses and many accounts
converging on one mailbox.
A wake refused by the §9.1 running-server cap returns 503, but the
per-server cooldown was recorded before the cap check ran. A player
held because the cluster was momentarily full would then also have to
wait out the wake cooldown once a slot freed, even though their refused
wake never actually flipped desiredState.
Split cooldownLimiter.allow into allowed (peek, no record) and record
(commit). Both wake paths now consult allowed for the 429, then call
record only after SetDesiredState succeeds — so neither a 503
at_capacity nor a SetDesiredState error consumes the cooldown. The
split is safe against the running cap, which counts CRD truth via
ListServers and is independent of the limiter.
`docker build` failed twice over because .dockerignore excluded two trees the
image actually needs. The panel stage's `COPY panel/ ./` hit `"/panel": not
found`, and even past that the Go build would fail: the root felis package
//go:embeds deploy/bootstrap.sh and deploy/crd/*.yaml, which `COPY . .` dropped
along with the excluded deploy/.
The stale header comment claimed only internal/store/migrations was embedded,
which is what licensed the over-broad exclusions. Rewrite it to name all three
embedded trees (migrations, deploy assets, panel static) and warn against
re-adding panel/, deploy/, or internal/ without re-checking the go:embed list.
Tighten node_modules -> **/node_modules so a working-tree build no longer drags
panel/node_modules over the Linux modules npm ci installs in the panel stage.
When a staff account already exists, the break-glass console now opens on a
thin top-level menu (menuModel) where account operations are peers rather than
tails of one wizard: provision/reset the Owner, or add an Operator. A fresh
machine with no Owner skips the menu and goes straight to Owner bootstrap, since
minting an Operator first would create a staff account the login gate rejects.
The Operator path reuses ownerModel via a bgOperation discriminator. It is
insert-only (performAddOperator -> InsertOperator), wraps a duplicate username as
api.ErrConflict and routes back to the provision form for a retry rather than
tearing down, and deliberately never flips the global local_auth toggle the way
the Owner thread does. The post-exit summary and audit trail distinguish the two
outcomes (isOperator); only the Owner provision claims local-password login was
enabled.
Tests cover the operator-model defaults, path selection (insert vs upsert and
the local-auth gate), conflict-retry versus generic teardown, isOperator
propagation, and the root menu routing for both fresh and admin-present
machines.
Add docs/troubleshooting.md covering the common failure modes the spec
implies, grounded in the actual control-plane code paths:
- Stuck Starting (PodNotReady / RconSecretUnavailable / RconNotReachable)
and the deliberate absence of a Starting->Failed timeout.
- Failed reachable only via InvalidSpec on a malformed spec.storage.size,
plus the stale status.endpoint=direct caveat after a failure.
- Routing via status.endpoint direct/fallback and the empty fallbackServer
pitfall; wake 403/429/503 gate order.
- online-mode coupling and Velocity's offline-mode routing refusal.
- Cloudflare Access 401/403, nil-Keyfunc fail-closed, audience checks,
the absence of an issuer check, and local-session gating.
- Internal service-token (FELIS_SERVICE_TOKEN) rejection path.
- link/claim error codes, Kaniko build denials (SA-by-absence RBAC,
default-deny egress, internal-registry push gate), and the registry
DNS contract.
- Reaper backup-before-delete invariant and false-delete vectors.
- Unimplemented idle auto-stop, permanently-zero players.online, the
inert CRD fields, and the always-survives world PVC behaviour.
Each item is labelled with its evidence grade (GO-TESTED / CODE-ONLY /
INTEGRATION-ONLY / INERT) so operators know what is verified versus
asserted.
Adversarial cross-check of the three §28 diagrams against the wake, claim
and link code paths surfaced two fidelity drifts:
- The wake/status/join-event lanes used abbreviated /internal/... paths;
the registered internal-face routes carry the /api/v1 prefix (api.go),
matching the convention the claim and link diagrams already use.
- The /link diagram showed the game posting {mc_uuid, auth_source}, but no
shipped in-game caller sends auth_source — LinkClient posts {mc_uuid}
and the API defaults auth_source to mojang server-side.
The claim diagram already matched the code (verbatim atomic UPDATE,
404/409/200 mapping) and is unchanged.
Wire the fourth mandated §23 metric to a real producer. The histogram
spans two reconcile passes, so anchor and observation must persist in
status:
- Add status.startRequestedAt, set once on the first Starting reconcile
of a start attempt and cleared on Stopped so the next start re-anchors.
- Observe felis_start_duration_seconds exactly when readiness is first
reached (ReadySignalAt - StartRequestedAt), guarded so a server that
reaches ready without a Starting pass records nothing.
- Mirror the field into the deepcopy and the structural CRD schema so the
apiserver does not prune it on patchStatus round-trips.
- Promote prometheus/client_golang and client_model to direct deps now
that the operator and its tests import them.
Tests drive a step clock through Starting -> Running asserting the exact
observed duration, and through Running -> Stopped asserting the metric is
observed once and the anchor clears.
A per-object reconcile cannot maintain felis_servers_total (spec §23): it
sees one server per call, so it could never Set a correct fleet-wide gauge
and inc/dec on transitions would drift on any missed event. Add a snapshot
producer instead.
metrics.SyncServerGauge Resets the GaugeVec then Sets one child per state,
so a state that drains to zero reports 0 rather than a stale last value.
operator.GaugeSyncer is a manager.Runnable that periodically Lists the
fleet and republishes from it, defaulting an unset desiredState to Stopped.
SyncOnce is exercised end-to-end against a fake client (List, default,
republish); the ticker loop in Start is the only untested I/O edge.
Wire the build subsystem to the felis_image_build_failures_total counter
(spec §23). It advances at the two terminal-failure producers: finishAt
(the Sync JobFailed/JobUnknown verdict — a kaniko failure or a CRITICAL
CVE from trivy's --exit-code 1) and Submit's job-creation bypass path,
which records its failure directly without going through finishAt.
Cancellations and successful builds are deliberately not counted.
A delta-asserting test exercises both Inc sites plus a successful-build
negative control that proves the StatusFailed guard discriminates rather
than firing on every terminal write, all over the existing in-memory
Store/Jobs fakes.