Setup previously called runner.StartConnector (`cloudflared service install`)
while the TUI applyCloudflareEdge separately installs cloudflared-felis.service
for the same tunnel from the same config -- two managed services serving one
tunnel from one connector config.
Drop StartConnector from cfsetup: running a connector is a host-specific side
effect (systemd/launchd/Windows service) that belongs with the caller, not in
this host- and domain-agnostic package whose documented side effects are tunnel
creation, DNS routing, and the Access app/policy calls. installCloudflaredService
in the host layer stays the single connector installer, so the routed-but-dead
1033 is still closed; RouteDNS --overwrite-dns still closes the stale-DNS 1033.
Setup created the tunnel, routed DNS, and wrote config.yml, but nothing
installed or started a connector for it. A one-click run therefore left the
tunnel routed-but-dead: every web hostname returned Cloudflare error 1033
(tunnel has no connector) even though the config was correct on disk.
Add a StartConnector step to the Runner seam, invoked right after the config
is written (and gated on ConfigPath, so a caller wanting only the Access
config is not forced to install a service). The ExecRunner implementation
runs `cloudflared --config <path> service install`, which installs and starts
a managed system service (systemd/launchd/Windows), and is idempotent on an
already-installed service. The orchestration — connector started, and only
after its config exists — is unit-tested against the fake Runner; the actual
service install is INTEGRATION-ONLY.
Together with the RouteDNS --overwrite-dns fix, this closes both distinct
paths to a 1033 half-state from a fresh setup: a stale DNS binding and a
missing connector.
RouteDNS ran `cloudflared tunnel route dns` without --overwrite-dns and
swallowed the resulting "record already exists" error as success. When a
hostname already had a CNAME from an earlier tunnel that was deleted and
recreated, the record stayed bound to the dead tunnel: the setup reported
the hostname "routed" while it kept returning Cloudflare error 1033 (the
tunnel it pointed at has no connector).
Pass --overwrite-dns so the record is repointed at the tunnel just created,
making the route idempotent and correct on every re-run, and drop the
now-unnecessary "already exists" swallow. INTEGRATION-ONLY (ExecRunner
shells out to the real cloudflared binary).
Wrap github.com/go-webauthn/webauthn behind the api.PasskeyVerifier
seam so the api package stays free of go-webauthn types. The adapter
covers the credential-creation ceremony only (BeginRegistration /
CreateCredential); the login/assertion path is a deferred slice.
Ceremony state crosses the seam as opaque marshaled SessionData, the
attestation as an io.Reader, and the verified result as a plain
VerifiedCredential. SessionData carries no expiry so the challenge
row's TTL stays the single liveness authority. New rejects an empty
RP id or origin list so a misconfigured deployment fails at
construction rather than minting unverifiable challenges.
Tests drive a real relying party against a virtual authenticator
(descope/virtualwebauthn): a full creation round-trip plus adversarial
guards proving origin-mismatch and user-mismatch are rejected and
already-bound credentials are excluded.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14), web app face.
An already-authenticated principal binds a passkey to their own account and
manages the credentials they have bound; email-OTP stays the fallback factor.
- four account routes: POST register/begin mints a credential-creation
challenge, POST register/finish verifies the attestation against the
server-stashed SessionData and binds the credential, GET/DELETE credentials
list and unbind the caller's OWN passkeys. App-tier, principal-scoped (the
body never names a user).
- PasskeyVerifier seam keeps go-webauthn out of this package: ceremony state
crosses as opaque bytes, attestation as an io.Reader, result as a plain
VerifiedCredential. A nil verifier makes begin/finish report 503 so the
authenticated boundary is exercised before the real verifier is wired in.
- the view never leaks the public key; credential_id collisions map to 409.
- OpenAPI: the four paths plus the PasskeyCredential schema, keeping the
served-routes parity gate green.
Scope: ENROLLMENT only. The passkey login/assertion path (proving a passkey
from an unauthenticated state) is deferred; every ceremony here rides on a
known principal.
Tests: handler + challenge state machine against a fake repo and a fake
verifier (no real attestation crypto, no SQL). The decisive assertion is the
session-data round-trip -- the finish body carries no challenge, so the only
path for the stashed blob into FinishRegistration is store-stash then consume,
proving the challenge is server-held and never client-echoed. Also covers
supersede-on-begin, single-use, expiry, 503-unavailable, 409-already-bound,
owner-scoped list/delete, and external-only face separation.
Phase 6 WebAuthn bind, enrollment-only slice (spec section 14). Adds the data
layer an already-authenticated principal needs to bind and manage passkeys:
- migration 0007: webauthn_credentials (one bound passkey per row, public
attestation material only) and webauthn_challenges (server-stashed ceremony
state between begin and finish, single-use via consumed_at). Both rows are
bound to a known user_id; there is no usernameless login lookup, since the
assertion/login path is a deferred slice.
- PasskeyCredential type and five Repo methods (create/consume challenge,
create/list/delete credential) with the PG semantics the handlers rely on:
supersede-prior-live on begin, expiry-before-consume single-use on finish,
credential_id UNIQUE -> ErrConflict, owner-scoped delete -> ErrNotFound.
- ErrPasskeyChallengeInvalid sentinel for a missing/expired/consumed ceremony.
cooldownLimiter began as the wake-only throttle; the OTP-start hardening
reused it via the atomic reserve/release. Its type comment still called it
a per-server wake limiter and justified the per-replica behaviour as
"acceptable because the operator reconcile is idempotent" -- true for wake,
false for OTP, whose every admitted send is a non-idempotent email.
Rewrite the comment to describe the shared per-key limiter and record the
honest KNOWN-LIMITATION: the atomic reserve/release closes the
intra-replica concurrent burst, but the in-memory map throttles per
replica, so cross-replica bounding still needs a shared store. No
behaviour change.
The email-OTP resend cooldown checked the window with a peek (allowed)
and only recorded it after delivery. For OTP that throttle is the sole
defense and each admitted send is a real, non-idempotent email, so a
burst of truly concurrent starts all passed the peek before any recorded
and every one mailed: N concurrent starts bombed a mailbox with N codes.
Add an atomic reserve/release pair to cooldownLimiter: reserve checks and
records the window in one critical section under the mutex, so a
concurrent burst yields exactly one winner; release rolls a reservation
back only if it is still the current one, so a slow failing caller never
clobbers a newer holder. handleEmailOTPStart now reserves both the
principal and the recipient key up front and defers a rollback that frees
both windows on any mint, create, or delivery error — preserving the old
"a failed send does not consume the cooldown" property, now race-free.
The wake path keeps allowed→record: its real gate is the running cap and
its side effect (SetDesiredState) is idempotent, so the peek gap is
harmless there.
Tests: a frozen-clock gate-mailer fires 8 concurrent starts for one
victim from one principal and asserts exactly one mail and one 202; a
flaky-mailer test proves a failed delivery releases the window so an
immediate retry in the same instant is admitted.
A Running, ready server always reported 0/0 players: markRunningReady
never wrote Status.Players, and markStopped only cleared it. The panel
therefore showed an empty tally for live servers.
Extend the readiness probe to also sample the player count. Prober.Probe
now returns a PlayerCount{Online, Max}: RconProber still gates readiness
on Dial+auth, then runs a best-effort `list` and parses the vanilla
reply ("There are N of a max of M players online"). A failed or
unparseable tally is swallowed (0/0) so it never blocks readiness. The
reconciler threads the count into markRunningReady, which writes
Status.Players; markStopped still resets it to zero.
handleEmailOTPStart minted and mailed a code on every call, so an
authenticated caller could drive unbounded mail to any address they
typed — an email-bomb primitive against arbitrary mailboxes.
Add a separate otpLimiter (its own sync.Once and map, distinct from the
wake limiter) and throttle each send on two keys before anything is
minted: the caller (user:<id>) and the recipient (email:<lower>). A
refused send mints no code and mails nothing; both cooldowns are
recorded only after delivery succeeds, mirroring the wake path so a
failed mint or delivery never consumes the throttle. The two-key design
stops both one account fanning out across addresses and many accounts
converging on one mailbox.
A wake refused by the §9.1 running-server cap returns 503, but the
per-server cooldown was recorded before the cap check ran. A player
held because the cluster was momentarily full would then also have to
wait out the wake cooldown once a slot freed, even though their refused
wake never actually flipped desiredState.
Split cooldownLimiter.allow into allowed (peek, no record) and record
(commit). Both wake paths now consult allowed for the 429, then call
record only after SetDesiredState succeeds — so neither a 503
at_capacity nor a SetDesiredState error consumes the cooldown. The
split is safe against the running cap, which counts CRD truth via
ListServers and is independent of the limiter.
Wire the fourth mandated §23 metric to a real producer. The histogram
spans two reconcile passes, so anchor and observation must persist in
status:
- Add status.startRequestedAt, set once on the first Starting reconcile
of a start attempt and cleared on Stopped so the next start re-anchors.
- Observe felis_start_duration_seconds exactly when readiness is first
reached (ReadySignalAt - StartRequestedAt), guarded so a server that
reaches ready without a Starting pass records nothing.
- Mirror the field into the deepcopy and the structural CRD schema so the
apiserver does not prune it on patchStatus round-trips.
- Promote prometheus/client_golang and client_model to direct deps now
that the operator and its tests import them.
Tests drive a step clock through Starting -> Running asserting the exact
observed duration, and through Running -> Stopped asserting the metric is
observed once and the anchor clears.
A per-object reconcile cannot maintain felis_servers_total (spec §23): it
sees one server per call, so it could never Set a correct fleet-wide gauge
and inc/dec on transitions would drift on any missed event. Add a snapshot
producer instead.
metrics.SyncServerGauge Resets the GaugeVec then Sets one child per state,
so a state that drains to zero reports 0 rather than a stale last value.
operator.GaugeSyncer is a manager.Runnable that periodically Lists the
fleet and republishes from it, defaulting an unset desiredState to Stopped.
SyncOnce is exercised end-to-end against a fake client (List, default,
republish); the ticker loop in Start is the only untested I/O edge.
Wire the build subsystem to the felis_image_build_failures_total counter
(spec §23). It advances at the two terminal-failure producers: finishAt
(the Sync JobFailed/JobUnknown verdict — a kaniko failure or a CRITICAL
CVE from trivy's --exit-code 1) and Submit's job-creation bypass path,
which records its failure directly without going through finishAt.
Cancellations and successful builds are deliberately not counted.
A delta-asserting test exercises both Inc sites plus a successful-build
negative control that proves the StatusFailed guard discriminates rather
than firing on every terminal write, all over the existing in-memory
Store/Jobs fakes.
Introduce internal/metrics exposing the four metric families spec §23
mandates at minimum: felis_servers_total (gauge by desired state),
felis_start_duration_seconds (histogram with Minecraft cold-start
buckets), felis_image_build_failures_total and
felis_reaper_worlds_deleted_total (counters). Collectors are
package-level vars so any subsystem records without an import cycle;
Register wires them into a prometheus.Registerer and is idempotent.
Wire registration into the operator against controller-runtime's global
Registry, so /metrics on the manager's existing metrics endpoint carries
the felis_* families. Instrument the reaper to increment
felis_reaper_worlds_deleted_total in lockstep with Summary.WorldsReaped,
at the one point a world's PVC has actually been deleted.
Add the insert-only Operator-creation path to the break-glass console
(felis breakGlass). An Operator is an additional staff admin: role=admin
with must_change_password=true, identical in shape to the Owner, since
Felis has no separate operator DB role (migration 0003).
Unlike the Owner upsert, provisioning is insert-only -- a username already
taken returns ErrConflict (ON CONFLICT DO NOTHING + zero RowsAffected)
rather than silently resetting a live account, so adding an Operator can
never clobber the Owner's or another Operator's credential. A typed
password is used as-is; an empty one is replaced with a generated
one-time credential returned for display. Operator-add does not touch
local_auth_enabled -- that global gate belongs to the Owner thread alone.
Accountability is recorded best-effort under a break_glass.operator_create
audit action, written only after a successful provision.
The TUI menu router that reaches this path is deferred; this lands the
fully unit-testable logic layer (provisionOperator, performAddOperator,
auditAddOperator) with the PGRepo insert kept integration-only.
QR scan-to-login is a device-code grant where the QR encodes the existing
short-lived account-link code (spec §B3 player game-login). velocity mints a
code in-game, renders it as a QR, the player scans it on a phone already signed
in to the panel, and that web session's verify writes the durable account_links
row bound to that user. The only new verifiable surface that flow needs is the
completion poll velocity calls to learn the link landed and admit the player.
Add GET /api/v1/internal/account/link/status/{mc_uuid}: a read-only, internal
handleLinkStatus keyed by the verified mc_uuid velocity already holds. It reuses
the existing UserByMCUUID, so it adds no migration and no mutation to the
load-bearing VerifyLinkCode; ErrNotFound maps to {linked:false} (pending /
not-yet-scanned), a hit to {linked:true, user_id}. Keying on the public UUID and
not the scanned code means the read carries no guessing surface and needs no
attempt cap — the internal face already gates it to service callers, and the poll
consumes nothing so a velocity restart re-polls safely.
QR render, limbo collision routing, in-game admit, and the reclaim
inherit-disambiguation stay CODE-ONLY (Java/Velocity) and are labeled as such;
this endpoint reports link completion only.
Document the route in openapi.yaml (x-felis-face internal, x-felis-tier service)
so the parity gate holds, and cover it with a hermetic vertical that proves the
poll reflects the durable link only after the external verify and binds the
verifier's id, plus unknown-uuid, idempotency, and internal-only face separation.
- Convert CRLF to LF across Go, panel, and plugin files
- Add Cloudflare API token template URL to breakGlass TUI edge intro
- Verify API token in cfsetup before creating tunnel, DNS, or Access app
Add internal/cfsetup, the verifiable core of an optional one-click
Cloudflare Tunnel + Access provisioning flow for the SysAdmin edge
(spec §14). It is domain-agnostic (every FQDN is composed from the
configured root_domain) and IdP-agnostic (any valid Access JWT aud is
accepted, whichever IdP fronts it), so a SysAdmin who brings their own
domain or Zero-Trust scheme stays fully supported.
The load-bearing safety property is a fail-closed guard on the
recommended Access policy. validateFailClosed is an allowlist that
refuses any policy that could be public: a bypass/non-allow decision, an
empty include, an "everyone" include not narrowed by a constraining
require (include rules are OR, so "everyone" beside an identity is still
public), or any include rule it cannot positively recognize as a scoped
identity. Setup runs the guard before any side effect, so a public
policy aborts the run with nothing created.
The tunnel ingress routes only the web hostnames to the local panel
origin and terminates in the mandatory fail-shut 404 catch-all; the raw
game host is never proxied. Gating preconditions (cloudflared present,
tunnel login completed, API token) are hard checks with no side effects
on failure.
The actual cloudflared exec, DNS routing, and Access API calls live in
runner.go and are integration-only: they require the operator's own live
Cloudflare account and interactive browser consent, which cannot be
unit-tested. The policy guard, ingress generation, request bodies, and
gating are unit-tested.
When the configured third-party Yggdrasil and the official Mojang service
issue the same username under different UUIDs, the non-genuine squatter is
displaced in favour of the real Mojang owner (正版优先). This adds the
Go-verifiable data layer of that flow on the internal (velocity) face.
- migration 0006: username_blacklist (barred squatter UUIDs) and
player_data_holds (the displaced account's 30-day data stash), both keyed
by mc_uuid so the genuine Mojang player — identical username, different
UUID — is never caught by the bar.
- POST /api/v1/internal/player/reclaim bars the squatter UUID and stashes
its data in one transaction (all-or-nothing). It is idempotent on a
retried callback and returns the hold's effective expiry — the first
reclaim's window, never a fresh now()+30d — so the rejected player is told
the truth about how long their data is kept.
- GET /api/v1/internal/player/blacklist/{mc_uuid} is the login-gate check
velocity calls to reject a barred squatter before admitting them.
Scope: velocity collision-routing, the limbo prompt, the authlib
dual-backend and the data-inherit flow are code-only (Java plus a QR-bound
device session a row cannot express) and are not part of this slice. Unit
tests cover the handlers and the in-memory repo contract; the Postgres SQL
path is exercised by integration only.
Capture which Yggdrasil authenticated an in-game UUID when a link code is
minted (spec §10 dual-Yggdrasil) and copy it onto the durable account_links
row at verify. The value originates in-game — the web verify side never sees
the authentication — so it threads through account_link_codes, mirroring how
mc_uuid (not user_id) lives on a code.
- migration 0005: add link_auth_source enum + auth_source column on both
account_link_codes and account_links; DEFAULT 'mojang' backfills existing
rows and sets the Mojang-priority default for a mint that omits the field
- mint validates an explicit auth_source (unknown value -> 400); verify
surfaces it in the 200 body and refreshes it on idempotent re-verify
Forced web onboarding proves a player controls an email before it is
bound to their account. POST /api/v1/account/email/start mints a random
6-digit code, mails it (or logs it server-side when no Mailer is wired —
the demo has no SMTP), and POST /api/v1/account/email/verify redeems it,
flipping users.email_verified in the same transaction that consumes the
code.
Brute force is bounded two ways: a 10-minute TTL and a 5-attempt cap,
both enforced in the repo so the fake and Postgres agree. Only the
sha-256 of the code is stored; the digits live only in the email. Both
routes are app-tier external — verifying your own email is scoped to the
principal, never names another user.
Root is machine authority, not a human identity, so `felis breakGlass`
now also records WHICH SysAdmin broke the glass. Even under
`sudo felis breakGlass` an account and password are entered in the TUI;
the root gate is necessary but no longer sufficient for accountability.
The console resolves one of three modes up front and audits the
difference:
- bootstrap (no staff account exists yet): the typed credential mints
the first Owner; the act is attributed to the OS user ($SUDO_USER,
else root) and recorded verified:false.
- recovery (an admin already exists): the operator authenticates as an
existing admin via bcrypt; the verified identity is the accountable
actor and the row is recorded verified:true.
- root override (the typed credential did not verify): a deliberate
OVERRIDE token proceeds under local-root authority, attributed to the
OS user and recorded verified:false. Break-glass never refuses -
recovering when no admin password can be produced is its whole job.
Attribution is best-effort, not proof (whoever runs this is root and can
edit Postgres directly); the audit row is honest about which it is.
- internal/api: AuditEntry gains an optional jsonb Payload (nil maps to
SQL NULL, so existing callers are unaffected); PGRepo.Audit writes it
and a new PGRepo.AdminExists drives the bootstrap-vs-recovery switch.
- the accountability row is written the instant the credential changes,
before local auth is enabled, so a failed toggle write can never leave
a reset credential with no "who did it" record.
- local_auth_enabled is now one exported api.LocalAuthEnabledKey shared
by the break-glass writer and the per-request reader, replacing two
drifting copies of the literal.
- break-glass password entry reuses the panel's 8-72-byte rule so a
credential set here is never later rejected by web change-password.
Covered by Go unit tests over a fake owner store: auth match/non-match,
the three audit modes and their payloads, that a dead audit sink does
not fail the recovery, that the audit precedes the toggle write, and a
headless drive of the TUI state machine asserting no credential reaches
provisioning without a verified admin or an explicit OVERRIDE.
Add username+password login for Owner/Operator staff accounts on
op.console, the primary web login when Zero Trust is not in front of the
API. Three handlers form the whole surface: login mints a server-side
session cookie, logout revokes it idempotently, and change-password
re-verifies the current password before rotating the hash and clearing
must_change_password.
- Session cookies are HttpOnly+Secure+SameSite=Lax, host-only, stored
server-side as a SHA-256 hash with a 12h TTL.
- Login is anti-enumeration: every failure runs a uniform bcrypt compare
against a dummy hash and returns the same vague error.
- Credential-bearing writes require Content-Type: application/json,
returning 415 otherwise, to close the cross-site form-POST forgery
vector as a belt to the SameSite cookie.
- Local auth fails closed: login is rejected unless local_auth_enabled
is set, so a Zero-Trust-only deployment never accepts a local password.
- Extend the users table with a nullable password_hash and
must_change_password; staff are role=admin rows with a hash, players
are role=user rows with hash NULL.
- /me now reports must_change_password so the panel can force a
first-login change.
Covered by Go unit tests (handlers, content-type guard, anti-enumeration,
forced-change lockdown) and the OpenAPI route-parity gate.
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them.
The dual-faced felis-api: internal (service) and external (public/app/admin) routes behind a Zero-Trust guard. Includes the access domain (whitelist, ban, and LuckPerms permission/group control over the owner-gated RCON path), the modpack submission endpoints, and the admin-tier SysAdmin fleet read. Structured access fields are charset-validated before assembly so no field can splice a second RCON command.
A user-directed extension over the build subsystem: an uploaded modpack stays in pending_review and is never built until an admin approves. Approval is a single-winner compare-and-swap that hands off to the image-build Job, keeping the mandatory vulnerability scan in front of any push.
Foundational libraries: deterministic resource naming, the RCON client, the Postgres store with embedded SQL migrations, configuration loading, and container image-build helpers.