Felis never actually sent mail: OTP codes for onboarding, email login and
op-login were only written to the felis-api log behind a "demo has no SMTP"
limitation, and the Settings/SMTP flow those comments promised was never
built. Combined with the bootstrap Owner's address being recorded unverified
(87279a1), op-login start always took the anti-enumeration neutral branch and
minted a fake request_id, so the in-game approve inevitably answered "No
pending operator sign-in with that code".
Give the codes a real delivery path, configured in felis.toml rather than a
web settings page so config keeps a single source of truth:
- config: new [smtp] table (host, port defaulting to 587, from, username,
password_ref). Validation requires a plausible from address and a sane
port; the password itself never enters the config file.
- internal/mail (new): stdlib net/smtp mailer implementing the api.OTPMailer
seam. Port 465 dials implicit TLS, other ports upgrade via STARTTLS when
advertised; AUTH only when a username is configured (PlainAuth itself
refuses plaintext, so the password cannot leak to a TLS-less relay).
Ping() proves reachability and credentials without sending mail. The
message shape (CRLF, Q-encoded bilingual subject) is pinned by test.
- platform: felis-smtp Secret constants and an optional FELIS_SMTP_PASSWORD
env var on the felis-api Deployment, mirroring felis-uploads-s3.
- cmd/felis api: construct the real mailer when [smtp] is configured; keep
the log fallback otherwise and say so at startup. Warn when a username is
set but the credentials env is empty.
- setup TUI: "e" on the summary/status screen opens the email form (host,
port, from, optional auth). Apply order: Ping preflight, [smtp] into both
host and pod config files, felis-smtp Secret piped to kubectl via stdin,
config Secret, felis-api rollout. A failed preflight leaves the install
untouched. SMTP is deliberately not a wizard rail step: first-run stays
mail-less by design, and the passkey minted at onboarding is the pre-SMTP
owner credential.
Also make PGRepo.UserByEmail match case-insensitively (lower(email) =
lower($1)), honoring the interface contract and the users_verified_email_
unique partial index; the fake repo already matched with EqualFold.
Existing installs need the felis-api Deployment manifest re-applied (e.g. a
bootstrap re-run) before the new env var exists; a rollout restart alone
cannot add it.
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.
Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).
Manifest-level fix; the live packet path is pending real-cluster verification.
InternalAPIBaseURL builds the felis-api internal-face DNS from SAAPI and the internal port for cross-namespace callers (the login limbo). The service-token Secret name/key now reference the shared naming constants so the Deployment wiring and the operator's login-pod injection cannot drift.
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them.