Files
Felis/internal/api/errors.go
T
flyemoji b3989fa4af fix(mail): prove SMTP deliverability before saving, and stop losing the relay
A live install passed the SMTP setup screen and then failed every one-time
code with a bare `internal error`. Four separate defects had to line up for
that, and each is fixed here.

The relay was configured with `from = noreply@<domain-A>` on an account
authenticated as `<user>@<domain-B>`. Providers that validate sender identity
— Fastmail among them — answer MAIL FROM with an unconditional 250 and only
refuse at end-of-DATA. Ping stopped at NOOP, so it never saw the refusal: the
wizard reported success, wrote the config, rolled felis-api, and every OTP
afterwards died at w.Close().

Ping now runs the same transaction a real code takes — connect, (STARTTLS,)
AUTH, MAIL FROM, RCPT TO, DATA — delivering one self-test message to the From
address, and SendOTP and Ping share deliver() so the check cannot drift from
the thing it checks. The self-test recipient cannot cause a false negative:
an authenticated submission relay accepts RCPT for any destination by
definition, while the sender identity it does validate is exactly what we
want tested. The setup screen now says a message will be sent, names the
address it went to, and warns that From must be an address the account is
allowed to send as.

A relay refusal also answered 500 `internal`, which reads as a broken panel
and sends the operator hunting through handler code instead of their [smtp]
block. It is now 502 `mail_undeliverable`, mapped inside deliverOTP so all
four doors that mail a code (onboarding, email login, op-login, migrate
step-up) answer alike. The relay's own text stays out of the response — it
can name the SMTP account, and these routes are reachable by any signed-in
player — and goes to the log instead.

writeError logged nothing when it collapsed an unmapped error to 500, so an
operator holding an `internal error` had nothing to grep for and diagnosis
degraded into guessing against a live install. It now logs the method, path,
wrapped chain and the same request_id the caller is shown.

Finally, write_felis_toml regenerated the config wholesale and never emitted
[smtp], so re-running the installer — the documented way to update felis-api —
silently erased a working relay and reverted OTP delivery to the no-Mailer
path, logging codes instead of sending them. It now carries the block forward,
cached on first read because the host toml is clobbered before the pod toml is
written. Same defect family as the root_domain loss fixed in ecbeb20: a
generated file holding a hand-set value with no carry-forward.

Tests cover the case a MAIL FROM probe cannot see: a fake relay that answers
250 to MAIL FROM and 550 at end-of-DATA must fail both Ping and SendOTP, and
the 502 must carry a distinct machine code without leaking the relay's text.
2026-07-21 00:11:40 +09:00

130 lines
6.7 KiB
Go

package api
import (
"encoding/json"
"errors"
"fmt"
"log"
"net/http"
)
// Sentinel errors the repository and cluster layers return so handlers can map
// domain outcomes onto HTTP status codes without leaking driver details.
var (
// ErrNotFound means the requested server / record does not exist.
ErrNotFound = errors.New("not found")
// ErrConflict means an atomic precondition failed (e.g. claim lost the race).
ErrConflict = errors.New("conflict")
// ErrLinkCodeInvalid means an account-link code is unknown or expired (spec
// §10). It is a client error (the verify endpoint exists; the code is bad), so
// handlers map it to 400, not 404.
ErrLinkCodeInvalid = errors.New("link code invalid or expired")
// ErrConsoleUnavailable means the RCON write channel could not be reached —
// the dial timed out, was refused, or the password was rejected (spec §8).
// Because readiness IS an RCON probe (spec §141: phase=Running ⟺ RCON
// answers), a reachable failure here is a transient/racy "the server isn't
// actually up", not a server bug. Handlers map it to 503, not 500, so the
// caller is told to wake/retry rather than shown an opaque internal error.
ErrConsoleUnavailable = errors.New("server console is unavailable")
// ErrOTPInvalid means an email one-time code is unknown, expired, already
// consumed, or did not match (spec §B2 onboarding). Like ErrLinkCodeInvalid it
// is a client error — the verify endpoint exists; the code is bad — so handlers
// map it to 400, not 404. A wrong-but-not-yet-locked guess collapses to it too,
// so the response never distinguishes "no such code" from "wrong digits".
ErrOTPInvalid = errors.New("email code invalid or expired")
// ErrOTPLocked means the live email code has exhausted its attempt budget: too
// many wrong guesses (spec §B2). It is distinct from ErrOTPInvalid so handlers
// can answer 429 (back off / request a new code) rather than inviting another
// guess against a code that will never accept one.
ErrOTPLocked = errors.New("email code locked: too many attempts")
// ErrPasskeyChallengeInvalid means a passkey enrollment ceremony cannot be
// finished: there is no live (unconsumed, unexpired) challenge for the caller and
// purpose (Phase 6 WebAuthn bind). Like ErrOTPInvalid it is a client error — the
// finish endpoint exists; the ceremony state is gone (never begun, already
// consumed, or expired) — so handlers map it to 400, not 404.
ErrPasskeyChallengeInvalid = errors.New("passkey challenge invalid or expired")
// ErrPlayerBindForbidden means a public Bind-Code redemption resolved to a STAFF
// account (role=admin), which the player-console bootstrap refuses (console-tier
// access model). Operators authenticate at op.console behind Zero Trust, never via
// the account-less console.<root_domain> door, so the public bootstrap provably
// never mints a session for an admin identity. It is distinct from ErrConflict so
// the handler answers 403 (wrong door) rather than 409 (already linked).
ErrPlayerBindForbidden = errors.New("bind code belongs to a staff account")
// ErrEmailTaken means a verified email would collide with another account's
// already-verified address (spec §B email-first login foundation, migration 0010).
// VerifyEmailOTP returns it — WITHOUT consuming the code, since the address, not
// the code, is the problem — when a DIFFERENT user has already proven the same
// address case-insensitively. It is the clean, application-level counterpart of
// the users_verified_email_unique index: a sequential double-verify meets this
// guard and gets a 409 instead of a raw unique-violation 500. Distinct from
// ErrConflict so the message can name the cause (the email is spoken for).
ErrEmailTaken = errors.New("email already verified on another account")
// ErrTooManyDiscoverableChallenges means the non-user-keyed discoverable ("usernameless")
// login challenge store is at its hard cap of live rows (task #40, migration 0013).
// Unlike the user-keyed enrollment/login challenges — which self-bound via a per-user
// supersede — a from-zero begin has no principal to key a fair per-caller limit on, so the
// table is capped globally and a begin over the cap is refused. Distinct from the other
// sentinels so the handler answers 429 (a transient "too busy, retry" — the cap self-clears
// as challenges expire), never a 400 that invites an immediate retry.
ErrTooManyDiscoverableChallenges = errors.New("too many discoverable login challenges in flight")
)
// apiError is a handler-level error carrying an HTTP status and a stable,
// machine-readable code. The error envelope matches the platform convention:
//
// {"error": {"code": "...", "message": "...", "request_id": "..."}}
type apiError struct {
status int
code string
msg string
}
func (e *apiError) Error() string { return e.msg }
// newError builds an apiError with a formatted message.
func newError(status int, code, format string, a ...any) *apiError {
return &apiError{status: status, code: code, msg: fmt.Sprintf(format, a...)}
}
// Common errors reused across handlers.
var (
errUnauthorized = newError(http.StatusUnauthorized, "unauthorized", "authentication required")
errForbidden = newError(http.StatusForbidden, "forbidden", "not permitted")
errBadRequest = newError(http.StatusBadRequest, "bad_request", "invalid request")
)
// writeJSON writes v as an indented JSON body with the given status.
func writeJSON(w http.ResponseWriter, status int, v any) {
w.Header().Set("Content-Type", "application/json; charset=utf-8")
w.WriteHeader(status)
enc := json.NewEncoder(w)
enc.SetEscapeHTML(false)
_ = enc.Encode(v)
}
// writeError renders err as the standard error envelope. Non-apiError values
// collapse to a 500 so driver/internal details never reach the client.
//
// That collapse is deliberately lossy on the wire and deliberately NOT lossy in
// the log. Everything the client is denied — the driver message, the wrapped
// chain, the handler that produced it — is written to stderr first, keyed by the
// same request_id the caller is shown. Without that line an operator holding a
// "internal error" has nothing to grep for, and diagnosis degrades into guessing
// against a live install; it cost a full debugging session to learn that once.
func writeError(w http.ResponseWriter, r *http.Request, err error) {
var ae *apiError
if !errors.As(err, &ae) {
log.Printf("api: %s %s: unmapped error (request_id=%s): %v",
r.Method, r.URL.Path, requestIDFromContext(r.Context()), err)
ae = newError(http.StatusInternalServerError, "internal", "internal error")
}
body := map[string]any{
"error": map[string]string{
"code": ae.code,
"message": ae.msg,
"request_id": requestIDFromContext(r.Context()),
},
}
writeJSON(w, ae.status, body)
}