Loading deploy/alerts/felis-alerts.yaml +25 −1 Changes for deploy/alerts/felis-alerts.yaml: 25 added lines, 1 removed line. Original line number Diff line number Diff line Loading @@ -6,7 +6,9 @@ # - felis-operator pod :8080/metrics → felis_servers_total, felis_start_duration_seconds # - felis-api internal :8081/metrics → felis_image_build_failures_total, # felis_mail_total, felis_rate_limited_total, # felis_auth_otp_lockouts_total # felis_auth_otp_lockouts_total, # felis_auth_failures_total, # felis_audit_write_failures_total # - node-exporter textfile collector → felis_db_backup_* (felis-db-backup.timer) # node_* / kube_* series come from node-exporter / kube-state-metrics. groups: Loading Loading @@ -142,3 +144,25 @@ groups: The audit log names the account (action auth.otp.locked); the owner was mailed. Unless they fumbled codes, someone is guessing at it (troubleshooting §17). - alert: FelisSignInFailures expr: sum(increase(felis_auth_failures_total[15m])) > 30 for: 5m labels: severity: warning annotations: summary: "over 30 refused sign-ins in 15 minutes" description: >- Wrong codes, unknown addresses or bad passkey assertions well above people mistyping: someone is guessing or enumerating. `sum by (door, reason) (increase(felis_auth_failures_total[15m]))` shows where; the audit rows (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17). - alert: FelisAuditWriteFailing expr: increase(felis_audit_write_failures_total[15m]) > 0 labels: severity: warning annotations: summary: "felis-api failed to write audit rows" description: >- The actions went through but their audit rows were lost. The felis-api log names each lost row (`audit: lost ...`); the usual cause is PostgreSQL being unreachable or out of disk. deploy/alerts/felis-alerts_test.yml +55 −0 Changes for deploy/alerts/felis-alerts_test.yml: 55 added lines, 0 removed lines. Original line number Diff line number Diff line Loading @@ -236,3 +236,58 @@ tests: - eval_time: 90m alertname: FelisOTPAccountLocked exp_alerts: [] - name: sign-in failure rate interval: 1m input_series: # Two doors failing at 3/min between them from t=0. - series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}' values: '0+2x40' - series: 'felis_auth_failures_total{door="op_login",reason="no_account",job="felis-api"}' values: '0+1x40' alert_rule_test: - eval_time: 8m alertname: FelisSignInFailures exp_alerts: [] - eval_time: 25m alertname: FelisSignInFailures exp_alerts: - exp_labels: severity: warning exp_annotations: summary: "over 30 refused sign-ins in 15 minutes" description: >- Wrong codes, unknown addresses or bad passkey assertions well above people mistyping: someone is guessing or enumerating. `sum by (door, reason) (increase(felis_auth_failures_total[15m]))` shows where; the audit rows (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17). - name: people mistyping stays quiet interval: 1m input_series: - series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}' values: '0 0 1 1 2 2 3 3 4 4 5x30' alert_rule_test: - eval_time: 30m alertname: FelisSignInFailures exp_alerts: [] - name: audit rows lost interval: 1m input_series: - series: 'felis_audit_write_failures_total{job="felis-api",instance="api-0"}' values: '0 0 0 2x20' alert_rule_test: - eval_time: 2m alertname: FelisAuditWriteFailing exp_alerts: [] - eval_time: 5m alertname: FelisAuditWriteFailing exp_alerts: - exp_labels: severity: warning job: felis-api instance: api-0 exp_annotations: summary: "felis-api failed to write audit rows" description: >- The actions went through but their audit rows were lost. The felis-api log names each lost row (`audit: lost ...`); the usual cause is PostgreSQL being unreachable or out of disk. deploy/alerts/felis-prometheusrule.yaml +22 −0 Changes for deploy/alerts/felis-prometheusrule.yaml: 22 added lines, 0 removed lines. Original line number Diff line number Diff line Loading @@ -144,3 +144,25 @@ spec: The audit log names the account (action auth.otp.locked); the owner was mailed. Unless they fumbled codes, someone is guessing at it (troubleshooting §17). - alert: FelisSignInFailures expr: sum(increase(felis_auth_failures_total[15m])) > 30 for: 5m labels: severity: warning annotations: summary: "over 30 refused sign-ins in 15 minutes" description: >- Wrong codes, unknown addresses or bad passkey assertions well above people mistyping: someone is guessing or enumerating. `sum by (door, reason) (increase(felis_auth_failures_total[15m]))` shows where; the audit rows (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17). - alert: FelisAuditWriteFailing expr: increase(felis_audit_write_failures_total[15m]) > 0 labels: severity: warning annotations: summary: "felis-api failed to write audit rows" description: >- The actions went through but their audit rows were lost. The felis-api log names each lost row (`audit: lost ...`); the usual cause is PostgreSQL being unreachable or out of disk. docs/troubleshooting.md +31 −4 Changes for docs/troubleshooting.md: 31 added lines, 4 removed lines. Original line number Diff line number Diff line Loading @@ -933,7 +933,9 @@ The series come from two processes: - `felis-api` internal face `:8081/metrics` (Service `felis-api-internal`) — `felis_image_build_failures_total`, and the sign-in series of §17 (`felis_mail_total`, `felis_rate_limited_total`, `felis_auth_otp_lockouts_total`). Unauthenticated like the probes; `felis_auth_otp_lockouts_total`, `felis_auth_failures_total`, `felis_sessions_revoked_total`, `felis_audit_write_failures_total`). Unauthenticated like the probes; ClusterIP-only, and the external face never serves it. - `felis_reaper_worlds_deleted_total` is produced inside the one-shot reaper CronJob, which exits long before any scrape interval — without a pushgateway Loading @@ -945,8 +947,8 @@ The series come from two processes: `deploy/alerts/` ships ready-made rules: build failures, slow starts, node disk/memory thresholds, the kubelet `DiskPressure` condition, control-plane database backup freshness (§16; needs node-exporter's textfile collector), and sign-in abuse: the mail budget, relay failures, throttled floods and account code locks (§17). sign-in abuse: the mail budget, relay failures, throttled floods, account code locks and the refused sign-in rate, plus lost audit rows (§17). - Plain Prometheus: add `felis-alerts.yaml` to `rule_files`. Check and unit-test it standalone with `promtool check rules felis-alerts.yaml` and Loading Loading @@ -1171,7 +1173,7 @@ Skips the pre-migration snapshot (`migrate up -no-backup`). The installer warns loudly when it is set. Use it only when the snapshot cannot work and you have another backup, e.g. an external database newer than the host's `pg_dump`. ## 17. Sign-in refused with 429, mail budget, account code locks ## 17. Sign-in refused with 429, mail budget, account code locks, failed sign-ins The public sign-in doors (`/api/v1/auth/*` except logout and the op-login status poll) have three limits of their own. Each answers 429 with a Loading Loading @@ -1232,6 +1234,29 @@ sudo -u postgres psql felis -c \ "DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '<name>');" ``` ### Who tried: the audit trail Every refused sign-in writes an `auth.<door>.failed` row (payload `reason`: `bad_code`, `no_account`, `staff_account`, `not_staff`, `bad_assertion`, ...) and counts in `felis_auth_failures_total{door,reason}`; `FelisSignInFailures` fires above 30 in 15 minutes. The first refusal of each throttled burst writes `auth.rate_limited` with the source address. Rows carry `actor_user_id` (the account, the column to attribute by), `client_ip` (the same address the limit keys on) and `user_agent`. `actor` is display text: a verified email or the username, never an address the caller set without verifying. ```sh sudo -u postgres psql felis -c " SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason FROM audit_logs WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour' ORDER BY created_at DESC LIMIT 50;" ``` A failed audit write does not fail the action; it logs `audit: lost ...` in `felis-api` and counts in `felis_audit_write_failures_total` (`FelisAuditWriteFailing`). The cause is almost always PostgreSQL (§16). ### Optional: a Cloudflare rate limiting rule in front The limits above live in the API, so they hold on any edge. Behind Cloudflare Loading Loading @@ -1274,3 +1299,5 @@ for 10 seconds (the Free plan's limits). | Sign-in 429 `rate_limited` for everyone at once | §17 | | 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 | | Right code refused; `otp_account_locked` / `FelisOTPAccountLocked` | §17 | | `FelisSignInFailures` / who is guessing, from where | §17 | | `FelisAuditWriteFailing` | §17 | internal/api/api_test.go +6 −1 Changes for internal/api/api_test.go: 6 added lines, 1 removed line. Original line number Diff line number Diff line Loading @@ -47,6 +47,7 @@ type fakeRepo struct { serverResources map[string]ResourceSpec resourceUpdates map[string]ResourceSpec audits []AuditEntry failAudit error // Audit fails with it (a store outage) joins []string // create-server seeding (spec §15) seeded map[string]bool // name -> servers row exists Loading Loading @@ -745,6 +746,9 @@ func (f *fakeRepo) SeedServer(_ context.Context, name, subdomain string, _, _, _ } func (f *fakeRepo) Ping(_ context.Context) error { return f.pingErr } func (f *fakeRepo) Audit(_ context.Context, e AuditEntry) error { if f.failAudit != nil { return f.failAudit } f.audits = append(f.audits, e) return nil } Loading Loading @@ -857,7 +861,8 @@ func (f *fakeRepo) SessionUser(_ context.Context, tokenHash string, now time.Tim return nil, ErrNotFound } return &SessionedUser{ ID: u.ID, Email: u.Email, Role: u.Role, ID: u.ID, Username: u.Username, Email: u.Email, Role: u.Role, EmailVerified: u.EmailVerified, }, nil } } Loading Loading
deploy/alerts/felis-alerts.yaml +25 −1 Changes for deploy/alerts/felis-alerts.yaml: 25 added lines, 1 removed line. Original line number Diff line number Diff line Loading @@ -6,7 +6,9 @@ # - felis-operator pod :8080/metrics → felis_servers_total, felis_start_duration_seconds # - felis-api internal :8081/metrics → felis_image_build_failures_total, # felis_mail_total, felis_rate_limited_total, # felis_auth_otp_lockouts_total # felis_auth_otp_lockouts_total, # felis_auth_failures_total, # felis_audit_write_failures_total # - node-exporter textfile collector → felis_db_backup_* (felis-db-backup.timer) # node_* / kube_* series come from node-exporter / kube-state-metrics. groups: Loading Loading @@ -142,3 +144,25 @@ groups: The audit log names the account (action auth.otp.locked); the owner was mailed. Unless they fumbled codes, someone is guessing at it (troubleshooting §17). - alert: FelisSignInFailures expr: sum(increase(felis_auth_failures_total[15m])) > 30 for: 5m labels: severity: warning annotations: summary: "over 30 refused sign-ins in 15 minutes" description: >- Wrong codes, unknown addresses or bad passkey assertions well above people mistyping: someone is guessing or enumerating. `sum by (door, reason) (increase(felis_auth_failures_total[15m]))` shows where; the audit rows (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17). - alert: FelisAuditWriteFailing expr: increase(felis_audit_write_failures_total[15m]) > 0 labels: severity: warning annotations: summary: "felis-api failed to write audit rows" description: >- The actions went through but their audit rows were lost. The felis-api log names each lost row (`audit: lost ...`); the usual cause is PostgreSQL being unreachable or out of disk.
deploy/alerts/felis-alerts_test.yml +55 −0 Changes for deploy/alerts/felis-alerts_test.yml: 55 added lines, 0 removed lines. Original line number Diff line number Diff line Loading @@ -236,3 +236,58 @@ tests: - eval_time: 90m alertname: FelisOTPAccountLocked exp_alerts: [] - name: sign-in failure rate interval: 1m input_series: # Two doors failing at 3/min between them from t=0. - series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}' values: '0+2x40' - series: 'felis_auth_failures_total{door="op_login",reason="no_account",job="felis-api"}' values: '0+1x40' alert_rule_test: - eval_time: 8m alertname: FelisSignInFailures exp_alerts: [] - eval_time: 25m alertname: FelisSignInFailures exp_alerts: - exp_labels: severity: warning exp_annotations: summary: "over 30 refused sign-ins in 15 minutes" description: >- Wrong codes, unknown addresses or bad passkey assertions well above people mistyping: someone is guessing or enumerating. `sum by (door, reason) (increase(felis_auth_failures_total[15m]))` shows where; the audit rows (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17). - name: people mistyping stays quiet interval: 1m input_series: - series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}' values: '0 0 1 1 2 2 3 3 4 4 5x30' alert_rule_test: - eval_time: 30m alertname: FelisSignInFailures exp_alerts: [] - name: audit rows lost interval: 1m input_series: - series: 'felis_audit_write_failures_total{job="felis-api",instance="api-0"}' values: '0 0 0 2x20' alert_rule_test: - eval_time: 2m alertname: FelisAuditWriteFailing exp_alerts: [] - eval_time: 5m alertname: FelisAuditWriteFailing exp_alerts: - exp_labels: severity: warning job: felis-api instance: api-0 exp_annotations: summary: "felis-api failed to write audit rows" description: >- The actions went through but their audit rows were lost. The felis-api log names each lost row (`audit: lost ...`); the usual cause is PostgreSQL being unreachable or out of disk.
deploy/alerts/felis-prometheusrule.yaml +22 −0 Changes for deploy/alerts/felis-prometheusrule.yaml: 22 added lines, 0 removed lines. Original line number Diff line number Diff line Loading @@ -144,3 +144,25 @@ spec: The audit log names the account (action auth.otp.locked); the owner was mailed. Unless they fumbled codes, someone is guessing at it (troubleshooting §17). - alert: FelisSignInFailures expr: sum(increase(felis_auth_failures_total[15m])) > 30 for: 5m labels: severity: warning annotations: summary: "over 30 refused sign-ins in 15 minutes" description: >- Wrong codes, unknown addresses or bad passkey assertions well above people mistyping: someone is guessing or enumerating. `sum by (door, reason) (increase(felis_auth_failures_total[15m]))` shows where; the audit rows (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17). - alert: FelisAuditWriteFailing expr: increase(felis_audit_write_failures_total[15m]) > 0 labels: severity: warning annotations: summary: "felis-api failed to write audit rows" description: >- The actions went through but their audit rows were lost. The felis-api log names each lost row (`audit: lost ...`); the usual cause is PostgreSQL being unreachable or out of disk.
docs/troubleshooting.md +31 −4 Changes for docs/troubleshooting.md: 31 added lines, 4 removed lines. Original line number Diff line number Diff line Loading @@ -933,7 +933,9 @@ The series come from two processes: - `felis-api` internal face `:8081/metrics` (Service `felis-api-internal`) — `felis_image_build_failures_total`, and the sign-in series of §17 (`felis_mail_total`, `felis_rate_limited_total`, `felis_auth_otp_lockouts_total`). Unauthenticated like the probes; `felis_auth_otp_lockouts_total`, `felis_auth_failures_total`, `felis_sessions_revoked_total`, `felis_audit_write_failures_total`). Unauthenticated like the probes; ClusterIP-only, and the external face never serves it. - `felis_reaper_worlds_deleted_total` is produced inside the one-shot reaper CronJob, which exits long before any scrape interval — without a pushgateway Loading @@ -945,8 +947,8 @@ The series come from two processes: `deploy/alerts/` ships ready-made rules: build failures, slow starts, node disk/memory thresholds, the kubelet `DiskPressure` condition, control-plane database backup freshness (§16; needs node-exporter's textfile collector), and sign-in abuse: the mail budget, relay failures, throttled floods and account code locks (§17). sign-in abuse: the mail budget, relay failures, throttled floods, account code locks and the refused sign-in rate, plus lost audit rows (§17). - Plain Prometheus: add `felis-alerts.yaml` to `rule_files`. Check and unit-test it standalone with `promtool check rules felis-alerts.yaml` and Loading Loading @@ -1171,7 +1173,7 @@ Skips the pre-migration snapshot (`migrate up -no-backup`). The installer warns loudly when it is set. Use it only when the snapshot cannot work and you have another backup, e.g. an external database newer than the host's `pg_dump`. ## 17. Sign-in refused with 429, mail budget, account code locks ## 17. Sign-in refused with 429, mail budget, account code locks, failed sign-ins The public sign-in doors (`/api/v1/auth/*` except logout and the op-login status poll) have three limits of their own. Each answers 429 with a Loading Loading @@ -1232,6 +1234,29 @@ sudo -u postgres psql felis -c \ "DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '<name>');" ``` ### Who tried: the audit trail Every refused sign-in writes an `auth.<door>.failed` row (payload `reason`: `bad_code`, `no_account`, `staff_account`, `not_staff`, `bad_assertion`, ...) and counts in `felis_auth_failures_total{door,reason}`; `FelisSignInFailures` fires above 30 in 15 minutes. The first refusal of each throttled burst writes `auth.rate_limited` with the source address. Rows carry `actor_user_id` (the account, the column to attribute by), `client_ip` (the same address the limit keys on) and `user_agent`. `actor` is display text: a verified email or the username, never an address the caller set without verifying. ```sh sudo -u postgres psql felis -c " SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason FROM audit_logs WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour' ORDER BY created_at DESC LIMIT 50;" ``` A failed audit write does not fail the action; it logs `audit: lost ...` in `felis-api` and counts in `felis_audit_write_failures_total` (`FelisAuditWriteFailing`). The cause is almost always PostgreSQL (§16). ### Optional: a Cloudflare rate limiting rule in front The limits above live in the API, so they hold on any edge. Behind Cloudflare Loading Loading @@ -1274,3 +1299,5 @@ for 10 seconds (the Free plan's limits). | Sign-in 429 `rate_limited` for everyone at once | §17 | | 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 | | Right code refused; `otp_account_locked` / `FelisOTPAccountLocked` | §17 | | `FelisSignInFailures` / who is guessing, from where | §17 | | `FelisAuditWriteFailing` | §17 |
internal/api/api_test.go +6 −1 Changes for internal/api/api_test.go: 6 added lines, 1 removed line. Original line number Diff line number Diff line Loading @@ -47,6 +47,7 @@ type fakeRepo struct { serverResources map[string]ResourceSpec resourceUpdates map[string]ResourceSpec audits []AuditEntry failAudit error // Audit fails with it (a store outage) joins []string // create-server seeding (spec §15) seeded map[string]bool // name -> servers row exists Loading Loading @@ -745,6 +746,9 @@ func (f *fakeRepo) SeedServer(_ context.Context, name, subdomain string, _, _, _ } func (f *fakeRepo) Ping(_ context.Context) error { return f.pingErr } func (f *fakeRepo) Audit(_ context.Context, e AuditEntry) error { if f.failAudit != nil { return f.failAudit } f.audits = append(f.audits, e) return nil } Loading Loading @@ -857,7 +861,8 @@ func (f *fakeRepo) SessionUser(_ context.Context, tokenHash string, now time.Tim return nil, ErrNotFound } return &SessionedUser{ ID: u.ID, Email: u.Email, Role: u.Role, ID: u.ID, Username: u.Username, Email: u.Email, Role: u.Role, EmailVerified: u.EmailVerified, }, nil } } Loading