Unverified Commit 857c73a6 authored by Lemon-miaow's avatar Lemon-miaow
Browse files

fix(audit): 审计按账号 id 归属并记录来源 IP/UA,登录失败与限速入审计和指标,写入失败计数告警

parent c4e4953f
Loading
Loading
Loading
Loading
+25 −1
Changes for deploy/alerts/felis-alerts.yaml: 25 added lines, 1 removed line.
Original line number Diff line number Diff line
@@ -6,7 +6,9 @@
#   - felis-operator pod :8080/metrics      → felis_servers_total, felis_start_duration_seconds
#   - felis-api internal :8081/metrics      → felis_image_build_failures_total,
#                                             felis_mail_total, felis_rate_limited_total,
#                                             felis_auth_otp_lockouts_total
#                                             felis_auth_otp_lockouts_total,
#                                             felis_auth_failures_total,
#                                             felis_audit_write_failures_total
#   - node-exporter textfile collector      → felis_db_backup_* (felis-db-backup.timer)
# node_* / kube_* series come from node-exporter / kube-state-metrics.
groups:
@@ -142,3 +144,25 @@ groups:
            The audit log names the account (action auth.otp.locked); the owner was
            mailed. Unless they fumbled codes, someone is guessing at it
            (troubleshooting §17).
      - alert: FelisSignInFailures
        expr: sum(increase(felis_auth_failures_total[15m])) > 30
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "over 30 refused sign-ins in 15 minutes"
          description: >-
            Wrong codes, unknown addresses or bad passkey assertions well above people
            mistyping: someone is guessing or enumerating. `sum by (door, reason)
            (increase(felis_auth_failures_total[15m]))` shows where; the audit rows
            (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
      - alert: FelisAuditWriteFailing
        expr: increase(felis_audit_write_failures_total[15m]) > 0
        labels:
          severity: warning
        annotations:
          summary: "felis-api failed to write audit rows"
          description: >-
            The actions went through but their audit rows were lost. The felis-api log
            names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
            being unreachable or out of disk.
+55 −0
Changes for deploy/alerts/felis-alerts_test.yml: 55 added lines, 0 removed lines.
Original line number Diff line number Diff line
@@ -236,3 +236,58 @@ tests:
      - eval_time: 90m
        alertname: FelisOTPAccountLocked
        exp_alerts: []
  - name: sign-in failure rate
    interval: 1m
    input_series:
      # Two doors failing at 3/min between them from t=0.
      - series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}'
        values: '0+2x40'
      - series: 'felis_auth_failures_total{door="op_login",reason="no_account",job="felis-api"}'
        values: '0+1x40'
    alert_rule_test:
      - eval_time: 8m
        alertname: FelisSignInFailures
        exp_alerts: []
      - eval_time: 25m
        alertname: FelisSignInFailures
        exp_alerts:
          - exp_labels:
              severity: warning
            exp_annotations:
              summary: "over 30 refused sign-ins in 15 minutes"
              description: >-
                Wrong codes, unknown addresses or bad passkey assertions well above people
                mistyping: someone is guessing or enumerating. `sum by (door, reason)
                (increase(felis_auth_failures_total[15m]))` shows where; the audit rows
                (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
  - name: people mistyping stays quiet
    interval: 1m
    input_series:
      - series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}'
        values: '0 0 1 1 2 2 3 3 4 4 5x30'
    alert_rule_test:
      - eval_time: 30m
        alertname: FelisSignInFailures
        exp_alerts: []
  - name: audit rows lost
    interval: 1m
    input_series:
      - series: 'felis_audit_write_failures_total{job="felis-api",instance="api-0"}'
        values: '0 0 0 2x20'
    alert_rule_test:
      - eval_time: 2m
        alertname: FelisAuditWriteFailing
        exp_alerts: []
      - eval_time: 5m
        alertname: FelisAuditWriteFailing
        exp_alerts:
          - exp_labels:
              severity: warning
              job: felis-api
              instance: api-0
            exp_annotations:
              summary: "felis-api failed to write audit rows"
              description: >-
                The actions went through but their audit rows were lost. The felis-api log
                names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
                being unreachable or out of disk.
+22 −0
Changes for deploy/alerts/felis-prometheusrule.yaml: 22 added lines, 0 removed lines.
Original line number Diff line number Diff line
@@ -144,3 +144,25 @@ spec:
              The audit log names the account (action auth.otp.locked); the owner was
              mailed. Unless they fumbled codes, someone is guessing at it
              (troubleshooting §17).
        - alert: FelisSignInFailures
          expr: sum(increase(felis_auth_failures_total[15m])) > 30
          for: 5m
          labels:
            severity: warning
          annotations:
            summary: "over 30 refused sign-ins in 15 minutes"
            description: >-
              Wrong codes, unknown addresses or bad passkey assertions well above people
              mistyping: someone is guessing or enumerating. `sum by (door, reason)
              (increase(felis_auth_failures_total[15m]))` shows where; the audit rows
              (action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
        - alert: FelisAuditWriteFailing
          expr: increase(felis_audit_write_failures_total[15m]) > 0
          labels:
            severity: warning
          annotations:
            summary: "felis-api failed to write audit rows"
            description: >-
              The actions went through but their audit rows were lost. The felis-api log
              names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
              being unreachable or out of disk.
+31 −4
Changes for docs/troubleshooting.md: 31 added lines, 4 removed lines.
Original line number Diff line number Diff line
@@ -933,7 +933,9 @@ The series come from two processes:
- `felis-api` internal face `:8081/metrics` (Service `felis-api-internal`) —
  `felis_image_build_failures_total`, and the sign-in series of §17
  (`felis_mail_total`, `felis_rate_limited_total`,
  `felis_auth_otp_lockouts_total`). Unauthenticated like the probes;
  `felis_auth_otp_lockouts_total`, `felis_auth_failures_total`,
  `felis_sessions_revoked_total`, `felis_audit_write_failures_total`).
  Unauthenticated like the probes;
  ClusterIP-only, and the external face never serves it.
- `felis_reaper_worlds_deleted_total` is produced inside the one-shot reaper
  CronJob, which exits long before any scrape interval — without a pushgateway
@@ -945,8 +947,8 @@ The series come from two processes:
`deploy/alerts/` ships ready-made rules: build failures, slow starts, node
disk/memory thresholds, the kubelet `DiskPressure` condition, control-plane
database backup freshness (§16; needs node-exporter's textfile collector), and
sign-in abuse: the mail budget, relay failures, throttled floods and account
code locks (§17).
sign-in abuse: the mail budget, relay failures, throttled floods, account
code locks and the refused sign-in rate, plus lost audit rows (§17).

- Plain Prometheus: add `felis-alerts.yaml` to `rule_files`. Check and unit-test
  it standalone with `promtool check rules felis-alerts.yaml` and
@@ -1171,7 +1173,7 @@ Skips the pre-migration snapshot (`migrate up -no-backup`). The installer warns
loudly when it is set. Use it only when the snapshot cannot work and you have
another backup, e.g. an external database newer than the host's `pg_dump`.

## 17. Sign-in refused with 429, mail budget, account code locks
## 17. Sign-in refused with 429, mail budget, account code locks, failed sign-ins

The public sign-in doors (`/api/v1/auth/*` except logout and the op-login
status poll) have three limits of their own. Each answers 429 with a
@@ -1232,6 +1234,29 @@ sudo -u postgres psql felis -c \
  "DELETE FROM otp_failure_windows WHERE user_id = (SELECT id FROM users WHERE username = '<name>');"
```

### Who tried: the audit trail

Every refused sign-in writes an `auth.<door>.failed` row (payload `reason`:
`bad_code`, `no_account`, `staff_account`, `not_staff`, `bad_assertion`, ...)
and counts in `felis_auth_failures_total{door,reason}`; `FelisSignInFailures`
fires above 30 in 15 minutes. The first refusal of each throttled burst writes
`auth.rate_limited` with the source address. Rows carry `actor_user_id` (the
account, the column to attribute by), `client_ip` (the same address the limit
keys on) and `user_agent`. `actor` is display text: a verified email or the
username, never an address the caller set without verifying.

```sh
sudo -u postgres psql felis -c "
  SELECT created_at, action, actor, client_ip, payload->>'reason' AS reason
  FROM audit_logs
  WHERE action LIKE 'auth.%' AND created_at > now() - interval '1 hour'
  ORDER BY created_at DESC LIMIT 50;"
```

A failed audit write does not fail the action; it logs `audit: lost ...` in
`felis-api` and counts in `felis_audit_write_failures_total`
(`FelisAuditWriteFailing`). The cause is almost always PostgreSQL (§16).

### Optional: a Cloudflare rate limiting rule in front

The limits above live in the API, so they hold on any edge. Behind Cloudflare
@@ -1274,3 +1299,5 @@ for 10 seconds (the Free plan's limits).
| Sign-in 429 `rate_limited` for everyone at once | §17 |
| 429 `mail_rate_limited` / `FelisMailBudgetExhausted` | §17 |
| Right code refused; `otp_account_locked` / `FelisOTPAccountLocked` | §17 |
| `FelisSignInFailures` / who is guessing, from where | §17 |
| `FelisAuditWriteFailing` | §17 |
+6 −1
Changes for internal/api/api_test.go: 6 added lines, 1 removed line.
Original line number Diff line number Diff line
@@ -47,6 +47,7 @@ type fakeRepo struct {
	serverResources map[string]ResourceSpec
	resourceUpdates map[string]ResourceSpec
	audits          []AuditEntry
	failAudit       error // Audit fails with it (a store outage)
	joins           []string
	// create-server seeding (spec §15)
	seeded  map[string]bool   // name -> servers row exists
@@ -745,6 +746,9 @@ func (f *fakeRepo) SeedServer(_ context.Context, name, subdomain string, _, _, _
}
func (f *fakeRepo) Ping(_ context.Context) error { return f.pingErr }
func (f *fakeRepo) Audit(_ context.Context, e AuditEntry) error {
	if f.failAudit != nil {
		return f.failAudit
	}
	f.audits = append(f.audits, e)
	return nil
}
@@ -857,7 +861,8 @@ func (f *fakeRepo) SessionUser(_ context.Context, tokenHash string, now time.Tim
				return nil, ErrNotFound
			}
			return &SessionedUser{
				ID: u.ID, Email: u.Email, Role: u.Role,
				ID: u.ID, Username: u.Username, Email: u.Email, Role: u.Role,
				EmailVerified: u.EmailVerified,
			}, nil
		}
	}
Loading