# Felis alert rules — plain Prometheus format (also the promtool-tested source # for felis-prometheusrule.yaml). See docs/troubleshooting.md §14 for scraping # and loading instructions. These are for a deployment that brings its own # Prometheus; every install already runs `felis watchdog` on the host, which # checks the same conditions without one and mails the owners (§14). # # felis_* series come from these processes: # - felis-operator-metrics Service :8080 → felis_servers_total, felis_server_phase, # felis_start_duration_seconds, # felis_build_info{component="operator"}, # controller_runtime_*, workqueue_* # - felis-api-internal Service :8081 → felis_image_build_failures_total, # felis_mail_total, felis_rate_limited_total, # felis_auth_otp_lockouts_total, # felis_auth_failures_total, # felis_audit_write_failures_total, # felis_build_info{component="api"} # - node-exporter textfile collector → felis_db_backup_* (felis-db-backup.timer) # node_* / kube_* series come from node-exporter / kube-state-metrics. groups: - name: felis.rules rules: - alert: FelisImageBuildFailures expr: increase(felis_image_build_failures_total[6h]) > 0 for: 5m labels: severity: warning annotations: summary: "modpack/image build failed in the last 6h" description: >- felis_image_build_failures_total increased. Inspect the failed build Job (kubectl logs -n felis-build job/); the same error text is on GET /api/v1/images/build/{id} and in the submitter's row in the panel. - alert: FelisSlowServerStarts expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300 for: 15m labels: severity: warning annotations: summary: "p90 server start time exceeds 5 minutes" description: >- Starts regularly take over five minutes (felis_start_duration_seconds, observed when readiness is first reached). A start that never completes records nothing — cross-check desiredState=Running servers with no ready phase (troubleshooting §1). - name: felis.platform.rules rules: - alert: FelisOperatorDown expr: absent(felis_build_info{component="operator"}) for: 10m labels: severity: critical annotations: summary: "felis-operator is down or not scraped" description: >- No felis_build_info{component="operator"} series for 10 minutes. Without the operator no server starts, stops or recovers. Check `kubectl -n felis get deploy felis-operator` and its log; if the pod is healthy, the felis-operator-metrics Service is not being scraped (troubleshooting §14). - alert: FelisAPIDown expr: absent(felis_build_info{component="api"}) for: 10m labels: severity: critical annotations: summary: "felis-api is down or not scraped" description: >- No felis_build_info{component="api"} series for 10 minutes. The panel, sign-in and the proxy's player lookups all go through felis-api. Check `kubectl -n felis get deploy felis-api` and its log; if the pod is healthy, the felis-api-internal Service is not being scraped (troubleshooting §14). - alert: FelisLoginGateDown expr: felis_server_phase{role="login",desired="Running",phase!="Running"} == 1 for: 10m labels: severity: critical annotations: summary: "the login gate {{ $labels.server }} is {{ $labels.phase }}" description: >- Every player connection passes through the login server first, so no one can join. The MinecraftServer's conditions carry the reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}` (troubleshooting §1, §2). - alert: FelisSystemServerDown expr: felis_server_phase{role!="",role!="login",desired="Running",phase!="Running"} == 1 for: 10m labels: severity: warning annotations: summary: "system server {{ $labels.server }} ({{ $labels.role }}) is {{ $labels.phase }}" description: >- Players who sign in are sent to the lobby; while it is down they stay at the gate. `kubectl -n minecraft describe minecraftserver {{ $labels.server }}` shows the reason (troubleshooting §1, §2). - alert: FelisServerFailed expr: felis_server_phase{role="",phase="Failed"} == 1 for: 5m labels: severity: warning annotations: summary: "server {{ $labels.server }} is Failed" description: >- The operator gave up on this server (a crash loop, an image that will not pull, a world volume that will not mount). Its conditions carry the reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}` (troubleshooting §2). - alert: FelisReconcileErrors expr: sum(increase(controller_runtime_reconcile_errors_total{controller="minecraftserver"}[15m])) > 10 for: 5m labels: severity: warning annotations: summary: "the operator failed over 10 reconciles in 15 minutes" description: >- Server changes are being retried instead of applied. The felis-operator log names each failing server and its error. - alert: FelisReconcileStuck expr: max(workqueue_longest_running_processor_seconds{name="minecraftserver"}) > 300 for: 5m labels: severity: critical annotations: summary: "an operator reconcile has been running for over 5 minutes" description: >- Each reconcile is bounded at 3 minutes, so this one is ignoring its deadline and holding a worker. The liveness probe restarts the operator once a pass passes 10 minutes; the log from before the restart shows where it hung. - name: felis.jobs.rules rules: - alert: FelisWorldJobFailed expr: kube_job_failed{namespace="minecraft",condition="true"} == 1 labels: severity: warning annotations: summary: "Job {{ $labels.job_name }} failed" description: >- A world backup, restore or reaper run failed; after a failed backup that world's newest archive is older than planned. `kubectl -n minecraft logs job/{{ $labels.job_name }}` has the error (troubleshooting §10). - alert: FelisReaperStale expr: time() - kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"} > 26 * 3600 for: 10m labels: severity: warning annotations: summary: "the world reaper has not succeeded in over 26h" description: >- felis-reaper runs daily; idle worlds are neither backed up nor reclaimed while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp` lists its runs, and the newest one's log shows why (troubleshooting §10). - name: felis.node.rules rules: - alert: FelisNodeDiskSpaceLow expr: >- node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"} / node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15 for: 15m labels: severity: warning annotations: summary: "node filesystem {{ $labels.mountpoint }} below 15% available" description: >- Sustained disk pressure evicts game pods and garbage-collects images (troubleshooting §13b). Free space before kubelet raises DiskPressure. - alert: FelisNodeDiskPressure expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1 for: 5m labels: severity: critical annotations: summary: "kubelet reports DiskPressure on {{ $labels.node }}" description: >- The eviction chain is in progress: control-plane pods hold system-cluster-critical and survive, game pods do not. Free disk now (troubleshooting §13b). - alert: FelisNodeMemoryLow expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10 for: 15m labels: severity: warning annotations: summary: "node memory available below 10% for 15m" description: >- PostgreSQL, the control plane, the registry and game servers share one node; sustained memory pressure risks OOM kills. - name: felis.backup.rules rules: - alert: FelisDBBackupStale expr: time() - max(felis_db_backup_last_success_timestamp_seconds) > 26 * 3600 for: 10m labels: severity: critical annotations: summary: "no control-plane database backup in over 26h" description: >- felis-db-backup.timer runs daily; the newest bundle is more than a day old. Read `journalctl -u felis-db-backup` on the host, then take one now with `sudo felis db backup` (troubleshooting §16). - alert: FelisDBBackupMetricMissing expr: absent(felis_db_backup_last_success_timestamp_seconds) for: 2h labels: severity: warning annotations: summary: "database backup freshness is not being scraped" description: >- No felis_db_backup_last_success_timestamp_seconds series, so FelisDBBackupStale cannot fire. Point node-exporter's --collector.textfile.directory at the directory of FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector) (troubleshooting §16). - name: felis.auth.rules rules: - alert: FelisMailBudgetExhausted expr: sum(increase(felis_mail_total{result="throttled"}[15m])) > 0 labels: severity: warning annotations: summary: "the install-wide mail budget refused mail" description: >- felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is spent, and every sign-in code is refused with 429 mail_rate_limited until it refills. Check felis_rate_limited_total for a flood before raising the budget (troubleshooting §17). - alert: FelisMailDeliveryFailing expr: sum(increase(felis_mail_total{result="failed"}[15m])) > 0 labels: severity: warning annotations: summary: "the SMTP relay refused mail in the last 15m" description: >- felis_mail_total{result="failed"} increased: sign-in codes are not being delivered (502 mail_undeliverable). The relay's reason is in the felis-api log (troubleshooting §17). - alert: FelisSignInFlood expr: sum(rate(felis_rate_limited_total{scope="auth_door"}[5m])) * 60 > 10 for: 10m labels: severity: warning annotations: summary: "sign-in doors refusing over 10 requests a minute" description: >- The per-address sign-in limit has been refusing callers for 10 minutes. A script is hammering the auth doors; if real users report rate_limited at once instead, [auth] client_ip_header is missing and everyone shares the proxy's address (troubleshooting §17). - alert: FelisOTPAccountLocked expr: sum by (purpose) (increase(felis_auth_otp_lockouts_total[1h])) > 0 labels: severity: warning annotations: summary: "an account's email-code sign-in locked after 10 wrong codes" description: >- Someone entered 10 wrong codes for one account within 24h ({{ $labels.purpose }}). The audit log names the account (action auth.otp.locked); the owner was mailed. Unless they fumbled codes, someone is guessing at it (troubleshooting §17). - alert: FelisSignInFailures expr: sum(increase(felis_auth_failures_total[15m])) > 30 for: 5m labels: severity: warning annotations: summary: "over 30 refused sign-ins in 15 minutes" description: >- Wrong codes, unknown addresses or bad passkey assertions well above people mistyping: someone is guessing or enumerating. `sum by (door, reason) (increase(felis_auth_failures_total[15m]))` shows where; the audit rows (action auth..failed) carry each caller's client_ip (troubleshooting §17). - alert: FelisAuditWriteFailing expr: increase(felis_audit_write_failures_total[15m]) > 0 labels: severity: warning annotations: summary: "felis-api failed to write audit rows" description: >- The actions went through but their audit rows were lost. The felis-api log names each lost row (`audit: lost ...`); the usual cause is PostgreSQL being unreachable or out of disk.