145 lines
6.7 KiB
YAML
145 lines
6.7 KiB
YAML
# Felis alert rules — plain Prometheus format (also the promtool-tested source
|
|
# for felis-prometheusrule.yaml). See docs/troubleshooting.md §14 for scraping
|
|
# and loading instructions.
|
|
#
|
|
# felis_* series come from two processes:
|
|
# - felis-operator pod :8080/metrics → felis_servers_total, felis_start_duration_seconds
|
|
# - felis-api internal :8081/metrics → felis_image_build_failures_total,
|
|
# felis_mail_total, felis_rate_limited_total,
|
|
# felis_auth_otp_lockouts_total
|
|
# - node-exporter textfile collector → felis_db_backup_* (felis-db-backup.timer)
|
|
# node_* / kube_* series come from node-exporter / kube-state-metrics.
|
|
groups:
|
|
- name: felis.rules
|
|
rules:
|
|
- alert: FelisImageBuildFailures
|
|
expr: increase(felis_image_build_failures_total[6h]) > 0
|
|
for: 5m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "modpack/image build failed in the last 6h"
|
|
description: >-
|
|
felis_image_build_failures_total increased. Inspect the failed build Job
|
|
(kubectl logs -n felis-build job/<build-job>); the same error text is on
|
|
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
|
|
- alert: FelisSlowServerStarts
|
|
expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300
|
|
for: 15m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "p90 server start time exceeds 5 minutes"
|
|
description: >-
|
|
Starts regularly take over five minutes (felis_start_duration_seconds,
|
|
observed when readiness is first reached). A start that never completes
|
|
records nothing — cross-check desiredState=Running servers with no ready
|
|
phase (troubleshooting §1).
|
|
- name: felis.node.rules
|
|
rules:
|
|
- alert: FelisNodeDiskSpaceLow
|
|
expr: >-
|
|
node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
|
|
/ node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15
|
|
for: 15m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "node filesystem {{ $labels.mountpoint }} below 15% available"
|
|
description: >-
|
|
Sustained disk pressure evicts game pods and garbage-collects images
|
|
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
|
|
- alert: FelisNodeDiskPressure
|
|
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
|
|
for: 5m
|
|
labels:
|
|
severity: critical
|
|
annotations:
|
|
summary: "kubelet reports DiskPressure on {{ $labels.node }}"
|
|
description: >-
|
|
The eviction chain is in progress: control-plane pods hold
|
|
system-cluster-critical and survive, game pods do not. Free disk now
|
|
(troubleshooting §13b).
|
|
- alert: FelisNodeMemoryLow
|
|
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
|
|
for: 15m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "node memory available below 10% for 15m"
|
|
description: >-
|
|
PostgreSQL, the control plane, the registry and game servers share one
|
|
node; sustained memory pressure risks OOM kills.
|
|
- name: felis.backup.rules
|
|
rules:
|
|
- alert: FelisDBBackupStale
|
|
expr: time() - max(felis_db_backup_last_success_timestamp_seconds) > 26 * 3600
|
|
for: 10m
|
|
labels:
|
|
severity: critical
|
|
annotations:
|
|
summary: "no control-plane database backup in over 26h"
|
|
description: >-
|
|
felis-db-backup.timer runs daily; the newest bundle is more than a day
|
|
old. Read `journalctl -u felis-db-backup` on the host, then take one now
|
|
with `sudo felis db backup` (troubleshooting §16).
|
|
- alert: FelisDBBackupMetricMissing
|
|
expr: absent(felis_db_backup_last_success_timestamp_seconds)
|
|
for: 2h
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "database backup freshness is not being scraped"
|
|
description: >-
|
|
No felis_db_backup_last_success_timestamp_seconds series, so
|
|
FelisDBBackupStale cannot fire. Point node-exporter's
|
|
--collector.textfile.directory at the directory of
|
|
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
|
|
(troubleshooting §16).
|
|
- name: felis.auth.rules
|
|
rules:
|
|
- alert: FelisMailBudgetExhausted
|
|
expr: sum(increase(felis_mail_total{result="throttled"}[15m])) > 0
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "the install-wide mail budget refused mail"
|
|
description: >-
|
|
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
|
|
spent, and every sign-in code is refused with 429 mail_rate_limited until
|
|
it refills. Check felis_rate_limited_total for a flood before raising the
|
|
budget (troubleshooting §17).
|
|
- alert: FelisMailDeliveryFailing
|
|
expr: sum(increase(felis_mail_total{result="failed"}[15m])) > 0
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "the SMTP relay refused mail in the last 15m"
|
|
description: >-
|
|
felis_mail_total{result="failed"} increased: sign-in codes are not being
|
|
delivered (502 mail_undeliverable). The relay's reason is in the
|
|
felis-api log (troubleshooting §17).
|
|
- alert: FelisSignInFlood
|
|
expr: sum(rate(felis_rate_limited_total{scope="auth_door"}[5m])) * 60 > 10
|
|
for: 10m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "sign-in doors refusing over 10 requests a minute"
|
|
description: >-
|
|
The per-address sign-in limit has been refusing callers for 10 minutes.
|
|
A script is hammering the auth doors; if real users report rate_limited
|
|
at once instead, [auth] client_ip_header is missing and everyone shares
|
|
the proxy's address (troubleshooting §17).
|
|
- alert: FelisOTPAccountLocked
|
|
expr: sum by (purpose) (increase(felis_auth_otp_lockouts_total[1h])) > 0
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "an account's email-code sign-in locked after 10 wrong codes"
|
|
description: >-
|
|
Someone entered 10 wrong codes for one account within 24h ({{ $labels.purpose }}).
|
|
The audit log names the account (action auth.otp.locked); the owner was
|
|
mailed. Unless they fumbled codes, someone is guessing at it
|
|
(troubleshooting §17).
|