Files
Felis/deploy/alerts/felis-alerts_test.yml
T

239 lines
9.6 KiB
YAML

# promtool unit tests: `promtool test rules felis-alerts_test.yml`
# Proves every shipped rule actually fires on its target condition (and stays
# silent before it).
rule_files:
- felis-alerts.yaml
evaluation_interval: 1m
tests:
- name: build failure and slow starts
interval: 1m
input_series:
# counter: quiet for 5m, then one failure per step.
- series: 'felis_image_build_failures_total'
values: '0x5 1x15'
# histogram: all observations land in the (300,600] bucket.
- series: 'felis_start_duration_seconds_bucket{le="120"}'
values: '0x22'
- series: 'felis_start_duration_seconds_bucket{le="300"}'
values: '0x22'
- series: 'felis_start_duration_seconds_bucket{le="600"}'
values: '0+10x21'
- series: 'felis_start_duration_seconds_bucket{le="+Inf"}'
values: '0+10x21'
alert_rule_test:
- eval_time: 2m
alertname: FelisImageBuildFailures
exp_alerts: []
- eval_time: 20m
alertname: FelisImageBuildFailures
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "modpack/image build failed in the last 6h"
description: >-
felis_image_build_failures_total increased. Inspect the failed build Job
(kubectl logs -n felis-build job/<build-job>); the same error text is on
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
- eval_time: 20m
alertname: FelisSlowServerStarts
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "p90 server start time exceeds 5 minutes"
description: >-
Starts regularly take over five minutes (felis_start_duration_seconds,
observed when readiness is first reached). A start that never completes
records nothing — cross-check desiredState=Running servers with no ready
phase (troubleshooting §1).
- name: node disk and memory thresholds
interval: 1m
input_series:
- series: 'node_filesystem_avail_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
values: '10x26'
- series: 'node_filesystem_size_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
values: '100x26'
- series: 'kube_node_status_condition{condition="DiskPressure",node="n1",status="true"}'
values: '0x4 1x22'
- series: 'node_memory_MemAvailable_bytes{instance="node1",job="node-exporter"}'
values: '5x26'
- series: 'node_memory_MemTotal_bytes{instance="node1",job="node-exporter"}'
values: '100x26'
alert_rule_test:
- eval_time: 2m
alertname: FelisNodeDiskPressure
exp_alerts: []
- eval_time: 20m
alertname: FelisNodeDiskSpaceLow
exp_alerts:
- exp_labels:
device: /dev/vda1
fstype: xfs
instance: node1
job: node-exporter
mountpoint: /
severity: warning
exp_annotations:
summary: "node filesystem / below 15% available"
description: >-
Sustained disk pressure evicts game pods and garbage-collects images
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
- eval_time: 20m
alertname: FelisNodeDiskPressure
exp_alerts:
- exp_labels:
condition: DiskPressure
node: n1
status: "true"
severity: critical
exp_annotations:
summary: "kubelet reports DiskPressure on n1"
description: >-
The eviction chain is in progress: control-plane pods hold
system-cluster-critical and survive, game pods do not. Free disk now
(troubleshooting §13b).
- eval_time: 20m
alertname: FelisNodeMemoryLow
exp_alerts:
- exp_labels:
instance: node1
job: node-exporter
severity: warning
exp_annotations:
summary: "node memory available below 10% for 15m"
description: >-
PostgreSQL, the control plane, the registry and game servers share one
node; sustained memory pressure risks OOM kills.
- name: database backup freshness
interval: 1m
input_series:
# The newest bundle was taken at t=0 and none since.
- series: 'felis_db_backup_last_success_timestamp_seconds{instance="node1",job="node-exporter",label="daily"}'
values: '0x1630'
alert_rule_test:
- eval_time: 25h
alertname: FelisDBBackupStale
exp_alerts: []
- eval_time: 27h
alertname: FelisDBBackupStale
exp_alerts:
- exp_labels:
severity: critical
exp_annotations:
summary: "no control-plane database backup in over 26h"
description: >-
felis-db-backup.timer runs daily; the newest bundle is more than a day
old. Read `journalctl -u felis-db-backup` on the host, then take one now
with `sudo felis db backup` (troubleshooting §16).
- eval_time: 27h
alertname: FelisDBBackupMetricMissing
exp_alerts: []
- name: database backup freshness not scraped
interval: 1m
input_series:
- series: 'up{job="node-exporter"}'
values: '1x200'
alert_rule_test:
- eval_time: 1h
alertname: FelisDBBackupMetricMissing
exp_alerts: []
- eval_time: 3h
alertname: FelisDBBackupMetricMissing
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "database backup freshness is not being scraped"
description: >-
No felis_db_backup_last_success_timestamp_seconds series, so
FelisDBBackupStale cannot fire. Point node-exporter's
--collector.textfile.directory at the directory of
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
(troubleshooting §16).
- name: sign-in mail budget and relay
interval: 1m
input_series:
# Created at zero on start; the budget refuses one mail at t=3m.
- series: 'felis_mail_total{kind="otp",result="throttled",job="felis-api"}'
values: '0 0 0 1x30'
- series: 'felis_mail_total{kind="otp",result="failed",job="felis-api"}'
values: '0x33'
alert_rule_test:
- eval_time: 2m
alertname: FelisMailBudgetExhausted
exp_alerts: []
- eval_time: 5m
alertname: FelisMailBudgetExhausted
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "the install-wide mail budget refused mail"
description: >-
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
spent, and every sign-in code is refused with 429 mail_rate_limited until
it refills. Check felis_rate_limited_total for a flood before raising the
budget (troubleshooting §17).
- eval_time: 5m
alertname: FelisMailDeliveryFailing
exp_alerts: []
- name: sign-in flood
interval: 1m
input_series:
# 30 refusals a minute from t=0; a lone refused script at 2/min stays quiet.
- series: 'felis_rate_limited_total{scope="auth_door",job="felis-api"}'
values: '0+30x40'
alert_rule_test:
- eval_time: 10m
alertname: FelisSignInFlood
exp_alerts: []
- eval_time: 20m
alertname: FelisSignInFlood
exp_alerts:
- exp_labels:
severity: warning
exp_annotations:
summary: "sign-in doors refusing over 10 requests a minute"
description: >-
The per-address sign-in limit has been refusing callers for 10 minutes.
A script is hammering the auth doors; if real users report rate_limited
at once instead, [auth] client_ip_header is missing and everyone shares
the proxy's address (troubleshooting §17).
- name: sign-in trickle stays quiet
interval: 1m
input_series:
- series: 'felis_rate_limited_total{scope="auth_door",job="felis-api"}'
values: '0+2x40'
alert_rule_test:
- eval_time: 30m
alertname: FelisSignInFlood
exp_alerts: []
- name: account email-code lock
interval: 1m
input_series:
- series: 'felis_auth_otp_lockouts_total{purpose="login_email",job="felis-api"}'
values: '0 0 1x90'
- series: 'felis_auth_otp_lockouts_total{purpose="op_login",job="felis-api"}'
values: '0x92'
alert_rule_test:
- eval_time: 1m
alertname: FelisOTPAccountLocked
exp_alerts: []
- eval_time: 10m
alertname: FelisOTPAccountLocked
exp_alerts:
- exp_labels:
severity: warning
purpose: login_email
exp_annotations:
summary: "an account's email-code sign-in locked after 10 wrong codes"
description: >-
Someone entered 10 wrong codes for one account within 24h (login_email).
The audit log names the account (action auth.otp.locked); the owner was
mailed. Unless they fumbled codes, someone is guessing at it
(troubleshooting §17).
- eval_time: 90m
alertname: FelisOTPAccountLocked
exp_alerts: []