feat(watchdog): 主机侧巡检定时器按异常邮件通知平台所有者,operator 增加 phase 与 build_info 指标、卡死存活探针与告警规则
This commit is contained in:
26 files changed
+2686
-34
No files matched your search
@@ -1,14 +1,20 @@
|
||||
# Felis alert rules — plain Prometheus format (also the promtool-tested source
|
||||
# for felis-prometheusrule.yaml). See docs/troubleshooting.md §14 for scraping
|
||||
# and loading instructions.
|
||||
# and loading instructions. These are for a deployment that brings its own
|
||||
# Prometheus; every install already runs `felis watchdog` on the host, which
|
||||
# checks the same conditions without one and mails the owners (§14).
|
||||
#
|
||||
# felis_* series come from two processes:
|
||||
# - felis-operator pod :8080/metrics → felis_servers_total, felis_start_duration_seconds
|
||||
# - felis-api internal :8081/metrics → felis_image_build_failures_total,
|
||||
# felis_* series come from these processes:
|
||||
# - felis-operator-metrics Service :8080 → felis_servers_total, felis_server_phase,
|
||||
# felis_start_duration_seconds,
|
||||
# felis_build_info{component="operator"},
|
||||
# controller_runtime_*, workqueue_*
|
||||
# - felis-api-internal Service :8081 → felis_image_build_failures_total,
|
||||
# felis_mail_total, felis_rate_limited_total,
|
||||
# felis_auth_otp_lockouts_total,
|
||||
# felis_auth_failures_total,
|
||||
# felis_audit_write_failures_total
|
||||
# felis_audit_write_failures_total,
|
||||
# felis_build_info{component="api"}
|
||||
# - node-exporter textfile collector → felis_db_backup_* (felis-db-backup.timer)
|
||||
# node_* / kube_* series come from node-exporter / kube-state-metrics.
|
||||
groups:
|
||||
@@ -37,6 +43,116 @@ groups:
|
||||
observed when readiness is first reached). A start that never completes
|
||||
records nothing — cross-check desiredState=Running servers with no ready
|
||||
phase (troubleshooting §1).
|
||||
- name: felis.platform.rules
|
||||
rules:
|
||||
- alert: FelisOperatorDown
|
||||
expr: absent(felis_build_info{component="operator"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-operator is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="operator"} series for 10 minutes. Without
|
||||
the operator no server starts, stops or recovers. Check
|
||||
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
|
||||
healthy, the felis-operator-metrics Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisAPIDown
|
||||
expr: absent(felis_build_info{component="api"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-api is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="api"} series for 10 minutes. The panel,
|
||||
sign-in and the proxy's player lookups all go through felis-api. Check
|
||||
`kubectl -n felis get deploy felis-api` and its log; if the pod is
|
||||
healthy, the felis-api-internal Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisLoginGateDown
|
||||
expr: felis_server_phase{role="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "the login gate {{ $labels.server }} is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Every player connection passes through the login server first, so no one
|
||||
can join. The MinecraftServer's conditions carry the reason:
|
||||
`kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §1, §2).
|
||||
- alert: FelisSystemServerDown
|
||||
expr: felis_server_phase{role!="",role!="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "system server {{ $labels.server }} ({{ $labels.role }}) is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Players who sign in are sent to the lobby; while it is down they stay at
|
||||
the gate. `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
shows the reason (troubleshooting §1, §2).
|
||||
- alert: FelisServerFailed
|
||||
expr: felis_server_phase{role="",phase="Failed"} == 1
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "server {{ $labels.server }} is Failed"
|
||||
description: >-
|
||||
The operator gave up on this server (a crash loop, an image that will not
|
||||
pull, a world volume that will not mount). Its conditions carry the
|
||||
reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §2).
|
||||
- alert: FelisReconcileErrors
|
||||
expr: sum(increase(controller_runtime_reconcile_errors_total{controller="minecraftserver"}[15m])) > 10
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the operator failed over 10 reconciles in 15 minutes"
|
||||
description: >-
|
||||
Server changes are being retried instead of applied. The felis-operator
|
||||
log names each failing server and its error.
|
||||
- alert: FelisReconcileStuck
|
||||
expr: max(workqueue_longest_running_processor_seconds{name="minecraftserver"}) > 300
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "an operator reconcile has been running for over 5 minutes"
|
||||
description: >-
|
||||
Each reconcile is bounded at 3 minutes, so this one is ignoring its
|
||||
deadline and holding a worker. The liveness probe restarts the operator
|
||||
once a pass passes 10 minutes; the log from before the restart shows
|
||||
where it hung.
|
||||
- name: felis.jobs.rules
|
||||
rules:
|
||||
- alert: FelisWorldJobFailed
|
||||
expr: kube_job_failed{namespace="minecraft",condition="true"} == 1
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Job {{ $labels.job_name }} failed"
|
||||
description: >-
|
||||
A world backup, restore or reaper run failed; after a failed backup that
|
||||
world's newest archive is older than planned.
|
||||
`kubectl -n minecraft logs job/{{ $labels.job_name }}` has the error
|
||||
(troubleshooting §10).
|
||||
- alert: FelisReaperStale
|
||||
expr: time() - kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"} > 26 * 3600
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the world reaper has not succeeded in over 26h"
|
||||
description: >-
|
||||
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
|
||||
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
|
||||
lists its runs, and the newest one's log
|
||||
shows why (troubleshooting §10).
|
||||
- name: felis.node.rules
|
||||
rules:
|
||||
- alert: FelisNodeDiskSpaceLow
|
||||
|
||||
@@ -291,3 +291,219 @@ tests:
|
||||
The actions went through but their audit rows were lost. The felis-api log
|
||||
names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
|
||||
being unreachable or out of disk.
|
||||
- name: operator and api presence
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_build_info{component="operator",version="v1",job="felis-operator",instance="op-0"}'
|
||||
values: '1x30'
|
||||
# felis-api stops being scraped after 5m; the series goes stale 5m later.
|
||||
- series: 'felis_build_info{component="api",version="v1",job="felis-api",instance="api-0"}'
|
||||
values: '1x5'
|
||||
alert_rule_test:
|
||||
- eval_time: 25m
|
||||
alertname: FelisOperatorDown
|
||||
exp_alerts: []
|
||||
- eval_time: 15m
|
||||
alertname: FelisAPIDown
|
||||
exp_alerts: []
|
||||
- eval_time: 25m
|
||||
alertname: FelisAPIDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
component: api
|
||||
exp_annotations:
|
||||
summary: "felis-api is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="api"} series for 10 minutes. The panel,
|
||||
sign-in and the proxy's player lookups all go through felis-api. Check
|
||||
`kubectl -n felis get deploy felis-api` and its log; if the pod is
|
||||
healthy, the felis-api-internal Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- name: operator never scraped
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_build_info{component="api",version="v1",job="felis-api",instance="api-0"}'
|
||||
values: '1x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 5m
|
||||
alertname: FelisOperatorDown
|
||||
exp_alerts: []
|
||||
- eval_time: 15m
|
||||
alertname: FelisOperatorDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
component: operator
|
||||
exp_annotations:
|
||||
summary: "felis-operator is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="operator"} series for 10 minutes. Without
|
||||
the operator no server starts, stops or recovers. Check
|
||||
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
|
||||
healthy, the felis-operator-metrics Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- name: system and user servers down
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_server_phase{server="login",role="login",phase="Starting",desired="Running"}'
|
||||
values: '1x20'
|
||||
- series: 'felis_server_phase{server="lobby",role="lobby",phase="Failed",desired="Running"}'
|
||||
values: '1x20'
|
||||
# Stopped on purpose: not an outage.
|
||||
- series: 'felis_server_phase{server="lobby2",role="lobby",phase="Stopped",desired="Stopped"}'
|
||||
values: '1x20'
|
||||
# A user server carries no role label (the operator publishes role="").
|
||||
- series: 'felis_server_phase{server="survival",phase="Failed",desired="Running"}'
|
||||
values: '1x20'
|
||||
- series: 'felis_server_phase{server="creative",phase="Running",desired="Running"}'
|
||||
values: '1x20'
|
||||
alert_rule_test:
|
||||
- eval_time: 9m
|
||||
alertname: FelisLoginGateDown
|
||||
exp_alerts: []
|
||||
- eval_time: 11m
|
||||
alertname: FelisLoginGateDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
server: login
|
||||
role: login
|
||||
phase: Starting
|
||||
desired: Running
|
||||
exp_annotations:
|
||||
summary: "the login gate login is Starting"
|
||||
description: >-
|
||||
Every player connection passes through the login server first, so no one
|
||||
can join. The MinecraftServer's conditions carry the reason:
|
||||
`kubectl -n minecraft describe minecraftserver login`
|
||||
(troubleshooting §1, §2).
|
||||
- eval_time: 11m
|
||||
alertname: FelisSystemServerDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
server: lobby
|
||||
role: lobby
|
||||
phase: Failed
|
||||
desired: Running
|
||||
exp_annotations:
|
||||
summary: "system server lobby (lobby) is Failed"
|
||||
description: >-
|
||||
Players who sign in are sent to the lobby; while it is down they stay at
|
||||
the gate. `kubectl -n minecraft describe minecraftserver lobby`
|
||||
shows the reason (troubleshooting §1, §2).
|
||||
- eval_time: 3m
|
||||
alertname: FelisServerFailed
|
||||
exp_alerts: []
|
||||
- eval_time: 6m
|
||||
alertname: FelisServerFailed
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
server: survival
|
||||
phase: Failed
|
||||
desired: Running
|
||||
exp_annotations:
|
||||
summary: "server survival is Failed"
|
||||
description: >-
|
||||
The operator gave up on this server (a crash loop, an image that will not
|
||||
pull, a world volume that will not mount). Its conditions carry the
|
||||
reason: `kubectl -n minecraft describe minecraftserver survival`
|
||||
(troubleshooting §2).
|
||||
- name: operator reconcile errors and a stuck pass
|
||||
interval: 1m
|
||||
input_series:
|
||||
# Two failed reconciles a minute from 6m on.
|
||||
- series: 'controller_runtime_reconcile_errors_total{controller="minecraftserver",job="felis-operator"}'
|
||||
values: '0x5 0+2x20'
|
||||
# One pass that started at 5m and never returns.
|
||||
- series: 'workqueue_longest_running_processor_seconds{name="minecraftserver",controller="minecraftserver",job="felis-operator"}'
|
||||
values: '0x5 60+60x20'
|
||||
alert_rule_test:
|
||||
- eval_time: 5m
|
||||
alertname: FelisReconcileErrors
|
||||
exp_alerts: []
|
||||
- eval_time: 25m
|
||||
alertname: FelisReconcileErrors
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "the operator failed over 10 reconciles in 15 minutes"
|
||||
description: >-
|
||||
Server changes are being retried instead of applied. The felis-operator
|
||||
log names each failing server and its error.
|
||||
- eval_time: 12m
|
||||
alertname: FelisReconcileStuck
|
||||
exp_alerts: []
|
||||
- eval_time: 20m
|
||||
alertname: FelisReconcileStuck
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
exp_annotations:
|
||||
summary: "an operator reconcile has been running for over 5 minutes"
|
||||
description: >-
|
||||
Each reconcile is bounded at 3 minutes, so this one is ignoring its
|
||||
deadline and holding a worker. The liveness probe restarts the operator
|
||||
once a pass passes 10 minutes; the log from before the restart shows
|
||||
where it hung.
|
||||
- name: world job failures and a late reaper
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'kube_job_failed{namespace="minecraft",job_name="backup-survival-abc",condition="true"}'
|
||||
values: '0x2 1x10'
|
||||
- series: 'kube_job_failed{namespace="minecraft",job_name="backup-survival-abc",condition="false"}'
|
||||
values: '1x2 0x10'
|
||||
# A build Job in another namespace is FelisImageBuildFailures' business.
|
||||
- series: 'kube_job_failed{namespace="felis-build",job_name="build-x",condition="true"}'
|
||||
values: '1x12'
|
||||
# Evaluation starts at the epoch, so "over a day ago" is a negative timestamp.
|
||||
- series: 'kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"}'
|
||||
values: '-100000x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 1m
|
||||
alertname: FelisWorldJobFailed
|
||||
exp_alerts: []
|
||||
- eval_time: 5m
|
||||
alertname: FelisWorldJobFailed
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
namespace: minecraft
|
||||
job_name: backup-survival-abc
|
||||
condition: "true"
|
||||
exp_annotations:
|
||||
summary: "Job backup-survival-abc failed"
|
||||
description: >-
|
||||
A world backup, restore or reaper run failed; after a failed backup that
|
||||
world's newest archive is older than planned.
|
||||
`kubectl -n minecraft logs job/backup-survival-abc` has the error
|
||||
(troubleshooting §10).
|
||||
- eval_time: 5m
|
||||
alertname: FelisReaperStale
|
||||
exp_alerts: []
|
||||
- eval_time: 15m
|
||||
alertname: FelisReaperStale
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
namespace: minecraft
|
||||
cronjob: felis-reaper
|
||||
exp_annotations:
|
||||
summary: "the world reaper has not succeeded in over 26h"
|
||||
description: >-
|
||||
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
|
||||
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
|
||||
lists its runs, and the newest one's log
|
||||
shows why (troubleshooting §10).
|
||||
- name: a reaper that ran yesterday stays quiet
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"}'
|
||||
values: '-50000x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 25m
|
||||
alertname: FelisReaperStale
|
||||
exp_alerts: []
|
||||
@@ -37,6 +37,116 @@ spec:
|
||||
observed when readiness is first reached). A start that never completes
|
||||
records nothing — cross-check desiredState=Running servers with no ready
|
||||
phase (troubleshooting §1).
|
||||
- name: felis.platform.rules
|
||||
rules:
|
||||
- alert: FelisOperatorDown
|
||||
expr: absent(felis_build_info{component="operator"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-operator is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="operator"} series for 10 minutes. Without
|
||||
the operator no server starts, stops or recovers. Check
|
||||
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
|
||||
healthy, the felis-operator-metrics Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisAPIDown
|
||||
expr: absent(felis_build_info{component="api"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-api is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="api"} series for 10 minutes. The panel,
|
||||
sign-in and the proxy's player lookups all go through felis-api. Check
|
||||
`kubectl -n felis get deploy felis-api` and its log; if the pod is
|
||||
healthy, the felis-api-internal Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisLoginGateDown
|
||||
expr: felis_server_phase{role="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "the login gate {{ $labels.server }} is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Every player connection passes through the login server first, so no one
|
||||
can join. The MinecraftServer's conditions carry the reason:
|
||||
`kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §1, §2).
|
||||
- alert: FelisSystemServerDown
|
||||
expr: felis_server_phase{role!="",role!="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "system server {{ $labels.server }} ({{ $labels.role }}) is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Players who sign in are sent to the lobby; while it is down they stay at
|
||||
the gate. `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
shows the reason (troubleshooting §1, §2).
|
||||
- alert: FelisServerFailed
|
||||
expr: felis_server_phase{role="",phase="Failed"} == 1
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "server {{ $labels.server }} is Failed"
|
||||
description: >-
|
||||
The operator gave up on this server (a crash loop, an image that will not
|
||||
pull, a world volume that will not mount). Its conditions carry the
|
||||
reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §2).
|
||||
- alert: FelisReconcileErrors
|
||||
expr: sum(increase(controller_runtime_reconcile_errors_total{controller="minecraftserver"}[15m])) > 10
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the operator failed over 10 reconciles in 15 minutes"
|
||||
description: >-
|
||||
Server changes are being retried instead of applied. The felis-operator
|
||||
log names each failing server and its error.
|
||||
- alert: FelisReconcileStuck
|
||||
expr: max(workqueue_longest_running_processor_seconds{name="minecraftserver"}) > 300
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "an operator reconcile has been running for over 5 minutes"
|
||||
description: >-
|
||||
Each reconcile is bounded at 3 minutes, so this one is ignoring its
|
||||
deadline and holding a worker. The liveness probe restarts the operator
|
||||
once a pass passes 10 minutes; the log from before the restart shows
|
||||
where it hung.
|
||||
- name: felis.jobs.rules
|
||||
rules:
|
||||
- alert: FelisWorldJobFailed
|
||||
expr: kube_job_failed{namespace="minecraft",condition="true"} == 1
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Job {{ $labels.job_name }} failed"
|
||||
description: >-
|
||||
A world backup, restore or reaper run failed; after a failed backup that
|
||||
world's newest archive is older than planned.
|
||||
`kubectl -n minecraft logs job/{{ $labels.job_name }}` has the error
|
||||
(troubleshooting §10).
|
||||
- alert: FelisReaperStale
|
||||
expr: time() - kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"} > 26 * 3600
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the world reaper has not succeeded in over 26h"
|
||||
description: >-
|
||||
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
|
||||
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
|
||||
lists its runs, and the newest one's log
|
||||
shows why (troubleshooting §10).
|
||||
- name: felis.node.rules
|
||||
rules:
|
||||
- alert: FelisNodeDiskSpaceLow
|
||||
|
||||
@@ -252,6 +252,13 @@ GOROOT_DIR="/opt/felis/go"
|
||||
NANO_SERVICE="/etc/systemd/system/felis-nano.service"
|
||||
DB_BACKUP_SERVICE="/etc/systemd/system/felis-db-backup.service"
|
||||
DB_BACKUP_TIMER="/etc/systemd/system/felis-db-backup.timer"
|
||||
WATCHDOG_SERVICE="/etc/systemd/system/felis-watchdog.service"
|
||||
WATCHDOG_TIMER="/etc/systemd/system/felis-watchdog.timer"
|
||||
WATCHDOG_STATE="/var/lib/felis/watchdog/state.json"
|
||||
# While this marker holds a future Unix time, felis watchdog mails nothing: an install
|
||||
# restarts the control plane and the system servers on purpose. cleanup removes it; the
|
||||
# time in it is the backstop for an installer killed before its EXIT trap runs.
|
||||
WATCHDOG_QUIET_FILE="/run/felis/watchdog-quiet-until"
|
||||
VELOCITY_DIR="/opt/felis/velocity"
|
||||
VELOCITY_USER="felis-velocity"
|
||||
VELOCITY_SERVICE="/etc/systemd/system/felis-velocity.service"
|
||||
@@ -336,6 +343,9 @@ cleanup() {
|
||||
for path in "${TEMP_PATHS[@]-}"; do
|
||||
[ -n "$path" ] && rm -rf -- "$path" || true
|
||||
done
|
||||
# A failed install leaves something broken the owners should hear about, so the
|
||||
# watchdog speaks again the moment the installer exits, however it exits.
|
||||
rm -f -- "$WATCHDOG_QUIET_FILE" 2>/dev/null || true
|
||||
}
|
||||
|
||||
remember_temp() { TEMP_PATHS+=("$1"); }
|
||||
@@ -2415,6 +2425,60 @@ run_migrations() {
|
||||
ok "migrations applied"
|
||||
}
|
||||
|
||||
# Silence felis watchdog's mail for the rest of this install (see WATCHDOG_QUIET_FILE).
|
||||
# Two hours covers a slow source build; cleanup lifts it as soon as the installer exits.
|
||||
quiet_watchdog() {
|
||||
install -d -m 0755 "$(dirname "$WATCHDOG_QUIET_FILE")"
|
||||
printf '%s\n' "$(( $(date +%s) + 7200 ))" > "$WATCHDOG_QUIET_FILE"
|
||||
}
|
||||
|
||||
# The platform watchdog: every two minutes it checks the control plane, the login gate,
|
||||
# the fleet, PostgreSQL, the game proxy, the database backups and the host's disks and
|
||||
# memory, and mails the owners (their verified addresses, over the [smtp] relay) what
|
||||
# has stayed wrong long enough to matter. It runs on the host so a k3s that is down is
|
||||
# still reported. The first run happens now, so a broken unit shows up in this install.
|
||||
install_watchdog_timer() {
|
||||
local disks="/,/var/lib/rancher/k3s,/var/lib/postgresql,/var/lib/felis" path
|
||||
for path in "$FELIS_WORLDS_HOST_PATH" "$FELIS_ARCHIVE_LOCAL_PATH" "$FELIS_DB_BACKUP_DIR"; do
|
||||
if [ -n "$path" ]; then disks="${disks},${path}"; fi
|
||||
done
|
||||
install -d -m 0700 "$(dirname "$WATCHDOG_STATE")"
|
||||
cat > "$WATCHDOG_SERVICE" <<EOF
|
||||
[Unit]
|
||||
Description=Felis platform watchdog (health checks, owner alert mail)
|
||||
After=network-online.target k3s.service postgresql.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=${HOST_BIN} watchdog -config ${STATE_DIR}/felis.host.toml -state ${WATCHDOG_STATE} -quiet-file ${WATCHDOG_QUIET_FILE} -backup-dir ${FELIS_DB_BACKUP_DIR} -proxy-addr 127.0.0.1:${FELIS_GAME_PORT} -disk-paths ${disks}
|
||||
TimeoutStartSec=3min
|
||||
Nice=5
|
||||
PrivateTmp=yes
|
||||
NoNewPrivileges=yes
|
||||
ProtectSystem=full
|
||||
EOF
|
||||
cat > "$WATCHDOG_TIMER" <<EOF
|
||||
[Unit]
|
||||
Description=Felis platform watchdog, every two minutes
|
||||
|
||||
[Timer]
|
||||
OnBootSec=3min
|
||||
OnUnitActiveSec=2min
|
||||
AccuracySec=15s
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
EOF
|
||||
systemctl daemon-reload
|
||||
systemctl enable --now felis-watchdog.timer
|
||||
if systemctl start felis-watchdog.service; then
|
||||
ok "watchdog: checks every 2 minutes and mails the owners' verified addresses (journalctl -u felis-watchdog)"
|
||||
else
|
||||
journalctl -u felis-watchdog.service -n 20 --no-pager >&2 || true
|
||||
warn "the first watchdog run failed (log above); nothing will be mailed until it runs: sudo systemctl start felis-watchdog.service"
|
||||
fi
|
||||
}
|
||||
|
||||
# The daily database backup. The first run happens now, so a broken pipeline (pg_dump
|
||||
# missing, directory unwritable) shows up in this install rather than in the first
|
||||
# restore someone needs.
|
||||
@@ -3019,6 +3083,7 @@ main() {
|
||||
main_nano
|
||||
return
|
||||
fi
|
||||
quiet_watchdog
|
||||
pause_package_background_timers
|
||||
detect_node_ip
|
||||
ensure_swap
|
||||
@@ -3073,6 +3138,8 @@ main() {
|
||||
install_velocity
|
||||
# After deploy_bundle: the bundle's MinecraftServer export reads the cluster.
|
||||
install_db_backup_timer
|
||||
# Last: its first run should see the platform as this install leaves it.
|
||||
install_watchdog_timer
|
||||
mark_bootstrap_done
|
||||
summary
|
||||
}
|
||||
|
||||
@@ -1152,6 +1152,75 @@ expect "a failed first backup shows its log" "JOURNAL: pg_dump: connection refus
|
||||
expect "a failed first backup is a loud warning" "WARN: the first database backup failed" "$out"
|
||||
rm -rf "$tdir"
|
||||
|
||||
wblock="$(awk '/^install_watchdog_timer\(\) \{/,/^}/' "$BS")"
|
||||
[ -n "$wblock" ] || { echo "FAIL: no install_watchdog_timer found in $BS"; exit 1; }
|
||||
[ "$(printf '%s\n' "$wblock" | wc -l)" -lt 60 ] \
|
||||
|| { echo "FAIL: the extracted block is not install_watchdog_timer -- did its closing brace move?"; exit 1; }
|
||||
qblock="$(awk '/^quiet_watchdog\(\) \{/,/^}/' "$BS")"
|
||||
[ -n "$qblock" ] || { echo "FAIL: no quiet_watchdog found in $BS"; exit 1; }
|
||||
|
||||
tdir="$(mktemp -d)"
|
||||
run_watchdog_timer() { # $1: exit status of the first run, $2: FELIS_WORLDS_HOST_PATH
|
||||
FIRST="$1" FELIS_WORLDS_HOST_PATH="$2" WATCHDOG_SERVICE="$tdir/felis-watchdog.service" WATCHDOG_TIMER="$tdir/felis-watchdog.timer" \
|
||||
WATCHDOG_STATE="$tdir/watchdog/state.json" WATCHDOG_QUIET_FILE=/run/felis/watchdog-quiet-until \
|
||||
FELIS_DB_BACKUP_DIR=/var/lib/felis/db-backups FELIS_ARCHIVE_LOCAL_PATH=/var/lib/felis/archives FELIS_GAME_PORT=25577 \
|
||||
HOST_BIN=/usr/local/bin/felis STATE_DIR=/etc/felis bash -c '
|
||||
set -Eeuo pipefail
|
||||
ok() { printf "OK: %s\n" "$*"; }; warn() { printf "WARN: %s\n" "$*"; }
|
||||
systemctl() { printf "SYSTEMCTL: %s\n" "$*"; [ "$1" != start ] || return "$FIRST"; }
|
||||
journalctl() { printf "JOURNAL: parse /etc/felis/felis.host.toml\n"; }
|
||||
'"$wblock"'
|
||||
install_watchdog_timer' 2>&1
|
||||
}
|
||||
|
||||
out="$(run_watchdog_timer 0 "")"
|
||||
unit="$(cat "$tdir/felis-watchdog.service")"
|
||||
timer="$(cat "$tdir/felis-watchdog.timer")"
|
||||
expect "the watchdog runs the host binary against the host config, dialing the proxy's port" \
|
||||
"ExecStart=/usr/local/bin/felis watchdog -config /etc/felis/felis.host.toml -state $tdir/watchdog/state.json -quiet-file /run/felis/watchdog-quiet-until -backup-dir /var/lib/felis/db-backups -proxy-addr 127.0.0.1:25577 -disk-paths /,/var/lib/rancher/k3s,/var/lib/postgresql,/var/lib/felis,/var/lib/felis/archives,/var/lib/felis/db-backups" "$unit"
|
||||
expect "a wedged run is killed before the next one is due twice over" "TimeoutStartSec=3min" "$unit"
|
||||
expect "the watchdog runs every two minutes" "OnUnitActiveSec=2min" "$timer"
|
||||
expect "the watchdog starts soon after boot" "OnBootSec=3min" "$timer"
|
||||
expect "the watchdog timer is enabled" "SYSTEMCTL: enable --now felis-watchdog.timer" "$out"
|
||||
expect "the first watchdog run happens during the install" "SYSTEMCTL: start felis-watchdog.service" "$out"
|
||||
expect "a working first run is reported" "OK: watchdog: checks every 2 minutes" "$out"
|
||||
if [ "$(stat -c %a "$tdir/watchdog" 2>/dev/null || stat -f %Lp "$tdir/watchdog")" = 700 ]; then
|
||||
echo "PASS the watchdog state directory is private (it caches the relay password)"
|
||||
else
|
||||
echo "FAIL the watchdog state directory must be 0700"; fails=$((fails + 1))
|
||||
fi
|
||||
|
||||
out="$(run_watchdog_timer 0 /srv/worlds)"
|
||||
expect "a custom worlds root is watched for free space" "-disk-paths /,/var/lib/rancher/k3s,/var/lib/postgresql,/var/lib/felis,/srv/worlds," "$(cat "$tdir/felis-watchdog.service")"
|
||||
|
||||
out="$(run_watchdog_timer 1 "")"
|
||||
expect "a failed first watchdog run shows its log" "JOURNAL: parse /etc/felis/felis.host.toml" "$out"
|
||||
expect "a failed first watchdog run is a loud warning" "WARN: the first watchdog run failed" "$out"
|
||||
|
||||
out="$(WATCHDOG_QUIET_FILE="$tdir/run/quiet" bash -c '
|
||||
set -Eeuo pipefail
|
||||
'"$qblock"'
|
||||
quiet_watchdog; cat "$WATCHDOG_QUIET_FILE"; date +%s' 2>&1)"
|
||||
until_ts="$(printf '%s\n' "$out" | sed -n 1p)"; now_ts="$(printf '%s\n' "$out" | sed -n 2p)"
|
||||
if [ -n "$until_ts" ] && [ "$((until_ts - now_ts))" -ge 3600 ] && [ "$((until_ts - now_ts))" -le 7200 ]; then
|
||||
echo "PASS the install quiets the watchdog for a bounded while"
|
||||
else
|
||||
echo "FAIL quiet_watchdog wrote '$until_ts' at $now_ts, want now+1h..2h"; fails=$((fails + 1))
|
||||
fi
|
||||
rm -rf "$tdir"
|
||||
|
||||
cblock="$(awk '/^cleanup\(\) \{/,/^}/' "$BS")"
|
||||
case "$cblock" in
|
||||
*'rm -f -- "$WATCHDOG_QUIET_FILE"'*) echo "PASS the installer lifts the watchdog's quiet period on exit" ;;
|
||||
*) echo "FAIL cleanup must remove WATCHDOG_QUIET_FILE, or a failed install stays silent"; fails=$((fails + 1)) ;;
|
||||
esac
|
||||
main_block="$(awk '/^main\(\) \{/,/^}/' "$BS")"
|
||||
case "$main_block" in
|
||||
*quiet_watchdog*pause_package_background_timers*install_db_backup_timer*install_watchdog_timer*mark_bootstrap_done*)
|
||||
echo "PASS main quiets the watchdog first and installs it last" ;;
|
||||
*) echo "FAIL main must call quiet_watchdog before any restart and install_watchdog_timer after the backup timer"; fails=$((fails + 1)) ;;
|
||||
esac
|
||||
|
||||
# ---------------------------------------------------------------------------------------
|
||||
if [ "$fails" -eq 0 ]; then
|
||||
echo "ALL PASS"
|
||||
|
||||
Reference in new issue
Block a user