fix(bootstrap): kubeconfig 仅 root 可读、固定 k3s 节点名、开启 NTP 与持久日志,DHCP 地址告警并交给 watchdog 监测

This commit is contained in:
Lemon-miaow committed 2026-09-26 01:25:48 +08:00
1 parent 298a1e635d
commit 6a4c30b3f1
4 files changed
+660 -8

No files matched your search

+50 -3
View File
@@ -8,7 +8,9 @@ see to the code path that emitted it.
## How to read this document
Each entry is **symptom → likely cause → where to look → fix**. Signals are
Each entry is **symptom → likely cause → where to look → fix**. The `kubectl`
commands run as root on the node (`sudo -i`, or `sudo k3s kubectl …`): the admin
kubeconfig `/etc/rancher/k3s/k3s.yaml` is readable by root only (§13c). Signals are
graded for how far the in-repo Go test suite proves the behaviour:
- **[GO-TESTED]** — a hermetic `*_test.go` exercises this exact path; the
@@ -1331,6 +1333,48 @@ the registry), or for a single image
deliberate second copy on the node; treat it as the recovery path, not as free
space.
## 13c. The host's address, name or clock changed
An install is bound to the address it was made on. `bootstrap.sh` writes that
address into the database connection string, `pg_hba.conf`, the network
policies, the panel certificate and the default `<ip>.nip.io` root domain, and
nothing re-addresses a live install. When the host loses the address (a DHCP
lease that came back different, a moved VM), felis-api cannot reach PostgreSQL
and the panel stops answering on its old name. The watchdog reports it as
`host-address` (critical, after 5 minutes). The installer warns at install time
when the address is a DHCP lease.
Remedy: give the host its old address back, either as a DHCP reservation on
the router or as a static address (`nmcli con mod <con> ipv4.method manual
ipv4.addresses <ip>/<prefix> ipv4.gateway <gw> ipv4.dns <dns> && nmcli con up
<con>` on Rocky), then restart felis-api (`sudo k3s kubectl -n felis rollout restart
deploy/felis-api`) and the proxy (`sudo systemctl restart felis-velocity`). Moving an install
to a new address is a reinstall onto a restored backup (docs/operations.md §5).
The node **name** is pinned. Every local-path volume (worlds, registry,
uploads, backups) is bound to its node by name, and k3s takes the name from the
hostname unless told otherwise, so renaming the host used to bring k3s back as
a second, empty node with every volume Pending on the old one. The installer
pins the name in `/etc/rancher/k3s/config.yaml.d/50-felis.yaml`
(`node-name:`); a hostname change is then harmless. Check the pin with `sudo
k3s kubectl get node -o jsonpath='{.items[0].metadata.annotations.k3s\.io/node-args}'`.
That file also sets `write-kubeconfig-mode: "0600"`: the admin kubeconfig
`/etc/rancher/k3s/k3s.yaml` is cluster-admin and readable by root only, so
use `sudo k3s kubectl` (or `sudo -E kubectl`).
The **clock** must be kept by NTP. Sign-in codes and sessions expire by it,
S3 refuses off-site uploads signed more than 15 minutes off, and certificate
checks fail on a clock far off. The installer turns NTP on (`timedatectl
set-ntp true`, installing chrony where there is no client to enable) unless
`FELIS_MANAGE_TIME_SYNC=0`. The watchdog reports an unsynchronized clock as
`clock` (warning, after 30 minutes). Check with `timedatectl` (want `System
clock synchronized: yes` and `NTP service: active`) and `chronyc sources`;
a firewall that drops outbound UDP 123 keeps it unsynchronized.
The system journal is persistent (`/etc/systemd/journald.conf.d/50-felis.conf`,
capped by `FELIS_JOURNAL_MAX_USE`, default 1G), so `journalctl -b -1` shows the
boot before a reboot.
---
## 14. Health alerts, and metrics for diagnosis (spec §23)
@@ -1359,6 +1403,8 @@ Every two minutes the host checks:
| Newest control-plane database backup over 26h old, or none (§16) | 10 min | critical |
| A watched filesystem below 15% free (below 5%: critical) | 15 min (5 min) | warning |
| Host memory available below 10% | 15 min | warning |
| The host no longer holds the address the install was made on (§13c) | 5 min | critical |
| The system clock is not synchronized by NTP (§13c) | 30 min | warning |
How it mails:
@@ -1388,8 +1434,9 @@ A healthy run logs `every check passed`. Otherwise it logs one line per
finding, and the mail's subject once one is sent.
A `-dry-run` from a shell uses the command's defaults, and those do not include
the game-proxy check. The unit carries `-proxy-addr 127.0.0.1:<game port>` and
the disk list the install chose. `systemctl cat felis-watchdog` shows both.
the game-proxy or the address check. The unit carries `-proxy-addr
127.0.0.1:<game port>`, `-node-ip <install address>` and the disk list the
install chose. `systemctl cat felis-watchdog` shows them.
### Metrics