fix(dbbackup): 备份新鲜度按最新的 daily 包算,手动或异地包不再掩盖停掉的定时任务

This commit is contained in:
Lemon-miaow committed 2026-09-27 16:40:50 +08:00
1 parent 46f99a7104
commit 840d339760
18 files changed
+262 -59

No files matched your search

+13 -3
View File
@@ -382,9 +382,19 @@ components:
Why the bundle lacks the MinecraftServer objects (the cluster did not
answer the export), when it does. A restore from it brings back the
database but no servers.
daily_at:
type: string
format: date-time
description: >
When the newest daily bundle (felis-db-backup.timer) in dir was
written, as of this record; absent when dir held none. Equal to
at when this record is a daily one.
stale:
type: boolean
description: True when there is no record or it is older than max_age_seconds.
description: >
True when there is no record, no daily bundle, or the newest daily
bundle is older than max_age_seconds. A newer manual, pre-migrate or
off-site bundle leaves it as it is: the daily timer has still stopped.
max_age_seconds:
type: integer
format: int64
@@ -3684,8 +3694,8 @@ paths:
description: >-
What the host's felis-db-backup.timer (or a manual `felis db backup`)
last recorded in platform_settings. last is null before the first
backup; stale is true then, and whenever the newest backup is older than
max_age_seconds. Read-only: backups run on the host, never through the API.
backup; stale is true then, and whenever the newest daily backup
(last.daily_at) is missing or older than max_age_seconds. Read-only: backups run on the host, never through the API.
x-felis-face: [external]
x-felis-tier: admin
security: [{ sessionCookie: [] }]
+1 -1
View File
@@ -690,7 +690,7 @@ production install:
- **Rehearse the rebuild** once on a spare VM: troubleshooting.md §16 "Rebuild on a new
host", every step but 8 (take-over) and 11 (the tunnel), then its checks: sign in with
an email code, restore one world and join it. `felis offsite status` and `felis db check` exit
non-zero when the copy or the newest bundle is stale; wire them into your monitoring,
non-zero when the copy or the newest daily bundle is stale; wire them into your monitoring,
or rely on the watchdog's mail.
### Moving to another host (planned)
+10 -5
View File
@@ -1665,7 +1665,7 @@ Every two minutes the host checks:
| Node `NotReady`, or kubelet reports Disk/Memory/PID pressure (§13b) | 2–5 min | critical |
| PostgreSQL unreachable | 3 min | critical |
| The game proxy (`felis-velocity`) refuses connections on the game port | 3 min | critical |
| Newest control-plane database backup over 26h old, or none (§16) | 10 min | critical |
| Newest daily control-plane database backup over 26h old, or none (§16); a newer manual or off-site bundle leaves it standing | 10 min | critical |
| A watched filesystem below 15% free (below 5%: critical) | 15 min (5 min) | warning |
| Host memory available below 10% | 15 min | warning |
| The host no longer holds the address the install was made on (§13c) | 5 min | critical |
@@ -2187,13 +2187,18 @@ Installer knobs: `FELIS_DB_BACKUP_DIR`, `FELIS_DB_BACKUP_KEEP`,
### Is the newest backup fresh?
Three places answer, all with the same 26 h limit:
Four places answer, all with the same 26 h limit on the newest **daily**
bundle, the one `felis-db-backup.timer` writes. A manual, `pre-migrate` or
`offsite` bundle taken since counts for a restore and leaves the alarm
standing: the timer has still stopped, and that bundle only ages from here.
- The panel: **管理 → 维护与备份** shows the newest backup, its kind and size,
and turns red with the fix commands when it is missing or overdue (read from
the `db_backup_last` platform setting each backup writes).
and turns red with the fix commands when the daily one is missing or overdue
(read from the `db_backup_last` platform setting each backup writes; its
`daily_at` is the newest daily bundle on disk when it was written).
- `sudo felis db check` exits 1 with the reason; `sudo felis db list` shows every
bundle with its age.
- The watchdog mails the owners (§14).
- Prometheus: `FelisDBBackupStale` (critical) and `FelisDBBackupMetricMissing`
(warning) in `deploy/alerts/`. They read
`felis_db_backup_last_success_timestamp_seconds`, which each daily run writes
@@ -2207,7 +2212,7 @@ When a backup is overdue:
```
sudo systemctl status felis-db-backup.timer # enabled? next run?
sudo journalctl -u felis-db-backup -n 50 --no-pager # why the last run failed
sudo felis db backup # take one now (label manual)
sudo systemctl start felis-db-backup.service # run the daily backup now; clears the alarm
```
**A bundle without the MinecraftServer objects.** When the cluster does not