From d601abb09bc34ff0dee3dac8f0004c40d2301eff Mon Sep 17 00:00:00 2001 From: Lemon-miaow Date: Sat, 26 Sep 2026 09:41:38 +0800 Subject: [PATCH] =?UTF-8?q?docs(operations):=20felis-api=20=E9=95=BF?= =?UTF-8?q?=E6=97=B6=E9=97=B4=E5=81=9C=E6=91=86=E6=97=B6=E6=97=A5=E5=BF=97?= =?UTF-8?q?=E4=B8=8E=E5=91=8A=E8=AD=A6=E7=9A=84=E5=AE=9E=E6=B5=8B=E6=97=B6?= =?UTF-8?q?=E9=97=B4=E7=BA=BF?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/operations.md | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/docs/operations.md b/docs/operations.md index 0458449..ef49937 100644 --- a/docs/operations.md +++ b/docs/operations.md @@ -87,6 +87,17 @@ the window: `/opt/felis/velocity/plugins/felis-link/last-servers.json`, the last server list the API answered with, until a refresh succeeds (every 15 s). +A longer outage reads like this in the logs [VM-VERIFIED]. The drill scaled felis-api +to 0 for about 8 minutes on the reference VM. + +- The proxy logged `server list refresh failed ... keeping current registrations` 11 s + in, then `still failing: 22 failed attempts over 304 s` at the 5-minute mark. +- The watchdog found `deployment/felis-api` critical on its first run after the scale. + It raised the alert on the first run past 5 minutes, at about 7 minutes; with no + `[smtp]` that is logged only (`journalctl -u felis-watchdog`). +- The proxy logged `server list refresh recovered after 32 failed attempts over 469 s` + as soon as the new pod was Available. + ## 2. Sizing ### What the platform itself uses