feat(metrics): publish felis_servers_total from a fleet snapshot

A per-object reconcile cannot maintain felis_servers_total (spec §23): it
sees one server per call, so it could never Set a correct fleet-wide gauge
and inc/dec on transitions would drift on any missed event. Add a snapshot
producer instead.

metrics.SyncServerGauge Resets the GaugeVec then Sets one child per state,
so a state that drains to zero reports 0 rather than a stale last value.
operator.GaugeSyncer is a manager.Runnable that periodically Lists the
fleet and republishes from it, defaulting an unset desiredState to Stopped.

SyncOnce is exercised end-to-end against a fake client (List, default,
republish); the ticker loop in Start is the only untested I/O edge.
This commit is contained in:
flyemoji committed 2026-06-30 12:52:53 +09:00
1 parent 2a93a9e0cb
commit 79eae7f669
5 files changed
+194

No files matched your search

+9
View File
@@ -79,6 +79,15 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
return 1
}
// Republish felis_servers_total from a periodic full List of the fleet. A
// per-object reconcile can never maintain a fleet-wide gauge correctly, so a
// snapshot Runnable owns it; it shares the manager's cached client and stops
// with the manager.
if err := mgr.Add(&operator.GaugeSyncer{Client: mgr.GetClient()}); err != nil {
fmt.Fprintf(stderr, "felis operator: add gauge syncer: %v\n", err)
return 1
}
if err := mgr.Start(ctrl.SetupSignalHandler()); err != nil {
fmt.Fprintf(stderr, "felis operator: manager exited: %v\n", err)
return 1