Loading
Sign in before continuing.
fix(operator): end the start attempt on success — stale anchor caused false StartupTimeout
Found live while validating the RCON-secret heal: a server that had already recovered to Ready was marked Failed(StartupTimeout) minutes later, the moment an unrelated pod rollout briefly dropped readyReplicas. The anchor (status.startRequestedAt) was never cleared on success, so its 300s budget kept ticking under a healthy server and any later blip spent it. markRunningReady now clears the anchor: every start-or-recovery attempt gets its own budget. It also flips ConditionProvisioned back to True — markFailed sets it False and nothing ever reset it, leaving a permanent failure flag on recovered servers that every conditions consumer would read. Unit tests pin both: anchor cleared on Ready, Provisioned recovers from Failed to Running.