fix(restore): replace a finished Job so retries enqueue; replicate felis-config

An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
This commit is contained in:
Lemon-miaow committed 2026-09-22 17:50:04 +08:00
1 parent fd0794d04d
commit 90ccbfede4
6 files changed
+134 -27

No files matched your search

+4 -2
View File
@@ -73,8 +73,10 @@ func TestAPIRole_CreatesJobsInBothNamespaces(t *testing.T) {
if mc.Namespace != "minecraft" {
t.Errorf("felis-api minecraft Role namespace = %q, want minecraft", mc.Namespace)
}
if !hasRule(mc, "batch", "jobs", "create") {
t.Error("felis-api (minecraft) must have batch/jobs:create for the restore Job")
for _, v := range []string{"create", "get", "delete"} {
if !hasRule(mc, "batch", "jobs", v) {
t.Errorf("felis-api (minecraft) must have batch/jobs:%s for the restore-Job lifecycle", v)
}
}
build := roleByName(t, rbac.Roles, "felis-api-builds")