fix(api): refuse backup/restore before a missing world volume
A server whose world PVC does not exist yet (never started) or no longer exists (the world was already reaped) accepted the backup/restore POST, answered 202, and the Job sat Pending on the missing claim until its deadline with nothing recorded anywhere — a silent no-op from the operator's seat. The live drill on the reaped `resolvecheck` world reproduced exactly that. Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume with "start it once to create it, then retry". The felis-api Role gains the matching get-only PVC grant — the first live run surfaced the missing RBAC as a 403 behind a 500, so the fix ships with it. Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one (which has a world) still backs up through the new gate end to end.
This commit is contained in:
8 files changed
+127
-3
No files matched your search
@@ -9,6 +9,16 @@ import (
|
||||
"felis.lolicon.best/internal/naming"
|
||||
)
|
||||
|
||||
// errNoWorldVolume is the shared 409 for backup and restore when the server's
|
||||
// world PVC does not exist: the Job would only hang Pending on the missing
|
||||
// claim — invisible to the caller and to the backups list — so the handlers
|
||||
// refuse up front. Starting the server once (which creates the claim via the
|
||||
// StatefulSet volumeClaimTemplate) unlocks both ops.
|
||||
func errNoWorldVolume() error {
|
||||
return newError(http.StatusConflict, "no_world_volume",
|
||||
"this server has no world volume yet — start it once to create it, then retry")
|
||||
}
|
||||
|
||||
// handleListBackups lists the world backups visible to the caller (spec §7 GET
|
||||
// /api/v1/backups; world_backups in §22). It is app-tier: an admin sees every
|
||||
// present backup; a regular user sees only the backups of worlds they formerly
|
||||
@@ -154,6 +164,19 @@ func (a *API) handleRestoreBackup(w http.ResponseWriter, r *http.Request) {
|
||||
return
|
||||
}
|
||||
|
||||
// World-volume gate: the restore Job mounts the world PVC read-write to unpack
|
||||
// the archive into it, so a missing claim means a Pod stuck Pending — a 202
|
||||
// "restoring" with nothing ever written. Same refusal as the backup face
|
||||
// (shared errNoWorldVolume): the operator starts the server once to create the
|
||||
// claim, then restores into it.
|
||||
if exists, err := a.Cluster.WorldVolumeExists(r.Context(), name); err != nil {
|
||||
writeError(w, r, err)
|
||||
return
|
||||
} else if !exists {
|
||||
writeError(w, r, errNoWorldVolume())
|
||||
return
|
||||
}
|
||||
|
||||
// Restorer is optional: when unwired the endpoint reports 503 rather than
|
||||
// panicking, so the authorization boundary above is exercised even before the
|
||||
// restore-Job executor is wired (see Restorer).
|
||||
@@ -285,6 +308,18 @@ func (a *API) enqueueBackup(w http.ResponseWriter, r *http.Request, name string,
|
||||
return
|
||||
}
|
||||
|
||||
// World-volume gate: the Job mounts the world PVC by claim name, and a missing
|
||||
// claim would leave its Pod Pending — a 202 "backing_up" with nothing ever
|
||||
// recorded anywhere. A never-started or already-reaped server is refused with
|
||||
// the same specificity as the stopped gate.
|
||||
if exists, err := a.Cluster.WorldVolumeExists(r.Context(), name); err != nil {
|
||||
writeError(w, r, err)
|
||||
return
|
||||
} else if !exists {
|
||||
writeError(w, r, errNoWorldVolume())
|
||||
return
|
||||
}
|
||||
|
||||
// Backuper is optional: when unwired the endpoint reports 503 rather than
|
||||
// panicking, so the authorization boundary above is exercised even before the
|
||||
// backup-Job executor is wired (see Backuper).
|
||||
|
||||
Reference in new issue
Block a user