fix(api): refuse backup/restore before a missing world volume

A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
This commit is contained in:
Lemon-miaow committed 2026-09-23 07:49:00 +08:00
1 parent 55d515d41f
commit 508a1c02da
8 files changed
+127 -3

No files matched your search

+35
View File
@@ -9,6 +9,16 @@ import (
"felis.lolicon.best/internal/naming"
)
// errNoWorldVolume is the shared 409 for backup and restore when the server's
// world PVC does not exist: the Job would only hang Pending on the missing
// claim — invisible to the caller and to the backups list — so the handlers
// refuse up front. Starting the server once (which creates the claim via the
// StatefulSet volumeClaimTemplate) unlocks both ops.
func errNoWorldVolume() error {
return newError(http.StatusConflict, "no_world_volume",
"this server has no world volume yet — start it once to create it, then retry")
}
// handleListBackups lists the world backups visible to the caller (spec §7 GET
// /api/v1/backups; world_backups in §22). It is app-tier: an admin sees every
// present backup; a regular user sees only the backups of worlds they formerly
@@ -154,6 +164,19 @@ func (a *API) handleRestoreBackup(w http.ResponseWriter, r *http.Request) {
return
}
// World-volume gate: the restore Job mounts the world PVC read-write to unpack
// the archive into it, so a missing claim means a Pod stuck Pending — a 202
// "restoring" with nothing ever written. Same refusal as the backup face
// (shared errNoWorldVolume): the operator starts the server once to create the
// claim, then restores into it.
if exists, err := a.Cluster.WorldVolumeExists(r.Context(), name); err != nil {
writeError(w, r, err)
return
} else if !exists {
writeError(w, r, errNoWorldVolume())
return
}
// Restorer is optional: when unwired the endpoint reports 503 rather than
// panicking, so the authorization boundary above is exercised even before the
// restore-Job executor is wired (see Restorer).
@@ -285,6 +308,18 @@ func (a *API) enqueueBackup(w http.ResponseWriter, r *http.Request, name string,
return
}
// World-volume gate: the Job mounts the world PVC by claim name, and a missing
// claim would leave its Pod Pending — a 202 "backing_up" with nothing ever
// recorded anywhere. A never-started or already-reaped server is refused with
// the same specificity as the stopped gate.
if exists, err := a.Cluster.WorldVolumeExists(r.Context(), name); err != nil {
writeError(w, r, err)
return
} else if !exists {
writeError(w, r, errNoWorldVolume())
return
}
// Backuper is optional: when unwired the endpoint reports 503 rather than
// panicking, so the authorization boundary above is exercised even before the
// backup-Job executor is wired (see Backuper).