fix(api): refuse backup/restore before a missing world volume

A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
This commit is contained in:
Lemon-miaow committed 2026-09-23 07:49:00 +08:00
1 parent 55d515d41f
commit 508a1c02da
8 files changed
+127 -3

No files matched your search

+6
View File
@@ -81,6 +81,12 @@ type Cluster interface {
// GetServer reads one MinecraftServer's lifecycle view, or ErrNotFound.
GetServer(ctx context.Context, name string) (*ServerInfo, error)
// WorldVolumeExists reports whether the server's world PVC exists in the
// server namespace. A server that never started — or whose world the
// retention reaper already archived and deleted — has no claim, and a
// backup/restore Job would hang Pending on the missing volume with nothing
// ever recorded, so both handlers refuse those up front.
WorldVolumeExists(ctx context.Context, name string) (bool, error)
// GetBySubdomain finds the MinecraftServer whose spec.subdomain matches, or
// ErrNotFound.
GetBySubdomain(ctx context.Context, subdomain string) (*ServerInfo, error)