feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)

Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
This commit is contained in:
flyemoji committed 2026-07-07 10:04:30 +09:00
1 parent fad48ff21d
commit 7a7c0d53ab
16 files changed
+1396 -1

No files matched your search

+43
View File
@@ -0,0 +1,43 @@
package backupjob
import (
"context"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"sigs.k8s.io/controller-runtime/pkg/client"
)
// K8sJobs is the production Jobs backed by a controller-runtime client. It creates
// the on-demand world-backup Job — nothing more: the backup Job is one-shot and
// self-cleaning (ttlSecondsAfterFinished), so there is no phase or cancel seam, and
// thus no config to hold. Every backup parameter arrives in the JobParams the
// Backuper builds from its own (defaulted) Config. The cluster-bootstrap objects
// (the weak felis-restore SA it reuses) are installed once by the deployment
// manifests, not per backup, so this binding never creates them. It is
// integration-tested against a live cluster, not the hermetic unit suite.
type K8sJobs struct {
c client.Client
}
// NewK8sJobs builds a Jobs over c.
func NewK8sJobs(c client.Client) *K8sJobs {
return &K8sJobs{c: c}
}
// CreateBackupJob renders and applies the backup Job. Its name is a deterministic
// function of the server (BackupJobName), so a concurrent backup of the same server
// collides on Create; that collision is mapped to ErrAlreadyExists, which the
// Backuper treats as success (idempotent enqueue).
func (k *K8sJobs) CreateBackupJob(ctx context.Context, p JobParams) error {
job, err := BackupJob(p)
if err != nil {
return err
}
if err := k.c.Create(ctx, job); err != nil {
if apierrors.IsAlreadyExists(err) {
return ErrAlreadyExists
}
return err
}
return nil
}