Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.
felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).
Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
26 lines
1.4 KiB
Go
26 lines
1.4 KiB
Go
package api
|
|
|
|
import "context"
|
|
|
|
// Backuper starts an on-demand world backup (spec §18/§19 WorldArchiver, run on
|
|
// demand rather than on the reaper's daily schedule) — the "back up before I touch
|
|
// it" lever behind POST /api/v1/servers/{name}/backup. Like Restorer it only STARTS
|
|
// the work: tarring the world PVC into the archive store is pod-filesystem work
|
|
// felis-api cannot do in-process — the world PVC is RWO and owned by the operator's
|
|
// StatefulSet, so the API has nothing to mount at request time. The production
|
|
// implementation therefore hands off to a backup Job (internal/backupjob), which
|
|
// self-records the world_backups row like the reaper. The call returns once the
|
|
// backup is enqueued, so the handler answers 202 (backing_up).
|
|
//
|
|
// formerOwner is the current owner recorded on the backup row so it can later be
|
|
// restored (empty when an admin backs up an unowned server). It returns ErrNotFound
|
|
// if the server is unknown to the execution backend; any other error maps to 500.
|
|
//
|
|
// It is an interface so the handler is tested against a fake; the production executor
|
|
// (internal/backupjob.Backuper) is integration-only, and until it is wired the
|
|
// API.Backuper is nil so POST /servers/{name}/backup reports 503 — the backup
|
|
// authorization boundary is exercised without shipping a stub that cannot run.
|
|
type Backuper interface {
|
|
Backup(ctx context.Context, serverName, formerOwner string) error
|
|
}
|