Files
Felis/docs/changes/2026-07-07-on-demand-world-backup.md
T
flyemoji 7a7c0d53ab feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
2026-07-07 10:04:30 +09:00

7.4 KiB

On-demand world backup (§B4 break-glass "Sync"; felis-api endpoint + Job executor)

  • Type: feature (addition)
  • Date: 2026-07-07
  • Area: internal/backupjob (new pkg), internal/api, cmd/felis, docs/openapi.yaml — Go, oracle-verified
  • Commit: pending
  • Task: #31 Phase B4 break-glass ops — the "Sync" operation, resolved with the user as immediate/on-demand world backup. Per the user's "两者都要" decision this is built in two halves: (this change) the felis-api endpoint that does the real backup-Job orchestration, and (a follow-up) a break-glass menu peer that calls it while the API is alive.

What it does

Adds POST /api/v1/servers/{name}/backup: an owner or admin snapshots a stopped server's world into the archive store on demand, recorded as a first-class world_backups row (reason manual) — restorable later by the existing restore path and expired by the reaper's retention pass, so it never leaks as an orphan archive.

The backup runs asynchronously as a one-shot Kubernetes Job (the new internal/backupjob package), mirroring how restore and image builds hand off to Jobs. The handler answers 202 backing_up.

Why

felis-api cannot archive a world in-process: the world PVC is RWO and owned by the operator's StatefulSet, so the API has nothing to mount at request time — the same constraint that already makes internal/restore a Job. The break-glass console (which runs direct-to-Postgres) likewise lacks the deployment coordinates (FELIS_IMAGE, FELIS_BACKUP_PVC) needed to render the Job. Both point to the same home: the orchestration belongs in felis-api, which holds those coordinates; other callers invoke the endpoint.

Design decisions

  • Backup Job self-records its world_backups row. Unlike the restore Job — which is deliberately DB-blind because it processes a potentially poisoned archive — the backup Job does mount the felis config Secret and inserts its own backup row, exactly like the reaper (the only other component holding both a world mount and the database). This avoids the archive-then-async-record split that would otherwise leak orphan archives on a crash. The security review for that one departure lives in internal/backupjob/jobspec.go and is frozen by jobspec_test.go. Rationale: a backup only reads a world the operator already owns and tars it (bytes, never executed), so restore's poisoned-input threat does not apply; its blast radius (DB + two PVCs) is a strict subset of the reaper's, and it never deletes a PVC nor calls the K8s API (SA token stays un-mounted).
  • World mounted read-only, backup PVC read-write — the mirror image of restore.
  • Stopped-gate (409 not_stopped). The world PVC is RWO and held by a running server, so a backup Job cannot double-mount it; the handler refuses unless the server is fully stopped (info.Ready || DesiredState != Stopped). This also guarantees a quiescent, non-torn archive. Mirrors handleRestoreBackup's gate.
  • Authorization is restore's front half, minus the former-owner match. Backup is initiated by the current owner and records their ownership, so there is no prior owner's data to leak — the leak guard that restore needs does not apply here. An admin may back up an unowned (released) world; the recorded former owner is then empty, exactly as the reaper records for an unowned reap.
  • Unique Job name per request. Each backup Job is named backup-<server>-<rand>, not a deterministic backup-<server>. A deterministic name would collide with a just-finished Job still inside its TTLSecondsAfterFinished window (10m), and the AlreadyExists → 202 path would then silently produce no archive — the exact window a user (or the console "立即备份" button) retries in. Unique names make every request produce its own archive; ErrAlreadyExists remains only as a defensive no-op on the ~impossible suffix collision. Ceiling (documented in backup.go): two truly simultaneous taps may schedule two backup Pods — both mount the world PVC read-only, so neither corrupts anything; single-flight-on-running is the upgrade path if a double-tap storm ever appears.
  • One retention clock. The entrypoint reuses the reaper's reaperConfig derivation so a manual backup expires on the same schedule as an inactivity backup — one policy, not two. The "bk-"+hex id scheme also matches, so manual and inactivity backups are indistinguishable downstream.
  • Fail-safe on record failure. If the row insert fails, the entrypoint deletes the just-written archive so a failed backup leaves no unrecorded bytes.
  • Optional executor, honest 503. Wired only when FELIS_IMAGE + FELIS_BACKUP_PVC are supplied (same gate as restore); otherwise API.Backuper is nil and the endpoint returns 503 backup_unavailable, so the authorization boundary is exercised before the Job executor is deployable.

Files

File Change
internal/backupjob/jobspec.go new — BackupJob renderer + BackupJobName; weak SA, token off, hardened container, world RO / backup RW, config-Secret mount
internal/backupjob/backup.go new — Backuper (idempotent enqueue) + Config/withDefaults
internal/backupjob/k8sjobs.go new — controller-runtime CreateBackupJob (AlreadyExists → idempotent)
internal/backupjob/jobspec_test.go new — freezes the Job's security shape incl. the deliberate config-Secret mount
internal/backupjob/backup_test.go new — asserts each Backup call mints a unique Job name (repeat-tap must not silently no-op)
cmd/felis/backup.go new — felis backup in-Pod entrypoint: archive + self-record + orphan-cleanup
cmd/felis/run.go dispatch case "backup" + usage line
internal/api/backuper.go new — the narrow Backuper port
internal/api/handlers_backups.go +handleBackupNow
internal/api/api.go Backuper field + POST /servers/{name}/backup route
internal/api/handlers_backup_now_test.go new — fakeBackuper + handler subtests
internal/api/backuper_wire_test.go new — compile-time Backuper = (*backupjob.Backuper)(nil)
cmd/felis/api.go wire backuper under the FELIS_IMAGE+FELIS_BACKUP_PVC gate; backupConfig helper
docs/openapi.yaml document the backupNow operation

Verification

WSL oracle (go1.26.4, authoritative for Go):

go build ./...  &&  go vet ./...  &&  go test ./...   → ALL GREEN

The internal/api OpenAPI served-route contract test (TestOpenAPIMatchesServedRoutes) initially failed — the new route was served but undocumented — and passes after adding the backupNow operation to docs/openapi.yaml. internal/backupjob and the new handler subtests pass. The controller-runtime K8sJobs binding is integration-only (needs a live cluster) and is exercised only by the interface conformance test.

Self-review outcome

  • ponytail (over-engineering): the backup Job is a near-mirror of the restore Job, not a shared parameterization — deliberate, because its security shape differs (it holds DB creds) and must be asserted independently, not hidden behind a shared knob. No speculative config; Config.withDefaults fills only real deployment values.
  • correctness: the RWO stopped-gate and the self-recording atomicity were traced to the reaper and restore before writing; the former-owner asymmetry vs restore is justified above.