feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)
Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.
felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).
Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
This commit is contained in:
16 files changed
+1396
-1
No files matched your search
@@ -0,0 +1,115 @@
|
||||
# On-demand world backup (§B4 break-glass "Sync"; felis-api endpoint + Job executor)
|
||||
|
||||
- **Type:** feature (addition)
|
||||
- **Date:** 2026-07-07
|
||||
- **Area:** `internal/backupjob` (new pkg), `internal/api`, `cmd/felis`, `docs/openapi.yaml` — Go, oracle-verified
|
||||
- **Commit:** `pending`
|
||||
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation, resolved with the user as **immediate/on-demand world backup**. Per the user's "两者都要" decision this is built in two halves: **(this change) the felis-api endpoint that does the real backup-Job orchestration**, and (a follow-up) a break-glass menu peer that calls it while the API is alive.
|
||||
|
||||
## What it does
|
||||
|
||||
Adds `POST /api/v1/servers/{name}/backup`: an owner or admin snapshots a **stopped**
|
||||
server's world into the archive store on demand, recorded as a first-class
|
||||
`world_backups` row (reason `manual`) — restorable later by the existing restore path
|
||||
and expired by the reaper's retention pass, so it never leaks as an orphan archive.
|
||||
|
||||
The backup runs asynchronously as a one-shot Kubernetes Job (the new
|
||||
`internal/backupjob` package), mirroring how restore and image builds hand off to
|
||||
Jobs. The handler answers **202 `backing_up`**.
|
||||
|
||||
## Why
|
||||
|
||||
felis-api cannot archive a world in-process: the world PVC is **RWO** and owned by the
|
||||
operator's StatefulSet, so the API has nothing to mount at request time — the same
|
||||
constraint that already makes `internal/restore` a Job. The break-glass console (which
|
||||
runs direct-to-Postgres) likewise lacks the deployment coordinates (`FELIS_IMAGE`,
|
||||
`FELIS_BACKUP_PVC`) needed to render the Job. Both point to the same home: the
|
||||
orchestration belongs in felis-api, which holds those coordinates; other callers
|
||||
invoke the endpoint.
|
||||
|
||||
## Design decisions
|
||||
|
||||
- **Backup Job self-records its `world_backups` row.** Unlike the restore Job — which
|
||||
is deliberately DB-blind because it processes a potentially poisoned archive — the
|
||||
backup Job **does** mount the felis config Secret and inserts its own backup row,
|
||||
exactly like the reaper (the only other component holding both a world mount and the
|
||||
database). This avoids the archive-then-async-record split that would otherwise leak
|
||||
orphan archives on a crash. The security review for that one departure lives in
|
||||
`internal/backupjob/jobspec.go` and is frozen by `jobspec_test.go`. Rationale: a
|
||||
backup only **reads** a world the operator already owns and tars it (bytes, never
|
||||
executed), so restore's poisoned-input threat does not apply; its blast radius (DB +
|
||||
two PVCs) is a strict subset of the reaper's, and it never deletes a PVC nor calls
|
||||
the K8s API (SA token stays un-mounted).
|
||||
- **World mounted read-only, backup PVC read-write** — the mirror image of restore.
|
||||
- **Stopped-gate (409 `not_stopped`).** The world PVC is RWO and held by a running
|
||||
server, so a backup Job cannot double-mount it; the handler refuses unless the server
|
||||
is fully stopped (`info.Ready || DesiredState != Stopped`). This also guarantees a
|
||||
quiescent, non-torn archive. Mirrors `handleRestoreBackup`'s gate.
|
||||
- **Authorization is restore's front half, minus the former-owner match.** Backup is
|
||||
initiated by the **current** owner and records **their** ownership, so there is no
|
||||
prior owner's data to leak — the leak guard that restore needs does not apply here.
|
||||
An admin may back up an unowned (released) world; the recorded former owner is then
|
||||
empty, exactly as the reaper records for an unowned reap.
|
||||
- **Unique Job name per request.** Each backup Job is named `backup-<server>-<rand>`,
|
||||
not a deterministic `backup-<server>`. A deterministic name would collide with a
|
||||
just-finished Job still inside its `TTLSecondsAfterFinished` window (10m), and the
|
||||
`AlreadyExists → 202` path would then silently produce **no** archive — the exact
|
||||
window a user (or the console "立即备份" button) retries in. Unique names make every
|
||||
request produce its own archive; `ErrAlreadyExists` remains only as a defensive
|
||||
no-op on the ~impossible suffix collision. Ceiling (documented in `backup.go`): two
|
||||
truly simultaneous taps may schedule two backup Pods — both mount the world PVC
|
||||
read-only, so neither corrupts anything; single-flight-on-running is the upgrade
|
||||
path if a double-tap storm ever appears.
|
||||
- **One retention clock.** The entrypoint reuses the reaper's `reaperConfig` derivation
|
||||
so a manual backup expires on the same schedule as an inactivity backup — one policy,
|
||||
not two. The `"bk-"+hex` id scheme also matches, so manual and inactivity backups are
|
||||
indistinguishable downstream.
|
||||
- **Fail-safe on record failure.** If the row insert fails, the entrypoint deletes the
|
||||
just-written archive so a failed backup leaves no unrecorded bytes.
|
||||
- **Optional executor, honest 503.** Wired only when `FELIS_IMAGE` + `FELIS_BACKUP_PVC`
|
||||
are supplied (same gate as restore); otherwise `API.Backuper` is nil and the endpoint
|
||||
returns 503 `backup_unavailable`, so the authorization boundary is exercised before
|
||||
the Job executor is deployable.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Change |
|
||||
|---|---|
|
||||
| `internal/backupjob/jobspec.go` | **new** — `BackupJob` renderer + `BackupJobName`; weak SA, token off, hardened container, world RO / backup RW, config-Secret mount |
|
||||
| `internal/backupjob/backup.go` | **new** — `Backuper` (idempotent enqueue) + `Config`/`withDefaults` |
|
||||
| `internal/backupjob/k8sjobs.go` | **new** — controller-runtime `CreateBackupJob` (AlreadyExists → idempotent) |
|
||||
| `internal/backupjob/jobspec_test.go` | **new** — freezes the Job's security shape incl. the deliberate config-Secret mount |
|
||||
| `internal/backupjob/backup_test.go` | **new** — asserts each `Backup` call mints a unique Job name (repeat-tap must not silently no-op) |
|
||||
| `cmd/felis/backup.go` | **new** — `felis backup` in-Pod entrypoint: archive + self-record + orphan-cleanup |
|
||||
| `cmd/felis/run.go` | dispatch `case "backup"` + usage line |
|
||||
| `internal/api/backuper.go` | **new** — the narrow `Backuper` port |
|
||||
| `internal/api/handlers_backups.go` | **+`handleBackupNow`** |
|
||||
| `internal/api/api.go` | `Backuper` field + `POST /servers/{name}/backup` route |
|
||||
| `internal/api/handlers_backup_now_test.go` | **new** — `fakeBackuper` + handler subtests |
|
||||
| `internal/api/backuper_wire_test.go` | **new** — compile-time `Backuper = (*backupjob.Backuper)(nil)` |
|
||||
| `cmd/felis/api.go` | wire `backuper` under the `FELIS_IMAGE`+`FELIS_BACKUP_PVC` gate; `backupConfig` helper |
|
||||
| `docs/openapi.yaml` | document the `backupNow` operation |
|
||||
|
||||
## Verification
|
||||
|
||||
WSL oracle (go1.26.4, authoritative for Go):
|
||||
|
||||
```
|
||||
go build ./... && go vet ./... && go test ./... → ALL GREEN
|
||||
```
|
||||
|
||||
The `internal/api` OpenAPI served-route contract test (`TestOpenAPIMatchesServedRoutes`)
|
||||
initially failed — the new route was served but undocumented — and passes after adding
|
||||
the `backupNow` operation to `docs/openapi.yaml`. `internal/backupjob` and the new
|
||||
handler subtests pass. The controller-runtime `K8sJobs` binding is integration-only
|
||||
(needs a live cluster) and is exercised only by the interface conformance test.
|
||||
|
||||
## Self-review outcome
|
||||
|
||||
- **ponytail (over-engineering):** the backup Job is a near-mirror of the restore Job,
|
||||
not a shared parameterization — deliberate, because its security shape differs (it
|
||||
holds DB creds) and must be asserted independently, not hidden behind a shared knob.
|
||||
No speculative config; `Config.withDefaults` fills only real deployment values.
|
||||
- **correctness:** the RWO stopped-gate and the self-recording atomicity were traced to
|
||||
the reaper and restore before writing; the former-owner asymmetry vs restore is
|
||||
justified above.
|
||||
@@ -23,7 +23,9 @@ not yet committed.
|
||||
|
||||
## Pending (built + verified, not yet committed)
|
||||
|
||||
_None — the break-glass halt (`c2ee21a`) and `/felis migrate` (`c1aa38b`) landed in the ledger below._
|
||||
| Change | Detail doc | Status |
|
||||
|---|---|---|
|
||||
| On-demand world backup — felis-api `POST /servers/{name}/backup` + backup-Job executor (§B4 "Sync", half 1 of 2) | [2026-07-07-on-demand-world-backup.md](2026-07-07-on-demand-world-backup.md) | oracle-green; commit `pending` |
|
||||
|
||||
## Committed change ledger
|
||||
|
||||
|
||||
Reference in new issue
Block a user