Unverified Commit 7a7c0d53 authored by Minseong Choi's avatar Minseong Choi 💬
Browse files

feat(api): add on-demand world backup endpoint and Job executor (§B4 Sync)

Add POST /api/v1/servers/{name}/backup: an owner or admin snapshots a
stopped server's world into the archive store on demand, recorded as a
first-class world_backups row (reason `manual`) — restorable by the
existing restore path and expired by the reaper's retention pass, so it
never leaks as an orphan archive. This is the break-glass "Sync" op,
resolved as immediate/on-demand backup.

felis-api cannot archive in-process (the world PVC is RWO, held by the
operator StatefulSet), so the work hands off to a one-shot Kubernetes Job
(new internal/backupjob) that mounts the world PVC read-only and the
backup PVC read-write, plus the felis config Secret so it self-records
its row atomically like the reaper. The Pod mirrors restore's weak-SA
isolation (SA token un-mounted, non-root, read-only rootfs, drop ALL);
the one reviewed departure is that config-Secret mount, frozen by
jobspec_test.go. Handler answers 202 backing_up; gated on the server
being Stopped (RWO world PVC), owner-or-admin, and FELIS_IMAGE +
FELIS_BACKUP_PVC being wired (else 503 backup_unavailable).

Each request mints a unique Job name (backup-<server>-<rand>) so a repeat
on-demand backup produces a fresh archive rather than colliding with a
just-finished Job still inside its TTL window and silently no-op'ing the
retry.
parent fad48ff2
Loading
Loading
Loading
Loading
+30 −0
Changes for cmd/felis/api.go: 30 added lines, 0 removed lines.
Original line number Diff line number Diff line
@@ -13,6 +13,7 @@ import (

	"felis.lolicon.best/internal/api"
	"felis.lolicon.best/internal/apis/felis/v1alpha1"
	"felis.lolicon.best/internal/backupjob"
	"felis.lolicon.best/internal/build"
	"felis.lolicon.best/internal/config"
	"felis.lolicon.best/internal/panel"
@@ -162,6 +163,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
		fmt.Fprintln(stderr, "felis api: restore executor disabled (needs FELIS_IMAGE and FELIS_BACKUP_PVC) — restore endpoint returns 503")
	}

	// On-demand backup subsystem (spec §18/§19 WorldArchiver, run on demand). Its
	// backup Job mirrors the restore Job's weak-SA isolation but additionally mounts
	// the config Secret so it self-records the world_backups row (see internal/
	// backupjob). It needs the same deployment-specific values as restore, so it is
	// wired under the same gate; otherwise the Backuper is left nil and the backup
	// endpoint honestly returns 503.
	var backuper api.Backuper
	if felisImage != "" && backupPVC != "" {
		backuper = &backupjob.Backuper{Jobs: backupjob.NewK8sJobs(cl), Config: backupConfig(cfg, felisImage, backupPVC)}
	} else {
		fmt.Fprintln(stderr, "felis api: backup executor disabled (needs FELIS_IMAGE and FELIS_BACKUP_PVC) — backup endpoint returns 503")
	}

	// One PGRepo instance backs both the handlers and the session verifier: the
	// SessionAuth that fronts the external face reads sessions/users/settings from
	// the same store the auth handlers write to, so a login and the next request
@@ -179,6 +193,7 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
		Internal:    api.BearerTokenAuth{Token: token},
		Builder:     builder,
		Restorer:    restorer,
		Backuper:    backuper,
		Submissions: submissions,
		// The external face is fronted by SessionAuth: it prefers a local-password
		// session cookie and otherwise delegates to the Cloudflare-Access JWT verifier,
@@ -359,6 +374,21 @@ func restoreConfig(cfg *config.Config, image, backupPVC string) restore.Config {
	}
}

// backupConfig builds the on-demand backup executor's config from felis.toml plus
// the deployment-supplied image and backup PVC. BackupRoot mirrors restoreConfig —
// it MUST equal [archive] local_path so the recorded ref resolves the same way a
// later restore Job mounts it. ConfigSecret/ConfigMount are left to backupjob's
// defaults (the control-plane manifest names), which is the Secret this backup Job
// mounts to self-record its world_backups row.
func backupConfig(cfg *config.Config, image, backupPVC string) backupjob.Config {
	return backupjob.Config{
		Namespace:  cfg.K8s.Namespace,
		Image:      image,
		BackupPVC:  backupPVC,
		BackupRoot: cfg.Archive.LocalPath,
	}
}

// reconcileBuilds polls unfinished builds on an interval and advances any whose
// Job has reached a terminal phase. It exits when ctx is cancelled.
func reconcileBuilds(ctx context.Context, b *build.Builder, stderr io.Writer) {

cmd/felis/backup.go

0 → 100644
+123 −0
Changes for cmd/felis/backup.go: 123 added lines, 0 removed lines.
Original line number Diff line number Diff line
package main

import (
	"crypto/rand"
	"encoding/hex"
	"flag"
	"fmt"
	"io"
	"time"

	"felis.lolicon.best/internal/backup"
	"felis.lolicon.best/internal/config"
	"felis.lolicon.best/internal/naming"
	"felis.lolicon.best/internal/reaper"
	"felis.lolicon.best/internal/store"
	ctrl "sigs.k8s.io/controller-runtime"
)

// cmdBackup is the in-Pod entrypoint the on-demand backup Job runs. internal/backupjob
// renders a Pod whose command is `/usr/local/bin/felis backup`. It tars the mounted
// world into the archive store AND records the world_backups row, then exits — it is
// NOT a user-facing command and is never invoked by hand.
//
// Unlike `felis restore`, this command DOES hold database credentials (via the mounted
// config Secret) and calls config.Load: a backup must record its row atomically with
// the archive, exactly like the reaper — otherwise a completed archive would leak as an
// orphan file the retention pass never expires. The security review for that departure
// lives in internal/backupjob/jobspec.go. The world is mounted directly at --worlds-root
// (single-PVC mount, like restore), so the archiver's resolver returns that root for any
// PVC; the archive is written into the backup PVC at cfg.Archive.LocalPath.
func cmdBackup(args []string, stdout, stderr io.Writer) int {
	fs := flag.NewFlagSet("backup", flag.ContinueOnError)
	fs.SetOutput(stderr)
	cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
	server := fs.String("server", "", "server name whose world is being backed up")
	formerOwner := fs.String("former-owner", "", "owner recorded on the backup row (empty for an unowned server)")
	worldsRoot := fs.String("worlds-root", "/world", "mount path of the world PVC being archived")
	if err := fs.Parse(args); err != nil {
		return 2
	}
	if *server == "" {
		fmt.Fprintln(stderr, "felis backup: --server is required")
		return 2
	}

	cfg, err := config.Load(*cfgPath)
	if err != nil {
		fmt.Fprintf(stderr, "felis backup: %v\n", err)
		return 1
	}
	if cfg.Archive.Store != "tarLocal" {
		fmt.Fprintf(stderr, "felis backup: archive store %q is not implemented in this build (only tarLocal)\n", cfg.Archive.Store)
		return 1
	}
	// Reuse the reaper's retention derivation so an on-demand backup expires on the
	// same clock as an inactivity backup — one retention policy, not two.
	rcfg, err := reaperConfig(cfg)
	if err != nil {
		fmt.Fprintf(stderr, "felis backup: %v\n", err)
		return 1
	}

	// The world PVC is mounted directly at worldsRoot; the resolver returns it for
	// any target, exactly as in cmdRestore. This is the same TarLocal the reaper
	// writes archives with.
	archiver := &backup.TarLocal{
		BackupRoot: cfg.Archive.LocalPath,
		Resolve: func(string) (string, error) {
			return *worldsRoot, nil
		},
	}

	ctx := ctrl.SetupSignalHandler()

	ref, size, err := archiver.Archive(ctx, *server, naming.WorldPVCName(*server))
	if err != nil {
		fmt.Fprintf(stderr, "felis backup: archive: %v\n", err)
		return 1
	}

	drv, err := store.Open(ctx, cfg.Database.URL)
	if err != nil {
		fmt.Fprintf(stderr, "felis backup: open database: %v\n", err)
		return 1
	}
	defer drv.Close()

	rec := reaper.BackupRecord{
		ID:          newBackupID(),
		ServerName:  *server,
		FormerOwner: *formerOwner,
		BackupRef:   string(ref),
		SizeBytes:   size,
		Reason:      "manual",
		ExpiresAt:   time.Now().Add(rcfg.Retention),
	}
	if err := reaper.NewPGStore(drv.DB()).InsertBackup(ctx, rec); err != nil {
		// The archive is written but unrecorded — an orphan the retention pass would
		// never expire. Delete it so a failed backup leaves no leaked bytes, mirroring
		// the reaper's archive-then-record atomicity.
		if delErr := archiver.Delete(ctx, ref); delErr != nil {
			fmt.Fprintf(stderr, "felis backup: record failed (%v) AND orphan archive %s could not be removed: %v\n", err, ref, delErr)
			return 1
		}
		fmt.Fprintf(stderr, "felis backup: record failed, orphan archive removed: %v\n", err)
		return 1
	}

	fmt.Fprintf(stdout, "felis backup: server=%s archived %d bytes to %s (backup %s)\n", *server, size, ref, rec.ID)
	return 0
}

// newBackupID mints a world_backups primary key, matching the reaper's "bk-"+hex
// scheme so a manual and an inactivity backup are indistinguishable downstream.
func newBackupID() string {
	var b [16]byte
	if _, err := rand.Read(b[:]); err != nil {
		// crypto/rand failure is fatal and unrecoverable; a time-based fallback would
		// be a weaker ID for no benefit. ponytail: panic is the honest failure here.
		panic("felis backup: crypto/rand: " + err.Error())
	}
	return "bk-" + hex.EncodeToString(b[:])
}
+3 −0
Changes for cmd/felis/run.go: 3 added lines, 0 removed lines.
Original line number Diff line number Diff line
@@ -16,6 +16,7 @@ Commands:
  api               Run the felis-api HTTP server
  reaper            Run the world reaper / backup batch
  restore           Extract a world archive into a world volume (internal Job entrypoint)
  backup            Archive a world into the backup store and record it (internal Job entrypoint)
  manifests         Render the control-plane RBAC + NetworkPolicy install bundle as YAML
  apply             Create a MinecraftServer CRD (direct K8s write; use -f server.json)
  setup             Run host bootstrap + first-run setup console (TUI; requires root/sudo)
@@ -43,6 +44,8 @@ func run(args []string, stdout, stderr io.Writer) int {
		return cmdReaper(rest, stdout, stderr)
	case "restore":
		return cmdRestore(rest, stdout, stderr)
	case "backup":
		return cmdBackup(rest, stdout, stderr)
	case "manifests":
		return cmdManifests(rest, stdout, stderr)
	case "apply":
+115 −0
Changes for docs/changes/2026-07-07-on-demand-world-backup.md: 115 added lines, 0 removed lines.
Original line number Diff line number Diff line
# On-demand world backup (§B4 break-glass "Sync"; felis-api endpoint + Job executor)

- **Type:** feature (addition)
- **Date:** 2026-07-07
- **Area:** `internal/backupjob` (new pkg), `internal/api`, `cmd/felis`, `docs/openapi.yaml` — Go, oracle-verified
- **Commit:** `pending`
- **Task:** #31 Phase B4 break-glass ops — the "Sync" operation, resolved with the user as **immediate/on-demand world backup**. Per the user's "两者都要" decision this is built in two halves: **(this change) the felis-api endpoint that does the real backup-Job orchestration**, and (a follow-up) a break-glass menu peer that calls it while the API is alive.

## What it does

Adds `POST /api/v1/servers/{name}/backup`: an owner or admin snapshots a **stopped**
server's world into the archive store on demand, recorded as a first-class
`world_backups` row (reason `manual`) — restorable later by the existing restore path
and expired by the reaper's retention pass, so it never leaks as an orphan archive.

The backup runs asynchronously as a one-shot Kubernetes Job (the new
`internal/backupjob` package), mirroring how restore and image builds hand off to
Jobs. The handler answers **202 `backing_up`**.

## Why

felis-api cannot archive a world in-process: the world PVC is **RWO** and owned by the
operator's StatefulSet, so the API has nothing to mount at request time — the same
constraint that already makes `internal/restore` a Job. The break-glass console (which
runs direct-to-Postgres) likewise lacks the deployment coordinates (`FELIS_IMAGE`,
`FELIS_BACKUP_PVC`) needed to render the Job. Both point to the same home: the
orchestration belongs in felis-api, which holds those coordinates; other callers
invoke the endpoint.

## Design decisions

- **Backup Job self-records its `world_backups` row.** Unlike the restore Job — which
  is deliberately DB-blind because it processes a potentially poisoned archive — the
  backup Job **does** mount the felis config Secret and inserts its own backup row,
  exactly like the reaper (the only other component holding both a world mount and the
  database). This avoids the archive-then-async-record split that would otherwise leak
  orphan archives on a crash. The security review for that one departure lives in
  `internal/backupjob/jobspec.go` and is frozen by `jobspec_test.go`. Rationale: a
  backup only **reads** a world the operator already owns and tars it (bytes, never
  executed), so restore's poisoned-input threat does not apply; its blast radius (DB +
  two PVCs) is a strict subset of the reaper's, and it never deletes a PVC nor calls
  the K8s API (SA token stays un-mounted).
- **World mounted read-only, backup PVC read-write** — the mirror image of restore.
- **Stopped-gate (409 `not_stopped`).** The world PVC is RWO and held by a running
  server, so a backup Job cannot double-mount it; the handler refuses unless the server
  is fully stopped (`info.Ready || DesiredState != Stopped`). This also guarantees a
  quiescent, non-torn archive. Mirrors `handleRestoreBackup`'s gate.
- **Authorization is restore's front half, minus the former-owner match.** Backup is
  initiated by the **current** owner and records **their** ownership, so there is no
  prior owner's data to leak — the leak guard that restore needs does not apply here.
  An admin may back up an unowned (released) world; the recorded former owner is then
  empty, exactly as the reaper records for an unowned reap.
- **Unique Job name per request.** Each backup Job is named `backup-<server>-<rand>`,
  not a deterministic `backup-<server>`. A deterministic name would collide with a
  just-finished Job still inside its `TTLSecondsAfterFinished` window (10m), and the
  `AlreadyExists → 202` path would then silently produce **no** archive — the exact
  window a user (or the console "立即备份" button) retries in. Unique names make every
  request produce its own archive; `ErrAlreadyExists` remains only as a defensive
  no-op on the ~impossible suffix collision. Ceiling (documented in `backup.go`): two
  truly simultaneous taps may schedule two backup Pods — both mount the world PVC
  read-only, so neither corrupts anything; single-flight-on-running is the upgrade
  path if a double-tap storm ever appears.
- **One retention clock.** The entrypoint reuses the reaper's `reaperConfig` derivation
  so a manual backup expires on the same schedule as an inactivity backup — one policy,
  not two. The `"bk-"+hex` id scheme also matches, so manual and inactivity backups are
  indistinguishable downstream.
- **Fail-safe on record failure.** If the row insert fails, the entrypoint deletes the
  just-written archive so a failed backup leaves no unrecorded bytes.
- **Optional executor, honest 503.** Wired only when `FELIS_IMAGE` + `FELIS_BACKUP_PVC`
  are supplied (same gate as restore); otherwise `API.Backuper` is nil and the endpoint
  returns 503 `backup_unavailable`, so the authorization boundary is exercised before
  the Job executor is deployable.

## Files

| File | Change |
|---|---|
| `internal/backupjob/jobspec.go` | **new** — `BackupJob` renderer + `BackupJobName`; weak SA, token off, hardened container, world RO / backup RW, config-Secret mount |
| `internal/backupjob/backup.go` | **new** — `Backuper` (idempotent enqueue) + `Config`/`withDefaults` |
| `internal/backupjob/k8sjobs.go` | **new** — controller-runtime `CreateBackupJob` (AlreadyExists → idempotent) |
| `internal/backupjob/jobspec_test.go` | **new** — freezes the Job's security shape incl. the deliberate config-Secret mount |
| `internal/backupjob/backup_test.go` | **new** — asserts each `Backup` call mints a unique Job name (repeat-tap must not silently no-op) |
| `cmd/felis/backup.go` | **new** — `felis backup` in-Pod entrypoint: archive + self-record + orphan-cleanup |
| `cmd/felis/run.go` | dispatch `case "backup"` + usage line |
| `internal/api/backuper.go` | **new** — the narrow `Backuper` port |
| `internal/api/handlers_backups.go` | **+`handleBackupNow`** |
| `internal/api/api.go` | `Backuper` field + `POST /servers/{name}/backup` route |
| `internal/api/handlers_backup_now_test.go` | **new** — `fakeBackuper` + handler subtests |
| `internal/api/backuper_wire_test.go` | **new** — compile-time `Backuper = (*backupjob.Backuper)(nil)` |
| `cmd/felis/api.go` | wire `backuper` under the `FELIS_IMAGE`+`FELIS_BACKUP_PVC` gate; `backupConfig` helper |
| `docs/openapi.yaml` | document the `backupNow` operation |

## Verification

WSL oracle (go1.26.4, authoritative for Go):

```
go build ./...  &&  go vet ./...  &&  go test ./...   → ALL GREEN
```

The `internal/api` OpenAPI served-route contract test (`TestOpenAPIMatchesServedRoutes`)
initially failed — the new route was served but undocumented — and passes after adding
the `backupNow` operation to `docs/openapi.yaml`. `internal/backupjob` and the new
handler subtests pass. The controller-runtime `K8sJobs` binding is integration-only
(needs a live cluster) and is exercised only by the interface conformance test.

## Self-review outcome

- **ponytail (over-engineering):** the backup Job is a near-mirror of the restore Job,
  not a shared parameterization — deliberate, because its security shape differs (it
  holds DB creds) and must be asserted independently, not hidden behind a shared knob.
  No speculative config; `Config.withDefaults` fills only real deployment values.
- **correctness:** the RWO stopped-gate and the self-recording atomicity were traced to
  the reaper and restore before writing; the former-owner asymmetry vs restore is
  justified above.
+3 −1
Changes for docs/changes/INDEX.md: 3 added lines, 1 removed line.
Original line number Diff line number Diff line
@@ -23,7 +23,9 @@ not yet committed.

## Pending (built + verified, not yet committed)

_None — the break-glass halt (`c2ee21a`) and `/felis migrate` (`c1aa38b`) landed in the ledger below._
| Change | Detail doc | Status |
|---|---|---|
| On-demand world backup — felis-api `POST /servers/{name}/backup` + backup-Job executor (§B4 "Sync", half 1 of 2) | [2026-07-07-on-demand-world-backup.md](2026-07-07-on-demand-world-backup.md) | oracle-green; commit `pending` |

## Committed change ledger

Loading