Commit Graph
14 Commits
Author SHA1 Message Date
Lemon-miaow 0e08fa7ced fix(servers): 主人可放弃服务器、管理员可删除服务器,由 reaper 归档世界后释放或移除 2026-09-27 03:49:49 +08:00
Lemon-miaow b7d4275ca9 feat(operator): 状态迁移、超时、自动重建、空闲停机与 RCON Secret 创建写 Event 和结构化日志,RBAC 增加 events create/patch 2026-09-25 19:32:25 +08:00
Lemon-miaow a977e229ca fix(operator): RCON Secret 改走无缓存按名读取,去掉 Secret watch,RBAC 收窄为 secrets get/create,不再缓存全命名空间 Secret 2026-09-25 19:25:24 +08:00
Lemon-miaow 6ec1b2726c feat(operator): 游戏容器加 startup/liveness 探针,超时启动按 1/2/4 分钟退避重建 Pod 至多 3 次,Running 每 60s 重探且连续 3 次失败才降级,Failed 放缓重排 2026-09-25 19:19:57 +08:00
Lemon-miaow 1054e62fa9 feat(api): 集群列表读改走 informer 缓存,审核队列、我的提交与备份列表改为服务端分页筛选并一次批量查构建,面板三页跟进 2026-09-25 18:13:21 +08:00
Lemon-miaow 7766e8efa4 fix(reaper): 回收前先停服并持有世界维护锁再归档,锁丢失不删卷;无世界卷的无主服务器只重置时钟不再重复回收 2026-09-25 03:31:50 +08:00
Lemon-miaow 489eff4494 feat(restore): 恢复前默认为当前世界做安全快照并串接恢复 Job,快照失败则不恢复;并发恢复另一备份返回 409 2026-09-25 02:52:18 +08:00
Lemon-miaow abfe60d62d fix(api): wake 与回档/备份/改文件按服务器互斥 2026-09-24 14:39:02 +08:00
Lemon-miaow 508a1c02da fix(api): refuse backup/restore before a missing world volume
A server whose world PVC does not exist yet (never started) or no longer exists
(the world was already reaped) accepted the backup/restore POST, answered 202,
and the Job sat Pending on the missing claim until its deadline with nothing
recorded anywhere — a silent no-op from the operator's seat. The live drill on
the reaped `resolvecheck` world reproduced exactly that.

Both handlers now read the world PVC (Cluster.WorldVolumeExists, over the same
naming.WorldPVCName the Jobs mount) and answer a specific 409 no_world_volume
with "start it once to create it, then retry". The felis-api Role gains the
matching get-only PVC grant — the first live run surfaced the missing RBAC as a
403 behind a 500, so the fix ships with it.

Live (auditfix38): resolvecheck -> 409 no_world_volume on both faces; test-one
(which has a world) still backs up through the new gate end to end.
2026-09-23 07:49:00 +08:00
Lemon-miaow f650bf892a fix(operator): idle auto-stop couldn't write — patch the spec, and grant the patch
Two stacked blockers behind the frozen auto-stop, both found live after the
first two fixes let the timer finally tick:

- The stop used a whole-object Update while the same reconcile loop writes
  status; that risks clobbering a concurrent status write. Switch to the
  reaper's merge-patch pattern (spec.desiredState only; EmptySince is left for
  markStopped to clear).
- The operator Role never carried minecraftservers:patch, so the call failed
  closed with 403 (visible in the operator log as 'cannot update resource
  "minecraftservers"'). Grant patch and pin it in the RBAC scope test.

With all three layers fixed, the auto-stop path is: timer persists (schema),
wake-up fires (requeue), spec write allowed (RBAC).
2026-09-23 04:08:58 +08:00
Lemon-miaow fd33fd05e1 fix(install): backups exist on a default install; retention resolves real world dirs (#6)
Three faces of one gap, all on the supported install path:

- Backup/restore answered 503 out of the box: nothing ever rendered the
  archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
  PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
  'felis manifests' names it by default (--backup-pvc= is the explicit
  no-store shape); bootstrap passes it through so the generated felis.toml
  [archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
  now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
  env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
  nothing deleted).
- Even when enabled it could not find a world on a stock install:
  resolveWorldDir now also resolves the exact local-path directory
  <pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
  glob, so a stale deleted PV's bytes can't be archived in place of the
  current world). Reaper Role gains persistentvolumeclaims:get (weaker
  than the delete it already held).

README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
2026-09-22 20:48:07 +08:00
Lemon-miaow ff7c57cf9c feat(api): expose async backup/restore job status (fixes #7)
Backup and restore only enqueue a cluster Job; a later failure left its
only trace in that Job object, invisible without kubectl. Add
GET /api/v1/servers/{name}/jobs (owner-or-admin) projecting the newest
20 managed Jobs (felis-backup / felis-restore) as
running|succeeded|failed with message and timestamps. Nil reader -> 503
jobs_unavailable, mirroring the backup/restore feature gates. RBAC gains
jobs:list; OpenAPI parity updated.
2026-09-22 20:36:36 +08:00
Lemon-miaow 90ccbfede4 fix(restore): replace a finished Job so retries enqueue; replicate felis-config
An E2E audit on a live install found that a FAILED restore held its
deterministic Job name for the rest of the 10-minute TTL, so the next
restore answered 202 'restoring' while nothing ran (ErrAlreadyExists was
treated as success unconditionally). K8sJobs now inspects the colliding
Job: in-flight still coalesces, finished (succeeded or failed) is
deleted and replaced. The minecraft-namespace Role gains jobs:get/delete
for exactly that replacement.

The same audit found the backup Job mounts the felis-config Secret but
the installer only provisions it in the control namespace, so every
backup Job stranded on FailedMount. felis setup now replicates it into
the minecraft namespace beside the service-token and forwarding
secrets.
2026-09-22 17:50:04 +08:00
flyemoji 47fcd90f75 feat(platform): add node orchestration and the felis entrypoint
The platform package that places servers across nodes and wires the operator, build, restore, and reaper subsystems, plus cmd/felis, the single binary that runs them.
2026-06-26 23:32:38 +09:00