Commit Graph
50 Commits
Author SHA1 Message Date
Lemon-miaow fe15b56074 docs(reaper): 说明回收前停服持锁、归档回读校验与存储清扫、无世界卷时的处理及运行汇总各计数 2026-09-25 03:34:54 +08:00
Lemon-miaow f8f112b8ca feat(api): 访问日志与 HTTP/运行时指标、安全响应头与面板 CSP、跨站写拦截、请求体读截止、SSE 定期重鉴权与续传、404/405/413 信封、status 按所有权裁剪、停机并发排空 2026-09-25 02:14:08 +08:00
Lemon-miaow 3ed6603201 feat(store): api/reaper/offsite/migrate 按版本集合校验 schema,库比二进制新时拒绝启动;bootstrap 迁移前失败回滚宿主二进制 2026-09-25 01:46:33 +08:00
Lemon-miaow b1678f78c8 feat(update): felis update 报告 JRE 与 PostgreSQL(含停更提示),bootstrap 支持 FELIS_UPGRADE_DEPS=1 升级 k3s/cloudflared 2026-09-25 01:35:52 +08:00
Lemon-miaow 3e61fde06d fix(bootstrap): 关闭 BuildKit 默认 attestation,无改动重跑不再重启 login/lobby;同版本重跑同时重启 registry gate 2026-09-25 00:43:02 +08:00
Lemon-miaow 82545548e7 feat(bootstrap): game stack 按 lock 文件固定构建并校验 sha256,基础镜像按 digest 固定,JRE 固定补丁版本 2026-09-25 00:24:02 +08:00
Lemon-miaow aace66e8d3 fix(bootstrap): 重跑只重启有变化的 Velocity、系统服与 PostgreSQL,无 firewalld 时用 nftables 收口 5432,PG 大版本不符拒绝启动 2026-09-25 00:06:47 +08:00
Lemon-miaow a4fa16ae69 feat(bootstrap): 控制面镜像按版本打 tag,升级后 rollout undo 可回到上一版本 2026-09-24 23:52:30 +08:00
Lemon-miaow 07eb682358 docs(upgrade): 列出安装器对下载物的校验与离线导入 registry 镜像的做法 2026-09-24 23:39:39 +08:00
Lemon-miaow 8296153a38 docs(build): 说明构建工具镜像、定时刷新与离线环境做法 2026-09-24 23:25:21 +08:00
Lemon-miaow 80c4ea8966 docs(registry): 说明 manifest 清理、GC 窗口、PVC 容量与上传上限 2026-09-24 23:10:40 +08:00
Lemon-miaow a27d76ae1d docs(build): 修正上下文存储与构建镜像拉取的过时描述 2026-09-24 22:26:29 +08:00
Lemon-miaow e836a73a8a fix(build): 构建并发上限与排队,构建命名空间加 ResourceQuota,SyncAll 逐个容错 2026-09-24 22:26:21 +08:00
Lemon-miaow e521cf5976 fix(submit): 审阅绑定上下文 sha256,批准须带摘要,构建只从内部 API 取上下文并校验字节 2026-09-24 22:24:50 +08:00
Lemon-miaow 13b65e19ec fix(build): 构建 pod 等出口策略生效再运行,加 seccomp、可选 user namespace 与磁盘上限,上下文解包限总字节与条目数 2026-09-24 20:42:14 +08:00
Lemon-miaow 9bda3a52fa fix(backup): 手动备份按服冷却、每服保留上限与独立保留期,容量驱逐不再删除回收世界的唯一副本 2026-09-24 19:58:59 +08:00
Lemon-miaow f69cdec9d4 fix(reaper): 有服务器处理失败或过期备份删不掉时以非零码退出,让 Job 失败告警触达运维 2026-09-24 19:34:09 +08:00
Lemon-miaow 22becd6859 fix(reaper): 认领时重置活跃时钟与预警,回收只复用本次认领后的 reaper 归档 2026-09-24 19:31:23 +08:00
Lemon-miaow 47890ca913 feat(offsite): 世界归档与数据库备份加密同步到异地 S3,reaper 确认异地副本后才删除世界 2026-09-24 19:25:16 +08:00
Lemon-miaow d17524cd67 feat(watchdog): 主机侧巡检定时器按异常邮件通知平台所有者,operator 增加 phase 与 build_info 指标、卡死存活探针与告警规则 2026-09-24 18:06:19 +08:00
Lemon-miaow 215bfd78d7 fix(images): 服务器镜像在创建时固定到仓库 digest,更换镜像需确认备份,安装器重建前先固定旧服并推送不可变版本标签 2026-09-24 16:57:38 +08:00
Lemon-miaow 857b6da886 fix(operator): 删除依赖 rcon-cli 的 preStop 死代码,缩容前由 operator 经 RCON 执行 save-all flush 2026-09-24 16:39:22 +08:00
Lemon-miaow 5521e498a9 fix(platform): minecraft 命名空间强制 PodSecurity baseline,reaper 世界根目录改走静态 hostPath PV 2026-09-24 16:28:12 +08:00
Lemon-miaow c1796bea17 feat(servers): 新建服务器默认空闲 10 分钟自动停服,面板可调,converge 回填旧服,系统服不休眠 2026-09-24 16:12:53 +08:00
Lemon-miaow bd909d5858 fix(operator): 读不出玩家数时暂停空闲自动停服,支持 1.12/Bukkit/EssentialsX 与颜色码 2026-09-24 16:06:55 +08:00
Lemon-miaow 857c73a66a fix(audit): 审计按账号 id 归属并记录来源 IP/UA,登录失败与限速入审计和指标,写入失败计数告警 2026-09-24 16:02:44 +08:00
Lemon-miaow c4e4953f3d fix(auth): 公开登录门按来源限速并设全站发信上限,冷却表定期清理 2026-09-24 15:51:42 +08:00
Lemon-miaow c7db7d4126 feat(db): 控制面 PG 定时备份、迁移前快照与原子恢复 2026-09-24 15:19:42 +08:00
Lemon-miaow abfe60d62d fix(api): wake 与回档/备份/改文件按服务器互斥 2026-09-24 14:39:02 +08:00
Lemon-miaow 7819e5de50 feat(netpol): 锁定游戏服出站并为 registry 加入站围栏 2026-09-24 14:25:17 +08:00
Lemon-miaow b9ebc872ad docs(troubleshooting): the [INERT] legend points at a field that no longer exists
The legend said "§12 lists the one field this still applies to", but the last
inert field (spec.storage.retainOnDelete) was removed rather than implemented
(§13), and §12 has said "every field below is read by a controller" since.
Reworded so a reader whose change looks ignored follows the condition question
instead of hunting for a dead field.
2026-09-24 10:35:35 +08:00
Lemon-miaow c57daaf861 feat(cli): add felis converge for fields a newer desired spec never delivered (tracker #1)
Provisioning is create-if-absent, so a field the desired spec gained after an
install (spec.rcon, spec.startup.healthHTTPPort, a derived env key) never
reaches the existing login/lobby CR while every re-run of setup reports
success — the reported 'configuration updates never reach an installed
deployment' symptom. converge is the explicit pass: it fills exactly the
zero-valued whitelist fields and the derived env (including a missing key,
which refreshDerivedEnv deliberately never adds), and never overwrites a
non-zero value. The timing stays with the operator because enabling RCON or
the HTTP readiness gate on a pre-listener image would wedge that server in
Starting until it was marked Failed.

Tests: fills predated fields while operator edits survive / non-zero values
left alone / absent + foreign + unset-image guards. usage table updated so the
router-parity test passes; troubleshooting gains §12b.
2026-09-24 10:35:14 +08:00
Lemon-miaow b14bacbfc6 fix(build): mirror the Trivy Java DB — jar-bearing builds failed closed at the scan gate (#72) 2026-09-24 03:11:01 +08:00
Lemon-miaow 397a400d57 fix(update): point the apply guidance at the installer, not felis setup (#54)
On a completed install 'felis setup' never re-runs the installer: its host-bootstrap phase only runs while an install marker is missing, so it opens the config console and moves no component. Live on the audit box, a clean 'felis setup' run left /opt/felis/velocity/velocity.jar's mtime and hash untouched while an installer re-run logged 'resolving the newest Velocity 3.5.1 build'. The 'felis update' guidance was wrong three ways accordingly: 'run: sudo felis setup' for panel/velocity/plugins, the 'felis setup is idempotent and re-runs the installer' trailer, and the felis-api-only exception block, whose scoping taught the same false model for velocity.

Point every planner-backed selector at the tested path -- re-running the installer (the README's install one-liner) -- and replace the scoped caveat with one trailer: the channel is not persisted (pass FELIS_VERSION_BOOTSTRAP=dev on a host that tracks main), the private repo's one-liner needs the README's token'd form, and 'felis setup is not this path'. troubleshooting.md SS15 drops the same false alternative and gains the channel caveat.

Gates: gofmt, go vet, go test ./..., deploy/bootstrap_test.sh all clean. Green live (v0.0.0+fix54 installed to /usr/local/bin over the fix52 backup, sha 0bd49467...): --panel and --velocity print the installer one-liner plus the single trailer, --mc stays command-free, --all prints the trailer once.
2026-09-23 21:01:32 +08:00
Lemon-miaow 72553cb414 docs(troubleshooting): the installer leaves docker stopped — start it before manual pushes
Both the §8e mirror recipe and the §13b re-mirror step run docker tag/push,
and a fresh install (or re-run) ends with the daemon stopped. One line each so
the runbook does not fail on 'Cannot connect to the Docker daemon'.
2026-09-23 19:35:13 +08:00
Lemon-miaow fa0e8d7d97 docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
2026-09-23 19:03:11 +08:00
Lemon-miaow 43df08b52a feat(alerts): ship Felis alert rules with promtool unit tests; document scraping & rules (troubleshooting §14) 2026-09-23 16:17:51 +08:00
Lemon-miaow ae6e9256c6 docs(troubleshooting): 8e — apply build-image overrides through the config Secret (restart alone does not) 2026-09-23 15:56:44 +08:00
Lemon-miaow 2010961d32 fix(workloads): world executors run as root so game-image worlds are readable
A live backup drill on test-one failed: 'tar walk: open
/world/world/level.dat: permission denied'. The world volume belongs to
the game image's own UID (root for every Paper image we ship), and Paper
saves level.dat mode 0600 — a fixed uid-1000 executor can neither read
it (backup/reaper archive) nor overwrite it (restore). The same identity
silently broke on-demand backups, restores, and the reaper for every
server that had saved once.

Run the backup Job, restore Job, file Job, and the reaper pod as root
with DAC_OVERRIDE on top of drop-ALL — the same owner-matching precedent
as the operator's forwarding-init container; DAC_OVERRIDE extends it to
game images whose UID is neither root nor ours. FSGroup is omitted when
zero so a root executor never chgrps the world volume. Shape tests
updated for the new identity.
2026-09-23 06:20:07 +08:00
Lemon-miaow 089d4f3a80 docs(audit): idle auto-stop postmortem — ledger #25 and self-checks in §11
Records the three stacked defects (schema pruning, no wake-up, missing RBAC
grant) with the live evidence trail, and turns §11 from a 'it is implemented'
note into a three-step self-check for the field. Deployed image note bumped to
auditfix22.
2026-09-23 04:14:11 +08:00
Lemon-miaow 8e7c7bbf24 fix(reaper): deliver pre-reap warnings for real — and never fake a delivery
The §18 warning path had no delivery channel at all: no Warner implementation
existed, `felis reaper` passed nil, and maybeWarn still stamped warned_3d_at/
warned_1d_at and counted `warned=N`. So every owned server was silently reaped
15 days after its last join with no notice, and the operator's only feedback
said warnings were sent. Two changes close that:

- Honest stamps: warned_* now records a DELIVERED notice. A nil Warner logs
  `warning suppressed — no warner wired` and does NOT stamp; a delivery error
  logs and retries on the next daily run (bounded by the warning window). The
  stamps are no longer burned by notices nobody received.

- A real channel: mail.SendNotice (the second and last message shape the mail
  package sends) plus a mailWarner that resolves the owner's VERIFIED email
  and mails the notice through the configured [smtp] relay. `felis reaper`
  wires it when [smtp] is set (same password_ref convention as felis-api) and
  prints exactly what happens when it is not.

Plumbing so the in-cluster CronJob can actually reach the relay: the reaper
pod gets the optional FELIS_SMTP_PASSWORD env (same Secret as felis-api), and
the "configure email" screen now refreshes the minecraft-namespace mirrors of
felis-smtp AND felis-config (a secretKeyRef is namespace-local, and the config
mirror is what carries [smtp] into the reaper's own config). `felis setup`'s
replica list gains felis-smtp for fresh installs.

Tests: the delivered/retried/suppressed matrix in internal/reaper (the old
"stamp advances on failure" contract is deliberately replaced), the notice
message shape, the warner's resolve/send/failure paths, and the CronJob's
optional-secret env. docs/troubleshooting.md §10 now states the real semantics.
2026-09-23 03:47:19 +08:00
Lemon-miaow 02fd2de502 fix(build): three drill-driven fixes so the lane actually completes on a starter node
The first live build (Kaniko v1.24, 4 vCPU / 5.5 GiB node) walked the new
transport end to end and hit three real defects, each invisible to unit tests:

- The Job requested its FULL limits (2 CPU / 4Gi per container), so the build
  Pod never scheduled on the platform's own starter node: FailedScheduling /
  Insufficient memory, Pending forever. Requests are now a small floor
  (250m / 512Mi, never above a configured cap) while the limits stay the
  safety caps.
- Kaniko re-copies the Dockerfile out of the context and chowns/chmods it to
  the source owner; a 65532-owned context (the distroless felis image uid)
  fails that under the pod's dropped capabilities ('copying dockerfile:
  chown /kaniko/Dockerfile: operation not permitted'). The fetch container
  now extracts as root — the uid Kaniko already runs as — so the copy
  succeeds; the pod was root by necessity regardless.
- Trivy's DB fetch is exactly what the build egress lock denies: the scan
  step failed closed on mirror.gcr.io. New [registry] trivy_db_repository
  renders --db-repository, and docs/troubleshooting.md §8e now carries the
  verified mirror recipe (docker pull/tag/push of aquasec/trivy-db:2 into the
  internal registry; --insecure already covers its plain HTTP).

Verified live after this batch: fetch initContainer streamed the blob through
the API + netpol + token, Kaniko built and pushed registry.felis.svc:5000/
user-uploads/sub-<id>:latest, and Trivy scanned against the mirrored DB.
2026-09-22 23:01:25 +08:00
Lemon-miaow 87a9f4eb25 feat(build)/docs: make executor images configurable; document the build lane's real seams (#9, #10)
- [registry] gains kaniko_image / trivy_image / build_cpu_limit /
  build_mem_limit overrides; empty keeps the compiled-in defaults. An
  air-gapped or mirrored install has no route to gcr.io/aquasec (the
  build egress policy allows only DNS + registry + package mirrors), so
  builds previously could not even start their executors.
- deferred-seams: the uploads-context entry now records WHY a mount is
  impossible (PVCs cannot cross namespaces) and that the s3 lane also
  lacks credentials in the build Pod — options captured for the real fix.
- troubleshooting 8e (executor ImagePullBackOff + the overrides),
  13b rewritten (verified eviction refusal, 5m pressure-transition,
  image-GC recovery), 15 (upgrade/rollback runbook for Recreate).
- Backup semantics decided and documented: a backup is the whole /data
  volume (worlds + config + plugins + cache) and a restore rolls all of
  it back — OpenAPI/README wording updated to match (same-tag images are
  still watched for regressions by the openapi parity gate).
2026-09-22 22:03:48 +08:00
Lemon-miaow 0a2d654e68 fix(platform): control plane runs system-cluster-critical, so eviction refuses it (#8)
Following the first shield attempt (custom class, value 1e6) a live drill
showed the limit: kubelet evicted the game pods and then the api,
operator and registry anyway — evicting them was never what reclaimed
the disk — and with the images containerd-only, the GC stage left
everything in ImagePullBackOff. A custom class cannot be raised past 1e9
(the API caps user-defined values), while kubelet's eviction refusal
needs >= 2e9, so the control plane now uses the built-in
system-cluster-critical.

Re-drilled: disk filled to 1.7G free -> login/lobby evicted, and kubelet
logged "cannot evict a critical pod" for felis-api/operator/registry,
which stayed Running throughout. Recovery facts now in troubleshooting
13b: the DiskPressure condition lingers ~5m after space is freed
(--eviction-pressure-transition-period), and game images GC'd while
their pods were evicted need the documented re-import (verified: 25s to
Running).
2026-09-22 21:55:54 +08:00
Lemon-miaow fe310743a2 fix(platform): give the control plane a PriorityClass eviction shield (#8)
A full disk made kubelet's node-pressure eviction pick control-plane pods
alongside game pods (both priority 0), and with the images existing only
in the node's containerd (air-gapped), losing the api meant a manual
image re-import. Every control-plane pod template (api/operator/reaper/
registry) now names the bundle's cluster-scoped felis-control-plane
PriorityClass: value 1,000,000, preemptionPolicy Never — eviction order
only, never preempting a running game server. The image-GC half is not
code-fixable on an air-gapped box; troubleshooting gains 13b with the
recovery path (re-run the installer to rebuild imports, or docker save |
k3s ctr images import - for one image).
2026-09-22 21:08:56 +08:00
Lemon-miaow 2b87a5a13b fix(install): grant the reaper traverse on the worlds root (#6)
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.

Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
2026-09-22 21:03:17 +08:00
flyemoji 23792d6251 fix(crd): remove spec.storage.retainOnDelete rather than leave it inert
The field validated, shipped in the CRD, and reached no controller. The world
PVC survives deletion unconditionally -- it is a StatefulSet VolumeClaimTemplate,
StatefulSet deletion does not cascade to template PVCs, and no finalizer exists
anywhere in the operator. So setting it true described what already happened,
and setting it false did nothing at all. False is the worse half: it reads as a
request to delete a world, and was silently ignored.

This departs from spec v4.1 §5, which asks for
"删除:finalizer 清 Service/STS/ConfigMap,PVC 按 retainOnDelete". Neither half
was ever built. Restoring that line means adding a finalizer whose other listed
duties -- Service, StatefulSet, ConfigMap -- ownerReference GC already performs,
so the only work it would newly do is delete worlds, on a path that does not
pass the reaper's verified-backup check. The reaper is the one thing in the
system allowed to destroy a world and it earns that by proving a backup first.
A second door without that check is not an improvement.

The spec is a frozen versioned document, so it is left alone and the departure
is recorded in troubleshooting.md §13, beside the behaviour it explains. §12
loses its inert row and its opening sentence, which existed to introduce this
one field: every field in that table is now read by a controller.

Deployed installs need nothing. A CR still carrying retainOnDelete keeps
working, because a v1 CRD prunes unknown keys on the next write and the
behaviour the field claimed to control was never conditional.

go build, go vet and go test ./... pass on Linux with zero failures; the CRD
still parses and storage keeps size and storageClassName.
2026-07-28 17:27:01 +09:00
flyemoji 3af5cc360c docs(troubleshooting): correct four fields the runbook documents as inert
Sections 11 and 12 told the operator that idle auto-stop, both startup
budgets, and the player tally are read by nobody. All four are read, and
§1 repeated the same claim in its strongest form: "the operator has no
start timeout ... loops forever".

  spec.idle.autoStopEnabled       reconciler.go:175
  spec.idle.emptySecondsBeforeStop  reconciler.go:175
  spec.startup.timeoutSeconds     reconciler.go:479, called at :126
  spec.startup.readinessTimeoutSeconds  reconciler.go:490, called at :157
  status.players.online           reconciler.go:413 (markRunningReady)

The repository already contained the disproof:
TestReconcileRunning_StartupTimeoutConvertsToFailed and
TestReconcileRunning_ReadinessTimeoutConvertsToFailed both assert the
escalation §1 said does not exist. The test §1 cited,
TestReconcileRunning_RconProbeFailureStaysStarting, only asserts that a
single failed probe does not flap the phase; that was read as "forever".

§11 was the costly one, because it misdiagnosed a configuration problem
as a missing feature. Idle auto-stop and the player tally both hang off
spec.rcon.enabled -- the tally is a by-product of the RCON readiness
probe (prober.go:63 runs `list`), and reconciler.go:175 carries the RCON
condition explicitly so a never-sampled zero cannot stop a server full of
people. Following the old text, an operator whose RCON was never enabled
would conclude the feature was unwritten and stop. The section now opens
with the jsonpath that reads spec.rcon.enabled.

Also separated the prober's fixed 5s dial timeout (prober.go:45) from
spec.startup.readinessTimeoutSeconds, which §1c conflated: the former
bounds one probe, the latter is a deadline for the whole start measured
from status.startRequestedAt.

spec.storage.retainOnDelete is the one field still genuinely inert, so
the [INERT] legend and §12 stay -- §12 now records the condition each
read field depends on instead of claiming none of them are read.
2026-07-28 13:49:07 +09:00
flyemoji 2ba994889e fix(platform): front the felis-api internal face on its own ClusterIP Service
The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.

Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).

Manifest-level fix; the live packet path is pending real-cluster verification.
2026-07-07 11:24:09 +09:00
flyemoji ac02c69612 docs(troubleshooting): add operator failure-mode checklist
Add docs/troubleshooting.md covering the common failure modes the spec
implies, grounded in the actual control-plane code paths:

- Stuck Starting (PodNotReady / RconSecretUnavailable / RconNotReachable)
  and the deliberate absence of a Starting->Failed timeout.
- Failed reachable only via InvalidSpec on a malformed spec.storage.size,
  plus the stale status.endpoint=direct caveat after a failure.
- Routing via status.endpoint direct/fallback and the empty fallbackServer
  pitfall; wake 403/429/503 gate order.
- online-mode coupling and Velocity's offline-mode routing refusal.
- Cloudflare Access 401/403, nil-Keyfunc fail-closed, audience checks,
  the absence of an issuer check, and local-session gating.
- Internal service-token (FELIS_SERVICE_TOKEN) rejection path.
- link/claim error codes, Kaniko build denials (SA-by-absence RBAC,
  default-deny egress, internal-registry push gate), and the registry
  DNS contract.
- Reaper backup-before-delete invariant and false-delete vectors.
- Unimplemented idle auto-stop, permanently-zero players.online, the
  inert CRD fields, and the always-survives world PVC behaviour.

Each item is labelled with its evidence grade (GO-TESTED / CODE-ONLY /
INTEGRATION-ONLY / INERT) so operators know what is verified versus
asserted.
2026-06-30 13:38:37 +09:00