- [registry] gains kaniko_image / trivy_image / build_cpu_limit / build_mem_limit overrides; empty keeps the compiled-in defaults. An air-gapped or mirrored install has no route to gcr.io/aquasec (the build egress policy allows only DNS + registry + package mirrors), so builds previously could not even start their executors. - deferred-seams: the uploads-context entry now records WHY a mount is impossible (PVCs cannot cross namespaces) and that the s3 lane also lacks credentials in the build Pod — options captured for the real fix. - troubleshooting 8e (executor ImagePullBackOff + the overrides), 13b rewritten (verified eviction refusal, 5m pressure-transition, image-GC recovery), 15 (upgrade/rollback runbook for Recreate). - Backup semantics decided and documented: a backup is the whole /data volume (worlds + config + plugins + cache) and a restore rolls all of it back — OpenAPI/README wording updated to match (same-tag images are still watched for regressions by the openapi parity gate).
This commit is contained in:
10 files changed
+165
-37
No files matched your search
+2
-2
@@ -19,14 +19,14 @@ IPv6-only 接入(`ssh -6 -i ~/.ssh/id_ed25519 root@fdb2:2c26:f4e4:0:21c:42ff:f
|
||||
| 19 | **reaper CronJob 渲染位置错误,永远无法调度**:`felis manifests` 把 CronJob 渲染到控制 ns,却引用 minecraft ns 的备份 PVC(Pod 不能跨 ns 挂 PVC:真机 `FailedScheduling: persistentvolumeclaim "felis-backups" not found`);修正 ns 后又发现 ServerAccount 也不能跨 ns 使用(`serviceaccount "felis-reaper" not found`)。而 reaper 的 Role/RoleBinding 本就在 minecraft ns | 真机三层取证(PVC/SA/调度)| CronJob 与 SA、RoleBinding subject 全部移到 `MinecraftNamespace`(提交 `e4f2cff`+`c839454`)。**修复后完整演练通过**(见下) |
|
||||
| 6 | **默认安装无备份能力 + 回收链路不可达/不可用**:(a) 没有任何环节渲染归档 PVC,`FELIS_BACKUP_PVC` 永远为空 → backup/restore 恒 503;(b) bootstrap 从不下传 reaper 三旗标 → 官方安装路径根本无法启用回收;(c) 即便启用,stock k3s 的 `<root>/<pvc>` 布局不存在,且 k3s storage root 是 `0700 root:root`、reaper pod 以 uid 1000 运行 → 真机 `lstat /worlds/…: permission denied`(fail-closed 跳过,但纯空转);(d) README 口径(“定时备份”)不实 | 真机全链路:默认渲染 PVC+env → 建服→marker→备份 202→jobs 端点 running→succeeded→改 marker→restore→读回原值 ✅;再建 resolvecheck 世界(20d idle,marker)→ CronJob 手动 Job:`evaluated=2 reaped=1`,归档含 marker、PVC+宿主目录回收、`world_backups` 得 `inactive_15d` 行、servers 行/CR 保留 ✅ | `platform/workloads.go` 渲染归档 PVC(minecraft ns、RWO 10Gi、默认 SC),`felis manifests --backup-pvc` 默认 `felis-backups`(`=` 空为显式关闭),bootstrap 统一下传 env/旗标并写 [archive] local_path;`resolveWorldDir` 新增精确 local-path 目录解析(读 PVC `spec.volumeName`,非 glob,杜绝陈旧 PV 目录顶替);ReaperRole 增 `pvc:get`;bootstrap 设 `FELIS_WORLDS_HOST_PATH` 时给 uid 1000 授 traverse(setfacl/o+x);README/故障手册改写。提交 `fd33fd0`+`2b87a5a` |
|
||||
| 8 | **磁盘打满灾难链**:DiskPressure → kubelet 驱逐控制面(无 PriorityClass 保护)→ 镜像被 GC(无外网、registry 空)→ 全部 ImagePullBackOff;释放后数分钟才恢复调度。恢复依赖人工 `docker save | k3s ctr images import -` | 三轮真机 drill(原缺陷复现 + 两轮修复验证):①自定义 1e6 类:游戏 pod 先走、控制面随后仍被驱逐(kubelet 日志逐条列出 ranked/evicted),随后镜像被 GC → ErrImagePull;②内建 `system-cluster-critical`(2e9):填盘至 1.7G free,kubelet 对 felis-api/operator/registry 全部报 **“cannot evict a critical pod”**,三者在整个 DiskPressure 期间保持 Running;游戏 pod(0)被驱逐;③释放磁盘后:`DiskPressure` 经 ~5 分钟(`eviction-pressure-transition-period`)转 False——即“恢复慢”的主因;被 GC 的游戏镜像按 runbook 重导入后 25s 恢复 ✅ | 控制面四类 pod 挂内建 `system-cluster-critical`(用户自定义类值上限 1e9,达不到 2e9 临界阈值;preemption 保留=管理面可调度,已记录权衡);troubleshooting §13b 固化事件链、5 分钟条件过渡与镜像恢复路径。提交见下 |
|
||||
| 9 | **升级策略 Recreate**:单副本 + Recreate:任何控制面升级=停机;坏升级(错 tag)中断约 95s 且需人工 `rollout undo`(无自动回滚) | 复核部署模板(Recreate + 单副本 + 无 leader election 的注释理由成立) | 不改策略(Recreate 是对无 leader election 的正确取舍),改为固化 runbook:troubleshooting §15 = 升级即重跑安装器;停机窗口=rollout 时长;坏镜像在安装器 180s 等待内以 `kubectl describe` 诊断呈现;回滚 `kubectl rollout undo`(镜像被 GC 时先按 §13b 重导入)。提交见下 |
|
||||
| 10 | **备份语义**:归档包含整个 /data(jar、libraries、cache),167MB;是否符合 world backup 定位待评估 | 真机 tar 清单(cache/、config/、eula.txt、server.properties、jar)+ 恢复语义(overlay+prune) | 评估结论:**整卷备份是正确语义**(回滚=整服状态回滚,含配置/插件),保留行为;文档化:troubleshooting §10「backup contains」+ OpenAPI/README 口径改为“整服数据卷”。提交见下 |
|
||||
|
||||
## 待决策台账(未修)
|
||||
|
||||
| # | 主题 | 说明 |
|
||||
|---|------|------|
|
||||
| 7 | 异步失败不可感知 | backup/restore 失败后无状态出口:restore 行不变、backup 无行;只有集群侧 Job/日志可查。建议状态字段或 `?failed` 查询 |
|
||||
| 9 | 升级策略 Recreate | 单副本 + Recreate:任何控制面升级=停机;坏升级(实测错 tag)服务中断约 95s 且需人工 `rollout undo`(无自动回滚)。建议 runbook/文档化 |
|
||||
| 10 | 备份语义 | 归档包含整个 /data(jar、libraries、cache),167MB;是否符合"world backup"定位待评估 |
|
||||
| 11 | PG 断连表现 | 会话查询失败报 401 而非 503(fail-closed 但误导;用户以为没登录) |
|
||||
| 12 | ready 门滞后 | 容器 Ready 后 6~10s 内 API 仍 409 not_running |
|
||||
| 13 | setup token 截断 | 43 字符 token + 长域名,80 列终端下 TUI 截断显示(复现:tmux 80 列) |
|
||||
|
||||
@@ -17,7 +17,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
|
||||
|
||||
- **即开即玩**:玩家尝试连接时自动唤醒服务器,空闲后自动休眠,像游戏主机一样省资源。
|
||||
- **Web 控制面板**:浏览器中查看服务器状态、在线玩家与资源用量,管理备份与恢复。
|
||||
- **备份与恢复**:一键把世界打包进集群内的归档库,支持从任意备份点回滚;默认安装就已启用(归档 PVC 与路径由安装器一并生成)。
|
||||
- **备份与恢复**:一键把整服数据(世界、配置、插件/模组,即整个 /data 卷)打包进集群内的归档库,支持从任意备份点回滚;默认安装就已启用(归档 PVC 与路径由安装器一并生成)。
|
||||
- **智慧回收(可选开启)**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间;安装时设置 `FELIS_WORLDS_HOST_PATH`(k3s 默认 `/var/lib/rancher/k3s/storage`)即启用每日回收,不设置则不删任何世界。
|
||||
- **多核心支持**:兼容 Paper、Fabric、Forge、NeoForge,经由 Velocity 代理统一入口。
|
||||
- **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建并部署。
|
||||
|
||||
+1
-1
@@ -17,7 +17,7 @@ Table of Contents
|
||||
|
||||
- **Wake on Join**: Servers start automatically when a player connects, and stop when idle — like hibernate for your server.
|
||||
- **Web Dashboard**: Monitor server status, online players, and resource usage from your browser, with backup and restore management.
|
||||
- **Backup & Restore**: One-click world snapshots into the cluster's archive store, with rollback from any backup point — enabled by default (the installer renders the archive PVC and its path).
|
||||
- **Backup & Restore**: One-click snapshots of a server's whole data volume (worlds, config, plugins/mods — the entire /data volume) into the cluster's archive store, with rollback from any backup point — enabled by default (the installer renders the archive PVC and its path).
|
||||
- **World Reaper** (opt in): Worlds idle for more than 15 days are automatically backed up and removed to free disk space. Enable it by setting `FELIS_WORLDS_HOST_PATH` at install time (on k3s: `/var/lib/rancher/k3s/storage`); without it, no world is ever deleted.
|
||||
- **Multi-core Support**: Compatible with Paper, Fabric, Forge, and NeoForge, federated behind a Velocity proxy.
|
||||
- **Modpack Submission**: Players submit custom modpacks; admin approval triggers automatic build and deployment.
|
||||
|
||||
@@ -402,6 +402,14 @@ func buildConfig(cfg *config.Config) build.Config {
|
||||
return build.Config{
|
||||
Namespace: cfg.Registry.BuildNamespace,
|
||||
RegistryURL: cfg.Registry.URL,
|
||||
// Empty overrides fall back to the build package's defaults, so an
|
||||
// install that has not imported kaniko/trivy keeps the compiled-in refs
|
||||
// (and fails loudly on pull rather than silently building with the wrong
|
||||
// image).
|
||||
KanikoImage: cfg.Registry.KanikoImage,
|
||||
TrivyImage: cfg.Registry.TrivyImage,
|
||||
CPULimit: cfg.Registry.BuildCPULimit,
|
||||
MemLimit: cfg.Registry.BuildMemLimit,
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -44,6 +44,32 @@ func TestAuthSourcesFromConfig(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestBuildConfig_ProjectsOverrides pins the [registry] overrides reaching the
|
||||
// build subsystem: unset fields must stay EMPTY (the build package's compiled-in
|
||||
// defaults apply there, not here), and set fields must pass through verbatim —
|
||||
// an air-gapped install points these at its imported mirrors.
|
||||
func TestBuildConfig_ProjectsOverrides(t *testing.T) {
|
||||
empty := buildConfig(&config.Config{})
|
||||
if empty.KanikoImage != "" || empty.TrivyImage != "" || empty.CPULimit != "" || empty.MemLimit != "" {
|
||||
t.Errorf("empty registry config must project empty overrides (defaults live in internal/build), got %+v", empty)
|
||||
}
|
||||
full := buildConfig(&config.Config{Registry: config.RegistryConfig{
|
||||
URL: "registry.felis.svc:5000",
|
||||
BuildNamespace: "felis-build",
|
||||
KanikoImage: "reg/kaniko:v1",
|
||||
TrivyImage: "reg/trivy:v1",
|
||||
BuildCPULimit: "1",
|
||||
BuildMemLimit: "2Gi",
|
||||
}})
|
||||
if full.KanikoImage != "reg/kaniko:v1" || full.TrivyImage != "reg/trivy:v1" ||
|
||||
full.CPULimit != "1" || full.MemLimit != "2Gi" {
|
||||
t.Errorf("registry overrides did not reach build.Config: %+v", full)
|
||||
}
|
||||
if full.Namespace != "felis-build" || full.RegistryURL != "registry.felis.svc:5000" {
|
||||
t.Errorf("namespace/registry url must keep projecting: %+v", full)
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every
|
||||
// felis-api listener: the shared factory must bound the header and idle phases
|
||||
// (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because
|
||||
|
||||
+13
-1
@@ -44,7 +44,19 @@ A grep across `*.md` and `*.go` returns both sets; only the Go ones are seams.
|
||||
apply is under way.
|
||||
- `internal/submit/blobstore.go:40` — the uploads PVC is mounted into felis-api but
|
||||
not into the Kaniko build Pod, so a submitted context is durable at the derived
|
||||
location without yet being readable by the build that consumes it.
|
||||
location without yet being readable by the build that consumes it. Audited
|
||||
2026-09-22: this is not a missing volume line — a PVC cannot cross namespaces
|
||||
(uploads live in the control namespace; build Pods run in `felis-build`), so the
|
||||
fix is a transport, not a mount. The `s3://` lane does not close it either: the
|
||||
build Job carries no AWS credentials (no env, and the weak SA's token is
|
||||
deliberately unmounted, so no IAM either). Options on the table: (a) object
|
||||
storage with credentials plumbed into the build Pod as a per-build Secret plus an
|
||||
egress allowance; (b) a context-handoff PVC/Job pair in `felis-build` fed from
|
||||
the API side; (c) a node-local path both sides mount (single-node only, and it
|
||||
hands an arbitrary Dockerfile a filesystem view — needs its own security review).
|
||||
Kaniko/Trivy images are external-only by default; `[registry] kaniko_image /
|
||||
trivy_image / build_cpu_limit / build_mem_limit` now override them for mirrored
|
||||
or air-gapped installs.
|
||||
|
||||
## Built; only its I/O is unverifiable from this repo
|
||||
|
||||
|
||||
+5
-4
@@ -2771,11 +2771,12 @@ paths:
|
||||
post:
|
||||
tags: [backups]
|
||||
operationId: backupNow
|
||||
summary: Back up a server's world on demand (owner-or-admin; server must be stopped).
|
||||
summary: Back up a server's data volume on demand (owner-or-admin; server must be stopped).
|
||||
description: >-
|
||||
Snapshots the server's world into the archive store as a first-class
|
||||
world_backups row (reason "manual"), restorable later like an inactivity
|
||||
backup. The world PVC is RWO and held by a running server, so the server must
|
||||
Snapshots the server's whole data volume (worlds, config, plugins/mods,
|
||||
jars, libraries — not just world folders) into the archive store as a
|
||||
first-class world_backups row (reason "manual"), restorable later like an
|
||||
inactivity backup. A restore replaces the volume with the archive. The world PVC is RWO and held by a running server, so the server must
|
||||
be fully stopped first (409 not_stopped otherwise). The backup runs
|
||||
asynchronously as a Job, so success is 202 (backing_up).
|
||||
x-felis-face: [external]
|
||||
|
||||
+85
-28
@@ -391,6 +391,32 @@ Inspect:
|
||||
kubectl logs -n felis-build job/<build-job>
|
||||
```
|
||||
|
||||
### 8e. Build Pods never start: executor images and air-gapped installs
|
||||
|
||||
The build Job runs Kaniko and Trivy from external registries by default
|
||||
(`gcr.io/kaniko-project/executor:latest`, `aquasec/trivy:latest`). On a box whose
|
||||
build namespace cannot reach those registries (the egress policy allows only
|
||||
DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box
|
||||
has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the
|
||||
build stays `building` until its deadline. Point the overrides at images the box
|
||||
CAN pull — typically imports into the node's containerd, pushed through the
|
||||
internal registry — in `felis.toml`:
|
||||
|
||||
```toml
|
||||
[registry]
|
||||
url = "registry.felis.svc:5000"
|
||||
build_namespace = "felis-build"
|
||||
kaniko_image = "registry.felis.svc:5000/mirror/kaniko:v1.23.2"
|
||||
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.58.1"
|
||||
build_cpu_limit = "2"
|
||||
build_mem_limit = "4Gi"
|
||||
```
|
||||
|
||||
then restart `felis-api` (it renders the Job from this config). Unset fields keep
|
||||
the defaults. Note the user-modpack context topologies are a separate,
|
||||
still-open seam (see `docs/deferred-seams.md`); this section only makes the
|
||||
executors reachable.
|
||||
|
||||
---
|
||||
|
||||
## 9. Registry push/pull failures (spec §15)
|
||||
@@ -426,6 +452,16 @@ reaps a world only when `now - last_active_at > 15d` (`inactive_15d`); the 15-da
|
||||
deadline is **hard-fixed in code** (only `warn_before` / `retention` /
|
||||
`max_local_bytes` are configurable from `felis.toml [archive]`).
|
||||
|
||||
### What a "backup" contains
|
||||
|
||||
A backup tars the server's ENTIRE data volume — the same volume the server mounts
|
||||
at `/data`: world folders, `server.properties`, plugins/mods, configs, jars,
|
||||
libraries, logs and cache, not just the `world/` directory. A restore replaces the
|
||||
volume's contents with the archive (files added since the backup are pruned), so a
|
||||
restore also rolls config/plugin changes back. Sizes are dominated by
|
||||
libraries/cache on stock Paper servers (~170MB for a fresh instance before any
|
||||
world growth) — do not size the archive PVC as if only world data were stored.
|
||||
|
||||
### The backup-before-delete invariant
|
||||
|
||||
The reap sequence (all [GO-TESTED] hermetically) preserves the world unless a
|
||||
@@ -583,8 +619,8 @@ A full disk is the one failure this platform cannot ride out by itself, because
|
||||
the images exist only in the node's containerd (air-gapped by design), so a
|
||||
GC'd image has no pull source.
|
||||
|
||||
Eviction. Every control-plane pod (api, operator, reaper, registry) runs under
|
||||
the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
|
||||
**Eviction.** Every control-plane pod (api, operator, reaper, registry) runs
|
||||
under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
|
||||
node-pressure eviction refuses to touch those pods — the log shows
|
||||
*"Eviction manager: cannot evict a critical pod"* for each of them — while
|
||||
game-server pods at the default priority 0 are evicted first. A drill that filled
|
||||
@@ -596,39 +632,32 @@ critical threshold. The built-in class allows preemption (its policy is fixed),
|
||||
so a control-plane pod that cannot fit may preempt a game pod — deliberate: the
|
||||
management plane must be placeable.
|
||||
|
||||
The pressure condition clears slowly. After you free space, the node can stay
|
||||
`DiskPressure:True` for up to ~5 minutes (`--eviction-pressure-transition-period`
|
||||
defaults to 5m, to stop the condition flapping); pods that need scheduling wait
|
||||
for it. This is the bulk of the "recovery takes minutes" observation, not a
|
||||
stuck node.
|
||||
**The pressure condition clears slowly.** After you free space, the node can stay
|
||||
`DiskPressure:True` for up to ~5 minutes
|
||||
(`--eviction-pressure-transition-period` defaults to 5m, to stop the condition
|
||||
flapping); pods that need scheduling wait for it. This is the bulk of the
|
||||
"recovery takes minutes" observation, not a stuck node.
|
||||
|
||||
But the *game* images can still be GC'd. If game pods were evicted, the kubelet
|
||||
may garbage-collect their images (unused > 2 minutes under imagefs pressure), and
|
||||
those pods then sit in `ImagePullBackOff` after recovery — re-import as above
|
||||
(`docker save felis-limbo:demo felis-lobby:demo | k3s ctr images import -`, then
|
||||
delete the stuck pods). Verified: both system servers returned to Running in
|
||||
~25s after the import.
|
||||
**The images may be gone.** If pods were evicted, the kubelet can garbage-collect
|
||||
their images (unused > 2 minutes under imagefs pressure). Those pods then sit in
|
||||
`ImagePullBackOff`/`ErrImagePull` for a tag that plainly exists —
|
||||
`k3s ctr images ls` shows it missing. Recovery:
|
||||
|
||||
Symptoms of the image-GC stage: pods stuck `ImagePullBackOff`/`ErrImagePull`
|
||||
with `kubectl describe pod` showing a pull attempt for a tag that plainly
|
||||
exists (`k3s ctr images ls` will show it missing — the kubelet GC removed it
|
||||
under imagefs pressure).
|
||||
|
||||
Recovery:
|
||||
|
||||
1. Free disk on the node (`df -h /var/lib/rancher`, the biggest consumers are
|
||||
1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are
|
||||
`k3s ctr images ls -q` and the world/backup PVCs under
|
||||
`/var/lib/rancher/k3s/storage`).
|
||||
2. Re-import the images by re-running the installer (it rebuilds imports from
|
||||
2. Re-import the images by re-running the installer (it rebuilds/re-imports from
|
||||
the local Docker store, which the kubelet GC does not touch):
|
||||
`curl -fsSL <installer URL> | sudo bash` (or `sudo felis setup`), then
|
||||
`kubectl -n felis rollout status deploy/felis-api`.
|
||||
3. Delete now-unschedulable stuck pods so they retry with the re-imported image.
|
||||
3. Delete the stuck pods so they retry against the re-imported image.
|
||||
|
||||
If the API itself is down and you only need the images back without a full
|
||||
installer run: `docker save felis:<tag> | k3s ctr images import -` restores one
|
||||
image from the Docker store (that store is deliberately a second copy; treat it
|
||||
as the recovery path, not as free space).
|
||||
For a single image without a full installer run:
|
||||
`docker save felis:<tag> | k3s ctr images import -` — the Docker store is
|
||||
deliberately a second copy; treat it as the recovery path, not as free space.
|
||||
Verified end to end in the drill: `docker save felis-limbo:demo
|
||||
felis-lobby:demo | k3s ctr images import -` plus pod deletion had both system
|
||||
servers Running ~25s later.
|
||||
|
||||
---
|
||||
|
||||
@@ -648,6 +677,32 @@ All four mandated metrics have real producers; scrape them when triaging:
|
||||
|
||||
---
|
||||
|
||||
## 15. Control-plane upgrades, and rolling back a bad one
|
||||
|
||||
There is no in-place updater: an upgrade is re-running the installer
|
||||
(`curl -fsSL <installer URL> | sudo bash`, or `sudo felis setup`), which
|
||||
rebuilds/re-imports the image and re-applies the bundle. Two properties of the
|
||||
control plane matter when you do:
|
||||
|
||||
- Both Deployments use strategy **Recreate** (single replica, no leader election:
|
||||
two overlapping instances would fight over the same cluster). An upgrade takes
|
||||
the panel/API down for the rollout window — seconds normally, longer if the new
|
||||
image still has to be imported.
|
||||
- If the new pod cannot start (bad tag, missing image), the installer's rollout
|
||||
wait fails after 180s and prints `kubectl describe` diagnostics: you see
|
||||
`ErrImagePull`/`ImagePullBackOff` there instead of a silent hang.
|
||||
|
||||
Roll back with:
|
||||
|
||||
```
|
||||
kubectl -n felis rollout undo deploy/felis-api
|
||||
kubectl -n felis rollout status deploy/felis-api
|
||||
```
|
||||
|
||||
(the same for `felis-operator` and `registry`). `rollout undo` returns to the
|
||||
previous ReplicaSet, whose image is normally still on the node; if it was GC'd
|
||||
(§13b), re-import it first.
|
||||
|
||||
## Quick reference: symptom → section
|
||||
|
||||
| Symptom | Section |
|
||||
@@ -663,10 +718,12 @@ All four mandated metrics have real producers; scrape them when triaging:
|
||||
| Local password login rejected | §5c |
|
||||
| Internal callers 401 (service token) | §6 |
|
||||
| Link/claim 400/409/412/403/404 | §7 |
|
||||
| Build push 400 / SA denied / egress hang / Failed | §8 |
|
||||
| Build push 400 / SA denied / egress hang / Failed / executor ImagePullBackOff | §8, §8e |
|
||||
| Registry push/pull unreachable | §9 |
|
||||
| World deleted unexpectedly / backup skipped | §10 |
|
||||
| Idle auto-stop not firing; player count 0 | §11 |
|
||||
| A config field seems ignored | §12 |
|
||||
| PVC left behind after delete | §13 |
|
||||
| Node out of disk; pods evicted / ImagePullBackOff | §13b |
|
||||
| Which metric to scrape | §14 |
|
||||
| Upgrade / roll back a bad control-plane image | §15 |
|
||||
@@ -119,6 +119,19 @@ type K8sConfig struct {
|
||||
type RegistryConfig struct {
|
||||
URL string `toml:"url"`
|
||||
BuildNamespace string `toml:"build_namespace"`
|
||||
// KanikoImage / TrivyImage / BuildCPULimit / BuildMemLimit override the
|
||||
// build subsystem's compiled-in defaults (gcr.io/kaniko-project/executor and
|
||||
// aquasec/trivy, 2 CPU / 4Gi per build container). The defaults assume the
|
||||
// build namespace can reach those registries; on an air-gapped or mirrored
|
||||
// install there IS no such reach (the build egress policy allows only DNS,
|
||||
// the internal registry and explicit package mirrors), so the operator must
|
||||
// point these at whatever their box can actually pull — typically images
|
||||
// imported into the node's containerd alongside the felis image. Empty keeps
|
||||
// the default.
|
||||
KanikoImage string `toml:"kaniko_image"`
|
||||
TrivyImage string `toml:"trivy_image"`
|
||||
BuildCPULimit string `toml:"build_cpu_limit"`
|
||||
BuildMemLimit string `toml:"build_mem_limit"`
|
||||
// UserUploadsContext is the object-store base under which a user-submitted
|
||||
// modpack's Kaniko build context is pinned. It belongs to the §16 build
|
||||
// subsystem's input domain (the build-context store), introduced by the
|
||||
|
||||
@@ -47,6 +47,10 @@ metallb_pool = "192.0.2.200-250"
|
||||
[registry]
|
||||
url = "registry.felis.svc:5000"
|
||||
build_namespace = "felis-build"
|
||||
kaniko_image = "registry.felis.svc:5000/mirror/kaniko:v1.23.2"
|
||||
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.58.1"
|
||||
build_cpu_limit = "1"
|
||||
build_mem_limit = "2Gi"
|
||||
|
||||
[archive]
|
||||
store = "tarLocal"
|
||||
@@ -84,6 +88,13 @@ func TestLoadValid(t *testing.T) {
|
||||
if cfg.Archive.S3.Bucket != "felis-backups" {
|
||||
t.Errorf("s3 bucket = %q", cfg.Archive.S3.Bucket)
|
||||
}
|
||||
if cfg.Registry.KanikoImage != "registry.felis.svc:5000/mirror/kaniko:v1.23.2" ||
|
||||
cfg.Registry.TrivyImage != "registry.felis.svc:5000/mirror/trivy:0.58.1" {
|
||||
t.Errorf("build image overrides = %q / %q", cfg.Registry.KanikoImage, cfg.Registry.TrivyImage)
|
||||
}
|
||||
if cfg.Registry.BuildCPULimit != "1" || cfg.Registry.BuildMemLimit != "2Gi" {
|
||||
t.Errorf("build resource overrides = %q / %q", cfg.Registry.BuildCPULimit, cfg.Registry.BuildMemLimit)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadAppliesDefaults(t *testing.T) {
|
||||
|
||||
Reference in new issue
Block a user