fix(bootstrap): 在源码构建后安装分布式节点防火墙

首次源码安装或旧发布版升级时,等待新二进制构建完成后调用 node firewall;worker 的源码安装也按所请求的版本重新构建。补充安装顺序回归检查和 main 接入说明,更新单机默认与多节点运维边界。

验证:Linux VM 上 bootstrap shell 测试全量通过;go test ./cmd/felis;shell 语法和 git diff --check。
This commit is contained in:
Lemon-miaow committed 2026-10-01 19:57:17 +08:00
1 parent adf563ffe3
commit b01aae77a2
4 files changed
+28 -13

No files matched your search

+8
View File
@@ -25,6 +25,14 @@ done
用包含此功能的 Felis 安装器在 A 重跑安装:`FELIS_DISTRIBUTED=1`、`FELIS_NODE_EXTERNAL_IP=<A 固定公网 IP>`、`FELIS_PEER_CIDRS="$PEERS"`,保留现有安装参数。该步骤会启用 `wireguard-native`、`flannel-external-ip`、NodeRestriction 和独立 agent token,并安装归档服务、最小 RBAC 和宿主机隔离规则。WireGuard 更换需要停服维护窗口。已有 worker 的对等地址列表也必须提前更新。
仅推送 main 不会自动发布 release。还未使用包含这些变更的发布资产时,在新版源码目录中以 root 执行下面的命令,显式从 main 构建;其他原安装参数继续保留。只更新安装脚本、仍使用默认 release 通道,可能下载到不支持分布式命令的旧二进制。
```bash
FELIS_REF=main FELIS_DISTRIBUTED=1 \
FELIS_NODE_EXTERNAL_IP=<A固定公网IP> \
FELIS_PEER_CIDRS="$PEERS" bash deploy/bootstrap.sh
```
直接生成部署清单时,增加:
```bash
+7 -8
View File
@@ -97,14 +97,13 @@ The installer also makes the system journal persistent (capped at
`FELIS_JOURNAL_MAX_USE`, default 1G; `FELIS_MANAGE_JOURNAL=0` skips it) and writes the
admin kubeconfig `/etc/rancher/k3s/k3s.yaml` root-only: run `sudo k3s kubectl`.
One node is the whole supported shape. A world volume is a ReadWriteOnce claim on the
node's local-path storage, so a game server's pod is pinned to the node that first
scheduled it and cannot move when that node fails; the operator and felis-api each run
as a single replica without leader election, so an upgrade or a node restart pauses
wakes and stops until their pod is back. Joining k3s agents to the cluster is untested
and gains no failover. A multi-node shape would need, at least, storage that can follow a
pod to another node and leader election in felis-operator (controller-runtime's
`LeaderElection`) so a second replica can stand by.
Single-node deployment remains the default. The opt-in [distributed mode](distributed.md)
keeps the sole API and operator on A and runs games on approved k3s agents. A world is
a ReadWriteOnce claim on its node's local-path storage; moving it requires an explicit
stopped migration through A's archive service. There is no automatic failover or
standby controller. An A restart pauses control operations until its workloads return;
a lost worker leaves its worlds on that node. Cross-node networking still requires the
three-machine acceptance described in the distributed runbook.
### Where the binary and the images come from