Wire the felis-api pod-image and Velocity jar version-gather seams #11

Closed
opened 2026-07-28 12:21:37 +09:00 by FLYEMOJ1 · 2 comments
FLYEMOJ1 commented 2026-07-28 12:21:37 +09:00 (Migrated from github.com)

问题

NewSysGatherer 只接好了 CLI 那一路(k3s / cloudflared --version)。felis-api 的 Pod 镜像版本和集群外 Velocity jar 的版本这两个 seam 是显式的 nil,对应组件会报 "gather seam not wired" 然后被 Runner 跳过。

这是有意为之、也记录在案的 —— 宁可明确报未接线,也不要拿一个零值或错版本去做判断。但结果是这两个组件在版本视图里始终缺席。

证据

Identified by Claude Opus 5 (claude-opus-5) while auditing the updater's integration seams on 2026-07-28.

internal/updater/gatherer_integration.go:23-32 的注释把待办和一个附带的坑一起写清楚了:

// The felis-api pod-image read and the off-cluster Velocity jar inspection are the
// remaining integration seams; they are deliberately left nil ... Wiring them — a k8s
// client read of the control-plane Deployment's image, and however the operator exposes
// the proxy host — is the next integration step, tracked in doc.go. When the jar seam
// lands, revisit versionFromJarName's "-SNAPSHOT" handling: it reports a snapshot build
// as its stable core, which is harmless while this seam is nil and Velocity is
// Notify-only, but would suppress a legitimate "a stable is now out" notice once a real
// current flows.

验收

两个组件在版本视图里报出真实版本,不再是 "gather seam not wired"。

备注

接 jar seam 的时候必须同时处理 versionFromJarName 的 -SNAPSHOT:它现在把快照构建汇报成对应的稳定版核心号。当前无害(seam 是 nil),一旦真的有版本流进来,它会把"稳定版已发布"的通知吃掉。这是个潜伏 bug,不是接线完就自然好的。

## 问题 `NewSysGatherer` 只接好了 CLI 那一路(k3s / cloudflared `--version`)。felis-api 的 Pod 镜像版本和集群外 Velocity jar 的版本这两个 seam 是显式的 nil,对应组件会报 "gather seam not wired" 然后被 Runner 跳过。 这是有意为之、也记录在案的 —— 宁可明确报未接线,也不要拿一个零值或错版本去做判断。但结果是这两个组件在版本视图里始终缺席。 ## 证据 Identified by Claude Opus 5 (claude-opus-5) while auditing the updater's integration seams on 2026-07-28. `internal/updater/gatherer_integration.go:23-32` 的注释把待办和一个附带的坑一起写清楚了: ```go // The felis-api pod-image read and the off-cluster Velocity jar inspection are the // remaining integration seams; they are deliberately left nil ... Wiring them — a k8s // client read of the control-plane Deployment's image, and however the operator exposes // the proxy host — is the next integration step, tracked in doc.go. When the jar seam // lands, revisit versionFromJarName's "-SNAPSHOT" handling: it reports a snapshot build // as its stable core, which is harmless while this seam is nil and Velocity is // Notify-only, but would suppress a legitimate "a stable is now out" notice once a real // current flows. ``` ## 验收 两个组件在版本视图里报出真实版本,不再是 "gather seam not wired"。 ## 备注 接 jar seam 的时候必须同时处理 `versionFromJarName` 的 `-SNAPSHOT`:它现在把快照构建汇报成对应的稳定版核心号。当前无害(seam 是 nil),一旦真的有版本流进来,它会把"稳定版已发布"的通知吃掉。这是个潜伏 bug,不是接线完就自然好的。
FLYEMOJ1 commented 2026-07-28 18:34:48 +09:00 (Migrated from github.com)

这条 issue 的正文是我写的,前提不准确 —— 和 #16 一样,先把事实摆出来,决定留给你。

正文说这两个 seam "是显式的 nil,对应组件会报 gather seam not wired 然后被 Runner 跳过"。对 NewSysGatherer 是对的,对实际发货的 felis update 不对:cmd/felis/update.go:118 用的是 updater.NewHostGatherer(...),不是 NewSysGatherer。NewSysGatherer 现在只剩测试(gatherer_test.go:202)和注释在引用它。

NewHostGatherer(internal/updater/gatherer_host.go,05cb8f6 引入)两格都答得上:

  • Velocity —— 已完成,而且比 issue 要求的强。 velocityJarVersion:92-113 读 jar 里 META-INF/MANIFEST.MF 的 Implementation-Version,那正是 Velocity 运行时用 getImplementationVersion 报自己版本的同一个字段;文件名只当兜底。bootstrap 把 jar 装成固定名 velocity.jar(gatherer_host.go:34-39),文件名解析本来就答不了,所以走 manifest 是必须的,不是可选的。
  • felis-api —— 答案换了个来源,是有意的。 hostGatherer.Current:68-76 用正在跑的这个二进制自己的 build stamp,不是 Deployment 的 image tag。理由写在 gatherer_host.go:19-27:bootstrap.sh 从同一份 checkout、同一次 git describe 同时产出 felis 镜像和 /usr/local/bin/felis,所以是同一个构件;读 Deployment 回答的是另一个问题("现在 rollout 的是什么")。而且它 fail-closed —— 未打戳的 go build 报 "dev",updates.Parse 直接报错而不是编一个 0.0.0。

所以 k8s Deployment 读这一格确实还没接,但它只对 in-cluster 那条路有意义,而那条路本身也还不存在(doc.go:50-54 把 CronJob 入口列在 "Still absent" 里)。issue 现在的验收条件"两个组件在版本视图里报出真实版本",在 felis update 上已经满足了。

备注里的 -SNAPSHOT 那条也需要修正一半。versionFromJarName 确实会吃掉快照标记 —— 正则是 \d+\.\d+\.\d+|\d+\.\d+(gatherer.go:101),velocity-3.4.0-SNAPSHOT-461.jar 抽出来就是 3.4.0。但现在跑的不是这条路:manifest 读到的原始字符串直接进 updates.Parse,而 version.go:67-71 明确把 - 之后的尾巴存进 Prerelease,没有丢。潜伏 bug 还在,只是范围缩到"手工放置的带版本号文件名 + manifest 读失败"这个兜底组合,不是 issue 里说的默认路径。

建议:要么按"k8s Deployment 读 + in-cluster CronJob"重写这条(那是真没做的部分,但它等的是 CronJob 入口,不是这条 issue),要么关掉。我倾向关掉,因为 issue 标题问的两件事,发货路径上都有答案了 —— 但这是你的决定。

Verified by Claude Opus 5 (claude-opus-5) on 2026-07-28: git grep -n 'NewSysGatherer\|NewHostGatherer' -- '*.go' 的全部结果里,只有 cmd/felis/update.go:118 是生产调用点。

这条 issue 的正文是我写的,前提不准确 —— 和 #16 一样,先把事实摆出来,决定留给你。 正文说这两个 seam "是显式的 nil,对应组件会报 gather seam not wired 然后被 Runner 跳过"。对 `NewSysGatherer` 是对的,对**实际发货的 `felis update`** 不对:`cmd/felis/update.go:118` 用的是 `updater.NewHostGatherer(...)`,不是 `NewSysGatherer`。`NewSysGatherer` 现在只剩测试(`gatherer_test.go:202`)和注释在引用它。 `NewHostGatherer`(`internal/updater/gatherer_host.go`,`05cb8f6` 引入)两格都答得上: - **Velocity —— 已完成,而且比 issue 要求的强。** `velocityJarVersion:92-113` 读 jar 里 `META-INF/MANIFEST.MF` 的 `Implementation-Version`,那正是 Velocity 运行时用 `getImplementationVersion` 报自己版本的同一个字段;文件名只当兜底。bootstrap 把 jar 装成固定名 `velocity.jar`(`gatherer_host.go:34-39`),文件名解析本来就答不了,所以走 manifest 是必须的,不是可选的。 - **felis-api —— 答案换了个来源,是有意的。** `hostGatherer.Current:68-76` 用**正在跑的这个二进制自己的 build stamp**,不是 Deployment 的 image tag。理由写在 `gatherer_host.go:19-27`:bootstrap.sh 从同一份 checkout、同一次 `git describe` 同时产出 felis 镜像和 `/usr/local/bin/felis`,所以是同一个构件;读 Deployment 回答的是另一个问题("现在 rollout 的是什么")。而且它 fail-closed —— 未打戳的 `go build` 报 "dev",`updates.Parse` 直接报错而不是编一个 0.0.0。 所以 k8s Deployment 读这一格**确实还没接**,但它只对 in-cluster 那条路有意义,而那条路本身也还不存在(`doc.go:50-54` 把 CronJob 入口列在 "Still absent" 里)。issue 现在的验收条件"两个组件在版本视图里报出真实版本",在 `felis update` 上已经满足了。 备注里的 `-SNAPSHOT` 那条也需要修正一半。`versionFromJarName` 确实会吃掉快照标记 —— 正则是 `\d+\.\d+\.\d+|\d+\.\d+`(`gatherer.go:101`),`velocity-3.4.0-SNAPSHOT-461.jar` 抽出来就是 `3.4.0`。但**现在跑的不是这条路**:manifest 读到的原始字符串直接进 `updates.Parse`,而 `version.go:67-71` 明确把 `-` 之后的尾巴存进 `Prerelease`,没有丢。潜伏 bug 还在,只是范围缩到"手工放置的带版本号文件名 + manifest 读失败"这个兜底组合,不是 issue 里说的默认路径。 建议:要么按"k8s Deployment 读 + in-cluster CronJob"重写这条(那是真没做的部分,但它等的是 CronJob 入口,不是这条 issue),要么关掉。我倾向关掉,因为 issue 标题问的两件事,发货路径上都有答案了 —— 但这是你的决定。 Verified by Claude Opus 5 (claude-opus-5) on 2026-07-28: `git grep -n 'NewSysGatherer\|NewHostGatherer' -- '*.go'` 的全部结果里,只有 `cmd/felis/update.go:118` 是生产调用点。
FLYEMOJ1 commented 2026-07-31 05:00:58 +09:00 (Migrated from github.com)

验收在发货路径上已满足,关闭。

felis update 的唯一生产装配点 cmd/felis/update.go:118 用 NewHostGatherer,两个组件都报真实版本:

  • velocity — 读安装 jar 的 META-INF/MANIFEST.MF Implementation-Version(internal/updater/gatherer_host.go:92-113),和代理运行时自报的是同一个字符串;文件名只作兜底。
  • felis-api — 读二进制自身的 build stamp,无 stamp 的本地构建 fail-closed,而不是编造一个让所有上游版本都像升级的 0.0.0(gatherer_host.go:66-76)。

NewSysGatherer 里那两个 nil seam 不删:internal/updater/doc.go:44-54 写明它们是 in-cluster 调用方(CronJob 入口,尚不存在、单独跟踪)的座位,gather error 优于错版本,gatherer_test.go:200-208 锁着这个 fail-closed 行为。

备注里的 -SNAPSHOT 丢失核实后只剩文件名兜底路径(internal/updater/gatherer.go:101-112,注释已写明 advisory-only、仅影响提示文案);manifest 主路径经 updates.Parse 保留 Prerelease。维持现状。

验收在发货路径上已满足,关闭。 `felis update` 的唯一生产装配点 `cmd/felis/update.go:118` 用 `NewHostGatherer`,两个组件都报真实版本: - **velocity** — 读安装 jar 的 `META-INF/MANIFEST.MF` `Implementation-Version`(`internal/updater/gatherer_host.go:92-113`),和代理运行时自报的是同一个字符串;文件名只作兜底。 - **felis-api** — 读二进制自身的 build stamp,无 stamp 的本地构建 fail-closed,而不是编造一个让所有上游版本都像升级的 0.0.0(`gatherer_host.go:66-76`)。 `NewSysGatherer` 里那两个 nil seam 不删:`internal/updater/doc.go:44-54` 写明它们是 in-cluster 调用方(CronJob 入口,尚不存在、单独跟踪)的座位,gather error 优于错版本,`gatherer_test.go:200-208` 锁着这个 fail-closed 行为。 备注里的 `-SNAPSHOT` 丢失核实后只剩文件名兜底路径(`internal/updater/gatherer.go:101-112`,注释已写明 advisory-only、仅影响提示文案);manifest 主路径经 `updates.Parse` 保留 Prerelease。维持现状。
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: FelisMC/Felis#11