feat(netpol): 锁定游戏服出站并为 registry 加入站围栏

This commit is contained in:
Lemon-miaow committed 2026-09-24 14:25:17 +08:00
1 parent 3424852a39
commit 7819e5de50
13 files changed
+620 -37

No files matched your search

+8 -8
View File
@@ -60,9 +60,9 @@ A grep across `*.md` and `*.go` returns both sets; only the Go ones are seams.
(recipe in docs/troubleshooting.md §8e); `trivy_java_db_repository` does the
same for the Java DB, which Trivy fetches so soon as the scanned image contains
a jar — i.e. for every real modpack build. Left unset on an egress-locked box
the scan step fails closed — Kaniko pushes, Trivy exits on the DB download —
which is the correct fail direction but leaves the build unfinished, so the
mirrors are part of a production build install.
the scan step fails closed — Trivy exits on the DB download before anything is
pushed — which is the correct fail direction but leaves every build unfinished,
so the mirrors are part of a production build install.
## Built; only its I/O is unverifiable from this repo
@@ -132,8 +132,8 @@ worth revisiting.
## Recorded outside the code
- `deploy/limbo/README.md:139` — no NetworkPolicy locks the minecraft-namespace
egress or the control-namespace ingress today, which is why the login pod reaches
`felis-api-internal:8081`. This is a conditional obligation rather than a seam: if
a future deployment adds either lock, it must also open that path. Spec v4.1 §21
asks for those policies; `cmd/felis/manifests.go` renders the game-port one.
- The minecraft-namespace egress is locked (`felis-server-egress`, DNS plus the
public internet with every private range and the node's own global addresses
excluded) and `felis-login-to-internal-api` opens the one platform path a game pod
needs — login → felis-api:8081. Any new in-cluster service a game server must call
needs its own allow policy next to that one (`internal/platform/netpol.go`).
+48 -14
View File
@@ -386,14 +386,21 @@ internal registry.
`reconcileBuilds` polls the Job; a Job reaching `Failed` is surfaced via
`writeBuildError` (JobPhase→Failed). [GO-TESTED for the mapping.] The underlying
cause — a kaniko build error or the **Trivy CRITICAL-CVE gate** failing the build
before push (spec §16) — is in the Job's pod logs and is [INTEGRATION-ONLY].
Inspect:
cause — a kaniko build error, the **Trivy CRITICAL-CVE gate** failing the build
(spec §16), or the final push — is in the Job's pod logs and is
[INTEGRATION-ONLY]. The pod runs `kaniko` (builds a tarball, never pushes) and
`trivy` (scans that tarball) as init containers, then `push` — so a CVE-rejected
image never reaches the registry. Inspect every step:
```
kubectl logs -n felis-build job/<build-job>
kubectl logs -n felis-build job/<build-job> --all-containers --prefix
```
A `push` that fails with `403` means the target repository is under `felis/` or
`mirror/` — the registry gate reserves those for the platform (§9); `401` means
the `felis-registry-push` Secret in `felis-build` is missing or stale (re-run the
installer).
### 8e. Build Pods never start: executor images and air-gapped installs
The build Job runs Kaniko and Trivy from external registries by default
@@ -422,9 +429,13 @@ build_mem_limit = "4Gi"
Mirror the executor images into the registry once. On the node itself, push
through the loopback hostPort the registry Deployment binds (docker treats
`127.0.0.1` as insecure by default; the installer leaves the daemon stopped, so
`sudo systemctl start docker` first):
`sudo systemctl start docker` first). The registry takes writes only from an
authenticated principal, and `mirror/` only from `platform`, so log in with the
platform token first:
```sh
kubectl -n felis get secret felis-registry-auth -o jsonpath='{.data.platform}' | base64 -d \
| docker login --username platform --password-stdin 127.0.0.1:5000
docker pull gcr.io/kaniko-project/executor:v1.24.0 # any versions you pin
docker pull aquasec/trivy:0.74.0
docker pull mirror.gcr.io/aquasec/trivy-java-db:1
@@ -434,6 +445,7 @@ docker tag mirror.gcr.io/aquasec/trivy-java-db:1 127.0.0.1:5000/mirror/trivy-ja
docker push 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0
docker push 127.0.0.1:5000/mirror/trivy:0.74.0
docker push 127.0.0.1:5000/mirror/trivy-java-db:1
docker logout 127.0.0.1:5000
```
From another machine, port-forward the registry instead (`kubectl -n felis
@@ -460,20 +472,19 @@ Unset fields keep the defaults.
`trivy_db_repository` is not optional on an egress-locked box. Trivy fetches its
vulnerability DB from `mirror.gcr.io`/`ghcr.io` unless told otherwise, and the
build egress policy denies those hosts — so the scan step fails closed
(`failed to download vulnerability DB`) and NO build ever completes, even though
Kaniko pushed the image. Mirror the DB into the internal registry once:
(`failed to download vulnerability DB`), nothing is pushed, and NO build ever
completes. Mirror the DB into the internal registry once:
```
# On the node (docker treats 127.0.0.1 as insecure by default), or through the
# port-forward above:
# port-forward above, logged in as platform (see the block above):
# docker pull mirror.gcr.io/aquasec/trivy-db:2
# docker tag mirror.gcr.io/aquasec/trivy-db:2 127.0.0.1:5000/mirror/trivy-db:2
# docker push 127.0.0.1:5000/mirror/trivy-db:2
```
The Job's Trivy container already runs with `--insecure`, so the internal
registry's plain HTTP works for the DB pull exactly as it does for the scanned
image. Re-mirror the tag periodically (Trivy refreshes the DB several times a
The Job's Trivy container runs with `--insecure`, so the internal registry's plain
HTTP works for the DB pull; reads need no credential. Re-mirror the tag periodically (Trivy refreshes the DB several times a
day upstream; a stale mirror only means stale CVE data, never a failed gate).
`trivy_java_db_repository` is the same story one step lazier: Trivy downloads
@@ -510,6 +521,27 @@ control namespace (or `--registry-namespace`):
control-plane pods' 256Mi — a live 475MB-layer push OOM-killed the 256Mi
template mid-upload (audit #46). Very large layers need headroom here, not
more CPU.
- **Write authorization:** registry:2 listens on the pod's loopback only; the
`registry-gate` sidecar (`felis registry-gate`, the felis image) owns the port and
the hostPort. Reads are anonymous — containerd, Kaniko and Trivy pull without a
credential — but an anonymous `GET /v2/` answers `401 Basic` so docker knows to
send the credential on a push. Every write needs HTTP basic auth against a token
in the `felis-registry-auth` Secret: `platform` may write anything (the
installer's own images, the `mirror/` DB copies); `build` (a build Job's `push`
container, via `felis-registry-push` in `felis-build`) may write anything outside
`felis/` and `mirror/` and may never delete. A missing Secret leaves the registry
read-only rather than down. The tokens persist in `/etc/felis/secrets.env`;
rotating one means editing it there and re-running the installer, then
`kubectl -n felis rollout restart deployment/registry` (the gate reads its tokens
at start).
- **Who can connect:** `felis-registry-ingress` admits only the `felis-build`
namespace to the registry pod. Node-local traffic (containerd pulls, the
installer's pushes through the hostPort) is always allowed by Kubernetes; game
servers cannot reach it at all (`felis-server-egress`).
- **GC pinning:** the registry pod's own images (registry:2 and the felis image
its gate runs) cannot be pulled from the registry they make up, so the installer
labels both `io.cri-containerd.pinned=pinned` in containerd and kubelet's image
GC never collects them. Check with `k3s ctr images ls | grep pinned`.
- **Selector quirk worth knowing:** the registry Service selector is only
`name + component=registry` — it deliberately lacks the
`part-of=felis-control-plane` label, so the registry is *invisible* to the
@@ -812,15 +844,17 @@ If a pull does NOT come back:
`/var/lib/rancher/k3s/storage`).
2. Check the registry: `kubectl -n felis get pods -l
app.kubernetes.io/component=registry` and, on the node,
`curl -s http://127.0.0.1:5000/v2/` (expect `{}`).
`curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:5000/healthz`
(expect `200`: the registry gate answers it only while registry:2 behind it
does). An anonymous `GET /v2/` answers `401` by design — see §9.
3. Check the mirror file: `/etc/rancher/k3s/registries.yaml` must map
`registry.felis.svc:5000` to `http://127.0.0.1:5000`. Missing or changed:
re-run the installer (it rewrites the file and restarts k3s only when the
content changed).
4. Re-mirror a tag the registry does not have (hand-built images were never
pushed): `sudo systemctl start docker` (the installer leaves the daemon
stopped), then `docker tag <ref> 127.0.0.1:5000/<repo>:<tag> && docker push
127.0.0.1:5000/<repo>:<tag>`.
stopped), log in as `platform` (§8e), then `docker tag <ref>
127.0.0.1:5000/<repo>:<tag> && docker push 127.0.0.1:5000/<repo>:<tag>`.
For an image that is in neither place, the old fallback still stands: re-run the
installer (it rebuilds/re-imports from the local Docker store AND mirrors into