docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path

- §8e: the executor-mirror recipe now pushes into the internal registry (the
  node's 127.0.0.1:5000, or a kubectl port-forward from another machine)
  instead of advising bare node-containerd imports — GC collects those and an
  air-gapped box cannot restore them.
- §13b: after an image GC the images come back on their own (registry + the
  registries.yaml mirror); keeps the operator checks (registry pod, mirror
  file, re-mirror a tag) and the old fallback for unmirrored images.
- §15: rollout undo no longer needs a manual re-import for installer-built tags.
- §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi
  registry memory floor (audit #46).
- deploy/{limbo,lobby}/README: manual image builds publish into the registry and
  point felis.toml at the registry ref.
This commit is contained in:
Lemon-miaow committed 2026-09-23 19:03:11 +08:00
1 parent 13d64e0000
commit fa0e8d7d97
3 files changed
+90 -35

No files matched your search

+10 -3
View File
@@ -90,14 +90,21 @@ docker build -f deploy/limbo/Dockerfile \
version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the
maven repo). maven repo).
Import into k3s and point config at it: Publish it into the cluster's registry and point config at it. On the node
itself (docker treats `127.0.0.1` as insecure by default):
``` ```
docker save felis-limbo:demo | sudo k3s ctr images import - docker tag felis-limbo:demo 127.0.0.1:5000/felis/limbo:demo
# felis.toml → [velocity] login_image = "felis-limbo:demo" docker push 127.0.0.1:5000/felis/limbo:demo
# felis.toml → [velocity] login_image = "registry.felis.svc:5000/felis/limbo:demo"
sudo felis setup sudo felis setup
``` ```
The registry keys a repository by the path after the host, so pushing through a
`kubectl -n felis port-forward svc/registry 5000:5000` from another machine is
equivalent. Hosting the image in the registry (rather than only importing it
into containerd) is what lets kubelet re-pull it after an image GC.
## Ports (handled for you) ## Ports (handled for you)
The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to
+6 -2
View File
@@ -31,8 +31,12 @@ docker build -f deploy/lobby/Dockerfile \
--build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \ --build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \
--build-arg PAPER_JAR_SHA256=<sha256 of that jar> \ --build-arg PAPER_JAR_SHA256=<sha256 of that jar> \
-t felis-lobby:demo . -t felis-lobby:demo .
docker save felis-lobby:demo | sudo k3s ctr images import - # Publish into the cluster's registry (on the node; docker treats 127.0.0.1 as
# felis.toml → [velocity] lobby_image = "felis-lobby:demo" # insecure by default — or through a `kubectl -n felis port-forward svc/registry
# 5000:5000`, which is equivalent: only the path after the host matters).
docker tag felis-lobby:demo 127.0.0.1:5000/felis/lobby:demo
docker push 127.0.0.1:5000/felis/lobby:demo
# felis.toml → [velocity] lobby_image = "registry.felis.svc:5000/felis/lobby:demo"
sudo felis setup sudo felis setup
``` ```
+74 -30
View File
@@ -398,21 +398,41 @@ The build Job runs Kaniko and Trivy from external registries by default
build namespace cannot reach those registries (the egress policy allows only build namespace cannot reach those registries (the egress policy allows only
DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box
has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the
build stays `building` until its deadline. Point the overrides at images the box build stays `building` until its deadline. Point the overrides at images **in
CAN pull — typically imports into the node's containerd, pushed through the the internal registry** — the one pull source that survives an image GC (a bare
internal registry — in `felis.toml`: node-containerd import does not: kubelet's image GC collects unused images under
disk pressure, and an air-gapped box then has nothing to restore them from) —
in `felis.toml`:
```toml ```toml
[registry] [registry]
url = "registry.felis.svc:5000" url = "registry.felis.svc:5000"
build_namespace = "felis-build" build_namespace = "felis-build"
kaniko_image = "registry.felis.svc:5000/mirror/kaniko:v1.23.2" kaniko_image = "registry.felis.svc:5000/mirror/kaniko-executor:v1.24.0"
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.58.1" trivy_image = "registry.felis.svc:5000/mirror/trivy:0.74.0"
trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2" trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2"
build_cpu_limit = "2" build_cpu_limit = "2"
build_mem_limit = "4Gi" build_mem_limit = "4Gi"
``` ```
Mirror the executor images into the registry once. On the node itself, push
through the loopback hostPort the registry Deployment binds (docker treats
`127.0.0.1` as insecure by default):
```sh
docker pull gcr.io/kaniko-project/executor:v1.24.0 # any versions you pin
docker pull aquasec/trivy:0.74.0
docker tag gcr.io/kaniko-project/executor:v1.24.0 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0
docker tag aquasec/trivy:0.74.0 127.0.0.1:5000/mirror/trivy:0.74.0
docker push 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0
docker push 127.0.0.1:5000/mirror/trivy:0.74.0
```
From another machine, port-forward the registry instead (`kubectl -n felis
port-forward svc/registry 5000:5000`) and push to `localhost:5000/...` — the
registry keys a repository by the path after the host, so pushes through either
door land in the same place the build Pods will pull from.
Put them in **both** `/etc/felis/felis.host.toml` (host-side CLI) and Put them in **both** `/etc/felis/felis.host.toml` (host-side CLI) and
`/etc/felis/felis.pod.toml` (the file rendered into the API's `felis-config` `/etc/felis/felis.pod.toml` (the file rendered into the API's `felis-config`
Secret — the two differ only in the database URL; the setup screens re-render Secret — the two differ only in the database URL; the setup screens re-render
@@ -436,12 +456,11 @@ build egress policy denies those hosts — so the scan step fails closed
Kaniko pushed the image. Mirror the DB into the internal registry once: Kaniko pushed the image. Mirror the DB into the internal registry once:
``` ```
# On a host with internet + docker access to the cluster's registry # On the node (docker treats 127.0.0.1 as insecure by default), or through the
# (add its address to the daemon's insecure-registries first; the registry # port-forward above:
# serves plain HTTP):
# docker pull mirror.gcr.io/aquasec/trivy-db:2 # docker pull mirror.gcr.io/aquasec/trivy-db:2
# docker tag mirror.gcr.io/aquasec/trivy-db:2 <registry-addr>:5000/mirror/trivy-db:2 # docker tag mirror.gcr.io/aquasec/trivy-db:2 127.0.0.1:5000/mirror/trivy-db:2
# docker push <registry-addr>:5000/mirror/trivy-db:2 # docker push 127.0.0.1:5000/mirror/trivy-db:2
``` ```
The Job's Trivy container already runs with `--insecure`, so the internal The Job's Trivy container already runs with `--insecure`, so the internal
@@ -466,6 +485,16 @@ control namespace (or `--registry-namespace`):
- **Storage:** PVC is RWO, `10Gi`, mounted at `/var/lib/registry`, **no - **Storage:** PVC is RWO, `10Gi`, mounted at `/var/lib/registry`, **no
`storageClassName`** → binds the cluster default class. If the cluster has no `storageClassName`** → binds the cluster default class. If the cluster has no
default StorageClass the PVC stays `Pending` and the registry never starts. default StorageClass the PVC stays `Pending` and the registry never starts.
- **Node-side pulls:** containerd cannot dial the Service VIP (the live stack
answered "Empty reply"), so the registry Deployment binds a loopback hostPort
(`127.0.0.1:<port>`) and the installer writes a `/etc/rancher/k3s/registries.yaml`
mirror relaying `registry.<ns>.svc:<port>` onto it. That pair is what lets
kubelet re-pull a garbage-collected image; both halves must survive together
(remove either and every pull after an image GC fails).
- **Memory:** the registry's limit is 2Gi, deliberately larger than the other
control-plane pods' 256Mi — a live 475MB-layer push OOM-killed the 256Mi
template mid-upload (audit #46). Very large layers need headroom here, not
more CPU.
- **Selector quirk worth knowing:** the registry Service selector is only - **Selector quirk worth knowing:** the registry Service selector is only
`name + component=registry` — it deliberately lacks the `name + component=registry` — it deliberately lacks the
`part-of=felis-control-plane` label, so the registry is *invisible* to the `part-of=felis-control-plane` label, so the registry is *invisible* to the
@@ -702,9 +731,11 @@ changes, because retention was never conditional in the first place.
## 13b. Node runs out of disk: what survives, and how to recover ## 13b. Node runs out of disk: what survives, and how to recover
A full disk is the one failure this platform cannot ride out by itself, because A full disk is the most destructive failure this stack sees: kubelet evicts game
the images exist only in the node's containerd (air-gapped by design), so a pods (the control plane is protected below), and its image GC then collects
GC'd image has no pull source. images nothing is running. The images have a pull source now — the in-cluster
registry — so they come back without an operator re-import; freeing space is
what completes the recovery.
**Eviction.** Every control-plane pod (api, operator, reaper, registry) runs **Eviction.** Every control-plane pod (api, operator, reaper, registry) runs
under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
@@ -725,26 +756,38 @@ management plane must be placeable.
flapping); pods that need scheduling wait for it. This is the bulk of the flapping); pods that need scheduling wait for it. This is the bulk of the
"recovery takes minutes" observation, not a stuck node. "recovery takes minutes" observation, not a stuck node.
**The images may be gone.** If pods were evicted, the kubelet can garbage-collect **The images may be gone — they come back on their own.** If pods were evicted,
their images (unused > 2 minutes under imagefs pressure). Those pods then sit in the kubelet can garbage-collect their images (unused > 2 minutes under imagefs
`ImagePullBackOff`/`ErrImagePull` for a tag that plainly exists — pressure). Every image this platform runs is ALSO hosted in the in-cluster
`k3s ctr images ls` shows it missing. Recovery: registry: the installer builds each one as `registry.<ns>.svc:5000/felis/…`
and mirrors it there, and the node's containerd is configured (a registries.yaml
mirror onto the registry's loopback hostPort) to relay those refs back through
it. So a GC'd image is re-pulled on the next attempt with no operator action —
delete the stuck pod to force an immediate retry (or wait out the backoff), and
the workload converges.
If a pull does NOT come back:
1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are 1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are
`k3s ctr images ls -q` and the world/backup PVCs under `k3s ctr images ls -q` and the world/backup PVCs under
`/var/lib/rancher/k3s/storage`). `/var/lib/rancher/k3s/storage`).
2. Re-import the images by re-running the installer (it rebuilds/re-imports from 2. Check the registry: `kubectl -n felis get pods -l
the local Docker store, which the kubelet GC does not touch): app.kubernetes.io/component=registry` and, on the node,
`curl -fsSL <installer URL> | sudo bash` (or `sudo felis setup`), then `curl -s http://127.0.0.1:5000/v2/` (expect `{}`).
`kubectl -n felis rollout status deploy/felis-api`. 3. Check the mirror file: `/etc/rancher/k3s/registries.yaml` must map
3. Delete the stuck pods so they retry against the re-imported image. `registry.felis.svc:5000` to `http://127.0.0.1:5000`. Missing or changed:
re-run the installer (it rewrites the file and restarts k3s only when the
content changed).
4. Re-mirror a tag the registry does not have (hand-built images were never
pushed): `docker tag <ref> 127.0.0.1:5000/<repo>:<tag> && docker push
127.0.0.1:5000/<repo>:<tag>`.
For a single image without a full installer run: For an image that is in neither place, the old fallback still stands: re-run the
`docker save felis:<tag> | k3s ctr images import -` — the Docker store is installer (it rebuilds/re-imports from the local Docker store AND mirrors into
deliberately a second copy; treat it as the recovery path, not as free space. the registry), or for a single image
Verified end to end in the drill: `docker save felis-limbo:demo `docker save felis:<tag> | k3s ctr images import -`. The Docker store remains a
felis-lobby:demo | k3s ctr images import -` plus pod deletion had both system deliberate second copy on the node; treat it as the recovery path, not as free
servers Running ~25s later. space.
--- ---
@@ -815,8 +858,9 @@ kubectl -n felis rollout status deploy/felis-api
``` ```
(the same for `felis-operator` and `registry`). `rollout undo` returns to the (the same for `felis-operator` and `registry`). `rollout undo` returns to the
previous ReplicaSet, whose image is normally still on the node; if it was GC'd previous ReplicaSet, whose image is normally still on the node; if the image GC
(§13b), re-import it first. collected it, the registry re-serves it automatically (§13b) for every tag the
installer built — only hand-built tags need a manual re-mirror.
## Quick reference: symptom → section ## Quick reference: symptom → section