docs(troubleshooting): 8e/9/13b/15 — registry-hosted images and the loopback pull path
- §8e: the executor-mirror recipe now pushes into the internal registry (the node's 127.0.0.1:5000, or a kubectl port-forward from another machine) instead of advising bare node-containerd imports — GC collects those and an air-gapped box cannot restore them. - §13b: after an image GC the images come back on their own (registry + the registries.yaml mirror); keeps the operator checks (registry pod, mirror file, re-mirror a tag) and the old fallback for unmirrored images. - §15: rollout undo no longer needs a manual re-import for installer-built tags. - §9: documents the loopback hostPort/mirror pair as one unit and the 2Gi registry memory floor (audit #46). - deploy/{limbo,lobby}/README: manual image builds publish into the registry and point felis.toml at the registry ref.
This commit is contained in:
3 files changed
+90
-35
No files matched your search
+10
-3
@@ -90,14 +90,21 @@ docker build -f deploy/limbo/Dockerfile \
|
|||||||
version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the
|
version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the
|
||||||
maven repo).
|
maven repo).
|
||||||
|
|
||||||
Import into k3s and point config at it:
|
Publish it into the cluster's registry and point config at it. On the node
|
||||||
|
itself (docker treats `127.0.0.1` as insecure by default):
|
||||||
|
|
||||||
```
|
```
|
||||||
docker save felis-limbo:demo | sudo k3s ctr images import -
|
docker tag felis-limbo:demo 127.0.0.1:5000/felis/limbo:demo
|
||||||
# felis.toml → [velocity] login_image = "felis-limbo:demo"
|
docker push 127.0.0.1:5000/felis/limbo:demo
|
||||||
|
# felis.toml → [velocity] login_image = "registry.felis.svc:5000/felis/limbo:demo"
|
||||||
sudo felis setup
|
sudo felis setup
|
||||||
```
|
```
|
||||||
|
|
||||||
|
The registry keys a repository by the path after the host, so pushing through a
|
||||||
|
`kubectl -n felis port-forward svc/registry 5000:5000` from another machine is
|
||||||
|
equivalent. Hosting the image in the registry (rather than only importing it
|
||||||
|
into containerd) is what lets kubelet re-pull it after an image GC.
|
||||||
|
|
||||||
## Ports (handled for you)
|
## Ports (handled for you)
|
||||||
|
|
||||||
The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to
|
The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to
|
||||||
|
|||||||
@@ -31,8 +31,12 @@ docker build -f deploy/lobby/Dockerfile \
|
|||||||
--build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \
|
--build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \
|
||||||
--build-arg PAPER_JAR_SHA256=<sha256 of that jar> \
|
--build-arg PAPER_JAR_SHA256=<sha256 of that jar> \
|
||||||
-t felis-lobby:demo .
|
-t felis-lobby:demo .
|
||||||
docker save felis-lobby:demo | sudo k3s ctr images import -
|
# Publish into the cluster's registry (on the node; docker treats 127.0.0.1 as
|
||||||
# felis.toml → [velocity] lobby_image = "felis-lobby:demo"
|
# insecure by default — or through a `kubectl -n felis port-forward svc/registry
|
||||||
|
# 5000:5000`, which is equivalent: only the path after the host matters).
|
||||||
|
docker tag felis-lobby:demo 127.0.0.1:5000/felis/lobby:demo
|
||||||
|
docker push 127.0.0.1:5000/felis/lobby:demo
|
||||||
|
# felis.toml → [velocity] lobby_image = "registry.felis.svc:5000/felis/lobby:demo"
|
||||||
sudo felis setup
|
sudo felis setup
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
+74
-30
@@ -398,21 +398,41 @@ The build Job runs Kaniko and Trivy from external registries by default
|
|||||||
build namespace cannot reach those registries (the egress policy allows only
|
build namespace cannot reach those registries (the egress policy allows only
|
||||||
DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box
|
DNS, the internal registry and `--package-cidr` mirrors — and an air-gapped box
|
||||||
has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the
|
has no route at all), the Pods sit in `ImagePullBackOff`/`ErrImagePull` and the
|
||||||
build stays `building` until its deadline. Point the overrides at images the box
|
build stays `building` until its deadline. Point the overrides at images **in
|
||||||
CAN pull — typically imports into the node's containerd, pushed through the
|
the internal registry** — the one pull source that survives an image GC (a bare
|
||||||
internal registry — in `felis.toml`:
|
node-containerd import does not: kubelet's image GC collects unused images under
|
||||||
|
disk pressure, and an air-gapped box then has nothing to restore them from) —
|
||||||
|
in `felis.toml`:
|
||||||
|
|
||||||
```toml
|
```toml
|
||||||
[registry]
|
[registry]
|
||||||
url = "registry.felis.svc:5000"
|
url = "registry.felis.svc:5000"
|
||||||
build_namespace = "felis-build"
|
build_namespace = "felis-build"
|
||||||
kaniko_image = "registry.felis.svc:5000/mirror/kaniko:v1.23.2"
|
kaniko_image = "registry.felis.svc:5000/mirror/kaniko-executor:v1.24.0"
|
||||||
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.58.1"
|
trivy_image = "registry.felis.svc:5000/mirror/trivy:0.74.0"
|
||||||
trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2"
|
trivy_db_repository = "registry.felis.svc:5000/mirror/trivy-db:2"
|
||||||
build_cpu_limit = "2"
|
build_cpu_limit = "2"
|
||||||
build_mem_limit = "4Gi"
|
build_mem_limit = "4Gi"
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Mirror the executor images into the registry once. On the node itself, push
|
||||||
|
through the loopback hostPort the registry Deployment binds (docker treats
|
||||||
|
`127.0.0.1` as insecure by default):
|
||||||
|
|
||||||
|
```sh
|
||||||
|
docker pull gcr.io/kaniko-project/executor:v1.24.0 # any versions you pin
|
||||||
|
docker pull aquasec/trivy:0.74.0
|
||||||
|
docker tag gcr.io/kaniko-project/executor:v1.24.0 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0
|
||||||
|
docker tag aquasec/trivy:0.74.0 127.0.0.1:5000/mirror/trivy:0.74.0
|
||||||
|
docker push 127.0.0.1:5000/mirror/kaniko-executor:v1.24.0
|
||||||
|
docker push 127.0.0.1:5000/mirror/trivy:0.74.0
|
||||||
|
```
|
||||||
|
|
||||||
|
From another machine, port-forward the registry instead (`kubectl -n felis
|
||||||
|
port-forward svc/registry 5000:5000`) and push to `localhost:5000/...` — the
|
||||||
|
registry keys a repository by the path after the host, so pushes through either
|
||||||
|
door land in the same place the build Pods will pull from.
|
||||||
|
|
||||||
Put them in **both** `/etc/felis/felis.host.toml` (host-side CLI) and
|
Put them in **both** `/etc/felis/felis.host.toml` (host-side CLI) and
|
||||||
`/etc/felis/felis.pod.toml` (the file rendered into the API's `felis-config`
|
`/etc/felis/felis.pod.toml` (the file rendered into the API's `felis-config`
|
||||||
Secret — the two differ only in the database URL; the setup screens re-render
|
Secret — the two differ only in the database URL; the setup screens re-render
|
||||||
@@ -436,12 +456,11 @@ build egress policy denies those hosts — so the scan step fails closed
|
|||||||
Kaniko pushed the image. Mirror the DB into the internal registry once:
|
Kaniko pushed the image. Mirror the DB into the internal registry once:
|
||||||
|
|
||||||
```
|
```
|
||||||
# On a host with internet + docker access to the cluster's registry
|
# On the node (docker treats 127.0.0.1 as insecure by default), or through the
|
||||||
# (add its address to the daemon's insecure-registries first; the registry
|
# port-forward above:
|
||||||
# serves plain HTTP):
|
|
||||||
# docker pull mirror.gcr.io/aquasec/trivy-db:2
|
# docker pull mirror.gcr.io/aquasec/trivy-db:2
|
||||||
# docker tag mirror.gcr.io/aquasec/trivy-db:2 <registry-addr>:5000/mirror/trivy-db:2
|
# docker tag mirror.gcr.io/aquasec/trivy-db:2 127.0.0.1:5000/mirror/trivy-db:2
|
||||||
# docker push <registry-addr>:5000/mirror/trivy-db:2
|
# docker push 127.0.0.1:5000/mirror/trivy-db:2
|
||||||
```
|
```
|
||||||
|
|
||||||
The Job's Trivy container already runs with `--insecure`, so the internal
|
The Job's Trivy container already runs with `--insecure`, so the internal
|
||||||
@@ -466,6 +485,16 @@ control namespace (or `--registry-namespace`):
|
|||||||
- **Storage:** PVC is RWO, `10Gi`, mounted at `/var/lib/registry`, **no
|
- **Storage:** PVC is RWO, `10Gi`, mounted at `/var/lib/registry`, **no
|
||||||
`storageClassName`** → binds the cluster default class. If the cluster has no
|
`storageClassName`** → binds the cluster default class. If the cluster has no
|
||||||
default StorageClass the PVC stays `Pending` and the registry never starts.
|
default StorageClass the PVC stays `Pending` and the registry never starts.
|
||||||
|
- **Node-side pulls:** containerd cannot dial the Service VIP (the live stack
|
||||||
|
answered "Empty reply"), so the registry Deployment binds a loopback hostPort
|
||||||
|
(`127.0.0.1:<port>`) and the installer writes a `/etc/rancher/k3s/registries.yaml`
|
||||||
|
mirror relaying `registry.<ns>.svc:<port>` onto it. That pair is what lets
|
||||||
|
kubelet re-pull a garbage-collected image; both halves must survive together
|
||||||
|
(remove either and every pull after an image GC fails).
|
||||||
|
- **Memory:** the registry's limit is 2Gi, deliberately larger than the other
|
||||||
|
control-plane pods' 256Mi — a live 475MB-layer push OOM-killed the 256Mi
|
||||||
|
template mid-upload (audit #46). Very large layers need headroom here, not
|
||||||
|
more CPU.
|
||||||
- **Selector quirk worth knowing:** the registry Service selector is only
|
- **Selector quirk worth knowing:** the registry Service selector is only
|
||||||
`name + component=registry` — it deliberately lacks the
|
`name + component=registry` — it deliberately lacks the
|
||||||
`part-of=felis-control-plane` label, so the registry is *invisible* to the
|
`part-of=felis-control-plane` label, so the registry is *invisible* to the
|
||||||
@@ -702,9 +731,11 @@ changes, because retention was never conditional in the first place.
|
|||||||
|
|
||||||
## 13b. Node runs out of disk: what survives, and how to recover
|
## 13b. Node runs out of disk: what survives, and how to recover
|
||||||
|
|
||||||
A full disk is the one failure this platform cannot ride out by itself, because
|
A full disk is the most destructive failure this stack sees: kubelet evicts game
|
||||||
the images exist only in the node's containerd (air-gapped by design), so a
|
pods (the control plane is protected below), and its image GC then collects
|
||||||
GC'd image has no pull source.
|
images nothing is running. The images have a pull source now — the in-cluster
|
||||||
|
registry — so they come back without an operator re-import; freeing space is
|
||||||
|
what completes the recovery.
|
||||||
|
|
||||||
**Eviction.** Every control-plane pod (api, operator, reaper, registry) runs
|
**Eviction.** Every control-plane pod (api, operator, reaper, registry) runs
|
||||||
under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
|
under the BUILT-IN `system-cluster-critical` PriorityClass (value 2e9). Kubelet's
|
||||||
@@ -725,26 +756,38 @@ management plane must be placeable.
|
|||||||
flapping); pods that need scheduling wait for it. This is the bulk of the
|
flapping); pods that need scheduling wait for it. This is the bulk of the
|
||||||
"recovery takes minutes" observation, not a stuck node.
|
"recovery takes minutes" observation, not a stuck node.
|
||||||
|
|
||||||
**The images may be gone.** If pods were evicted, the kubelet can garbage-collect
|
**The images may be gone — they come back on their own.** If pods were evicted,
|
||||||
their images (unused > 2 minutes under imagefs pressure). Those pods then sit in
|
the kubelet can garbage-collect their images (unused > 2 minutes under imagefs
|
||||||
`ImagePullBackOff`/`ErrImagePull` for a tag that plainly exists —
|
pressure). Every image this platform runs is ALSO hosted in the in-cluster
|
||||||
`k3s ctr images ls` shows it missing. Recovery:
|
registry: the installer builds each one as `registry.<ns>.svc:5000/felis/…`
|
||||||
|
and mirrors it there, and the node's containerd is configured (a registries.yaml
|
||||||
|
mirror onto the registry's loopback hostPort) to relay those refs back through
|
||||||
|
it. So a GC'd image is re-pulled on the next attempt with no operator action —
|
||||||
|
delete the stuck pod to force an immediate retry (or wait out the backoff), and
|
||||||
|
the workload converges.
|
||||||
|
|
||||||
|
If a pull does NOT come back:
|
||||||
|
|
||||||
1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are
|
1. Free disk on the node (`df -h /var/lib/rancher`; the biggest consumers are
|
||||||
`k3s ctr images ls -q` and the world/backup PVCs under
|
`k3s ctr images ls -q` and the world/backup PVCs under
|
||||||
`/var/lib/rancher/k3s/storage`).
|
`/var/lib/rancher/k3s/storage`).
|
||||||
2. Re-import the images by re-running the installer (it rebuilds/re-imports from
|
2. Check the registry: `kubectl -n felis get pods -l
|
||||||
the local Docker store, which the kubelet GC does not touch):
|
app.kubernetes.io/component=registry` and, on the node,
|
||||||
`curl -fsSL <installer URL> | sudo bash` (or `sudo felis setup`), then
|
`curl -s http://127.0.0.1:5000/v2/` (expect `{}`).
|
||||||
`kubectl -n felis rollout status deploy/felis-api`.
|
3. Check the mirror file: `/etc/rancher/k3s/registries.yaml` must map
|
||||||
3. Delete the stuck pods so they retry against the re-imported image.
|
`registry.felis.svc:5000` to `http://127.0.0.1:5000`. Missing or changed:
|
||||||
|
re-run the installer (it rewrites the file and restarts k3s only when the
|
||||||
|
content changed).
|
||||||
|
4. Re-mirror a tag the registry does not have (hand-built images were never
|
||||||
|
pushed): `docker tag <ref> 127.0.0.1:5000/<repo>:<tag> && docker push
|
||||||
|
127.0.0.1:5000/<repo>:<tag>`.
|
||||||
|
|
||||||
For a single image without a full installer run:
|
For an image that is in neither place, the old fallback still stands: re-run the
|
||||||
`docker save felis:<tag> | k3s ctr images import -` — the Docker store is
|
installer (it rebuilds/re-imports from the local Docker store AND mirrors into
|
||||||
deliberately a second copy; treat it as the recovery path, not as free space.
|
the registry), or for a single image
|
||||||
Verified end to end in the drill: `docker save felis-limbo:demo
|
`docker save felis:<tag> | k3s ctr images import -`. The Docker store remains a
|
||||||
felis-lobby:demo | k3s ctr images import -` plus pod deletion had both system
|
deliberate second copy on the node; treat it as the recovery path, not as free
|
||||||
servers Running ~25s later.
|
space.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -815,8 +858,9 @@ kubectl -n felis rollout status deploy/felis-api
|
|||||||
```
|
```
|
||||||
|
|
||||||
(the same for `felis-operator` and `registry`). `rollout undo` returns to the
|
(the same for `felis-operator` and `registry`). `rollout undo` returns to the
|
||||||
previous ReplicaSet, whose image is normally still on the node; if it was GC'd
|
previous ReplicaSet, whose image is normally still on the node; if the image GC
|
||||||
(§13b), re-import it first.
|
collected it, the registry re-serves it automatically (§13b) for every tag the
|
||||||
|
installer built — only hand-built tags need a manual re-mirror.
|
||||||
|
|
||||||
## Quick reference: symptom → section
|
## Quick reference: symptom → section
|
||||||
|
|
||||||
|
|||||||
Reference in new issue
Block a user