fix(platform): front the felis-api internal face on its own ClusterIP Service

The login limbo pod dials FELIS_API_BASE_URL = felis-api.<ns>.svc:8081 (the
internal face, service-token auth) to mint bind codes and poll link status, but
the only Service named felis-api is the external NodePort face and declares only
port 443. A Service answers only on its declared ports, so felis-api:8081 had no
backend and every login-pod internal call silently failed to connect.

Render a separate ClusterIP Service felis-api-internal for port 8081 and repoint
InternalAPIBaseURL at it. A second port on the NodePort Service is not an option:
Type=NodePort allocates a node port for every declared port with no per-port
opt-out, so it would publish the no-Zero-Trust internal face on every node's
external IP. A distinct ClusterIP Service keeps 8081 in-cluster only, reachable
by the login pod via DNS and by the on-node break-glass console via the
ClusterIP (exported as APIInternalServiceName / APIInternalPort).

Manifest-level fix; the live packet path is pending real-cluster verification.
This commit is contained in:
flyemoji committed 2026-07-07 11:24:09 +09:00
1 parent 73195ca48f
commit 2ba994889e
4 files changed
+111 -14

No files matched your search

+11
View File
@@ -280,6 +280,17 @@ The internal face (`--internal-addr :8081`, routes under
`/api/v1/internal/...`) is **never** Zero-Trust; it authenticates a single
service token via `Authorization: Bearer <token>`, compared in constant time.
In-cluster it is reached through the ClusterIP Service `felis-api-internal` (port
8081), which is separate from the external NodePort `felis-api` (443) precisely so
the no-Zero-Trust face is never exposed on a node. On the control-plane node the
break-glass console reaches it by resolving that Service's ClusterIP and dialing
`:8081`.
- **Internal calls fail to *connect* (not 401)** → the `felis-api-internal` Service
is missing or its selector no longer matches the api pods. `kubectl -n felis get
svc felis-api-internal` must show a ClusterIP with 8081; a bare `felis-api` name
serves only 443 and every internal call would hang/refuse.
- **All internal calls 401** → the token is unset or wrong. The API reads env
`FELIS_SERVICE_TOKEN`. If unset, startup logs: